Emotional Classification Method, Device, Equipment, Storage Medium and Program Product
By extracting visual feature of the target image and combining semantic features of emotional text, the similarity between the image and emotional text is calculated, and the problem of inaccurate emotional classification caused by excessive characters or complex scenes in the image is solved, and higher emotional classification accuracy is achieved.
Patent Information
- Application Number
- CN202311236970.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-09-23
AI Technical Summary
When using visual information to classify characters' emotions, there are too many characters in the image or complex scenes, resulting in inaccurate emotional classification.
By extracting visual features of the target image and combining the semantic features of multiple emotional texts, the similarity between the image and the emotional text is calculated, so as to select the target emotional text and use its corresponding emotional relationship category as the emotional classification of the characters in the image.
It effectively improves the accuracy of emotional classification of characters in the image, and improves the classification effect in complex scenes by combining visual information and semantic information.
Smart Images

Figure CN117726848B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer application technologies, and particularly to an emotion classification method, apparatus, device, storage medium, and program product. Background Art
[0002] In an image involving people, the emotional relationship of the people in the image can be recognized according to the interactions between the people in the image, scene information, etc., so as to assist people in better understanding the content of the image. Traditional emotion classification methods usually use visual information for relationship recognition. Specifically, first, the image and the pair of people with pre-discriminated relationships are input into a visual backbone network to calculate visual features, then the visual features of the people are extracted by Region of Interest Pooling (RoIPooling), and a graph convolutional network is used to perform feature information transfer on all the extracted people features to model the emotional relationship between the people. Finally, the features of the pair of people with pre-discriminated relationships are input into a classifier for emotion relationship classification.
[0003] However, in the process of using visual information for relationship recognition, if there are too many people in the image or the scene is complex, there will be a problem of inaccurate emotion classification of the people in the image. Therefore, how to improve the accuracy of emotion classification of the people in the image is a technical problem that needs to be solved urgently at present. Summary of the Invention
[0004] Embodiments of this application provide an emotion classification method, apparatus, device, storage medium, and program product, which can effectively improve the accuracy of emotion classification of the people in the image.
[0005] On the one hand, embodiments of this application provide an emotion classification method, which includes:
[0006] Extract features from a target image to obtain visual features of the target image;
[0007] Extract features from each of multiple emotion texts corresponding to the target image to obtain semantic features of each of the emotion texts;
[0008] Obtain the similarity between the target image and each of the emotion texts according to the visual features of the target image and the semantic features of each of the emotion texts;
[0009] Select a target emotion text from the multiple emotion texts according to the similarity between the target image and each of the emotion texts, and use the emotion relationship category corresponding to the target emotion text as the emotion classification of the people in the target image.
[0010] In one embodiment, the method further includes:
[0011] Perform visual vocabulary extraction on the target image to obtain the visual vocabulary of the target image;
[0012] Fuse the visual vocabulary of the target image with multiple preset emotional relationship categories respectively to obtain multiple emotional texts corresponding to the target image.
[0013] In one embodiment, the performing visual vocabulary extraction on the target image to obtain the visual vocabulary of the target image includes:
[0014] Obtain the similarity between each category feature in the text corpus and the target image;
[0015] Select target category features from the text corpus according to the similarity between each category feature and the target image;
[0016] Determine the target category features as the visual vocabulary of the target image.
[0017] In one embodiment, the obtaining the similarity between each category feature in the text corpus and the target image includes:
[0018] Traverse the text corpus in at least one dimension, and obtain the similarity between each category feature in the currently traversed text corpus and the target image;
[0019] The determining the target category features as the visual vocabulary of the target image includes:
[0020] After the traversal ends, determine the target category features selected from the text corpus in each dimension as the visual vocabulary of the target image.
[0021] In one embodiment, the performing feature extraction on the target image to obtain the visual features of the target image includes:
[0022] Determine the local area of the target person in the target image;
[0023] Perform feature extraction on the local area to obtain the visual features;
[0024] The taking the emotional relationship category corresponding to the target emotional text as the emotional classification of the person in the target image includes:
[0025] Take the emotional relationship category corresponding to the target emotional text as the emotional classification of the target person in the target image.
[0026] In one embodiment, the emotional classification of the person in the target image is determined by an emotional classification model; the training method of the emotional classification model includes:
[0027] Obtain training images and the class labels of the training images, where the class labels are used to indicate the emotion classification of the people in the training images;
[0028] Invoke an initial emotion classification model to extract features from the training images to obtain the visual features of the training images;
[0029] Extract features from each of the multiple training emotion texts corresponding to the training images to obtain the semantic features of each of the training emotion texts;
[0030] According to the visual features of the training images and the semantic features of each of the training emotion texts, obtain the similarity between the training images and each of the training emotion texts;
[0031] According to the similarity between the training images and each of the training emotion texts, select target training emotion texts from the multiple training emotion texts, and use the emotion relationship category corresponding to the target training emotion texts as the predicted emotion classification of the people in the training images;
[0032] Train the initial emotion classification model in the direction of reducing the difference between the predicted emotion classification and the emotion classification indicated by the class labels to obtain the emotion classification model.
[0033] In one embodiment, the multiple training emotion texts corresponding to the training images are obtained through an emotion text extraction model, and the training method of the emotion text extraction model includes:
[0034] Obtain the training images and multiple reference emotion texts of the training images;
[0035] Invoke an initial emotion text extraction model to extract visual vocabulary from the training images to obtain the visual vocabulary of the training images;
[0036] Fuse the visual vocabulary of the training images with multiple preset emotion relationship categories respectively to obtain multiple predicted emotion texts corresponding to the training images;
[0037] Train the initial emotion text extraction model in the direction of reducing the difference between the multiple predicted emotion texts and the corresponding reference emotion texts to obtain the emotion text extraction model.
[0038] On the other hand, an embodiment of the present application provides an emotion classification device, and the emotion classification device includes:
[0039] A feature extraction unit, configured to extract features from a target image to obtain the visual features of the target image;
[0040] The feature extraction unit is further configured to extract features from each of the multiple emotion texts corresponding to the target image to obtain semantic features of each of the emotion texts;
[0041] The similarity acquisition unit is configured to acquire the similarity between the target image and each of the emotion texts according to the visual features of the target image and the semantic features of each of the emotion texts;
[0042] The emotion classification unit is configured to select a target emotion text from the multiple emotion texts according to the similarity between the target image and each of the emotion texts, and use the emotion relationship category corresponding to the target emotion text as the emotion classification of the person in the target image.
[0043] On the other hand, an embodiment of the present application provides a computer device, including a processor, a storage device, and a communication interface. The processor, the storage device, and the communication interface are interconnected. Among them, the storage device is used to store a computer program that supports the computer device to execute the above method. The computer program includes program instructions. The processor is configured to call the program instructions to execute the following steps:
[0044] Extract features from the target image to obtain the visual features of the target image;
[0045] Extract features from each of the multiple emotion texts corresponding to the target image to obtain semantic features of each of the emotion texts;
[0046] According to the visual features of the target image and the semantic features of each of the emotion texts, acquire the similarity between the target image and each of the emotion texts;
[0047] According to the similarity between the target image and each of the emotion texts, select a target emotion text from the multiple emotion texts, and use the emotion relationship category corresponding to the target emotion text as the emotion classification of the person in the target image.
[0048] On the other hand, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute the above emotion classification method.
[0049] On the other hand, an embodiment of the present application provides a computer program product. The computer program product includes a computer program. The computer program is suitable for being loaded and executed by a processor to execute the above emotion classification method.
[0050] In the embodiments of the present application, by extracting features from each of the multiple emotion texts corresponding to the target image, the semantic features of each emotion text are obtained. According to the visual features of the target image and the semantic features of each emotion text, the similarity between the target image and each emotion text is obtained. According to the similarity between the target image and each emotion text, the emotion classification of the person in the target image is determined, which can make full use of the visual information and semantic information of the image, thereby effectively improving the accuracy of the emotion classification of the person in the image. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0052] Figure 1 is a schematic structural diagram of an emotion classification system provided by an embodiment of the present application;
[0053] Figure 2 is a schematic flowchart of an emotion classification method provided by an embodiment of the present application;
[0054] Figure 3 is a schematic flowchart of another emotion classification method provided by an embodiment of the present application;
[0055] Figure 4 is a schematic structural diagram of an emotion classification device provided by an embodiment of the present application;
[0056] Figure 5 is a schematic structural diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0058] The embodiments of the present application can perform emotion classification on the people in a single image. For example, a computer device can display a user interface. If a user wants to perform emotion classification on the people in a certain image, the user can click the emotion classification button in the user interface. Then, the computer device can respond to the user's click operation and display an image selection area in the user interface. The user can select the image on which emotion classification is to be performed in the image selection area. Among them, the image selected by the user can be an image stored in the local memory of the computer device, or an image stored in the cloud, or an image on the Internet, which is not specifically limited by the embodiments of the present application. Optionally, after the computer device displays the image selection area in the user interface, the user can drag the image on which emotion classification is to be performed to the image selection area. The computer device can respond to the user's selection operation or drag operation to obtain the image, and then perform emotion classification on the people in the image by using the emotion classification method provided by the embodiments of the present application.
[0059] Optionally, the embodiments of the present application can also perform emotion classification on the people in a video. For example, a computer device can display a video playback interface. The video playback interface can play any video, and the video playback interface can also include an emotion classification button. If a user wants to perform emotion classification on the people in a target video, the user can click on the target video. The computer device can respond to this click operation and display the target video in the video playback interface. Then the user can click the emotion classification button in the video playback interface to submit an emotion classification instruction for the target video in the video playback interface. The computer device can respond to the emotion classification instruction and perform emotion classification on the people in the target video. Optionally, when the computer device plays any video, it can detect whether emotion classification has been started. If emotion classification has been started, before playing any frame of the video, it can perform emotion classification on the people in that frame, and then when playing that frame, it can display the emotion classification of the people in that frame. Among them, whether emotion classification is started can be set by the user. For example, the video playback interface can include an emotion classification button. If the user clicks the emotion classification button when emotion classification has not been started, the computer device can respond to this click operation and start emotion classification; if the user clicks the emotion classification button after emotion classification has been started, the computer device can respond to this click operation and turn off emotion classification.
[0060] The emotion classification method provided by the embodiments of the present application can be applied in a client, a server or a computer device. The client or the server can include a video player, an emotion classification plugin, etc. The client or the server can be installed or integrated in a content publishing platform or a browser, and the content publishing platform or the browser can run on a computer device. The computer device includes but is not limited to a smart phone, a camera, a wearable device or a computer, etc.
[0061] In the specific embodiments of the present application, data related to users, such as images, etc., are involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of relevant data need to comply with local laws, regulations, and standards.
[0062] Please refer to Figure 1 , Figure 1 which is a schematic structural diagram of an emotion classification system provided by an embodiment of the present application. Exemplarily, an external text corpus (such as scene categories, types of human interactions, etc.) can be introduced, and through a pre-trained (Contrastive Language-Image Pre-Training, CLIP) model, the feature Encoder text (class_name) of each category in the text corpus is calculated for similarity with the input image Encoder visual (image), and the category with the highest similarity is selected as the visual vocabulary Visual_vocabs i of the image. For example, in this way, the visual vocabulary "home" of the input image is obtained from the place label library (such as the place365 scene dataset), and the visual vocabulary "watching TV" is obtained from the action label library (such as the kinetics action recognition dataset).
[0063] Furthermore, multiple emotion texts corresponding to the input image can be constructed. For example, multiple emotion relationship categories can be preset in advance. For example, the multiple preset emotion relationship categories can include "sorrow", "fear", "surprise", "happiness", "ecstasy", "rage", "vigilance", "hatred", etc. Then, a conditional text template P class , that is, an emotion text, can be constructed for each preset emotion relationship category. For example, according to the visual vocabulary "home" and "watching TV" and the emotion relationship category "happiness", the conditional text template "The background is in {home}, there is a {watching TV} behavior among the people, and the emotional relationship among the people is {happiness}" is constructed. Another example is that according to the visual vocabulary "home" and "watching TV" and the emotion relationship category "sorrow", the conditional text template "The background is in {home}, there is a {watching TV} behavior among the people, and the emotional relationship among the people is {sorrow}" is constructed, and so on.
[0064] Secondly, multiple constructed sentiment texts (such as T1, T2, T3) and the input image can be separately fed into the sentiment classification model to extract features, that is, feature extraction is performed on the input image to obtain the visual features of the input image, and feature extraction is performed on each sentiment text among the multiple sentiment texts corresponding to the input image to obtain the semantic features of each sentiment text. Then, according to the visual features of the input image and the semantic features of each sentiment text, the similarity between the input image and each sentiment text is obtained. According to the similarity between the input image and each sentiment text, a target sentiment text is selected from the multiple sentiment texts, and the sentiment relationship category corresponding to the target sentiment text is used as the sentiment classification of the person in the target image.
[0065] The sentiment classification solution provided by the embodiments of this application, by introducing an external text corpus (i.e., Figure 1 the sentiment-related external corpus in Figure 1 ), and the CLIP model (i.e., Figure 1 the CLIP text model in Figure 1 ), calculates the similarity between each category feature in the text corpus and the input image (i.e.,
[0066] the visual-text similarity calculation in
[0067] ), and obtains the visual vocabulary of the input image through the similarity calculation, thereby constructing the sentiment text corresponding to the input image (i.e., Figure 2 ), Figure 2 the sentiment relationship classification of the person in Figure 2The emotional classification scheme shown includes but is not limited to steps S201 to S205, where:
[0068] S201, extract features from the target image to obtain the visual features of the target image.
[0069] The visual features may include one or more of color features, shape features, texture features, and spatial position relationship features. The color feature is a global feature that describes the surface properties of the scenery corresponding to the image or image region. The shape feature is represented by contour features and region features. The contour feature mainly targets the outer boundary of the object, while the region feature relates to the entire shape region. The texture feature is a global feature that reflects the visual features of homogeneous phenomena in the image and embodies the arrangement attributes of the surface tissue structure with slow transformation or periodic changes on the object surface. The spatial relationship feature describes the spatial position relationship of the objects in the image.
[0070] Exemplarily, the texture features of the image can be extracted by the Local Binary Patterns (LBP) algorithm. The shape features of the image can be extracted by the Histogram of Oriented Gradient (HOG) algorithm. The spatial position relationship features of the image can be extracted by the Scale Invariant Feature Transform (SIFT) algorithm. The color features of the image can be extracted by color moments or color histograms, etc.
[0071] In one implementation, feature extraction can be performed on the entire image region of the target image to obtain all the visual features of the target image. Optionally, feature extraction can also be performed on a specified region of the target image to obtain partial visual features of the target image.
[0072] In one implementation, the local region of the target person in the target image can be determined, feature extraction is performed on the local region to obtain visual features, feature extraction is performed on each of the multiple emotional texts corresponding to the target image to obtain the semantic features of each emotional text, the similarity between the target image and each emotional text is obtained based on the visual features of the target image and the semantic features of each emotional text, the target emotional text is selected from the multiple emotional texts according to the similarity between the target image and each emotional text, and the emotional relationship category corresponding to the target emotional text is used as the emotional classification of the target person in the target image.
[0073] For example, if the target image includes multiple people and the user wants to perform emotion classification on a certain person or certain people in the target image, then the user can submit an emotion classification instruction for the target person. For example, the user can submit an emotion classification instruction for the target person by means of selecting a certain person or certain people in the target image through frame selection, and then the person selected by the user through frame selection can be determined as the target person. Another example is that the people in the target image can be recognized, the recognized people can be labeled, and the identifiers of the labeled people are displayed. The user can submit an emotion classification instruction for the target person by clicking on the identifier of a certain person or certain people, and then the person corresponding to the identifier clicked by the user can be determined as the target person. Then, feature extraction can be performed on each of the multiple emotion texts corresponding to the target image to obtain the semantic features of each emotion text. According to the visual features of the target image and the semantic features of each emotion text, the similarity between the target image and each emotion text is obtained. According to the similarity between the target image and each emotion text, a target emotion text is selected from the multiple emotion texts, and the emotion relationship category corresponding to the target emotion text is used as the emotion classification of the target person in the target image.
[0074] Optionally, the local area of the target person in the target image and the target area associated with the local area can be determined, feature extraction is performed on the local area and the target area to obtain visual features, feature extraction is performed on each of the multiple emotion texts corresponding to the target image to obtain the semantic features of each emotion text. According to the visual features of the target image and the semantic features of each emotion text, the similarity between the target image and each emotion text is obtained. According to the similarity between the target image and each emotion text, a target emotion text is selected from the multiple emotion texts, and the emotion relationship category corresponding to the target emotion text is used as the emotion classification of the target person in the target image.
[0075] Among them, the target area refers to the area used to assist in analyzing the emotion classification of the target person. For example, it may include the area where objects (such as a TV set, a sofa, etc.) are located in the target image, the area where other people whose distance from the target person is less than a preset distance threshold are located, the area where other people interacting with the target person are located, and so on.
[0076] S202. Feature extraction is performed on each of the multiple emotion texts corresponding to the target image to obtain the semantic features of each emotion text.
[0077] In one implementation, visual word extraction can be performed on the target image to obtain the visual words of the target image, and the visual words of the target image are respectively fused with multiple preset emotion relationship categories to obtain multiple emotion texts corresponding to the target image.
[0078] For example, assume that the visual vocabulary of the target image includes "home" and "watching TV". Based on the visual vocabulary "home" and "watching TV", and the emotional relationship category "happy", an emotional text can be constructed as "The background is in {home}, there is a behavior of {watching TV} among the characters, and the emotional relationship among the characters is {happy}". Another example, based on the visual vocabulary "home" and "watching TV", and the emotional relationship category "sorrowful", another emotional text can be constructed as "The background is in {home}, there is a behavior of {watching TV} among the characters, and the emotional relationship among the characters is {sorrowful}", and so on.
[0079] In one implementation, the similarity between each category feature in the text corpus and the target image can be obtained. Based on the similarity between each category feature and the target image, the target category feature is selected from the text corpus, and the target category feature is determined as the visual vocabulary of the target image.
[0080] Among them, the text corpus contains a large number of rich category features. The text corpus can include, for example, the place365 scene dataset and the kinetics action recognition dataset, etc. The similarity between each category feature in the place365 scene dataset and the kinetics action recognition dataset and the target image can be calculated. Based on the similarity between each category feature and the target image, the target category feature is selected from the text corpus, and the target category feature is determined as the visual vocabulary of the target image. For example, the category feature with a similarity greater than the preset ratio threshold to the target image can be determined as the target category feature. Another example, the category features can be sorted in descending order of similarity to the target image, and the top k category features are determined as the target category features, where k is a positive integer and k can be a preset value, such as 3 or 10, etc. Another example, the category feature with the largest similarity to the target image in each dataset can be determined as the target category feature. Exemplarily, the category feature with the largest similarity to the target image in the place365 scene dataset is "home", and the category feature with the largest similarity to the target image in the kinetics action recognition dataset is "watching TV", then the target category features can include "home" and "watching TV".
[0081] Optionally, since the CLIP model contains sufficient semantic information, the similarity between each category feature in the text corpus and the target image can be calculated through the text end network of the pre-trained CLIP model. Based on the similarity between each category feature and the target image, the target category feature is selected from the text corpus, and the target category feature is determined as the visual vocabulary of the target image.
[0082] In one implementation, it is possible to traverse a text corpus of at least one dimension, obtain the similarity between each category feature in the currently traversed text corpus and the target image, and select the target category features from the currently traversed text corpus according to the similarity between each category feature in the currently traversed text corpus and the target image. After the traversal is completed, the target category features selected from the text corpora of each dimension are determined as the visual vocabulary of the target image.
[0083] Among them, the text corpus may include text corpora of at least one dimension. For example, it may include a text corpus of the scene dimension (such as the place365 scene dataset), a text corpus of the action dimension (such as the kinetics action recognition dataset), a text corpus of the expression dimension, a text corpus of the object category dimension, a text corpus of the human pose dimension, and so on. Among them, the text corpus of the scene dimension is used to identify the scene of the target image. For example, the background of the target image is at home, or outdoors, or at a specific location (such as Shanghai Disney). The text corpus of the action dimension is used to identify the actions of the people in the target image. For example, the people in the target image are watching TV, walking, taking pictures, etc. The text corpus of the expression dimension is used to identify the expressions of the people in the target image. For example, a certain person in the target image is happy, sad or frightened. The text corpus of the object category dimension is used to identify the categories of the objects in the target image. For example, the target image includes a TV set, a sofa, etc. The text corpus of the human pose dimension is used to identify the poses of the people in the target image. For example, person A and person B in the target image are sitting on the sofa.
[0084] In one implementation, multiple emotion texts corresponding to the target image can be obtained through an emotion text extraction model. The training method of the emotion text extraction model may include: obtaining training images and multiple reference emotion texts of the training images; calling an initial emotion text extraction model to perform visual vocabulary extraction on the training images to obtain the visual vocabulary of the training images; fusing the visual vocabulary of the training images with multiple preset emotion relationship categories respectively to obtain multiple predicted emotion texts corresponding to the training images; training the initial emotion text extraction model in the direction of reducing the difference between the multiple predicted emotion texts and the corresponding reference emotion texts to obtain the emotion text extraction model. The emotion text extraction model trained by the above training method can accurately extract multiple emotion texts corresponding to the target image. It can be understood that the emotion text extraction model is called to perform visual vocabulary extraction on the target image to obtain the visual vocabulary of the target image; the visual vocabulary of the target image is fused with multiple preset emotion relationship categories respectively to obtain multiple predicted emotion texts corresponding to the target image.
[0085] S203. Obtain the similarity between the target image and each sentiment text according to the visual features of the target image and the semantic features of each sentiment text.
[0086] Optionally, the target image and multiple sentiment texts corresponding to the target image can be input into the sentiment classification model. The sentiment classification model extracts features from the target image to obtain the visual features of the target image, and extracts features from each sentiment text among the multiple sentiment texts corresponding to the target image to obtain the semantic features of each sentiment text. Then, according to the visual features of the target image and the semantic features of each sentiment text, obtain the similarity between the target image and each sentiment text. According to the similarity between the target image and each sentiment text, select the target sentiment text from the multiple sentiment texts, and use the sentiment relationship category corresponding to the target sentiment text as the sentiment classification of the person in the target image.
[0087] Among them, the sentiment classification model can include a vision and text network. For example, it can be a Convolutional Neural Networks (CNN) or a Long Short-Term Memory (LSTM), etc., which is not specifically limited by the embodiments of the present application.
[0088] In one implementation, the training method of the sentiment classification model can include: obtaining training images and the category labels of the training images, where the category labels are used to indicate the sentiment classification of the people in the training images; calling the initial sentiment classification model to extract features from the training images to obtain the visual features of the training images; extracting features from each training sentiment text among the multiple training sentiment texts corresponding to the training images to obtain the semantic features of each training sentiment text; obtaining the similarity between the training images and each training sentiment text according to the visual features of the training images and the semantic features of each training sentiment text; selecting the target training sentiment text from the multiple training sentiment texts according to the similarity between the training images and each training sentiment text, and using the sentiment relationship category corresponding to the target training sentiment text as the predicted sentiment classification of the people in the training images; training the initial sentiment classification model in the direction of reducing the difference between the predicted sentiment classification and the sentiment classification indicated by the category labels to obtain the sentiment classification model.
[0089] The sentiment classification model trained by the above training method can accurately classify the sentiment of the people in the target image.
[0090] In one implementation, the multiple training emotion texts corresponding to the training images are obtained through an emotion text extraction model. The training method of the emotion text extraction model may include: obtaining the training images and multiple reference emotion texts of the training images; calling an initial emotion text extraction model to perform visual vocabulary extraction on the training images to obtain the visual vocabulary of the training images; fusing the visual vocabulary of the training images with multiple preset emotion relationship categories respectively to obtain multiple predicted emotion texts corresponding to the training images; training the initial emotion text extraction model in the direction of reducing the difference between the multiple predicted emotion texts and the corresponding reference emotion texts to obtain the emotion text extraction model.
[0091] The emotion text extraction model trained by the above training method can accurately obtain multiple emotion texts corresponding to any image.
[0092] It can be understood that the emotion classification model and the emotion text extraction model can be two independent neural network models, or can be integrated into one neural network model, that is, this neural network model can not only construct multiple emotion texts corresponding to the target image, but also classify the emotion of the target person in the target image.
[0093] S204, select a target emotion text from multiple emotion texts according to the similarity between the target image and each emotion text.
[0094] For example, the emotion text with the highest similarity to the target image can be determined as the target emotion text. Another example is that the emotion text with a similarity greater than a preset similarity threshold to the target image can be determined as the target emotion text. Another example is that each emotion text can be sorted according to the size relationship of the similarity to the target image, and the emotion text ranked at the nth position is determined as the target emotion text, where n is a positive integer and n is a preset value, such as 1, or 2, or the last one, etc.
[0095] S205, use the emotion relationship category corresponding to the target emotion text as the emotion classification of the person in the target image.
[0096] Specifically, since the emotion text is obtained by fusing the visual vocabulary of the target image with the preset emotion relationship categories, the emotion relationship category corresponding to any emotion text is the emotion relationship category fused to obtain this emotion text. On this basis, after determining the target emotion text, the emotion relationship category corresponding to the target emotion text can be used as the emotion classification of the person in the target image.
[0097] In the embodiments of the present application, by extracting features from each of the multiple emotion texts corresponding to the target image, the semantic features of each emotion text are obtained. According to the visual features of the target image and the semantic features of each emotion text, the similarity between the target image and each emotion text is obtained. According to the similarity between the target image and each emotion text, the emotion classification of the person in the target image is determined, which can make full use of the visual information and semantic information of the image, thereby effectively improving the accuracy of emotion classification of the person in the image.
[0098] Based on the above description, please refer to Figure 3 , Figure 3 which is a schematic flowchart of another emotion classification method provided by the embodiments of the present application. This emotion classification method can be executed by a client, a server, or a computer device. As Figure 3 shown, the emotion classification method includes but is not limited to steps S301 to S312, where:
[0099] S301, obtain a training image and the class label of the training image, where the class label is used to indicate the emotion classification of the person in the training image.
[0100] S302, call the initial emotion classification model to extract features from the training image to obtain the visual features of the training image.
[0101] S303, extract features from each of the multiple training emotion texts corresponding to the training image to obtain the semantic features of each training emotion text.
[0102] In one implementation, the multiple training emotion texts corresponding to the training image are obtained through an emotion text extraction model. The training method of the emotion text extraction model may include: obtaining the training image and multiple reference emotion texts of the training image; calling the initial emotion text extraction model to perform visual vocabulary extraction on the training image to obtain the visual vocabulary of the training image; fusing the visual vocabulary of the training image with multiple preset emotion relationship categories respectively to obtain multiple predicted emotion texts corresponding to the training image; training the initial emotion text extraction model in the direction of reducing the difference between the multiple predicted emotion texts and the corresponding reference emotion texts to obtain the emotion text extraction model.
[0103] The emotion text extraction model trained by the above training method can accurately obtain multiple emotion texts corresponding to any image.
[0104] It can be understood that the sentiment classification model and the sentiment text extraction model can be two independent neural network models, or can be integrated to obtain a neural network model, that is, this neural network model can not only construct multiple sentiment texts corresponding to the target image, but also classify the sentiment of the target person in the target image.
[0105] S304. Obtain the similarity between the training image and each training sentiment text according to the visual features of the training image and the semantic features of each training sentiment text.
[0106] S305. Select a target training sentiment text from multiple training sentiment texts according to the similarity between the training image and each training sentiment text, and use the sentiment relationship category corresponding to the target training sentiment text as the predicted sentiment classification of the person in the training image.
[0107] S306. Train the initial sentiment classification model in the direction of reducing the difference between the predicted sentiment classification and the sentiment classification indicated by the category label to obtain a sentiment classification model.
[0108] The sentiment text extraction model trained by the above training method can accurately obtain multiple sentiment texts corresponding to any image.
[0109] S307. Extract visual words from the target image to obtain the visual words of the target image.
[0110] In one implementation, the similarity between each category feature in the text corpus and the target image can be obtained. According to the similarity between each category feature and the target image, a target category feature is selected from the text corpus, and the target category feature is determined as the visual word of the target image.
[0111] Among them, the text corpus contains a large number of rich category features. The text corpus can, for example, include the place365 scene dataset and the kinetics action recognition dataset, etc. The similarity between each category feature in the place365 scene dataset and the kinetics action recognition dataset and the target image can be calculated. According to the similarity between each category feature and the target image, the target category feature is selected from the text corpus, and the target category feature is determined as the visual vocabulary of the target image. For example, the category feature with a similarity greater than a preset ratio threshold to the target image can be determined as the target category feature. Another example is that the category features can be sorted in descending order of similarity to the target image, and the top k category features are determined as the target category features, where k is a positive integer and k can be a preset value, such as 3 or 10, etc. Another example is that the category feature with the highest similarity to the target image in each dataset can be determined as the target category feature. Exemplarily, the category feature with the highest similarity to the target image in the place365 scene dataset is "home", and the category feature with the highest similarity to the target image in the kinetics action recognition dataset is "watching TV", then the target category features can include "home" and "watching TV".
[0112] Optionally, since the CLIP model contains sufficient semantic information, the similarity between each category feature in the text corpus and the target image can be calculated through the text end network of the pre-trained CLIP model. According to the similarity between each category feature and the target image, the target category feature is selected from the text corpus, and the target category feature is determined as the visual vocabulary of the target image.
[0113] In one implementation, at least one dimension of the text corpus can be traversed to obtain the similarity between each category feature in the currently traversed text corpus and the target image. According to the similarity between each category feature in the currently traversed text corpus and the target image, the target category feature is selected from the currently traversed text corpus. After the traversal is completed, the target category features selected from the text corpus of each dimension are determined as the visual vocabulary of the target image.
[0114] Among them, the text corpus may include text corpora of at least one dimension. For example, it may include a text corpus of the scene dimension (such as the place365 scene dataset), a text corpus of the action dimension (such as the kinetics action recognition dataset), a text corpus of the expression dimension, a text corpus of the object category dimension, a text corpus of the human pose dimension, and so on. Among them, the text corpus of the scene dimension is used to identify the scene of the target image. For example, the background of the target image is at home, or outdoors, or at a specific location (such as Shanghai Disneyland), etc. The text corpus of the action dimension is used to identify the behavior of the people in the target image. For example, the people in the target image are watching TV, walking, taking pictures, etc. The text corpus of the expression dimension is used to identify the expression of the people in the target image. For example, a certain person in the target image is happy, sad, or frightened. The text corpus of the object category dimension is used to identify the category of the objects in the target image. For example, the target image includes a TV set, a sofa, etc. The text corpus of the human pose dimension is used to identify the pose of the people in the target image. For example, person A and person B in the target image are sitting on the sofa.
[0115] S308. Fuse the visual words of the target image with multiple preset emotional relationship categories respectively to obtain multiple emotional texts corresponding to the target image.
[0116] For example, assume that the visual words of the target image include "home" and "watching TV". According to the visual words "home" and "watching TV" and the emotional relationship category "happy", an emotional text can be constructed, which can be "The background is in {home}, there is a {watching TV} behavior among people, and the emotional relationship among people is {happy}". Another example, according to the visual words "home" and "watching TV" and the emotional relationship category "sorrowful", another emotional text can be constructed, which can be "The background is in {home}, there is a {watching TV} behavior among people, and the emotional relationship among people is {sorrowful}", and so on.
[0117] In one implementation, multiple emotional texts corresponding to the target image can be obtained through an emotional text extraction model. The training method of the emotional text extraction model may include: obtaining training images and multiple reference emotional texts of the training images; calling an initial emotional text extraction model to perform visual vocabulary extraction on the training images to obtain the visual vocabulary of the training images; fusing the visual vocabulary of the training images with multiple preset emotional relationship categories respectively to obtain multiple predicted emotional texts corresponding to the training images; training the initial emotional text extraction model in the direction of reducing the difference between the multiple predicted emotional texts and the corresponding reference emotional texts to obtain the emotional text extraction model. The emotional text extraction model trained through the above training method can accurately extract multiple emotional texts corresponding to the target image. It can be understood that the emotional text extraction model is called to perform visual vocabulary extraction on the target image to obtain the visual vocabulary of the target image; the visual vocabulary of the target image is fused with multiple preset emotional relationship categories respectively to obtain multiple predicted emotional texts corresponding to the target image.
[0118] S309, call an emotion classification model to perform feature extraction on the target image to obtain the visual features of the target image.
[0119] The visual features may include one or more of color features, shape features, texture features, and spatial position relationship features.
[0120] Exemplarily, the emotion classification model can extract the texture features of the image through the LBP algorithm. The emotion classification model can also extract the shape features of the image through the HOG algorithm. The emotion classification model can also extract the spatial position relationship features of the image through the SIFT algorithm. The emotion classification model can also extract the color features of the image through methods such as color moments or color histograms.
[0121] In one implementation, feature extraction can be performed on the entire image area of the target image to obtain all the visual features of the target image. Optionally, feature extraction can also be performed on a specified area of the target image to obtain partial visual features of the target image.
[0122] In one implementation, the local area of the target person in the target image can be determined, feature extraction is performed on the local area to obtain visual features, feature extraction is performed on each of the multiple emotional texts corresponding to the target image to obtain the semantic features of each emotional text, the similarity between the target image and each emotional text is obtained according to the visual features of the target image and the semantic features of each emotional text, the target emotional text is selected from the multiple emotional texts according to the similarity between the target image and each emotional text, and the emotional relationship category corresponding to the target emotional text is used as the emotion classification of the target person in the target image.
[0123] For example, if the target image includes multiple persons and the user wants to perform emotion classification on a certain person or certain persons in the target image, then the user can submit an emotion classification instruction for the target person. For example, the user can submit an emotion classification instruction for the target person by means of selecting a certain person or certain persons in the target image by drawing a box around them, and then the person or persons selected by the user can be determined as the target person. Also, for example, the persons in the target image can be recognized, the recognized persons can be labeled, and the identifiers of the labeled persons can be displayed. The user can submit an emotion classification instruction for the target person by clicking on the identifier of a certain person or certain persons, and then the person corresponding to the identifier clicked by the user can be determined as the target person. Then, feature extraction can be performed on each of the multiple emotion texts corresponding to the target image to obtain the semantic features of each emotion text. Based on the visual features of the target image and the semantic features of each emotion text, the similarity between the target image and each emotion text can be obtained. Based on the similarity between the target image and each emotion text, a target emotion text can be selected from the multiple emotion texts, and the emotion relationship category corresponding to the target emotion text can be used as the emotion classification of the target person in the target image.
[0124] Optionally, the local region of the target person in the target image and the target region associated with the local region can be determined, feature extraction can be performed on the local region and the target region to obtain visual features, feature extraction can be performed on each of the multiple emotion texts corresponding to the target image to obtain the semantic features of each emotion text. Based on the visual features of the target image and the semantic features of each emotion text, the similarity between the target image and each emotion text can be obtained. Based on the similarity between the target image and each emotion text, a target emotion text can be selected from the multiple emotion texts, and the emotion relationship category corresponding to the target emotion text can be used as the emotion classification of the target person in the target image.
[0125] Among them, the target region refers to the region used to assist in analyzing the emotion classification of the target person. For example, it can include the region where the objects (such as a TV set, a sofa, etc.) in the target image are located, the region where other persons whose distance from the target person is less than a preset distance threshold are located, the region where other persons interacting with the target person are located, and so on.
[0126] S310. Feature extraction is performed on each of the multiple emotion texts corresponding to the target image to obtain the semantic features of each emotion text.
[0127] S311. Based on the visual features of the target image and the semantic features of each emotion text, the similarity between the target image and each emotion text is obtained.
[0128] S312. Select a target emotion text from multiple emotion texts according to the similarity between the target image and each emotion text, and use the emotion relationship category corresponding to the target emotion text as the emotion classification of the person in the target image.
[0129] For example, the emotion text with the highest similarity to the target image can be determined as the target emotion text. Another example is that the emotion text with a similarity greater than a preset similarity threshold to the target image can be determined as the target emotion text. Another example is that each emotion text can be sorted according to the size relationship of the similarity to the target image, and the emotion text ranked at the nth position is determined as the target emotion text, where n is a positive integer and n is a preset value, such as 1, or 2, or the last one, etc.
[0130] In addition, since the emotion text is obtained by fusing the visual vocabulary of the target image with the preset emotion relationship categories, the emotion relationship category corresponding to any emotion text is the emotion relationship category obtained by fusion. On this basis, after determining the target emotion text, the emotion relationship category corresponding to the target emotion text can be used as the emotion classification of the person in the target image.
[0131] In the embodiments of the present application, by extracting the visual vocabulary of the target image, the visual vocabulary of the target image is obtained, and the visual vocabulary of the target image is respectively fused with multiple preset emotion relationship categories to obtain multiple emotion texts corresponding to the target image, which can maximize the semantic information valuable for judging the emotional relationship between people. In addition, according to the visual features of the target image and the semantic features of each emotion text, the similarity between the target image and each emotion text is obtained, and according to the similarity between the target image and each emotion text, the emotion classification of the person in the target image is determined, which can make full use of the semantic information of the text label (i.e., the emotion text) and the visual information of the target image, thereby effectively improving the accuracy of the emotion classification of the person in the target image.
[0132] The embodiments of the present application also provide a computer storage medium, in which program instructions are stored, and when the program instructions are executed, they are used to implement the corresponding methods described in the above embodiments.
[0133] Please refer to again Figure 4 , Figure 4 which is a schematic structural diagram of an emotion classification device provided by the embodiments of the present application.
[0134] In one implementation of the emotion classification device of the embodiments of the present application, the emotion classification device includes the following structure.
[0135] A feature extraction unit 401, configured to extract features from the target image to obtain the visual features of the target image;
[0136] The feature extraction unit 401 is further configured to extract features from each of the multiple emotion texts corresponding to the target image to obtain semantic features of each of the emotion texts;
[0137] The similarity acquisition unit 402 is configured to acquire the similarity between the target image and each of the emotion texts according to the visual features of the target image and the semantic features of each of the emotion texts;
[0138] The emotion classification unit 403 is configured to select a target emotion text from the multiple emotion texts according to the similarity between the target image and each of the emotion texts, and use the emotion relationship category corresponding to the target emotion text as the emotion classification of the person in the target image.
[0139] In one embodiment, the emotion classification device may further include a visual vocabulary extraction unit 404 and a fusion unit 405, where:
[0140] The visual vocabulary extraction unit 404 is configured to extract visual vocabulary from the target image to obtain the visual vocabulary of the target image;
[0141] The fusion unit 405 is configured to fuse the visual vocabulary of the target image with multiple preset emotion relationship categories respectively to obtain multiple emotion texts corresponding to the target image.
[0142] In one embodiment, when the visual vocabulary extraction unit 404 extracts visual vocabulary from the target image to obtain the visual vocabulary of the target image, it includes:
[0143] Acquire the similarity between each category feature in the text corpus and the target image;
[0144] Select a target category feature from the text corpus according to the similarity between each category feature and the target image;
[0145] Determine the target category feature as the visual vocabulary of the target image.
[0146] In one embodiment, when the similarity acquisition unit 402 acquires the similarity between each category feature in the text corpus and the target image, it includes:
[0147] Traverse the text corpus in at least one dimension to acquire the similarity between each category feature in the currently traversed text corpus and the target image;
[0148] The determining the target category feature as the visual vocabulary of the target image includes:
[0149] After the traversal ends, the target category features selected from the text corpora of each dimension are determined as the visual vocabulary of the target image.
[0150] In one embodiment, when the feature extraction unit 401 extracts features from a target image to obtain the visual features of the target image, it includes:
[0151] Determine the local area of the target person in the target image;
[0152] Extract features from the local area to obtain the visual features;
[0153] Regarding the emotional relationship category corresponding to the target emotional text as the emotional classification of the person in the target image includes:
[0154] Regarding the emotional relationship category corresponding to the target emotional text as the emotional classification of the target person in the target image.
[0155] In one embodiment, the emotional classification of the person in the target image is determined by an emotional classification model; the emotional classification device may further include a training image acquisition unit 406 and a training unit 407, where:
[0156] The training image acquisition unit 406 is configured to acquire training images and the category labels of the training images, where the category labels are used to indicate the emotional classification of the persons in the training images;
[0157] The feature extraction unit 401 is further configured to call an initial emotional classification model to extract features from the training images to obtain the visual features of the training images;
[0158] The feature extraction unit 401 is further configured to extract features from each of the multiple training emotional texts corresponding to the training images to obtain the semantic features of each of the training emotional texts;
[0159] The similarity acquisition unit 402 is further configured to obtain the similarity between the training images and each of the training emotional texts according to the visual features of the training images and the semantic features of each of the training emotional texts;
[0160] The emotional classification unit 403 is further configured to select a target training emotional text from the multiple training emotional texts according to the similarity between the training images and each of the training emotional texts, and regard the emotional relationship category corresponding to the target training emotional text as the predicted emotional classification of the person in the training image;
[0161] A training unit 407, configured to train the initial sentiment classification model in a direction of reducing the difference between the predicted sentiment classification and the sentiment classification indicated by the category label, so as to obtain the sentiment classification model.
[0162] In one embodiment, the multiple training sentiment texts corresponding to the training image are obtained through a sentiment text extraction model. The sentiment classification device may further include a visual vocabulary extraction unit 404 and a fusion unit 405, where:
[0163] The training image acquisition unit 406 is further configured to acquire the training image and multiple reference sentiment texts of the training image;
[0164] The visual vocabulary extraction unit 404 is configured to call an initial sentiment text extraction model to perform visual vocabulary extraction on the training image, so as to obtain the visual vocabulary of the training image;
[0165] The fusion unit 405 is configured to fuse the visual vocabulary of the training image with multiple preset sentiment relationship categories respectively, so as to obtain multiple predicted sentiment texts corresponding to the training image;
[0166] The training unit 407 is further configured to train the initial sentiment text extraction model in a direction of reducing the difference between the multiple predicted sentiment texts and the corresponding reference sentiment texts, so as to obtain the sentiment text extraction model.
[0167] In the embodiment of the present application, the feature extraction unit 401 extracts features of each sentiment text in the multiple sentiment texts corresponding to the target image to obtain semantic features of each sentiment text. The similarity acquisition unit 402 acquires the similarity between the target image and each sentiment text according to the visual features of the target image and the semantic features of each sentiment text. The sentiment classification unit 403 determines the sentiment classification of the person in the target image according to the similarity between the target image and each sentiment text, which can make full use of the visual information and semantic information of the image, thereby effectively improving the accuracy of the sentiment classification of the person in the image.
[0168] Please refer to Figure 5 , Figure 5 is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device in the embodiment of the present application includes structures such as a power supply module, and includes a processor 501, a storage device 502, and a communication interface 503. Data can be exchanged between the processor 501, the storage device 502, and the communication interface 503, and the corresponding three-dimensional human body reconstruction method is implemented by the processor 501.
[0169] The storage device 502 may include volatile memory, such as random-access memory (RAM); the storage device 502 may also include non-volatile memory, such as flash memory, solid-state drive (SSD), etc.; the storage device 502 may further include a combination of the above types of memories.
[0170] The processor 501 may be a central processing unit (CPU). The processor 501 may also be a combination of a CPU and a GPU. In a server, multiple CPUs and GPUs may be included as needed for corresponding sentiment classification. In one embodiment, the storage device 502 is used to store program instructions. The processor 501 may call the program instructions to implement various methods involved in the above embodiments of the present application.
[0171] In the first possible implementation manner, the processor 501 of the computer device calls the program instructions stored in the storage device 502 to extract features from the target image to obtain the visual features of the target image; extract features from each of the multiple sentiment texts corresponding to the target image to obtain the semantic features of each sentiment text; obtain the similarity between the target image and each sentiment text according to the visual features of the target image and the semantic features of each sentiment text; select a target sentiment text from the multiple sentiment texts according to the similarity between the target image and each sentiment text, and use the sentiment relationship category corresponding to the target sentiment text as the sentiment classification of the person in the target image.
[0172] In one embodiment, the processor 501 is further configured to perform the following operations:
[0173] Extract visual words from the target image to obtain the visual words of the target image;
[0174] Fuse the visual words of the target image with multiple preset sentiment relationship categories respectively to obtain multiple sentiment texts corresponding to the target image.
[0175] In one embodiment, when the processor 501 extracts visual words from the target image to obtain the visual words of the target image, the following operations may be performed:
[0176] Obtain the similarity between each category feature in the text corpus and the target image;
[0177] Select target category features from the text corpus according to the similarity between each category feature and the target image;
[0178] Determine the target category features as the visual vocabulary of the target image.
[0179] In one embodiment, when the processor 501 obtains the similarity between each category feature in the text corpus and the target image, the following operations may be performed:
[0180] Traverse the text corpus in at least one dimension, and obtain the similarity between each category feature in the currently traversed text corpus and the target image;
[0181] The determining the target category features as the visual vocabulary of the target image includes:
[0182] After the traversal is completed, determine the target category features selected from the text corpora in each dimension as the visual vocabulary of the target image.
[0183] In one embodiment, when the processor 501 extracts features from the target image to obtain the visual features of the target image, the following operations may be performed:
[0184] Determine the local area of the target person in the target image;
[0185] Extract features from the local area to obtain the visual features;
[0186] The taking the emotional relationship category corresponding to the target emotional text as the emotional classification of the person in the target image includes:
[0187] Take the emotional relationship category corresponding to the target emotional text as the emotional classification of the target person in the target image.
[0188] In one embodiment, the emotional classification of the person in the target image is determined by an emotional classification model; the processor 501 is further configured to perform the following operations:
[0189] Obtain training images and the category labels of the training images, where the category labels are used to indicate the emotional classification of the people in the training images;
[0190] Call an initial emotional classification model to extract features from the training images to obtain the visual features of the training images;
[0191] Extract features from each of the multiple training emotional texts corresponding to the training images to obtain the semantic features of each of the training emotional texts;
[0192] Obtain the similarity between the training image and each training emotion text according to the visual features of the training image and the semantic features of each training emotion text;
[0193] Select a target training emotion text from the multiple training emotion texts according to the similarity between the training image and each training emotion text, and use the emotion relationship category corresponding to the target training emotion text as the predicted emotion classification of the person in the training image;
[0194] Train the initial emotion classification model in the direction of reducing the difference between the predicted emotion classification and the emotion classification indicated by the category label to obtain the emotion classification model.
[0195] In one embodiment, the multiple training emotion texts corresponding to the training image are obtained through an emotion text extraction model, and the processor 501 is further configured to perform the following operations:
[0196] Obtain the training image and multiple reference emotion texts of the training image;
[0197] Call the initial emotion text extraction model to perform visual vocabulary extraction on the training image to obtain the visual vocabulary of the training image;
[0198] Fuse the visual vocabulary of the training image with multiple preset emotion relationship categories respectively to obtain multiple predicted emotion texts corresponding to the training image;
[0199] Train the initial emotion text extraction model in the direction of reducing the difference between the multiple predicted emotion texts and the corresponding reference emotion texts to obtain the emotion text extraction model.
[0200] In the embodiment of the present application, the processor 501 extracts the features of each emotion text in the multiple emotion texts corresponding to the target image to obtain the semantic features of each emotion text, obtains the similarity between the target image and each emotion text according to the visual features of the target image and the semantic features of each emotion text, and determines the emotion classification of the person in the target image according to the similarity between the target image and each emotion text, which can make full use of the visual information and semantic information of the image, thereby effectively improving the accuracy of the emotion classification of the person in the image.
[0201] Those of ordinary skill in the art can understand that to implement all or part of the processes in the above-described embodiment methods, it can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above various methods. Among them, the computer-readable storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc. The computer-readable storage medium mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of blockchain nodes, etc.
[0202] The foregoing disclosure only shows some embodiments of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Those of ordinary skill in the art can understand the implementation of all or part of the processes of the above embodiments, and the equivalent changes made according to the claims of the present application still fall within the scope covered by the present invention.
Claims
1. A method for sentiment classification, characterized in that, Including: Performing feature extraction on a target image to obtain visual features of the target image; Based on the similarity between each category feature in a text corpus and the target image, determining a target category feature as the visual vocabulary of the target image, where the text corpus includes text corpora of at least one dimension, and determining the target category features selected from the text corpora of each dimension as the visual vocabulary of the target image; Fusing the visual vocabulary of the target image with multiple preset emotional relationship categories respectively to obtain multiple emotional texts corresponding to the target image; Performing feature extraction on each emotional text among the multiple emotional texts corresponding to the target image to obtain semantic features of each emotional text; According to the visual features of the target image and the semantic features of each emotional text, obtaining the similarity between the target image and each emotional text; According to the similarity between the target image and each emotional text, selecting a target emotional text from the multiple emotional texts, and using the emotional relationship category corresponding to the target emotional text as the emotional classification of the person in the target image.
2. The method according to claim 1, wherein The step of determining a target category feature as the visual vocabulary of the target image based on the similarity between each category feature in a text corpus and the target image includes: For any dimension of the text corpus, obtaining the similarity between each category feature in the text corpus of that dimension and the target image; According to the similarity between each category feature and the target image, selecting a target category feature from the text corpus of that dimension; Determining the target category feature as the visual vocabulary of the target image.
3. The method according to claim 1, characterized in that The step of determining a target category feature as the visual vocabulary of the target image based on the similarity between each category feature in a text corpus and the target image includes: Traversing the at least one dimension of the text corpus, and obtaining the similarity between each category feature in the currently traversed text corpus and the target image; According to the similarity between each category feature and the target image, selecting a target category feature from the currently traversed text corpus; After the traversal is completed, determining the target category features selected from the text corpora of each dimension as the visual vocabulary of the target image.
4. The method according to claim 1, wherein The step of performing feature extraction on a target image to obtain visual features of the target image includes: Determining a local area of a target person in the target image; Performing feature extraction on the local area to obtain the visual features; The step of using the emotional relationship category corresponding to the target emotional text as the emotional classification of the person in the target image includes: Using the emotional relationship category corresponding to the target emotional text as the emotional classification of the target person in the target image.
5. The method according to claim 1, characterized in that, The emotional classification of the person in the target image is determined by an emotional classification model; The training method of the emotional classification model includes: Obtaining training images and category labels of the training images, where the category labels are used to indicate the emotional classification of the people in the training images; Call the initial emotion classification model to extract features from the training image to obtain the visual features of the training image; Extract features from each of the multiple training emotion texts corresponding to the training image to obtain the semantic features of each of the training emotion texts; According to the visual features of the training image and the semantic features of each of the training emotion texts, obtain the similarity between the training image and each of the training emotion texts; According to the similarity between the training image and each of the training emotion texts, select the target training emotion text from the multiple training emotion texts, and use the emotion relationship category corresponding to the target training emotion text as the predicted emotion classification of the person in the training image; Train the initial emotion classification model in the direction of reducing the difference between the predicted emotion classification and the emotion classification indicated by the category label to obtain the emotion classification model.
6. The method according to claim 5, wherein The multiple training emotion texts corresponding to the training image are obtained through an emotion text extraction model, and the training method of the emotion text extraction model includes: Obtain the training image and multiple reference emotion texts of the training image; Call the initial emotion text extraction model to extract visual vocabulary from the training image to obtain the visual vocabulary of the training image; Fuse the visual vocabulary of the training image with multiple preset emotion relationship categories respectively to obtain multiple predicted emotion texts corresponding to the training image; Train the initial emotion text extraction model in the direction of reducing the difference between the multiple predicted emotion texts and the corresponding reference emotion texts to obtain the emotion text extraction model.
7. An emotion classification device, characterized in that The device includes: A feature extraction unit for extracting features from a target image to obtain the visual features of the target image; A visual vocabulary extraction unit for determining the target category feature as the visual vocabulary of the target image based on the similarity between each category feature in the text corpus and the target image, where the text corpus includes text corpora of at least one dimension, and the target category features selected from the text corpora of each dimension are determined as the visual vocabulary of the target image; A fusion unit for fusing the visual vocabulary of the target image with multiple preset emotion relationship categories respectively to obtain multiple emotion texts corresponding to the target image; The feature extraction unit is further configured to extract features from each of the multiple emotion texts corresponding to the target image to obtain the semantic features of each of the emotion texts; A similarity acquisition unit for obtaining the similarity between the target image and each of the emotion texts according to the visual features of the target image and the semantic features of each of the emotion texts; An emotion classification unit for selecting a target emotion text from the multiple emotion texts according to the similarity between the target image and each of the emotion texts, and using the emotion relationship category corresponding to the target emotion text as the emotion classification of the person in the target image.
8. A computer device, characterized in that, The computer device includes a processor, a storage device, and a communication interface, and the processor, the storage device, and the communication interface are connected to each other, where: The storage device is used to store a computer program, and the computer program includes program instructions; The processor is used to call the program instructions to execute the sentiment classification method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a processor, they are used to execute the sentiment classification method according to any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes a computer program, and the computer program is adapted to be loaded and executed by a processor to execute the sentiment classification method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Image processing method and terminal equipment
CN115294150A
Emotion analysis method and device, electronic equipment and storage medium
CN116541520A