Emotion detection method and device, electronic equipment and storage medium
By extracting the media information and inputting it into multiple emotion detection models, and determining the emotion classification based on decision weights, the problem that a single model is difficult to accurately detect complex emotions is solved, and the accuracy and robustness of emotion detection are improved.
Patent Information
- Application Number
- CN202510100502.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, it is difficult to achieve ideal accuracy through a single emotion detection model, especially when processing complex emotions in media information.
A emotion detection method is proposed. By extracting the feature of the detection media information, inputting the feature information into a pre-trained multiple emotion detection models, and determining the target emotion classification based on the decision weight.
Through the combination of multiple emotion detection models, complex emotions in media information can be detected more comprehensively, biases during detection of a single model, and accuracy and robustness of emotion detection can be improved.
Smart Images

Figure CN119989095A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of information processing technology, and in particular to an emotion detection method, device, electronic device and storage medium. Background Art
[0002] In today's era of information explosion, people are exposed to a large amount of media information every day, which contains rich emotional information.
[0003] In the prior art, in order to detect emotions in media information, a single emotion detection model is usually used to determine the emotion type of the media information. However, in the process of implementing the present invention, it is found that the prior art has at least the following technical problems: due to the complexity of emotions contained in media information, it is often difficult to achieve ideal accuracy through a single emotion detection model. Summary of the invention
[0004] The embodiments of the present invention provide an emotion detection method, an apparatus, an electronic device and a storage medium to achieve the purpose of improving the accuracy and robustness of emotion detection.
[0005] According to one aspect of the present invention, there is provided an emotion detection method, comprising:
[0006] When the media information to be detected is obtained, a feature extraction operation is performed on the media information to be detected to obtain feature information of the media information to be detected;
[0007] Input the feature information into different pre-trained emotion detection models respectively to obtain output results of each emotion detection model; wherein the output results include the correlation degree value between the media information to be detected and each preset emotion type;
[0008] Based on the predetermined decision weights corresponding to each of the emotion detection models and the output results, a target emotion classification of the media information to be detected is determined.
[0009] According to another aspect of the present invention, there is provided an emotion detection device, the device comprising:
[0010] A feature extraction module is used to perform a feature extraction operation on the media information to be detected to obtain feature information of the media information to be detected when the media information to be detected is obtained;
[0011] An information input module, used to input the feature information into different pre-trained emotion detection models respectively, and obtain the output result of each emotion detection model; wherein the output result includes the correlation degree value between the media information to be detected and each preset emotion type;
[0012] The target emotion classification determination module is used to determine the target emotion classification of the media information to be detected based on the predetermined decision weights corresponding to each emotion detection model and the output results.
[0013] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0014] at least one processor; and
[0015] a memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the emotion detection method described in any embodiment of the present invention.
[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the emotion detection method described in any embodiment of the present invention when executed.
[0018] The technical solution of the embodiment of the present invention, when the media information to be detected is obtained, performs a feature extraction operation on the media information to be detected to obtain feature information of the media information to be detected; the feature information is respectively input into different pre-trained emotion detection models to obtain the output results of each emotion detection model, thereby performing emotion detection on the media information to be detected through multiple different emotion detection models, which is conducive to comprehensively detecting the complex emotions contained in the media information to be detected; and, based on the decision weights and output results corresponding to each predetermined emotion detection model, determines the target emotion classification of the media information to be detected. The technical solution of this embodiment determines the target emotion classification through decision weights combined with the output results of different emotion detection models, reduces the detection deviation generated when a single emotion detection model is used for detection, and is conducive to improving the accuracy and robustness of emotion detection.
[0019] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0021] Figure 1 is a flow chart of an emotion detection method provided according to an embodiment of the present invention;
[0022] Figure 2 is a flow chart of another emotion detection method provided according to an embodiment of the present invention;
[0023] Figure 3 is a structural schematic diagram of an emotion detection device provided according to an embodiment of the present invention;
[0024] Figure 4 It is a schematic diagram of the structure of an electronic device that implements the emotion detection method of an embodiment of the present invention. DETAILED DESCRIPTION
[0025] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "etc." and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products, or devices.
[0027] It should be noted that the collection, collection, updating, analysis, processing, use, transmission, storage and other aspects of user personal information involved in the technical solution of this disclosure are in compliance with the provisions of relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken for user personal information to prevent illegal access to user personal information data and maintain the security of user personal information and network security.
[0028] Figure 1 This is a flow chart of an emotion detection method provided according to an embodiment of the present invention. This embodiment is applicable to the case of detecting the emotion type reflected by the transmitted media information. The method can be executed by an emotion detection device, which can be implemented in the form of hardware and / or software.
[0029] like Figure 1 As shown, the method of this embodiment may specifically include:
[0030] S110 . When the media information to be detected is obtained, a feature extraction operation is performed on the media information to be detected to obtain feature information of the media information to be detected.
[0031] The media information to be detected is media information of a sentiment classification to be detected, and the media information to be detected may include at least one of text, image and video.
[0032] In order to facilitate feature extraction processing, when the media information to be detected is text, feature information of the media information to be detected can be directly extracted; when the media information to be detected is an image or video, the media information to be detected can be converted into text data. Specifically, feature extraction operation is performed on the media information to be detected to obtain feature information of the media information to be detected, including: when the media information to be detected is an image or video, the media information to be detected is converted into text data to be detected; preprocessing operation is performed on the text data to be detected, and the text data to be detected is updated based on the data after the preprocessing operation; and feature information of the updated text data to be detected is extracted.
[0033] In a specific implementation, when the media information to be detected is an image or video, the optical character recognition technology can be used to identify the text in the image or video, and the identified text constitutes the text data to be detected. When the media information to be detected is text, the media information to be detected can be directly used as the text data to be detected for subsequent processing.
[0034] To improve the accuracy of sentiment classification, preprocessing operations can be performed on the text data to be detected. Among them, the preprocessing operations include at least one of denoising processing, word segmentation processing, and stop word removal processing. Specifically, the text data to be detected can be first denoised, and the text data after denoising can be split into independent words to complete the word segmentation operation. To reduce interference with sentiment detection, stop words such as "de" and "le", punctuation marks, and special characters can be removed. Through the preprocessing operation, clean and effective text data is obtained, and the data after the preprocessing operation is updated as the text data to be detected. Finally, feature extraction is performed on the updated text data to be detected to obtain feature information.
[0035] In this embodiment, different information processing methods are provided for different forms of media information to be detected, which can extract features from different forms of media information to be detected; moreover, preprocessing operations are performed before feature extraction, so as to obtain clean and effective text data to be detected, which is beneficial to improving the accuracy of sentiment detection.
[0036] In this embodiment, extracting the feature information of the updated text data to be detected includes at least one of the following: extracting the feature information of the updated text data to be detected by the term frequency-inverse document frequency method; extracting the feature information of the updated text data to be detected based on a pre-determined bag-of-words model; extracting the feature information of the updated text data to be detected based on a pre-constructed word embedding model.
[0037] In this embodiment, three methods for extracting the feature information of the text data to be detected are provided. Optionally, the feature information of the updated text data to be detected can be extracted by the term frequency-inverse document frequency (TF-IDF) method. Specifically, the frequency of each word in the text data to be detected can be first determined, and then the inverse document frequency can be calculated. For each word, the term frequency and the inverse document frequency are multiplied to obtain the TF-IDF value of the word, and this TF-IDF value reflects the importance of the word in the text data to be detected. Based on the TF-IDF values, a matrix can be combined and used as the feature information of the text data to be detected. Or, a feature vector corresponding to the text data to be detected can be constructed through a bag-of-words model, and the vector obtained after normalizing this feature vector is used as the feature information of the text data to be detected. Or, the text data to be detected can be input into a pre-constructed word embedding model, and the information output by the word embedding model is used as the feature information of the text data to be detected.
[0038] In this embodiment, multiple methods for extracting the feature information of the text data to be detected are provided, thus improving the convenience of determining the feature information.
[0039] S120, inputting the feature information into different pre-trained emotion detection models respectively to obtain the output result of each emotion detection model.
[0040] The preset emotion type may include at least one of positive, negative and neutral.
[0041] It should be noted that the emotion detection model is used to detect the correlation value between each preset emotion type and the media information to be detected. The higher the correlation value, the higher the proportion of the preset emotion type contained in the media information to be detected; conversely, the lower the proportion of the preset emotion type contained in the media information to be detected.
[0042] Optionally, the sentiment detection model includes at least one of a script model, a dictionary model, a bidirectional transformer model, an efficient and accurate replacement word classification encoder model, an arrangement language model, and a review model. Exemplarily, the bidirectional transformer model includes a roberta-base model, the efficient and accurate replacement word classification encoder model can be an electra-large model, the arrangement language model can be an xlnet-base model, and the review model includes a dianping model.
[0043] In this embodiment, the feature information can be input into the script model, the dictionary model, the bidirectional transformer model, the efficient and accurate replacement word classification encoder model, the arrangement language model and the comment model respectively, and an output result is obtained for each emotion detection model. Optionally, the output result includes the correlation degree value between the media information to be detected and each preset emotion type. The correlation degree value can be expressed in the form of probability. For example, the feature information is input into the script model, and the output result of the script model includes: the first probability that the emotion reflected in the media information to be detected is classified as positive, the second probability that it is negative, and the third probability that it is neutral.
[0044] S130: Determine a target emotion classification of the media information to be detected based on a predetermined decision weight and output result corresponding to each emotion detection model.
[0045] In this embodiment, a decision weight may be pre-assigned to each emotion detection model, wherein the decision weight is used to reflect the importance of the emotion detection model; the greater the decision weight, the greater the role played by the emotion detection model corresponding to the decision weight in determining the target emotion classification; otherwise, the smaller the role played.
[0046] In a specific implementation, the target emotion classification of the media information to be detected can be determined based on the output results of the emotion detection model whose decision weight is greater than the preset weight value; or, for each emotion type, the decision weight of each emotion detection model can be used to perform a weighted sum of the correlation degree values of the preset emotion type output by each emotion detection model; based on the obtained sum value corresponding to each preset emotion type, the target emotion classification is determined; alternatively, the target emotion classification of the media information to be detected can be determined based on the output results of a preset number of emotion detection models with the largest decision weights.
[0047] In practical applications, alternative emotion types can be determined based on the output results of each emotion detection model. Specifically, for the output results of each emotion detection model, the preset emotion type with the largest correlation value is determined as the alternative emotion type. Among the alternative emotion types corresponding to each output result, the alternative emotion type with the largest number of repetitions is used as the target emotion type. For example, the output results of six emotion detection models give the alternative emotion types positive, positive, positive, positive, neutral and neutral, respectively. The number of repetitions of positive is 4, and the number of repetitions of neutral is 2. In this case, positive can be used as the target emotion type.
[0048] The technical solution of the embodiment of the present invention, when the media information to be detected is obtained, performs a feature extraction operation on the media information to be detected to obtain feature information of the media information to be detected; the feature information is respectively input into different pre-trained emotion detection models to obtain the output results of each emotion detection model, thereby performing emotion detection on the media information to be detected through multiple different emotion detection models, which is conducive to comprehensively detecting the complex emotions contained in the media information to be detected; and, based on the decision weights and output results corresponding to each predetermined emotion detection model, determines the target emotion classification of the media information to be detected. The technical solution of this embodiment determines the target emotion classification through decision weights combined with the output results of different emotion detection models, reduces the detection deviation generated when a single emotion detection model is used for detection, and is conducive to improving the accuracy and robustness of emotion detection.
[0049] Figure 2 is a flowchart of another emotion detection method provided according to an embodiment of the present invention. Based on the above embodiment, this embodiment further includes: determining the content length corresponding to the media information to be detected; based on the content length, determining the weight distribution method corresponding to the media information to be detected, and determining the decision weight of each emotion detection model according to the weight distribution method. The explanations of the terms that are the same or corresponding to the above embodiments are not repeated here. Figure 2 As shown, the method includes:
[0050] S210: When the media information to be detected is obtained, a feature extraction operation is performed on the media information to be detected to obtain feature information of the media information to be detected.
[0051] In practical applications, sentiment classification detection can be performed on the media information to be detected that meets the preset conditions. The preset conditions include at least one of the following: the source region of the media information to be detected is a preset region; the generation time of the media information to be detected is a preset time period; the belonging medium of the media information to be detected is a medium in a preset blacklist. Among them, the belonging medium can be at least one of a periodical, a newspaper, and a website. Through the preset conditions, this embodiment can screen out media information that meets the requirements of region, generation time, and belonging medium for detection, which is conducive to meeting the diverse needs of users.
[0052] In this embodiment, the preset emotion type includes at least one of positive, negative and neutral. Before performing feature extraction operation on the media information to be detected, it also includes: if the media information to be detected includes a title, extracting the title of the media information to be detected; comparing the title with the historical media information in the pre-stored historical media data set to determine the associated media information matching the title; performing feature extraction operation on the media information to be detected, including: if the standard emotion classification corresponding to the associated media information is negative, performing feature extraction operation on the media information to be detected.
[0053] The historical media data set includes historical media information and standard sentiment classification corresponding to the historical media information.
[0054] It should be noted that the media information to be detected includes information of the title and the text; or the media information to be detected only includes the text. Since the title can summarize the text and has less content, in order to improve the detection efficiency, the emotion type detection can be performed on the title before the feature extraction of the media information to be detected.
[0055] Specifically, in the case where the media information to be detected includes a title, the title of the media information to be detected can be extracted. In order to determine the sentiment classification corresponding to the title, the title can be compared with the historical media information in the historical media data set. Exemplarily, the degree of similarity between the title and the historical media information in the historical media data set can be determined. For example, the cosine similarity between the title and the title of the historical media information can be determined as the degree of similarity between the title and the historical media information. Alternatively, the pre-similarity between the title and the body of the historical media information can be determined as the degree of similarity between the title and the historical media information.
[0056] In a specific implementation, the historical media information with the highest degree of similarity to the title can be used as the associated media information, and the standard sentiment classification corresponding to the associated media information in the historical media data set can be used as the sentiment classification corresponding to the title. If the sentiment classification corresponding to the title is positive or neutral, the sentiment classification corresponding to the title can be directly used as the target sentiment classification of the media information to be detected, without performing a feature extraction operation on the media information to be detected. If the sentiment classification corresponding to the title is negative, in order to reduce the detection error of negative information, a feature extraction operation can be performed on the media information to be detected, and according to the above content, the extracted feature information can be input into different sentiment detection models to obtain the output results of each sentiment detection model, and then the target sentiment classification is determined based on the output results.
[0057] This embodiment first determines the sentiment classification corresponding to the title in the media information to be detected, and uses different methods to determine the target sentiment classification based on the different sentiment classifications corresponding to the titles, which is beneficial to improving the efficiency of determining the target sentiment classification and minimizing the detection error of negative media information.
[0058] S220, determining the content length corresponding to the media information to be detected; based on the content length, determining a weight distribution method corresponding to the media information to be detected, and determining the decision weight of each emotion detection model according to the weight distribution method.
[0059] Among them, the weight allocation method includes a long text allocation method and / or a short text allocation method; the long text allocation method is an allocation method that makes the decision weight of the first detection model greater than the decision weight of the second detection model; the short text allocation method is an allocation method that makes the decision weight of the second detection model greater than the decision weight of the first detection model; the first detection model is a pre-set emotion detection model corresponding to the long text, and the second detection model is a pre-set emotion detection model corresponding to the short text.
[0060] In practical applications, the detection accuracy of the emotion detection model is affected by the length of the text data corresponding to the media information to be detected, that is, the same emotion detection model has different detection accuracy for text data to be detected with different content lengths.
[0061] In order to improve the detection accuracy, the content length corresponding to the media information to be detected can be determined first, so as to detect the sentiment classification of the media information to be detected based on the content length. Specifically, the method for determining the content length corresponding to the media information to be detected is: when the media information to be detected is an image or a video, the media information to be detected is converted into text data to be detected, the number of characters contained in the text data to be detected is determined, and the number of characters is used as the content length corresponding to the media information to be detected. When the media information to be detected is text, the number of characters contained in the media information to be detected is directly determined, and the number of characters is used as the content length corresponding to the media information to be detected.
[0062] Further, based on the content length, the implementation method of determining the weight distribution method corresponding to the media information to be detected may include: when the content length is greater than a preset length threshold, the weight distribution method corresponding to the media information to be detected is determined as a long text distribution method; when the content length is less than or equal to the preset length threshold, the weight distribution method corresponding to the media information to be detected may be determined as a short text distribution method. Exemplarily, the preset length threshold may be 512 characters.
[0063] It should be noted that different sentiment detection models have different detection accuracy for texts of different lengths. Exemplarily, the review model, dictionary model and script model have higher detection accuracy for texts with longer content lengths; the bidirectional transformer model, efficient and accurate replacement word classification encoder model and arrangement language model have higher detection accuracy for texts with shorter content lengths, then the first detection model may include at least one of the review model, dictionary model and script model, and the second detection model includes one of the bidirectional transformer model, efficient and accurate replacement word classification encoder model and arrangement language model. By using a long text allocation method when the content length of the text corresponding to the media information to be detected is greater than a preset length threshold, the weight of the first detection model with higher detection accuracy for the long text is increased, so that the weight of the first detection model is higher than the weight of the second detection model, thereby improving the detection accuracy of the media information to be detected; and, when the content length corresponding to the media information to be detected is less than or equal to the preset length threshold, the weight of the second detection model is increased, so that the weight of the second detection model is higher than the weight of the first detection model, thereby improving the detection accuracy of the short text whose content length is less than or equal to the preset length threshold, that is, improving the detection accuracy of the media information to be detected.
[0064] S230, input the feature information into different pre-trained emotion detection models respectively to obtain the output result of each emotion detection model.
[0065] Optionally, the feature information is input into different pre-trained sentiment detection models to obtain the output results of each sentiment detection model, including: inputting the feature information into pre-trained script model, dictionary model, bidirectional transformer model, efficient and accurate replacement word classification encoder model, arrangement language model and comment model to obtain the output results corresponding to the script model, dictionary model, bidirectional transformer model, efficient and accurate replacement word classification encoder model, arrangement language model and comment model respectively.
[0066] It should be noted that the script model, dictionary model, bidirectional transformer model, efficient and accurate replacement word classification encoder model, permutation language model and comment model are all different types of models, which reduces the correlation between the various sentiment detection models and allows different sentiment detection models to learn or capture different features in the information, thereby improving the generalization ability, accuracy and efficiency of the sentiment detection model.
[0067] In this embodiment, the output results corresponding to the script model, the dictionary model, the bidirectional transformer model, the efficient and accurate replacement word classification encoder model, the arrangement language model and the comment model are all the correlation degree values of the media information to be detected being positive, negative and neutral, for example, the correlation degree value is a probability. The larger the correlation degree value, the higher the possibility that the media information to be detected is classified as the sentiment. For example, if the output result of the dictionary model is that the probability of positive is 0.1, the probability of negative is 0.5, and the probability of neutral is 0.4, it means that the dictionary model determines that the sentiment classification corresponding to the media information to be detected is negative.
[0068] This embodiment determines the output results through different emotion detection models, so that the various situations of emotion classification of the media information to be detected are comprehensively reflected through different output results, which is conducive to improving the generalization ability, accuracy and efficiency of the emotion detection model.
[0069] S240: Determine a target emotion classification of the media information to be detected based on the predetermined decision weights and output results corresponding to each emotion detection model.
[0070] In this embodiment, based on the decision weights and output results corresponding to each predetermined emotion detection model, the implementation method of determining the target emotion classification of the media information to be detected includes: for each preset emotion type, based on the decision weights corresponding to each predetermined emotion detection model, weighted averaging the correlation degree values corresponding to the preset emotion type in the output results of each emotion detection model, and using the obtained average value as the score of the preset emotion type; using the preset emotion type with the largest score as the target emotion classification of the media information to be detected.
[0071] In this embodiment, a score can be obtained for each preset emotion type, and the score is used to reflect the possibility that the media information to be detected is of the preset emotion type. The higher the score, the higher the possibility that the media information to be detected is of the preset emotion type; conversely, the lower the possibility that the media information to be detected is of the preset emotion type.
[0072] Exemplarily, the preset emotion types include positive, negative and neutral. A score is determined for each of the positive, negative and neutral. In order to more clearly illustrate the score determination process, take the positive as an example. For the preset emotion classification of positive, based on the output results of different emotion detection models, the probability that the emotion classification of the media information to be detected is positive determined by different emotion detection models is obtained. According to the decision weights corresponding to each predetermined emotion detection model, the probabilities that the emotion classification is positive in each output result are weighted and averaged, and the average value obtained is used as the score that the emotion classification of the media information to be detected is positive.
[0073] Furthermore, the maximum score indicates that the preset emotion type is most likely to be the emotion type contained in the media information to be detected, and thus the preset emotion type with the maximum score can be used as the target emotion classification of the media information to be detected.
[0074] This embodiment obtains the score of the preset emotion type by weighted averaging, thereby combining the detection results of different emotion detection models, which is conducive to improving the accuracy of the target emotion type.
[0075] Figure 3 1 is a schematic diagram of the structure of an emotion detection device provided according to an embodiment of the present invention, and the device is used to execute the emotion detection method provided in any of the above embodiments. The device and the emotion detection method of the above embodiments belong to the same inventive concept, and the details not described in detail in the embodiment of the emotion detection device can refer to the embodiment of the above emotion detection method. Figure 3 As shown, the device comprises:
[0076] The feature extraction module 10 is used to perform a feature extraction operation on the media information to be detected to obtain feature information of the media information to be detected when the media information to be detected is obtained;
[0077] The information input module 11 is used to input the feature information into different pre-trained emotion detection models to obtain the output results of each emotion detection model; wherein the output results include the correlation degree value between the media information to be detected and each preset emotion type;
[0078] The target emotion classification determination module 12 is used to determine the target emotion classification of the media information to be detected based on the predetermined decision weights and output results corresponding to each emotion detection model.
[0079] Based on any optional technical solution in the embodiment of the present invention, optionally, the feature extraction module 10 includes:
[0080] The information conversion submodule is used to convert the media information to be detected into text data to be detected when the media information to be detected is an image or a video;
[0081] A preprocessing submodule, used to perform a preprocessing operation on the text data to be detected, and update the text data to be detected based on the data after the preprocessing operation; wherein the preprocessing operation includes at least one of a denoising process, a word segmentation process, and a stop word removal process;
[0082] The information extraction submodule is used to extract the updated feature information of the text data to be detected.
[0083] Based on any optional technical solution in the embodiments of the present invention, optionally, the information extraction submodule includes at least one of the following:
[0084] A first extraction unit, used for extracting the updated feature information of the text data to be detected by using a word frequency inverse document frequency method;
[0085] A second extraction unit, used for extracting feature information of the updated text data to be detected based on a predetermined bag-of-words model;
[0086] The third extraction unit is used to extract the updated feature information of the text data to be detected based on the pre-built word embedding model.
[0087] On the basis of any optional technical solution in the embodiment of the present invention, optionally, the sentiment detection model includes at least one of a script model, a dictionary model, a bidirectional transformer model, an efficient and accurate replacement word classification encoder model, an arrangement language model and a comment model;
[0088] The information input module 11 includes:
[0089] The information input submodule is used to input feature information into the pre-trained script model, dictionary model, bidirectional transformer model, efficient and accurate replacement word classification encoder model, arrangement language model and comment model, and obtain the output results corresponding to the script model, dictionary model, bidirectional transformer model, efficient and accurate replacement word classification encoder model, arrangement language model and comment model respectively.
[0090] Based on any optional technical solution in the embodiment of the present invention, optionally, the target emotion classification determination module 12 includes:
[0091] A score determination submodule is used to, for each preset emotion type, perform weighted averaging of the association degree values corresponding to the preset emotion type in the output results of each emotion detection model based on the predetermined decision weights corresponding to each emotion detection model, and use the obtained average value as the score of the preset emotion type;
[0092] The target emotion classification determination submodule is used to select the preset emotion type with the largest score as the target emotion classification of the media information to be detected.
[0093] Based on any optional technical solution in the embodiment of the present invention, optionally, the preset emotion type includes at least one of positive, negative and neutral; and the device further includes:
[0094] A title extraction module, used for extracting the title of the media information to be detected before performing a feature extraction operation on the media information to be detected, if the media information to be detected contains a title;
[0095] An information comparison module is used to compare the title with the historical media information in the pre-stored historical media data set to determine the associated media information that matches the title; wherein the historical media data set includes the historical media information and the standard sentiment classification corresponding to the historical media information;
[0096] The feature extraction module 10 comprises:
[0097] The feature extraction submodule is used to perform feature extraction operations on the media information to be detected when the standard sentiment classification corresponding to the associated media information is negative.
[0098] Based on any optional technical solution in the embodiments of the present invention, optionally, the method further includes:
[0099] A content length determination module, used to determine the content length corresponding to the media information to be detected before determining the target emotion classification of the media information to be detected based on the decision weights and output results corresponding to each predetermined emotion detection model;
[0100] A decision weight determination module is used to determine a weight distribution method corresponding to the media information to be detected based on the content length, and determine the decision weight of each emotion detection model according to the weight distribution method;
[0101] Among them, the weight allocation method includes a long text allocation method and / or a short text allocation method; the long text allocation method is an allocation method that makes the decision weight of the first detection model greater than the decision weight of the second detection model; the short text allocation method is an allocation method that makes the decision weight of the second detection model greater than the decision weight of the first detection model; the first detection model is a pre-set emotion detection model corresponding to the long text, and the second detection model is a pre-set emotion detection model corresponding to the short text.
[0102] The technical solution of the embodiment of the present invention, when the media information to be detected is obtained, performs a feature extraction operation on the media information to be detected to obtain feature information of the media information to be detected; the feature information is respectively input into different pre-trained emotion detection models to obtain the output results of each emotion detection model, thereby performing emotion detection on the media information to be detected through multiple different emotion detection models, which is conducive to comprehensively detecting the complex emotions contained in the media information to be detected; and, based on the decision weights and output results corresponding to each predetermined emotion detection model, determines the target emotion classification of the media information to be detected. The technical solution of this embodiment determines the target emotion classification through decision weights combined with the output results of different emotion detection models, reduces the detection deviation generated when a single emotion detection model is used for detection, and is conducive to improving the accuracy and robustness of emotion detection.
[0103] It is worth noting that in the embodiment of the above-mentioned emotion detection device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.
[0104] Figure 4 : is a schematic diagram of the structure of an electronic device that implements the emotion detection method of an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0105] like Figure 4As shown, the electronic device 20 includes at least one processor 21, and a memory connected to the at least one processor 21, such as a read-only memory (ROM) 22, a random access memory (RAM) 23, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 21 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 22 or the computer program loaded from the storage unit 28 to the random access memory (RAM) 23. In RAM23, various programs and data required for the operation of the electronic device 20 can also be stored. The processor 21, ROM22 and RAM23 are connected to each other through a bus 24. An input / output (I / O) interface 25 is also connected to the bus 24.
[0106] A number of components in the electronic device 20 are connected to the I / O interface 25, including: an input unit 26, such as a keyboard, a mouse, etc.; an output unit 27, such as various types of displays, speakers, etc.; a storage unit 28, such as a disk, an optical disk, etc.; and a communication unit 29, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 29 allows the electronic device 20 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0107] The processor 21 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 21 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 21 performs the various methods and processes described above, such as the emotion detection method.
[0108] In some embodiments, the emotion detection method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 28. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 20 via the ROM 22 and / or the communication unit 29. When the computer program is loaded into the RAM 23 and executed by the processor 21, one or more steps of the emotion detection method described above may be performed. Alternatively, in other embodiments, the processor 21 may be configured to perform the emotion detection method in any other appropriate manner (e.g., by means of firmware).
[0109] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0110] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0111] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0112] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0113] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0114] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.
[0115] This embodiment also provides a computer program product, including a computer program, which, when executed by a processor, implements the emotion detection method provided in any embodiment of the present application.
[0116] In the process of implementation, the computer program product can be written in one or more programming languages or a combination thereof to perform the computer program code of the present invention, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect through the Internet).
[0117] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.
[0118] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for emotion detection, characterized in that: include: When the media information to be detected is obtained, a feature extraction operation is performed on the media information to be detected to obtain feature information of the media information to be detected; Input the feature information into different pre-trained emotion detection models respectively to obtain output results of each emotion detection model; wherein the output results include the correlation degree value between the media information to be detected and each preset emotion type; Based on the predetermined decision weights corresponding to each of the emotion detection models and the output results, a target emotion classification of the media information to be detected is determined.
2. The method according to claim 1, characterized in that The performing a feature extraction operation on the media information to be detected to obtain feature information of the media information to be detected includes: In the case where the media information to be detected is an image or a video, converting the media information to be detected into text data to be detected; Performing a preprocessing operation on the text data to be detected, and updating the text data to be detected based on the data after the preprocessing operation; wherein the preprocessing operation includes at least one of a denoising process, a word segmentation process, and a stop word removal process; Extract the updated feature information of the text data to be detected.
3. The method according to claim 2, characterized in that The step of extracting updated feature information of the text data to be detected includes at least one of the following: Extract the updated feature information of the text data to be detected by using the word frequency inverse document frequency method; Based on the predetermined bag-of-words model, extract the updated feature information of the text data to be detected; Based on the pre-built word embedding model, the feature information of the updated text data to be detected is extracted.
4. The method according to claim 1, characterized in that: The sentiment detection model includes at least one of a script model, a dictionary model, a bidirectional transformer model, an efficient and accurate replacement word classification encoder model, an arrangement language model, and a review model; The step of inputting the feature information into different pre-trained emotion detection models to obtain the output result of each emotion detection model includes: The feature information is input into a pre-trained script model, a dictionary model, a bidirectional transformer model, a high-efficiency and accurate replacement word classification encoder model, an arrangement language model and a review model, and the output results corresponding to the script model, the dictionary model, the bidirectional transformer model, the high-efficiency and accurate replacement word classification encoder model, the arrangement language model and the review model are obtained respectively.
5. The method according to claim 1, characterized in that: The step of determining a target emotion classification of the media information to be detected based on the predetermined decision weights corresponding to each emotion detection model and the output result includes: For each preset emotion type, based on the predetermined decision weight corresponding to each emotion detection model, weighted average is performed on the association degree values corresponding to the preset emotion type in the output results of each emotion detection model, and the obtained average value is used as the score of the preset emotion type; The preset emotion type with the largest score is used as the target emotion category of the media information to be detected.
6. The method according to claim 1, characterized in that The preset emotion type includes at least one of positive, negative and neutral; Before performing the feature extraction operation on the media information to be detected, the method further includes: In the case where the media information to be detected includes a title, extracting the title of the media information to be detected; Comparing the title with historical media information in a pre-stored historical media data set to determine associated media information matching the title; wherein the historical media data set includes the historical media information and a standard sentiment classification corresponding to the historical media information; The performing a feature extraction operation on the media information to be detected includes: When the standard sentiment classification corresponding to the associated media information is negative, a feature extraction operation is performed on the media information to be detected.
7. The method according to claim 1, characterized in that Before determining the target emotion classification of the media information to be detected based on the predetermined decision weights corresponding to each emotion detection model and the output results, the method further includes: Determine the content length corresponding to the media information to be detected; Based on the content length, determining a weight distribution method corresponding to the media information to be detected, and determining a decision weight of each of the emotion detection models according to the weight distribution method; Among them, the weight allocation method includes a long text allocation method and / or a short text allocation method; the long text allocation method is an allocation method that makes the decision weight of the first detection model greater than the decision weight of the second detection model; the short text allocation method is an allocation method that makes the decision weight of the second detection model greater than the decision weight of the first detection model; the first detection model is a pre-set emotion detection model corresponding to the long text, and the second detection model is a pre-set emotion detection model corresponding to the short text.
8. An emotion detection device, characterized in that: include: A feature extraction module is used to perform a feature extraction operation on the media information to be detected to obtain feature information of the media information to be detected when the media information to be detected is obtained; An information input module, used to input the feature information into different pre-trained emotion detection models respectively, and obtain the output result of each emotion detection model; wherein the output result includes the correlation degree value between the media information to be detected and each preset emotion type; The target emotion classification determination module is used to determine the target emotion classification of the media information to be detected based on the predetermined decision weights corresponding to each emotion detection model and the output results.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the emotion detection method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the emotion detection method described in any one of claims 1-7 when executed.