Emotion data capitalization method and system based on fine granularity emotion division
By processing multi-source text big data and building emotional data sets, the problem of difficulty in converting into emotional data assets in the existing technology is solved, efficient emotional data assetization is achieved, and the application value of data in various fields is enhanced.
Patent Information
- Application Number
- CN202510009943.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-03
AI Technical Summary
The existing technology is difficult to effectively transform multi-source text big data into emotional data assets with standardized, standardized, secure, and high reuse value, and cannot meet the practical application needs of governments, enterprises and research institutions.
By acquiring multi-source text big data, noise processing, standardization and anonymization are performed, fine-grained emotional data sets are constructed, and a general emotional data set is established based on the data set and asset-based processing is performed.
It realizes efficient transformation from original text data to emotional data assets, is practical and operable, and promotes the application and development of text data in various fields.
Smart Images

Figure CN120067290A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data assetization, and particularly to an emotion data assetization method and system based on fine-grained emotion division. Background Art
[0002] Currently, for the mining and analysis from the emotional perspective of text big data such as social media voices, netizen comments, product evaluations, complaints and reports, chat conversations, interactive suggestions, etc., there is a relatively urgent need for governments, enterprises, and research institutions. Converting text data into standardized, secure, highly reusable emotion data assets can not only meet the actual application needs of relevant institutions, but also conform to the current era background of the booming digital economy and the inevitable trend of data assetization. Summary of the Invention
[0003] The purpose of the present invention is to provide an emotion data assetization method and system based on fine-grained emotion division, aiming to solve the above problems in the prior art.
[0004] An embodiment of the present invention provides an emotion data assetization method based on fine-grained emotion division, including:
[0005] Obtaining multi-source text big data, and converting the text big data to obtain an initial data resource;
[0006] Constructing a fine-grained emotion data set based on the initial data resource, and using the fine-grained emotion data set as the first emotion data asset; and
[0007] Establishing a general emotion data set according to the first emotion data asset, using the general emotion data set as the second emotion data asset, and performing assetization processing on the first emotion data asset and the second emotion data asset.
[0008] An embodiment of the present invention provides an emotion data assetization system based on fine-grained emotion division, including:
[0009] A data module, configured to obtain multi-source text big data, and convert the text big data to obtain an initial data resource;
[0010] A first asset module, configured to construct a fine-grained emotion data set based on the initial data resource, and use the fine-grained emotion data set as the first emotion data asset; and
[0011] A second asset module, configured to establish a general emotion data set according to the first emotion data asset, use the general emotion data set as the second emotion data asset, and perform assetization processing on the first emotion data asset and the second emotion data asset.
[0012] An embodiment of the present invention further provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the above-mentioned emotional data assetization method based on fine-grained emotion division are implemented.
[0013] An embodiment of the present invention further provides a computer-readable storage medium, on which an implementation program for information transmission is stored. When the program is executed by a processor, the steps of the above-mentioned emotional data assetization method based on fine-grained emotion division are implemented.
[0014] The adoption of the embodiment of the present invention may include the following beneficial effects: The embodiment of the present invention proposes a method for efficiently converting text big data into emotional data assets. The entire process not only realizes the conversion from raw data to emotional data assets, but also has practical applicability and operability. The embodiment of the present invention will greatly promote the application and development of various types of text data in various fields. Description of the Drawings
[0015] In order to more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0016] Figure 1 It is a flowchart of the emotional data assetization method based on fine-grained emotion division according to an embodiment of the present invention;
[0017] Figure 2 It is a schematic diagram of the overall process according to an embodiment of the present invention;
[0018] Figure 3 It is a schematic diagram of the correspondence relationship between feature words and emotion types according to an embodiment of the present invention;
[0019] Figure 4 It is a schematic diagram of the important fields of data asset I according to an embodiment of the present invention;
[0020] Figure 5 It is a schematic diagram of the important fields of data asset II according to an embodiment of the present invention;
[0021] Figure 6 It is a schematic diagram of the emotional data assetization system based on fine-grained emotion division according to an embodiment of the present invention. Detailed Embodiments
[0022] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the following will clearly and completely describe the technical solutions in one or more embodiments of this specification in conjunction with the accompanying drawings in one or more embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this document.
[0023] Method Embodiment
[0024] According to an embodiment of the present invention, there is provided an emotional data assetization method based on fine-grained emotion division. Figure 1 It is a flowchart of the emotional data assetization method based on fine-grained emotion division according to an embodiment of the present invention. As Figure 1 shown, the emotional data assetization method based on fine-grained emotion division according to an embodiment of the present invention specifically includes:
[0025] Step S101, obtaining multi-source text big data, transforming the text big data to obtain initial data resources, specifically including:
[0026] Performing noise processing on the text big data, performing integrated normalization processing on the processed data according to standardized fields and converting it into a standard format, numbering the standardized data according to the subordinate relationship to generate corresponding identifiers, anonymizing the fields involving sensitive information in the data, encrypting and storing the anonymized data and adding permission management, and performing quality inspection and security inspection on the stored data in real time to obtain initial data resources;
[0027] Among them, the standardized fields include data content, data type, release time, release region, media source, like count, comment count, forward count, access count, favorite count, author name, and author gender;
[0028] The anonymization processing includes data desensitization, data generalization, data perturbation, and de-identification processing;
[0029] The quality inspection includes inspections of data format, integrity, accuracy, consistency, and uniqueness of the data, inspections of invalid data and duplicate data, and inspections of outliers;
[0030] The security inspection includes inspections of data encryption management, data access control, security vulnerability scanning and auditing, and data damage recovery;
[0031] Step S102: Construct a fine-grained sentiment dataset based on the initial data resource, and use the fine-grained sentiment dataset as the first sentiment data asset, specifically including:
[0032] Generate a semantic dependency graph for the initial data resource using the SDP and DEP semantic dependency algorithms. Search for language units in the semantic dependency graph through the depth-first search algorithm DFS and the breadth-first search algorithm BFS, and select the language units using the N-Gram language model according to the principle of maximizing joint probability to obtain the feature words to be encoded included in each piece of text after word segmentation. Encode the sentiment types of each piece of text according to the fine-grained sentiment knowledge graph, and obtain the sentiment feature words of each piece of text based on the sentiment type and its feature words to be encoded of each piece of text. Encode the sentiment attribute values of each piece of text through the sentiment feature words, and perform quality inspection and security inspection on the encoded sentiment data in real time to generate a fine-grained sentiment dataset;
[0033] The sentiment attribute values include the hit count, sentiment intensity, valence, sentiment arousal degree, and sentiment control sense of each sentiment in each piece of text;
[0034] Among them, encoding the sentiment attribute values of each piece of text through the sentiment feature words specifically includes:
[0035] Calculate the hit count of each sentiment in each piece of text using formula 1, and calculate the weighted hit count of interaction of each sentiment in each piece of text using formula 2; Obtain the sentiment intensity of each sentiment in each piece of text through any one of the calculation methods in formula 3, formula 4, or formula 5; Calculate the valence of each sentiment in each piece of text using formula 6; Obtain the sentiment arousal degree of each sentiment in each piece of text through any one of the calculation methods in formula 7, formula 8, or formula 9; Calculate the control type of each sentiment type in each piece of text using formula 10, and obtain the control sense of each sentiment in each piece of text through any one of the calculation methods in formula 11, formula 12, or formula 13;
[0036] count(k j ,x i )=y Formula 1;
[0037] S(k j ,x i )=L(k j ,x i )×count(k j ,x i ) Formula 2;
[0038]
[0039] intensity 2 (kj , x i ) = max({f(w t ) | t = 1, 2, 3, …, y}) Formula 4;
[0040]
[0041] valence(k j , x i ) = g(k j ) Formula 6;
[0042]
[0043]
[0044] arousal 3 (k j , x i ) = max({h(w t ) | t = 1, 2, 3, …, y}) Formula 9;
[0045] control nominal (k j , x i ) = m(k j ) Formula 10;
[0046]
[0047]
[0048] control ordinal,3 (k j , x i ) = m(k j ) × max({n(w t ) | t = 1, 2, 3, …, y}) Formula 13;
[0049] Among them, x i represents the i-th text among n texts. For text x i , the emotion k j (from k 1 , k 2 , … k 50 ) hits y feature words. S(k j , x i ) is the weighted hit number of a certain interaction volume of the emotion k j of the i-th text. count(k j , x i ) is the hit number of emotion feature words of emotion k j on the i-th text. L(kj ,x i ) is the weighted score of the interaction volume corresponding to this text, intensity(k j ,x i ) is the intensity of emotion k j on the i-th text. Subscripts 1, 2, and 3 are used to distinguish formula types. y is the hit count of emotion k j . f(w t ) is the intensity of a certain feature word, valence(k j ,x i ) represents the valence of emotion type k j on the i-th text. g(k j ) represents the valence of emotion type k j . arousal(k j ,x i ) represents the arousal level of emotion type k j on the i-th text. Subscripts 1, 2, and 3 are used to distinguish formula types. h(w t ) is the arousal level of a certain feature word, control nominal (k j ,x i ) represents the sense of emotion type control of emotion type k j on the i-th text. m(k j ) represents the control type of emotion type k j . control ordinal (k j ,x i ) represents the sense of control of emotion type k j on the i-th text. Subscripts 1, 2, and 3 are used to distinguish formula types. n(w t ) is the control degree of a certain feature word;
[0050] Step S103: Establish a general emotion dataset based on the first emotion data asset, use the general emotion dataset as the second emotion data asset, and perform assetization processing on the first emotion data asset and the second emotion data asset, specifically including:
[0051] Construct corresponding general emotion indicators based on the fine-grained emotion dataset and according to user needs, calculate each emotion attribute in the general emotion indicators based on the fine-grained emotion types and the relative contribution degrees of various types, and establish the general emotion dataset required by users using each emotion attribute;
[0052] The assetization processing includes emotion data asset registration, confirmation and empowerment of emotion data assets and corresponding trading entities, asset trading, and asset entry into the statement.
[0053] The following details the above technical solutions of the embodiments of the present invention in combination with the specific situation of the method for assetizing sentiment data based on fine-grained sentiment division in the embodiments of the present invention.
[0054] The embodiments of the present invention propose a solution for efficiently converting various types of text big data into sentiment data assets through four main steps. The first step is to resourceify the original data. Convert the original data into standard, standardized, and secure data resources, including data cleaning and preprocessing, and data quality and security inspection. The second step is to generate fine-grained sentiment data assets. Based on the fine-grained sentiment knowledge graph and model algorithms, encode the sentiment types and attributes of the data, and after data quality and security inspection, generate Data Asset 1: "Fine-grained Sentiment Dataset". The third step is to generate general sentiment data assets. According to the common application requirements of sentiment computing, such as "5 common sentiments", "satisfaction", "sense of security", "sentiment arousal degree", etc., use the algorithm model of general sentiment indicators to encode the general sentiment types and attributes of Data Asset 1 to meet the requirements of general sentiment indicators; after data quality and value evaluation, generate Data Asset 2: "General Sentiment Dataset". The fourth step is to market sentiment data assets. Specifically, it includes data asset registration, tabulation, and trading, etc.
[0055] The embodiments of the present invention use methods such as a fine-grained sentiment knowledge graph and classification models to convert sentiment data resources into data assets, and the overall steps are as Figure 2 shown. Specifically, it includes:
[0056] Step 1: Resourceify the original data
[0057] 1. Data cleaning
[0058] One is noise processing. Remove noise and outlier topics to achieve data validity, accuracy, and consistency. The topic can be a certain concept, event, policy, product, brand, etc.
[0059] The other is data standardization. Integrate and normalize multi-source text data, and at the same time convert the data into a standard format and standardize the data items. The data includes standardized fields such as "data content", "data type", "release time", "release region", "media source", "like count", "comment count", etc. (as shown in Table 1), and at the same time assign a data number to each data according to its subordinate relationship as its identifier, such as "total number", "main post number", etc.
[0060] Table 1 Field standardization rules
[0061]
[0062] Thirdly, anonymization and security. Anonymization processing is carried out on fields with sensitive personal information, including data desensitization, data generalization, data perturbation, and de-identification processing, etc. Fields such as "author name" and "author gender" are recoded, and the content only contains data types such as discrete numerical values that do not contain actual personal information; and through a preset string list, information such as irrelevant personal names and specific place names existing in the "data content" column is anonymized, effectively protecting personal privacy while retaining the useful information of the data, ensuring the value of the data in analysis and application.
[0063] The data is encrypted and stored and permission management is imposed, so that only authorized personnel can access the data, and relevant standards are complied with to ensure the compliance of data processing.
[0064] 2. Data quality inspection and security inspection
[0065] Data quality inspection and verification include checking the data format, as well as the integrity, accuracy, consistency, and uniqueness of the data, inspecting invalid data and duplicate data, and detecting outliers, etc. Security inspection includes checking and verifying data encryption management, data access control, security vulnerability scanning and auditing, data damage recovery, etc.
[0066] Based on the above two links, the messy raw data is transformed into standardized, anonymized, and secure data resources.
[0067] Step 2: Generate Data Asset I: "Fine-grained Sentiment Dataset"
[0068] 1. Sentiment type coding
[0069] First, intelligent word segmentation of the data is carried out by natural language processing. The entries of the fine-grained sentiment knowledge graph include feature words, abbreviations, expressions, symbols, sentence patterns and their combinations, hereinafter collectively referred to as feature words; the specific implementation includes the following sub-steps:
[0070] Using the SDP and DEP semantic dependency algorithms, the text data is mapped into a graph structure, and a semantic dependency graph is generated for the text data to be analyzed; among them, SDP (Semantic Dependency Parsing) and DEP (Dependency Parsing) are two important tasks in natural language processing.
[0071] Using two graph search algorithms, the depth-first DFS algorithm and the breadth-first BFS algorithm, to search for appropriate language units LU on the semantic dependency graph, and each language unit LU is a word segmentation result;
[0072] The N-Gram language model is adopted, and the language unit LU = {W 1 , W 2 , … W n} is selected according to the principle of maximizing the joint probability; P(W 1:n ) is the joint probability, which is specifically expressed as:
[0073]
[0074] Among them, W k is the feature word unit in the language unit group, k is the arrangement order of the feature word units, n is the upper bound of k, k = 1, 2, …, n; the relationship between each sentence and its language unit group satisfies the Markov relationship, and the language unit LU = {W 1 , W 2 , … W n} can be selected according to the principle of maximizing the joint probability; each feature word unit W k is not necessarily connected together.
[0075] After word segmentation according to the above steps, the feature words to be encoded contained in each text are obtained; then, each text data of the standardized data is encoded for the sentiment type, and the "feature word - sentiment type" correspondence of the single data, that is, the "hit situation" field, is obtained. This field has two possible values, 0 or 1. 0 means that the sentiment does not hit any feature words, and 1 means that at least one feature word is hit. As Figure 3 shown, the text in the figure hits k kinds of sentiments, and the number of feature words for each sentiment is y 1 , y 2 , …, y k . According to the fine-grained sentiment knowledge graph, the sentiment types hit by each text and the feature words to be encoded hit are obtained to get the sentiment feature words. This fine-grained sentiment knowledge graph contains 50 kinds of fine-grained sentiments, and the specific classification is as follows:
[0076] (1) Mood-related sentiments: Emotional experiences with low arousal level and long duration, including loneliness, numbness, calmness, depression, anxiety, decadence;
[0077] (2) Stress-related sentiments: Emotional experiences with high arousal level and limited duration, including happiness, fear, surprise, anger, sadness, excitement, relaxation, anxiety, nervousness, touch, alertness;
[0078] (3) Evaluative emotions: Emotional experiences generated after an individual evaluates an information event, which can be further classified into two categories according to whether they are directed towards oneself or the outside world. Among them, externally directed emotions include sympathy, indifference, ridicule, disgust, condemnation, resentment, suspicion, affection, satisfaction, praise, trust, contempt, gratitude, admiration, yearning, jealousy, helplessness, expectation, optimism, blessing, disappointment; internally directed emotions include depression, embarrassment, inferiority, shyness, guilt, pride, frustration, annoyance, palpitation, grievance, doubt, boredom.
[0079] 2. Emotional Attribute Value Encoding
[0080] Through the emotional feature words in the above steps, the hit count, intensity, valence, and arousal degree of each emotion in each text are obtained, and these features are collectively referred to as emotional attribute values.
[0081] Suppose there are n texts, for text x i (from x 1 , x 2 , … x n ) for emotion k j (from k 1 , k 2 , … k 50 ) hits y feature words, then the attribute values of emotion k of x i are calculated as follows: j A. Hit count (count)
[0082] count(k
[0083] , x j , x i ) = y (2);
[0084] That is, the sum of the number of emotion k j feature words.
[0085] In addition, there is a weighted hit count calculation method that takes into account the interaction volume of the text itself. The specific implementation method is as follows:
[0086] First, match the text to be analyzed with a fine-grained emotion dictionary to calculate the weighted score of the interaction volume of the text and the hit count of a certain emotion; secondly, calculate the weighted hit count of this interaction volume. The relevant formula is as follows:
[0087] S(k j , x i ) = L(k j , x i ) × count(k j , x i ) (3);
[0088] Among them, S(k j , x i) is the weighted hit count of a certain interaction volume for the i-th text, and this interaction volume can be behavioral indicators such as "likes", "forwards", "comments", etc. count(k j , the weighted hit count of a certain interaction volume for the i-th text, and this interaction volume can be behavioral indicators such as "likes", "forwards", "comments", etc. count(k j ,x i ) is the hit count of sentiment feature words of sentiment k j on the i-th text, L(k j ,x i ) is the weighted score of the interaction volume corresponding to this text, and this weighted score has the following calculation method:
[0089]
[0090] Among them, L(k j ,x i ) is the weighted score of the interaction volume of a single text, M is the adjustment coefficient, when M ≤ 0, the formula has no application value, when M > 140, it will cause the weighted result to exceed the reasonable range and approach the extreme, so the value range of M is specified as (0, 140]; Interaction is the interaction volume of a single text, Interaction > 0; a is the base of the logarithmic function and a > 1.
[0091] B. Sentiment intensity
[0092]
[0093] intensity 2 (k j ,x i ) = max({f(w t ) | t = 1, 2, 3, …, y}) (6);
[0094]
[0095] Among them, formula (5) is the arithmetic mean method, formula (6) is the maximum value method, formula (7) is the summation method, and users can choose any one or more of the above methods to calculate the sentiment intensity. Among them, intensity(k j ,x i ) is the sentiment intensity value of sentiment k j on the i-th text, the subscripts 1, 2, 3 are used to distinguish the formula types, y is the hit count of sentiment k j , and f(w t ) is the intensity of a certain feature word, which is defined in the knowledge graph. This intensity function f(w) accepts a sentiment feature word string and returns an integer in the range of [1, 5].
[0096] C. Valence
[0097] valence(k j ,x i ) = g(k j ) (8);
[0098] Each emotional feature word that is hit has an endogenous valence. g(k) accepts an emotional category string and returns -1, 0, or 1. 1 represents positive emotion, 0 represents neutral emotion, and -1 represents negative emotion; g(k) is defined in the knowledge graph; valence(k j ,x i ) represents the valence of emotional category k j in the i-th text, and g(k j ) represents the valence of emotional category k j .
[0099] D. Emotional arousal (arousal)
[0100]
[0101]
[0102] arousal 3 (k j ,x i ) = max({h(w t )|t = 1,2,3,…,y}) (11);
[0103] Among them, equation (9) is the summation method, equation (10) is the arithmetic mean method, and equation (11) is the maximum value method. The user can choose any one or more of the above methods to calculate the emotional arousal. arousal(k j ,x i ) represents the arousal of emotional category k j in the i-th text. The subscripts 1, 2, and 3 are used to distinguish the formula types. h(w t ) is the arousal of a certain feature word, and this function is defined in the knowledge graph. h(w t ) accepts an emotional category string and returns an integer in the range [1,5]. 1 represents extremely low arousal, and 5 represents extremely high arousal.
[0104] E. Emotional sense of control (control)
[0105] control nominal (k j ,x i ) = m(k j ) (12);
[0106]
[0107]
[0108] control ordinal,3 (k j ,x i ) = m(k j ) × max({n(w t ) | t = 1, 2, 3, …, y}) (15);
[0109] Equation (12) is the control sense of emotion types, and equations (13) to (15) are the emotion control senses; among them, equation (13) is the summation method, equation (14) is the arithmetic mean method, and equation (15) is the maximum value method. The user can choose any one or more of the above methods to calculate the emotion control sense. control nominal (k j ,x i ) represents the control sense of emotion type k j in the i-th text, and m(k j ) represents the control type of emotion type k j , which is defined in the knowledge graph. m(k) accepts a string of emotion types and returns -1, 0, or 1. 1 represents domination, 0 represents no domination relationship, and -1 represents being dominated; control ordinal (k j ,x i ) represents the control sense of emotion type k j in the i-th text. The subscripts 1, 2, and 3 are used to distinguish formula types. n(w t ) is the control degree of a certain feature word, which is defined in the knowledge graph. n(w t ) accepts a string of emotion types and returns an integer in the range [1, 5]. 1 represents a very low sense of domination (sense of being dominated), and 5 represents a very high sense of domination (sense of being dominated).
[0110] 3. Data Quality and Security Inspection
[0111] Quality inspection includes manual inspection of emotion category encoding (accuracy rate) and hit rate (the number of single data not recognized / the total number of data); security inspection includes access permission management, encrypted storage detection, and encrypted transmission detection.
[0112] In summary, after step 2, data asset I is obtained: fine-grained emotion dataset, as Figure 4 shown. The "hit situation" field corresponds to the result of emotion type encoding, and the other fields correspond to the results of emotion attribute encoding.
[0113] Step 3: Generate data asset II: "General Emotion Dataset"
[0114] Based on the fine-grained sentiment dataset obtained in Step 2, commonly used and general metrics in practical applications can also be quickly obtained, including but not limited to "5 common sentiments", "satisfaction", "sense of security", and "emotional arousal". The following introduces the calculation process of general sentiment metrics and several recommended metrics.
[0115] 1. Calculation of General Sentiment Metrics
[0116] Calculating general sentiment metrics requires considering the relative contribution degree of each sentiment attribute of the fine-grained sentiment types to this metric. That is, if there is a general sentiment metric Z, the values of its respective sentiment attributes are composed by the fine-grained sentiment set containing 50 elements according to their weights on Z. In addition to some commonly used metrics such as satisfaction and sense of security, users can also define other possible metrics independently. The specific algorithm for general sentiment metrics is as follows:
[0117] Suppose there are n texts, for text x i (from x 1 , x 2 , … x n ) for sentiment k j (j ∈ {1, 2, …, 50}), the respective sentiment attributes are defined by count(k j , x i ), S(k j , x j ), intensity(k j , x i ), arousal(k j , x i ) and control ordinal (k j , x i ). There are 5 attributes, so the respective sentiment attributes of the composite sentiment Z of x i are defined as the first-order tensor Z:
[0118]
[0119] Among them, Z is the set of the respective attributes of the general sentiment metric Z; Z is a first-order tensor, and both A and C are second-order tensors; A and C can be represented as 50×5 matrices, j ∈ {1, 2, 3,..., 50} and q ∈ {1, 2, 3, 4, 5}, the maximum value of j is the number of sentiment types, and the maximum value of q represents the number of non-categorical sentiment attributes. A reflects the values of the respective sentiment attributes of each sentiment type, and C reflects the weighted values of the respective sentiment attributes of each sentiment type in the general sentiment Z; Z can be represented as a column vector of size 5, which reflects the values of different sentiment attributes of the general sentiment Z.
[0120] To determine the value c of each element in C in the above formula jq : Without any prior assumptions, methods including but not limited to confirmatory factor analysis can be used to determine each c jq value.
[0121] When there are hypothesized relationships due to the application domain, the user can make prior assumptions about the relationships between the values of each c jq value, and the remaining unknown relationships can continue to be obtained by the above methods.
[0122] 2. General emotional attribute value encoding
[0123] After determining the coefficients and types of general emotions, it is necessary to calculate the relevant attributes of each piece of text. The following are several data sets obtained through attribute calculation in the embodiments of the present invention:
[0124] 2.1 General emotion: 5 common emotion data subsets
[0125] In the application, the emotional calculation process tends to classify emotions into 5 major categories: joy, sadness, anger, fear, and disgust (coarse granularity, referred to as common emotion types for short); classification calculation can be performed based on 50 fine-granularity emotion types to achieve attribute encoding of the five common emotions of the text. Some of the names of common emotion types are the same as those of fine-granularity emotion types. To distinguish them, the common emotion types are respectively named: 'joy', 'sadness', 'anger', 'fear', 'disgust' to distinguish from joy, sadness, anger, fear, and disgust in the fine-granularity emotion types.
[0126] For each piece of data, there are 5 common emotional indicators: 'joy', 'sadness', 'anger', 'fear', and 'disgust'. The specific calculation methods are as follows:
[0127]
[0128]
[0129]
[0130]
[0131]
[0132] Among them, Z 喜悦′ 、Z 悲伤′ 、Z 愤怒′ 、Z 恐惧′ 、Z 厌恶′They are all first-order tensors with a size of 5×1, representing the values of the emotional attributes of 5 common emotional indicators; Joy, Sorrow, Rage, Fear, Disgust, and A are all second-order tensors with a size of 50×5. Joy, Sorrow, Rage, Fear, Disgust represent the weighted values of 50 emotional categories for 5 common emotional indicators on 5 emotional attributes. A represents the values of 5 emotional attributes of 50 emotional categories. The elements in A have the following characteristics: For the 5 common emotions, except that the column values of the elements corresponding to the columns of these common emotions are 1, the values of other columns are all 0. The user can adjust the elements of this matrix according to their own definitions.
[0133] The above data forms Data Asset II: the subset of common emotional data of the general emotional data asset, which is a part of it.
[0134] 2.2 General Emotions: Subset of Satisfaction-Related Emotional Data
[0135] For each piece of data, there are three indicators: satisfaction, dissatisfaction, and satisfaction degree. The specific calculation methods are as follows:
[0136]
[0137]
[0138]
[0139] Among them, Satis and Unsatis are both first-order tensors with a size of 50×1, representing the weighted values of the hit counts of each emotional category for satisfaction and dissatisfaction. A is a second-order tensor with a size of 50×5, and the first row represents the hit counts of 50 emotional categories; for Satis and Unsatis, except for the columns corresponding to the elements of E satis and E unsatis the values of other columns are all 0. E satis ∈{satisfaction, admiration, touch, pride, happiness, love, belief, gratitude, praise, optimism, relaxation, calmness}; E unsatis ∈{sadness, frustration, depression, annoyance, loneliness, fear, decadence, palpitation, depression, inferiority, ridicule, doubt, condemnation, anxiety, jealousy, anger, anxiety, disappointment, grievance, disgust, resentment, helplessness, suspicion, numbness, contempt}.
[0140] The above data forms Data Asset II: the subset of satisfaction-related emotional data of the general emotional data asset, which is a part of it.
[0141] 2.3 General Emotions: Subset of Security-Related Emotional Data
[0142] For each piece of data, there are three indicators: sense of security, sense of insecurity, and the proportion of sense of security. The specific calculation methods are as follows:
[0143]
[0144]
[0145]
[0146] Among them, Secure and Insecure are both first-order tensors of size 50×1, representing the weighted values of the hit numbers of each emotional category for the sense of security and the sense of insecurity. A is a second-order tensor of size 50×5, and the first row represents the hit numbers of 50 emotional categories; for Secure and Insecure, except for the columns corresponding to the elements of E secure and E insecure the values of other columns are all 0, where E secure ∈{happy, relaxed, calm, touched, proud, satisfied, liked, grateful, optimistic}, E insecure ∈
[0147] {sad, fearful, alert, anxious, nervous, puzzled, annoyed, flustered, frustrated, down, anxious, sympathetic, disappointed, skeptical}.
[0148] The above data forms Data Asset II: a subset of the sense-of-security-related emotional data of the General Emotional Data Asset and is part of it.
[0149] 2.4 General Emotions: Subset of Emotional Arousal Data
[0150] For each piece of data, there are arousal values arousal({k 1 ,k 2 ,…,k 50}) of 50 fine-grained emotions. By classifying and summing the first-order tensor of size 50×1 according to the valence of the emotional categories corresponding to its elements, four indicators can be obtained for each piece of text: negative emotional arousal, positive emotional arousal, neutral emotional arousal, and the proportion of negative emotional arousal. The specific calculation methods are as follows:
[0151]
[0152]
[0153]
[0154]
[0155] Among them, neg, pos, and neu represent the set of negative sentiment categories, the set of positive sentiment categories, and the set of neutral sentiment categories respectively. The three are determined by valence(k t ,x i ) = g(k t ); arousal(k t ,x i ) represents the emotional arousal degree of emotion k t above the text x i .
[0156] The above data forms Data Asset II: the emotional arousal degree data subset of the general emotion data asset, which is a part of it.
[0157] 2.5 Other free customization
[0158] Users can also propose other possible emotion categories to construct new common emotion indicators.
[0159] 3. Data quality and security detection
[0160] It is consistent with the data quality and security detection content in Step 2. Quality inspection includes manual inspection of the general emotion category coding (accuracy rate) and hit rate (the number of single data not recognized / the total number of data); security inspection includes access permission management, encrypted storage detection, and encrypted transmission detection.
[0161] To sum up, after Step 3, Data Asset II: the general emotion dataset is obtained. In the example proposed in the embodiment of the present invention, the general emotion dataset has 4 recommended datasets (5 common emotion data subsets, satisfaction-related emotion data subset, security-related emotion data subset, emotional arousal degree data subset); in addition, based on Step 2, users can calculate new emotion dimensions according to the application definition to form other data subsets different from the above 4 recommended data subsets, as Figure 5 shown.
[0162] Step 4: Marketization of emotional data assets
[0163] After the dataset generated in the above steps is subject to compliance review by a third-party agency, it is registered in the data asset registration center. After determining the data ownership, it can enter the market for free trading; in addition, according to the needs of users, professional institutions can also evaluate the quality and value of the data assets to realize the inclusion of data assets in the financial statements.
[0164] System embodiment
[0165] According to the embodiment of the present invention, an emotional data assetization system based on fine-grained emotion classification is provided. Figure 6Schematic diagram of an emotional data assetization system based on fine-grained emotion division according to an embodiment of the present invention, as Figure 6 shown, the emotional data assetization system based on fine-grained emotion division according to an embodiment of the present invention specifically includes:
[0166] A data module 60, configured to obtain multi-source text big data, convert the text big data, and obtain initial data resources;
[0167] A first asset module 62, configured to construct a fine-grained emotion data set based on the initial data resources, and use the fine-grained emotion data set as the first emotional data asset; and
[0168] A second asset module 64, configured to establish a general emotion data set according to the first emotional data asset, use the general emotion data set as the second emotional data asset, and perform assetization processing on the first emotional data asset and the second emotional data asset.
[0169] The embodiment of the present invention is a system embodiment corresponding to the above method embodiment. The specific operations of each module can be understood with reference to the description of the method embodiment, and will not be elaborated here.
[0170] Device Embodiment 1
[0171] The embodiment of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps described in the method embodiment are implemented.
[0172] Device Embodiment 2
[0173] The embodiment of the present invention provides a computer-readable storage medium, on which an implementation program for information transmission is stored. When the program is executed by a processor, the steps described in the method embodiment are implemented.
[0174] The computer-readable storage medium described in this embodiment includes, but is not limited to: ROM, RAM, magnetic disk, optical disk, etc.
[0175] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for capitalizing emotional data based on fine-grained emotional segmentation, characterized in that include: Acquire multi-source text big data, transform the text big data, and obtain initial data resources; Building a fine-grained emotion dataset based on the initial data resource, and using the fine-grained emotion dataset as a first emotion data asset; as well as A universal emotion data set is established based on the first emotion data asset, the universal emotion data set is used as a second emotion data asset, and the first emotion data asset and the second emotion data asset are assetized.
2. The method according to claim 1, characterized in that The text big data is transformed to obtain initial data resources, which specifically include: Noise processing is performed on the text big data, integrated normalization processing is performed on the processed data according to standardized fields and converted into a standard format, data numbering is performed on the standardized data according to the subordinate relationship, corresponding identifiers are generated, and fields involving sensitive information in the data are anonymized, the anonymized data is encrypted and stored and permission management is added, and quality inspection and security inspection are performed on the stored data in real time to obtain initial data resources; The standardized fields include data content, data type, release time, release region, media source, number of likes, number of comments, number of reposts, number of visits, number of favorites, author name and author gender; The anonymization processing includes data desensitization, data generalization, data perturbation and de-identification processing; The quality check includes checking the data format and the completeness, accuracy, consistency and uniqueness of the data, checking invalid data and duplicate data, and checking abnormal values; The security inspection includes inspection of data encryption management, data access control, security vulnerability scanning and auditing, and data damage recovery.
3. The method according to claim 1, characterized in that Constructing a fine-grained sentiment dataset based on the initial data resources specifically includes: The SDP and DEP semantic dependency algorithms are used to generate a semantic dependency graph for the initial data resources, the language units in the semantic dependency graph are searched by the depth-first search algorithm DFS and the breadth-first search algorithm BFS, and the N-Gram language model is used to select the language units according to the principle of joint probability maximization to obtain the feature words to be encoded contained in each text after word segmentation, and the emotion type of each text is encoded according to the fine-grained emotion knowledge graph, and the emotion feature words of each text are obtained according to the emotion type of each text and its feature words to be encoded, the emotion attribute value of each text is encoded by the emotion feature words, and the emotion data obtained after encoding is quality checked and security checked in real time to generate a fine-grained emotion data set.
4. The method according to claim 3, characterized in that The emotional attribute values include the number of hits, emotional intensity, valence, emotional arousal and emotional control of each emotion in each text.
5. The method according to claim 4, characterized in that The emotional attribute value encoding of each text by using the emotional feature words specifically includes: The hit count of each emotion in each text is calculated using formula 1, and the interactive weighted hit count of each emotion in each text is calculated using formula 2; the emotional intensity of each emotion in each text is calculated using any one of formulas 3, 4, or 5; the valence of each emotion in each text is calculated using formula 6; the emotional arousal of each emotion in each text is calculated using any one of formulas 7, 8, or 9; the control type of each emotion type in each text is calculated using formula 10, and the sense of control of each emotion in each text is calculated using any one of formulas 11, 12, or 13; count(k j ,x i )=y formula 1; S(k j ,x i )=L(k j ,x i )×count(k j ,x i ) Formula 2; intensity2(k j ,x i )=max({f(w t )|t=1,2,3,…,y}) Formula 4; valence(k j ,x i )=g(k j ) Formula 6; arousal3(k j ,x i )=max({h(w t )|t=1,2,3,...,y}) in 9; control nominal (k j ,x i )=m(k j ) Formula 10; control ordina1,3 (k j ,x i )=m(k j )×max({n(w t )|t=1,2,3,…,y}) Formula 13; Among them, x i Represents the i-th text among n texts, for text x i Speaking of emotions j (from k1, k2, ... k 50 ) hits y feature words, S(k j ,x i ) is the sentiment k of the i-th text j The weighted hit count of a certain interaction volume, count(k j ,x j ) is the emotion k j The number of hits of sentiment feature words on the i-th text, L(k j ,x i ) is the weighted score of the interaction volume corresponding to the text, intensity(k j ,x i ) is the emotion k j The strength of the text on the i-th text. The subscripts 1, 2, and 3 are used to distinguish the formula type. y is the sentiment k j The number of hits, f(w t ) is the strength of a certain feature word, valence(k j ,x i ) represents the emotion type k j In the valence of the i-th text, g(k j ) represents the emotion type k j The potency, arousal(k j ,x i ) represents the emotion type k j In the arousal degree of the i-th text, the subscripts 1, 2, and 3 are used to distinguish the formula type, h(w t ) is the arousal degree of a certain feature word, control nominal (k j ,x i ) represents the emotion type k j The control sense of the emotional category in the i-th text, m(k j ) represents the emotion type k j Control type, control ordinal (k j ,x i ) represents the emotion type k j In the control sense of the i-th text, the subscripts 1, 2, and 3 are used to distinguish the formula type, n(w t ) is the control degree of a certain feature word.
6. The method according to claim 5, characterized in that Establishing a general emotion data set according to the first emotion data asset specifically includes: Based on the fine-grained emotion data set and according to user needs, a corresponding general emotion index is constructed, each emotion attribute in the general emotion index is calculated according to the fine-grained emotion types contained in the general emotion index and the relative contribution of each type, and the general emotion data set required by the user is established using the each emotion attribute.
7. The method according to claim 1, characterized in that The assetization process includes registration of emotional data assets, confirmation and empowerment of emotional data assets and corresponding transaction entities, asset transactions and inclusion of assets in the balance sheet.
8. A sentiment data assetization system based on fine-grained sentiment segmentation, characterized by include: A data module is used to obtain multi-source text big data and transform the text big data to obtain initial data resources; A first asset module, configured to construct a fine-grained emotion dataset based on the initial data resource, and use the fine-grained emotion dataset as a first emotion data asset; as well as The second asset module is used to establish a universal emotion data set based on the first emotion data asset, use the universal emotion data set as the second emotion data asset, and perform assetization processing on the first emotion data asset and the second emotion data asset.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method for assetizing emotional data based on fine-grained emotional segmentation as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores an implementation program for information transmission, and when the program is executed by a processor, the steps of the method for assetizing emotional data based on fine-grained emotional division as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Fine-grained emotion analysis method and device, computer equipment and storage medium
CN110516245A
Information processing method and device, electronic equipment and storage medium
CN114255071A
Policy announcement network comment sentiment analysis method, system and equipment
CN115238709A
Emotion knowledge graph construction method and device based on aspects
CN115391570A
Tourist online comment fine-grained sentiment analysis method and system
CN116737922A