A method and system for capitalizing emotional data based on fine-grained emotional segmentation

By performing noise processing and standardization on multi-source text big data, using semantic dependency graphs and N-Gram models to encode sentiment categories, and generating fine-grained and general sentiment datasets, the difficult problem of converting text big data into sentiment data assets is solved, and efficient data conversion and improved applicability are achieved.

CN120067290BActive Publication Date: 2025-09-19BEIJING COGNITIVE INSIGHT TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510009943.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-09-19
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

Existing technologies make it difficult to efficiently convert text big data such as social media and netizen comments into standardized, standard, and secure emotional data assets, and cannot meet the actual application needs of governments, enterprises, and research institutions.

Method used

By acquiring multi-source text big data, performing noise processing and standardization, the semantic dependency graph is generated using the SDP and DEP semantic dependency algorithms. Combining the depth-first search and breadth-first search algorithms, the N-Gram language model is used to encode sentiment categories and generate a fine-grained sentiment dataset. A general sentiment dataset is established based on general sentiment indicators, and finally assetized.

Benefits of technology

It has achieved the efficient conversion of text big data into emotional data assets, improved the applicability and operability of data, and promoted the application and development of text data in various fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067290B_ABST
    Figure CN120067290B_ABST
Patent Text Reader

Abstract

The present application provides a method and system for assetizing sentiment data based on fine-grained sentiment segmentation. The method according to the present application includes: acquiring multi-source text big data, transforming the text big data, and obtaining initial data resources; constructing a fine-grained sentiment data set based on the initial data resources, and using the fine-grained sentiment data set as the first sentiment data asset; and establishing a general sentiment data set based on the first sentiment data asset, using the general sentiment data set as the second sentiment data asset, and assetizing the first sentiment data asset and the second sentiment data asset. The present application proposes a method for efficiently transforming various types of text big data into sentiment data assets. The entire process not only realizes the transformation from raw data to sentiment data assets, but also has both practical applicability and operability, which will greatly promote the application and development of various types of text data in various fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data assetization, and in particular to a method and system for assetizing emotional data based on fine-grained emotional segmentation. Background Art

[0002] Currently, there is a pressing need for governments, businesses, and research institutions to conduct sentiment analysis on big data text, including social media posts, online comments, product reviews, complaints, chat conversations, and interactive suggestions. Converting text data into standardized, secure, and highly reusable sentiment data assets not only meets the practical application needs of relevant organizations but also aligns with the current era of booming digital economies and the inevitable trend of data assetization. Summary of the Invention

[0003] The purpose of the present invention is to provide a method and system for capitalizing emotional data based on fine-grained emotional segmentation, aiming to solve the above-mentioned problems in the prior art.

[0004] The embodiment of the present invention provides a method for capitalizing emotional data based on fine-grained emotional segmentation, including:

[0005] Acquire multi-source text big data, and transform the text big data to obtain initial data resources;

[0006] Building a fine-grained emotion dataset based on the initial data resource, and using the fine-grained emotion dataset as a first emotion data asset; and

[0007] A universal emotion data set is established based on the first emotion data asset, the universal emotion data set is used as the second emotion data asset, and assetization processing is performed on the first emotion data asset and the second emotion data asset.

[0008] The embodiment of the present invention provides a system for capitalizing emotional data based on fine-grained emotional segmentation, including:

[0009] A data module is used to obtain multi-source text big data and transform the text big data to obtain initial data resources;

[0010] a first asset module, configured to construct a fine-grained emotion dataset based on the initial data resource, and use the fine-grained emotion dataset as a first emotion data asset; and

[0011] The second asset module is configured to establish a universal emotion data set based on the first emotion data asset, use the universal emotion data set as a second emotion data asset, and perform assetization processing on the first emotion data asset and the second emotion data asset.

[0012] An embodiment of the present invention also provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps of the above-mentioned method for assetizing emotional data based on fine-grained emotional segmentation are implemented.

[0013] An embodiment of the present invention further provides a computer-readable storage medium, on which a program for implementing information transmission is stored. When the program is executed by a processor, the steps of the above-mentioned method for capitalizing emotional data based on fine-grained emotional division are implemented.

[0014] The following benefits can be achieved by implementing embodiments of the present invention: This invention proposes a method for efficiently converting text data into sentiment data assets. This entire process not only achieves the conversion from raw data to sentiment data assets, but also possesses both practical applicability and operability. This invention will greatly promote the application and development of various types of text data in various fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate one or more embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0016] Figure 1 This is a flow chart of a method for capitalizing emotional data based on fine-grained emotional segmentation according to an embodiment of the present invention;

[0017] Figure 2 It is a schematic diagram of the overall process of an embodiment of the present invention;

[0018] Figure 3 Schematic diagram of the correspondence between feature words and sentiment categories in an embodiment of the present invention;

[0019] Figure 4 This is a schematic diagram of important fields of data asset I according to an embodiment of the present invention;

[0020] Figure 5 This is a schematic diagram of important fields of data asset II according to an embodiment of the present invention;

[0021] Figure 6 Schematic diagram of an emotional data assetization system based on fine-grained emotional segmentation according to an embodiment of the present invention. DETAILED DESCRIPTION

[0022] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below in conjunction with the drawings in one or more embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this document.

[0023] Method Example

[0024] According to an embodiment of the present invention, a method for capitalizing emotional data based on fine-grained emotional segmentation is provided. Figure 1 This is a flow chart of a method for capitalizing emotional data based on fine-grained emotional segmentation according to an embodiment of the present invention. Figure 1 As shown, the method for capitalizing emotional data based on fine-grained emotional segmentation according to an embodiment of the present invention specifically includes:

[0025] Step S101, obtaining multi-source text big data, transforming the text big data to obtain initial data resources, specifically includes:

[0026] Noise processing is performed on the text big data, the processed data is integrated and normalized according to standardized fields and converted into a standard format, the standardized data is numbered according to affiliation, corresponding identifiers are generated, and fields involving sensitive information in the data are anonymized, the anonymized data is encrypted and stored and permission management is added, and the stored data is quality checked and security inspected in real time to obtain initial data resources;

[0027] The standardized fields include data content, data type, release time, release region, media source, number of likes, number of comments, number of reposts, number of visits, number of favorites, author name, and author gender;

[0028] The anonymization processing includes data desensitization, data generalization, data perturbation and de-identification processing;

[0029] The quality check includes checking the data format and the completeness, accuracy, consistency and uniqueness of the data, checking for invalid data and duplicate data, and checking for outliers;

[0030] The security inspection includes inspection of data encryption management, data access control, security vulnerability scanning and auditing, and data damage recovery;

[0031] Step S102, constructing a fine-grained emotion dataset based on the initial data resource and using the fine-grained emotion dataset as a first emotion data asset, specifically includes:

[0032] A semantic dependency graph is generated for the initial data resources using the SDP and DEP semantic dependency algorithms, the language units in the semantic dependency graph are searched using the depth-first search algorithm DFS and the breadth-first search algorithm BFS, and the language units are selected using the N-Gram language model according to the joint probability maximization principle, to obtain the feature words to be encoded contained in each text after word segmentation, to encode the emotion type of each text according to the fine-grained emotion knowledge graph, and to obtain the emotion feature words of each text according to the emotion type of each text and its feature words to be encoded, to encode the emotion attribute value of each text using the emotion feature words, and to perform quality inspection and security inspection on the emotion data obtained after encoding in real time, to generate a fine-grained emotion dataset;

[0033] The emotional attribute values ​​include the number of hits, emotional intensity, valence, emotional arousal and emotional control of each emotion in each text;

[0034] The emotional attribute value encoding of each text using the emotional feature words specifically includes:

[0035] The number of hits for each emotion in each text is calculated using formula 1, and the number of interaction-weighted hits for each emotion in each text is calculated using formula 2; the emotional intensity of each emotion in each text is calculated using any one of formulas 3, 4, or 5; the valence of each emotion in each text is calculated using formula 6; the emotional arousal of each emotion in each text is calculated using any one of formulas 7, 8, or 9; the control type of each emotion type in each text is calculated using formula 10, and the sense of control of each emotion in each text is calculated using any one of formulas 11, 12, or 13.

[0036] count(k j ,x i )=y formula 1;

[0037] S(k j ,x i )=L(k j ,x i )×count(k j ,x i ) Formula 2;

[0038]

[0039] intensity2(k j ,xi )=max({f(w t )|t=1,2,3,…,y}) Formula 4;

[0040]

[0041] valence(k j ,x i )=g(k j ) Formula 6;

[0042]

[0043]

[0044] arousal3(k j ,x i )=max({h(w t )|t=1,2,3,…,y}) Formula 9;

[0045] control nominal (k j ,x i )=m(k j ) Formula 10;

[0046]

[0047]

[0048] control ordinal,3 (k j ,x i )=m(k j )×max({n(w t )|t=1,2,3,…,y}) Formula 13;

[0049] Among them, x i Represents the i-th text in n texts, for text x i Speaking of emotions j (from k1, k2, ... k 50 ) hits y feature words, S(k j ,x i ) is the sentiment k of the i-th text j The weighted hit count of a certain interaction amount, count(k j ,x i ) is the emotion k j The number of sentiment feature words hit on the i-th text, L(k j ,x i) is the weighted score of the interaction volume corresponding to the text, intensity(k j ,x i ) is the emotion k j The strength of the i-th text, the subscripts 1, 2, and 3 are used to distinguish the formula type, and y is the sentiment k j The number of hits, f(w t ) is the strength of a certain feature word, valence(k j ,x i ) represents the emotion type k j In the valence of the i-th text, g(k j ) represents the emotion type k j The potency, arousal (k j ,x i ) represents the emotion type k j In the arousal degree of the i-th text, the subscripts 1, 2, and 3 are used to distinguish the formula type, h(w t ) is the arousal degree of a certain feature word, control nominal (k j ,x i ) represents the emotion type k j The emotional category control sense in the i-th text, m(k j ) represents the emotion type k j Control type, control ordinal (k j ,x i ) represents the emotion type k j In the control sense of the i-th text, the subscripts 1, 2, and 3 are used to distinguish the formula type, n(w t ) is the control degree of a certain feature word;

[0050] Step S103, establishing a universal emotion data set based on the first emotion data asset, using the universal emotion data set as a second emotion data asset, and performing assetization processing on the first emotion data asset and the second emotion data asset, specifically includes:

[0051] Based on the fine-grained emotion dataset and according to user needs, a corresponding general emotion index is constructed, each emotion attribute in the general emotion index is calculated according to the fine-grained emotion categories contained in the general emotion index and the relative contribution of each category, and the general emotion dataset required by the user is established using the emotion attributes;

[0052] The assetization process includes registration of emotional data assets, confirmation and empowerment of emotional data assets and corresponding transaction entities, asset transactions and asset inclusion.

[0053] The above technical solution of the embodiment of the present invention is described in detail below in conjunction with the specific situation of the method for capitalizing emotional data based on fine-grained emotional segmentation according to the embodiment of the present invention.

[0054] This embodiment of the present invention proposes a solution for efficiently converting various types of text big data into sentiment data assets through four main steps. The first step is raw data resource conversion. This involves converting raw data into standard, standardized, and secure data resources, including data cleaning and preprocessing, as well as data quality and security verification. The second step is generating fine-grained sentiment data assets. Based on a fine-grained sentiment knowledge graph and model algorithms, the data is encoded with sentiment types and attributes. After undergoing data quality and security verification, Data Asset 1, a "fine-grained sentiment dataset," is generated. The third step is generating general sentiment data assets. Based on common application requirements of sentiment computing, such as "five common emotions," "satisfaction," "security," and "emotional arousal," Data Asset 1 is encoded with general sentiment types and attributes using a general sentiment indicator algorithm model to meet the requirements of general sentiment indicators. After data quality and value assessment, Data Asset 2, a "general sentiment dataset," is generated. The fourth step is the commercialization of sentiment data assets. This specifically involves data asset registration, table entry, and transaction.

[0055] The embodiment of the present invention uses methods based on fine-grained emotional knowledge graphs, classification models, etc. to transform emotional data resources into data assets. The overall steps are as follows: Figure 2 Specifically including:

[0056] Step 1: Raw data resource

[0057] 1. Data cleaning

[0058] The first is noise processing. This involves removing noise and outlier topics to ensure data validity, accuracy, and consistency. Topics can be concepts, events, policies, products, brands, and so on.

[0059] The second is data standardization. Multi-source text data is integrated and normalized, and the data is converted into a standard format with standardized data items. The data includes standardized fields such as "data content," "data type," "release time," "release region," "media source," "number of likes," and "number of comments" (as shown in Table 1). Each piece of data is numbered according to its affiliation as its identifier, such as "total number," "main post number," etc.

[0060] Table 1 Field standardization rules

[0061]

[0062] The third is anonymization and security. Fields containing sensitive personal information are anonymized, including data desensitization, data generalization, data perturbation, and de-identification. Fields such as "Author Name" and "Author Gender" are recoded to contain only discrete numerical values ​​and other data types that do not contain actual personal information. Information such as names of people and specific places that are not related to the subject in the "Data Content" column is anonymized using a preset string list, effectively protecting personal privacy while retaining the useful information of the data and ensuring its value in analysis and application.

[0063] Data is encrypted and stored, and permission management is applied so that only authorized personnel can access the data. Relevant standards are followed to ensure the compliance of data processing.

[0064] 2. Data quality inspection and security inspection

[0065] Data quality inspection and verification include checking data format, data integrity, accuracy, consistency, and uniqueness, checking for invalid and duplicate data, detecting outliers, etc. Security inspection includes inspection and verification of data encryption management, data access control, security vulnerability scanning and auditing, and data damage recovery.

[0066] Based on the above two steps, the messy original data is transformed into standardized, anonymized and secure data resources.

[0067] Step 2: Generate Data Asset I: "Fine-Grained Sentiment Dataset"

[0068] 1. Emotional category coding

[0069] First, the data is segmented through natural language processing. The entries of the fine-grained sentiment knowledge graph include feature words, abbreviations, expressions, symbols, sentence patterns and their combinations, which are collectively referred to as feature words. The specific implementation includes the following sub-steps:

[0070] Using the SDP and DEP semantic dependency algorithms, text data is mapped into a graph structure to generate a semantic dependency graph for the text data to be analyzed. SDP (Semantic Dependency Parsing) and DEP (Dependency Parsing) are two important tasks in natural language processing.

[0071] Use two graph search algorithms, the depth-first DFS algorithm and the breadth-first BFS algorithm, to search for suitable language units LU on the semantic dependency graph. Each language unit LU is a word segmentation result.

[0072] Adopt N-Gram language model and select language unit LU={W1,W2,…W n}; P(W 1:n ) is the joint probability, which is specifically expressed as:

[0073]

[0074] Among them, W k is the characteristic word unit in the language unit group, k is the arrangement order of the characteristic word unit, n is the upper bound of k, k = 1, 2, ..., n; the relationship between each sentence and its language unit group satisfies the Markov relationship, and the language unit LU = {W1, W2, ...W n}; Each feature word unit W k They are not necessarily connected to each other.

[0075] After word segmentation according to the above steps, the feature words to be compiled are obtained for each text; then, the sentiment type is encoded for each text data of the standardized data to obtain the "feature word-sentiment type" correspondence of the single data, that is, the "hit situation" field. This field has two possible values ​​0 or 1. 0 means that the sentiment does not hit any feature word, and 1 means that at least one feature word is hit. Figure 3 As shown in the figure, the text has a total of k emotions, and the number of characteristic words for each emotion is y1, y2, ..., y k Based on the fine-grained emotional knowledge graph, we obtain the emotional category of each text and its corresponding feature words to be encoded, and obtain the emotional feature words. The fine-grained emotional knowledge graph contains 50 types of fine-grained emotions, which are specifically classified as follows:

[0076] (1) Mood: emotional experiences with low arousal and long duration, including loneliness, numbness, calmness, depression, anxiety, and decadence;

[0077] (2) Stress-related emotions: emotional experiences with high arousal and limited duration, including happiness, fear, surprise, anger, sadness, excitement, relaxation, anxiety, tension, emotion, and alertness;

[0078] (3) Evaluative emotions: The emotional experience generated by an individual after evaluating an information event can be further categorized as directed towards oneself or the outside world. Externally directed emotions include sympathy, indifference, ridicule, disgust, condemnation, resentment, questioning, love, satisfaction, praise, trust, contempt, gratitude, admiration, longing, jealousy, helplessness, expectation, optimism, wishing, and disappointment; internally directed emotions include depression, embarrassment, inferiority, shyness, guilt, pride, frustration, worry, panic, grievance, doubt, and boredom.

[0079] 2. Emotional attribute value encoding

[0080] The emotional feature words in the above steps are used to obtain the hit count, intensity, valence and arousal of each emotion in each text. These features are collectively referred to as emotional attribute values.

[0081] Assume there are n texts, for text x i (from x1, x2, ... x n ) to express emotions j (from k1, k2, ... k 50 ) hits y feature words, then x i Emotional k j The calculation method of each attribute value is as follows:

[0082] A. Hits (count)

[0083] count(k j ,x i )=y (2);

[0084] That is, emotion k j The sum of the number of feature words.

[0085] In addition, there is a method for calculating the weighted hit count that takes into account the amount of interaction with the text itself. The specific implementation is as follows:

[0086] First, the text to be analyzed is matched with a fine-grained sentiment dictionary to calculate the weighted score of the text's interaction count and the number of hits for a certain sentiment. Second, the weighted hit count for the interaction count is calculated. The formula involved is as follows:

[0087] S(k j ,x i )=L(k j ,x i )×count(k j ,x i ) (3);

[0088] Among them, S(k j ,x i ) is the sentiment k of the i-th text j The weighted hit count of a certain interaction amount, which can be a behavioral indicator such as "like", "forward", "comment", etc. count(k j ,x i ) is the emotion k j The number of sentiment feature words hit on the i-th text, L(k j ,x i ) is the weighted score of the interaction volume corresponding to the text, and the weighted score is calculated as follows:

[0089]

[0090] Among them, L(k j ,x i ) is the weighted score of the interaction volume of a single text, M is the adjustment coefficient, and when M≤0, the formula has no application value. When M>140, the weighted result exceeds the reasonable range and approaches the extreme. Therefore, the value range of M is specified as (0,140]; Interaction is the interaction volume of a single text, Interaction>0; a is the base of the logarithmic function and a>1.

[0091] B. Intensity

[0092]

[0093] intensity2(k j ,x i )=max({f(w t )|t=1,2,3,…,y}) (6);

[0094]

[0095] Among them, formula (5) is the arithmetic mean method, formula (6) is the maximum method, and formula (7) is the summation method. Users can choose any one or more of the above methods to calculate the intensity of emotion. j ,x i ) is the emotion k j The sentiment intensity value of the i-th text. The subscripts 1, 2, and 3 are used to distinguish the formula type. y is the sentiment k j The number of hits, f(w t ) is the strength of a certain feature word, which is defined in the knowledge graph. The strength function f(w) accepts a sentiment feature word string and returns an integer in the range of [1,5].

[0096] C. Valence

[0097] valence(k j ,x i )=g(k j ) (8);

[0098] Each hit sentiment feature word has an intrinsic valence. g(k) accepts a sentiment type string and feedback -1, 0 or 1, where 1 represents positive sentiment, 0 represents neutral sentiment, and -1 represents negative sentiment. g(k) is defined in the knowledge graph. valence(k j ,x i ) represents the emotion type k j In the valence of the i-th text, g(k j) represents the emotion type k j The potency of .

[0099] D. Arousal

[0100]

[0101]

[0102] arousal3(k j ,x i )=max({h(w t )|t=1,2,3,…,y}) (11);

[0103] Among them, formula (9) is the summation method, formula (10) is the arithmetic mean method, and formula (11) is the maximum value method. Users can choose any one or more of the above methods to calculate the emotional arousal. j ,x i ) represents the emotion type k j In the arousal degree of the i-th text, the subscripts 1, 2, and 3 are used to distinguish the formula type, h(w t ) is the arousal degree of a certain feature word. This function is defined in the knowledge graph. h(w t ) accepts an emotion type string and returns an integer in the range [1,5], where 1 represents very low arousal and 5 represents very high arousal.

[0104] E. Emotional control

[0105] control nominal (k j ,x i )=m(k j ) (12);

[0106]

[0107]

[0108] control ordinal,3 (k j ,x i )=m(k j )×max({n(w t )|t=1,2,3,…,y}) (15);

[0109] Formula (12) is the control sense of emotion type, and formulas (13) to (15) are the control sense of emotion; among them, formula (13) is the summation method, formula (14) is the arithmetic mean method, and formula (15) is the maximum value method. Users can choose any one or more of the above methods to calculate the control sense of emotion. nominal (k j ,x i ) represents the emotion type k j The emotional category control sense in the i-th text, m(k j ) represents the emotion type k j The control type of the function is defined in the knowledge graph. m(k) accepts a sentiment type string and feedback of -1, 0 or 1. 1 represents dominance, 0 represents no dominance relationship, and -1 represents a dominated relationship. ordinal (k j ,x i ) represents the emotion type k j In the control sense of the i-th text, the subscripts 1, 2, and 3 are used to distinguish the formula type, n(w t ) is the degree of control of a certain feature word. This function is defined in the knowledge graph. n(w t ) accepts a sentiment type string and returns an integer in the range [1,5], where 1 represents extremely low dominance (or being dominated) and 5 represents extremely high dominance (or being dominated).

[0110] 3. Data quality and security inspection

[0111] Quality inspection includes manual inspection of emotion category coding (accuracy) and hit rate (number of unrecognized single data items / total number of data items); security inspection includes access permission management, encrypted storage detection, and encrypted transmission detection.

[0112] In summary, after completing step 2, we obtain data asset I: fine-grained sentiment dataset, such as Figure 4 As shown, the "hit situation" field corresponds to the result of emotion category encoding, and the remaining fields correspond to the results of emotion attribute encoding.

[0113] Step 3: Generate Data Asset II: "General Sentiment Dataset"

[0114] Based on the fine-grained sentiment dataset obtained in step 2, we can quickly obtain commonly used and general indicators in practical applications, including but not limited to "five common emotions," "satisfaction," "sense of security," and "emotional arousal." The following describes the calculation process for common sentiment indicators and several recommended indicators.

[0115] 1. Calculation of general sentiment indicators

[0116] The calculation of a general sentiment index requires considering the relative contribution of each emotion attribute within each fine-grained emotion category to the index. Specifically, if a general sentiment index Z exists, its individual emotion attribute values ​​are composited from a 50-element fine-grained emotion set based on the weights of each emotion attribute within Z. In addition to commonly used indicators such as satisfaction and sense of security, users can also define other possible indicators. The specific algorithm for the general sentiment index is as follows:

[0117] Assume there are n texts, for text x i (from x1, x2, ... x n ) to express emotions j Each sentiment attribute of (j∈{1,2,…,50}) is represented by count(k j ,x i )、S(k j ,x j )、intensity(k j ,x i )、arousal(k j ,x i ) and control ordinal (k j ,x i ) definition, there are 5 attributes, then x i Each emotional attribute of the composite emotion Z is defined as a first-order tensor Z:

[0118]

[0119] Among them, Z is the set of attributes of the universal sentiment index Z; Z is a first-order tensor, and A and C are both second-order tensors; A and C can be represented as 50×5 matrices, j∈{1,2,3,...,50} and q∈{1,2,3,4,5}, the maximum value of j is the number of sentiment categories, and the maximum value of q represents the number of non-categorical sentiment attributes, A reflects the value of each sentiment attribute of each sentiment category, and C reflects the weighted value of each sentiment attribute of each sentiment category in the universal sentiment Z; Z can be represented as a column vector of size 5, which reflects the values ​​of different sentiment attributes of the universal sentiment Z.

[0120] In order to determine the value c of each element in C in the above formula jq : In the absence of any assumptions, methods including but not limited to confirmatory factor analysis can be used to determine the jq value.

[0121] When there is a hypothetical relationship due to the application field, the user can jq The relationship between the values ​​generates a priori assumptions, and the remaining unknown relationships can be further obtained by the above method.

[0122] 2. General emotional attribute value encoding

[0123] After determining the coefficients and types of general emotions, it is necessary to calculate the relevant attributes of each text. The following are several data sets obtained through attribute calculation in this embodiment of the present invention:

[0124] 2.1 General Sentiment: 5 Common Sentiment Data Subsets

[0125] In applications, the emotion calculation process tends to categorize emotions into five broad categories (coarse-grained, commonly referred to as common emotion categories): joy, sadness, anger, fear, and disgust. Classification and calculation can be performed based on 50 fine-grained emotion categories, enabling attribute encoding of the five common emotions in text. Some common emotion categories share the same names as fine-grained emotion categories. To distinguish them, the common emotion categories are named "joy," "sadness," "anger," "fear," and "disgust" to distinguish them from the fine-grained emotion categories of joy, sadness, anger, fear, and disgust.

[0126] For each piece of data, there are five commonly used emotion indicators: joy, sadness, anger, fear, and disgust. The specific calculation method is as follows:

[0127]

[0128]

[0129]

[0130]

[0131]

[0132] Among them, Z 喜悦′ 、Z 悲伤′ 、Z 愤怒′ 、Z 恐惧′ 、Z 厌恶′ Each is a first-order tensor of size 5×1, representing the values ​​of the emotional attributes of the five common emotion indicators. Joy, Sorrow, Rage, Fear, Disgust, and A are all second-order tensors of size 50×5. Joy, Sorrow, Rage, Fear, and Disgust represent the weighted values ​​of the five emotional attributes for the 50 emotion categories of the five common emotion indicators. A represents the values ​​of the five emotional attributes for the 50 emotion categories. The elements in A have the following characteristics: for the five common emotions, except for the columns corresponding to the elements related to these common emotions, the values ​​of all other columns are 0. Users can adjust the elements of this matrix according to their own definitions.

[0133] The above data forms a subset of commonly used emotional data of Data Asset II: General Emotional Data Asset, of which it is a part.

[0134] 2.2 General Sentiment: Satisfaction-Related Sentiment Data Subset

[0135] For each piece of data, there are three indicators: satisfaction, dissatisfaction, and satisfaction. The specific calculation method is as follows:

[0136]

[0137]

[0138]

[0139] Among them, Satis and Unsatis are both first-order tensors of size 50×1, representing the weighted values ​​of the hit counts of each emotion type for satisfaction and dissatisfaction. A is a second-order tensor of size 50×5, and the first row represents the hit counts of 50 emotion types. For Satis and Unsatis, except for the column number corresponding to E satis and E unsatis Except for the column containing the elements of E, the values ​​of other columns are all 0. satis ∈{satisfied, admired, moved, proud, happy, loved, believed, grateful, praised, optimistic, relaxed, calm}; E unsatis ∈{sadness, frustration, depression, worry, loneliness, fear, decadence, panic, depression, inferiority, ridicule, doubt, blame, anxiety, jealousy, anger, anxiety, disappointment, grievance, disgust, resentment, helplessness, questioning, numbness, contempt}.

[0140] The above data forms Data Asset II: a subset of satisfaction-related sentiment data of the general sentiment data asset, of which it is a part.

[0141] 2.3 General Emotion: Security-Related Emotion Data Subset

[0142] For each piece of data, there are three indicators: sense of security, insecurity, and the proportion of sense of security. The specific calculation method is as follows:

[0143]

[0144]

[0145]

[0146] Among them, Secure and Insecure are both first-order tensors of size 50×1, representing the weighted values ​​of the hit counts of each emotion type for security and insecurity. A is a second-order tensor of size 50×5, and the first row represents the hit counts of 50 emotion types. For Secure and Insecure, except for the column number corresponding to E secure and E insecure Except for the column of elements, the values ​​of other columns are all 0, where E secure ∈{happy, relaxed, calm, moved, proud, satisfied, fond, grateful, optimistic}, E insecure ∈

[0147] {sadness, fear, alertness, anxiety, tension, doubt, worry, panic, frustration, depression, anxiety, sympathy, disappointment, questioning}.

[0148] The above data forms Data Asset II: a subset of security-related emotional data of the general emotional data asset, of which it is a part.

[0149] 2.4 General Emotion: Emotional Arousal Data Subset

[0150] For each piece of data, there are 50 kinds of fine-grained emotional arousal values ​​arousal ({k1, k2, ..., k 50}), the first-order tensor of size 50×1 is classified and summed according to the valence of the emotion types corresponding to its elements. Four indicators, namely negative emotion arousal, positive emotion arousal, neutral emotion arousal, and negative emotion arousal ratio, can be obtained for each text. The specific calculation method is as follows:

[0151]

[0152]

[0153]

[0154]

[0155] Among them, neg, pos and neu represent the negative emotion category set, positive emotion category set and neutral emotion category set respectively, and the three are represented by valence(k t ,x i )=g(k t ) decision; arousal(k t ,x i ) indicates emotion k t In the text x i Emotional arousal above .

[0156] The above data forms Data Asset II: the emotional arousal data subset of the general emotional data asset, and is part of it.

[0157] 2.5 Other free customization

[0158] Users can also propose other possible types of emotions and construct new commonly used emotion indicators.

[0159] 3. Data quality and security testing

[0160] The data quality and security testing content is consistent with that in step 2. Quality inspection includes manual verification of general emotion category coding (accuracy) and hit rate (number of unrecognized single data items / total number of data items); security inspection includes access rights management, encrypted storage testing, and encrypted transmission testing.

[0161] In summary, after step 3 is completed, data asset II: general emotion dataset is obtained. In the example proposed in the embodiment of the present invention, the general emotion dataset has 4 recommended datasets (5 common emotion data subsets, satisfaction-related emotion data subsets, security-related emotion data subsets, and emotion arousal data subsets); in addition, based on step 2, users can calculate new emotion dimensions according to the application definition to form other data subsets different from the above 4 recommended data subsets, such as Figure 5 shown.

[0162] Step 4: Marketization of emotional data assets

[0163] The data set generated by the above steps will undergo compliance review by a third-party agency and then be registered at the Data Asset Registration Center. Once the data ownership is confirmed, it can enter the market for free trading. In addition, according to user needs, the quality and value of data assets can also be assessed by professional agencies to ensure that the data assets are included in the table.

[0164] System Example

[0165] According to an embodiment of the present invention, a sentiment data assetization system based on fine-grained sentiment segmentation is provided. Figure 6 Schematic diagram of the emotional data assetization system based on fine-grained emotional segmentation according to an embodiment of the present invention. Figure 6 As shown, the emotional data assetization system based on fine-grained emotional segmentation according to an embodiment of the present invention specifically includes:

[0166] Data module 60, used to obtain multi-source text big data, transform the text big data, and obtain initial data resources;

[0167] A first asset module 62 is configured to construct a fine-grained emotion dataset based on the initial data resource, and use the fine-grained emotion dataset as a first emotion data asset; and

[0168] The second asset module 64 is configured to create a universal emotion data set based on the first emotion data asset, use the universal emotion data set as a second emotion data asset, and perform assetization processing on the first emotion data asset and the second emotion data asset.

[0169] The embodiment of the present invention is a system embodiment corresponding to the above-mentioned method embodiment. The specific operations of each module can be understood by referring to the description of the method embodiment, which will not be repeated here.

[0170] Device Example 1

[0171] An embodiment of the present invention provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the steps described in the method embodiment when executed by the processor.

[0172] Device Example 2

[0173] An embodiment of the present invention provides a computer-readable storage medium, on which a program for implementing information transmission is stored. When the program is executed by a processor, the steps described in the method embodiment are implemented.

[0174] The computer-readable storage medium in this embodiment includes, but is not limited to, ROM, RAM, magnetic disk, or optical disk.

[0175] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for capitalizing emotional data based on fine-grained emotional segmentation, characterized by include: Acquire multi-source text big data, and transform the text big data to obtain initial data resources; Constructing a fine-grained emotion dataset based on the initial data resource and using the fine-grained emotion dataset as a first emotion data asset specifically includes: A semantic dependency graph is generated for the initial data resources using the SDP and DEP semantic dependency algorithms, the language units in the semantic dependency graph are searched using the depth-first search algorithm DFS and the breadth-first search algorithm BFS, and the language units are selected using the N-Gram language model according to the joint probability maximization principle to obtain the feature words to be encoded contained in each text after word segmentation, the emotion type of each text is encoded according to the fine-grained emotion knowledge graph, and the emotion feature words of each text are obtained according to the emotion type of each text and its feature words to be encoded, the emotion attribute value of each text is encoded using the emotion feature words, and the emotion data obtained after encoding is quality checked and security checked in real time to generate a fine-grained emotion data set; wherein the emotion attribute value includes the number of hits, emotion intensity, valence, emotion arousal and emotion control of each emotion in each text; Creating a universal emotion data set based on the first emotion data asset, using the universal emotion data set as a second emotion data asset, and performing assetization processing on the first emotion data asset and the second emotion data asset; The assetization process includes registration of emotional data assets, confirmation and empowerment of emotional data assets and corresponding transaction entities, asset transactions and asset inclusion.

2. The method according to claim 1, characterized in that The text big data is converted to obtain initial data resources, specifically including: Noise processing is performed on the text big data, the processed data is integrated and normalized according to standardized fields and converted into a standard format, the standardized data is numbered according to affiliation, corresponding identifiers are generated, and fields involving sensitive information in the data are anonymized, the anonymized data is encrypted and stored and permission management is added, and the stored data is quality checked and security inspected in real time to obtain initial data resources; The standardized fields include data content, data type, release time, release region, media source, number of likes, number of comments, number of reposts, number of visits, number of favorites, author name, and author gender; The anonymization processing includes data desensitization, data generalization, data perturbation and de-identification processing; The quality check includes checking the data format and the completeness, accuracy, consistency and uniqueness of the data, checking for invalid data and duplicate data, and checking for outliers; The security inspection includes inspection of data encryption management, data access control, security vulnerability scanning and auditing, and data damage recovery.

3. The method according to claim 1, characterized in that The emotional attribute value encoding of each text using the emotional feature words specifically includes: The number of hits for each emotion in each text is calculated using formula 1, and the number of interaction-weighted hits for each emotion in each text is calculated using formula 2; the emotional intensity of each emotion in each text is calculated using any one of formulas 3, 4, or 5; the valence of each emotion in each text is calculated using formula 6; the emotional arousal of each emotion in each text is calculated using any one of formulas 7, 8, or 9; the control type of each emotion type in each text is calculated using formula 10, and the sense of control of each emotion in each text is calculated using any one of formulas 11, 12, or 13. Formula 1: Formula 2: Formula 3: Formula 4: Formula 5: Formula 6: Formula 7: Formula 8: Formula 9; Formula 10; Formula 11; Formula 12; Formula 13; in, express The i-th text in the text, for the text Talk about emotions From Hit feature words indivual, Indicates the number of hit feature words, For the Sentiment of the text The weighted hit count of a certain interaction volume, For emotion In the The number of hits of sentiment feature words on the text, The weighted score for the interaction volume corresponding to the text. It's emotion In the The strength of the text, the subscripts 1, 2, 3 are used to distinguish the formula type, It's emotion The number of hits, is the strength of a certain feature word, Expressing emotion types In the Valence of the article Expressing emotion types The potency, Expressing emotion types In the The awakening degree of the text. The subscripts 1, 2, and 3 are used to distinguish the formula type. is the arousal degree of a certain feature word, Expressing emotion types In the The emotional type of the text controls the sense of Expressing emotion types The control type, Expressing emotion types In the The subscripts 1, 2, and 3 are used to distinguish the formula types. It is the degree of control of a certain feature word.

4. The method according to claim 3, characterized in that Establishing a general emotion dataset based on the first emotion data asset specifically includes: Based on the fine-grained emotion data set and according to user needs, a corresponding general emotion index is constructed, each emotion attribute in the general emotion index is calculated according to the fine-grained emotion types contained in the general emotion index and the relative contribution of each type, and the general emotion data set required by the user is established using the various emotion attributes.

5. A sentiment data assetization system based on fine-grained sentiment segmentation, characterized by include: A data module is used to obtain multi-source text big data and transform the text big data to obtain initial data resources; The first asset module is configured to construct a fine-grained emotion dataset based on the initial data resource and use the fine-grained emotion dataset as the first emotion data asset, specifically for: A semantic dependency graph is generated for the initial data resources using the SDP and DEP semantic dependency algorithms, the language units in the semantic dependency graph are searched using the depth-first search algorithm DFS and the breadth-first search algorithm BFS, and the language units are selected using the N-Gram language model according to the joint probability maximization principle to obtain the feature words to be encoded contained in each text after word segmentation, the emotion type of each text is encoded according to the fine-grained emotion knowledge graph, and the emotion feature words of each text are obtained according to the emotion type of each text and its feature words to be encoded, the emotion attribute value of each text is encoded using the emotion feature words, and the emotion data obtained after encoding is quality checked and security checked in real time to generate a fine-grained emotion data set; wherein the emotion attribute value includes the number of hits, emotion intensity, valence, emotion arousal and emotion control of each emotion in each text; a second asset module, configured to establish a universal emotion data set based on the first emotion data asset, use the universal emotion data set as a second emotion data asset, and perform assetization processing on the first emotion data asset and the second emotion data asset; The assetization process includes registration of emotional data assets, confirmation and empowerment of emotional data assets and corresponding transaction entities, asset transactions and asset inclusion.

6. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps of the method for capitalizing emotional data based on fine-grained emotional segmentation as described in any one of claims 1 to 4 are implemented.

7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores an implementation program for information transmission, and when the program is executed by a processor, the steps of the method for capitalizing emotional data based on fine-grained emotional division as described in any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Policy announcement network comment sentiment analysis method, system and equipment

    CN115238709A