Time domain secondary sentiment analysis method and system based on text symbolization

By constructing an emoticon tag library and a Chinese vocabulary library, text data is preprocessed, encoded and sequenced, and using deep learning models combined with a time-domain secondary screening mechanism, the problem of insufficient single-dimensional analysis and dynamic monitoring capabilities in the existing technology is solved, and the multi-dimensional, dynamic and refined emotion analysis is realized, and the accuracy and early warning capabilities of emotion analysis are improved.

CN120354850APending Publication Date: 2025-07-22XINHUANET CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510245538.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Existing emotion analysis technologies mainly focus on single-dimensional analysis, ignore the potential value of secondary emotions, lack dynamic monitoring capabilities, fail to make full use of the relationship between emoji and text emotions, find it difficult to identify emotional turning points, and cannot provide timely warnings.

Method used

By constructing an emoticon tag library and a Chinese vocabulary library, the text data is preprocessed, encoded and sequenced, and the deep learning model combined with the time domain secondary screening mechanism is used to explore potential emotions, track emotional inflection points and label them into the database.

Benefits of technology

It realizes multi-dimensional, dynamic and refined sentiment analysis, improves the accuracy and early warning capabilities of sentiment analysis, and enhances the traceability and scalability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354850A_ABST
    Figure CN120354850A_ABST
Patent Text Reader

Abstract

The invention relates to a time domain secondary sentiment analysis method and system based on text symbolization. According to the method, an expression label library and a Chinese vocabulary library are constructed, text data are subjected to preprocessing, encoding and sequence filling processing, a deep learning model is combined with a time domain secondary screening mechanism to mine potential emotions, and meanwhile, emotion inflection points are tracked and subjected to labeling storage; the technical problems of single-dimensional analysis, lack of dynamic monitoring ability and insufficient emoticon utilization in the existing emotion analysis technology are solved, multi-dimensional, dynamic and refined emotion analysis is achieved, the precision, depth and early warning ability of emotion analysis are remarkably improved, and the emotion analysis efficiency is improved. And meanwhile, the technical effects of traceability and expandability of the system are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet technologies, and in particular, to a time-domain secondary sentiment analysis method and system based on text symbolization. Background Art

[0002] With the rapid development of Internet technologies, social media platforms (such as Weibo, Twitter, etc.) have become important channels for people to express their opinions and emotions. On these platforms, users not only express emotions through text, but also widely use emojis to enhance the effect of emotional expression. Therefore, analyzing the sentiment tendency in social media texts is of great significance for understanding public sentiment, monitoring public opinion dynamics, and responding to emergencies.

[0003] The existing sentiment analysis technologies mainly focus on classifying the positive, negative, and neutral sentiments of texts, and there are no more fine-grained and diversified judgment criteria. These methods have the following limitations:

[0004] First of all, most sentiment analysis methods only focus on the main sentiment, ignoring the potential value of secondary sentiments, and it is difficult to comprehensively reflect complex emotional dynamics. Secondly, most of the existing methods are based on static text analysis, and fail to fully consider the changing trend of sentiment over time, unable to effectively capture the changing trend of sentiment over time, and it is difficult to achieve dynamic monitoring and early warning of sentiment. Moreover, emojis are an important part of emotional expression, but the existing methods often regard them as auxiliary information and fail to fully explore their association with text sentiment. In addition, in a dynamic social media environment, the change of sentiment is often accompanied by the emergence of emotional inflection points. The existing methods are difficult to accurately identify emotional inflection points and cannot provide timely early warnings for public opinion monitoring and crisis management.

[0005] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention

[0006] Embodiments of the present application provide a time-domain secondary sentiment analysis method and system based on text symbolization, so as to at least solve the technical problems of single-dimensional analysis, lack of dynamic monitoring ability, and insufficient utilization of emojis in the existing sentiment analysis technologies.

[0007] According to one aspect of an embodiment of the present application, a time-domain secondary sentiment analysis method based on text symbolization is provided, including: obtaining text data containing emoticons, and preprocessing the text data; constructing an emoticon tag library and a Chinese vocabulary library based on the preprocessed text data; wherein the emoticon tag library is used to correspond emoticons to specific emotion categories, and the Chinese vocabulary library is used to perform word segmentation and feature extraction on the preprocessed text data; encoding the preprocessed text data, and performing sequence filling on the encoded data; performing emoticon symbolization analysis on the sequence-filled data using a deep learning model, and mining potential emotions in combination with a time-domain secondary screening mechanism; wherein the deep learning model is trained based on the emoticon tag library and the Chinese vocabulary library; tracking and judging emotional turning points according to a data source trigger mechanism, and labeling the analysis results into a database.

[0008] Optionally, the text data is preprocessed, including: performing primary screening of the text data through a filtering rule library to remove redundant characters, topic tags and non-Chinese characters; performing deduplication processing on the text data that has passed the primary screening through a deduplication rule library; performing secondary processing on the deduplication processed text data through a symbol extraction library to locate and extract emoticons in the text; and performing tertiary adaptation on the secondary processed text data through a complex sample processing pool to convert one-to-many text into multiple sentences.

[0009] Optionally, the preprocessed text data is used to construct an emoticon tag library and a Chinese vocabulary library respectively, including: dividing the preprocessed text data into an emoticon set and a text data set; screening and counting the emoticon set to construct the emoticon tag library; analyzing and counting the text data set to construct the Chinese vocabulary library; and using a historical database to regularly update the Chinese vocabulary library by loading a training model.

[0010] Optionally, the preprocessed text data is encoded, including: filtering and screening the preprocessed text data to remove irrelevant characters and redundant information; performing sentence segmentation on the filtered and screened text data to split the text into independent vocabulary units; mapping each vocabulary unit in the segmented short text data to a unique digital code through a preset coding library; organizing the mapped digital codes in a list form, and representing the text data in the form of a one-hot encoding.

[0011] Optionally, the time domain secondary screening mechanism includes: tracking the initial weight value of emotional expressions in different window periods in the time domain space and recording them with labels; using a sliding time window to determine the secondary weight value of the next occurrence of emotional expressions, labeling it, and recording the increase with the initial value, screening out secondary amplified emotional expressions and marking them.

[0012] Optionally, the time domain secondary screening mechanism adopts a probability floating method in the time domain to predict potential emotions, including: calculating the probability value corresponding to each emotion label in the first time window; calculating the probability value corresponding to each emotion label in the second time window, and calculating the difference with the probability value in the first time window; screening out the emotion labels with increasing probability values and the largest increasing floating changes as secondary emotion labels, so as to identify potential emotion trends.

[0013] Optionally, sequence padding is performed on the encoded data, including: determining the maximum length of the data; padding the data whose length does not exceed the maximum length with 0 as the padding element; truncating the data whose length exceeds the maximum length and retaining the first maximum length elements.

[0014] Optionally, tracking and analyzing emotional inflection points is performed according to a data source trigger mechanism, and the analysis results are labeled and stored in a database, including: determining a data source trigger mechanism, the data source trigger mechanism including setting a monitored time period, a monitored subject or topic, and a monitored user account; collecting corresponding text data according to the determined data source trigger mechanism; performing emotional analysis on the collected text data to obtain emotional labels and probability values of the emotional labels; determining emotional inflection points according to the emotional labels and the probability values of the emotional labels; wherein the emotional inflection points are points where the probability values of the emotional labels change significantly; analyzing the determined emotional inflection points and analyzing the causes and trends of emotional changes; labeling the analysis results to generate labeled data, and storing the labeled data in a database.

[0015] According to another aspect of an embodiment of the present application, a time-domain secondary sentiment analysis system based on text symbolization is provided, comprising: a data acquisition and preprocessing module, used to acquire text data containing emoticons and preprocess the text data; a database construction module, used to construct an emoticon tag library and a Chinese vocabulary library based on the preprocessed text data; wherein the emoticon tag library is used to correspond emoticons to specific emotion categories, and the Chinese vocabulary library is used to perform word segmentation and feature extraction on the preprocessed text data; a data encoding and sequence filling module, used to encode the preprocessed text data, and to perform sequence filling on the encoded data; a sentiment analysis module, used to perform emoticon symbolization analysis on the sequence-filled data using a deep learning model, and to mine potential emotions in combination with a time-domain secondary screening mechanism; wherein the deep learning model is trained based on the emoticon tag library and the Chinese vocabulary library; an emotional inflection point tracking and judgment module, used to track and judge emotional inflection points according to a data source trigger mechanism, and to label and store the analysis results.

[0016] According to another aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor, and a memory storing a program, wherein the program comprises instructions, and when the instructions are executed by the processor, the processor executes the time-domain secondary sentiment analysis method based on text symbolization described above.

[0017] In an embodiment of the present application, by constructing an expression tag library and a Chinese vocabulary library, the text data is preprocessed, encoded and sequenced, and a deep learning model is used in combination with a time-domain secondary screening mechanism to mine potential emotions. At the same time, the emotional turning points are tracked and labeled for storage, thereby solving the technical problems of single-dimensional analysis, lack of dynamic monitoring capabilities, and insufficient utilization of emoticons in existing sentiment analysis technologies, achieving multi-dimensional, dynamic and refined sentiment analysis, significantly improving the accuracy, depth and early warning capabilities of sentiment analysis, and enhancing the technical effects of traceability and scalability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other embodiments can be obtained based on these drawings without creative work.

[0019] Figure 1 A flowchart of a time-domain secondary sentiment analysis method based on text symbolization provided in an embodiment of the present application;

[0020] Figure 2 A schematic diagram of a time-domain secondary sentiment analysis system based on text symbolization provided in an embodiment of the present application;

[0021] Figure 3 A flowchart of a method for time-domain secondary sentiment analysis of Internet short text information provided in an embodiment of the present application;

[0022] Figure 4 A flowchart of an information matrix module provided in an embodiment of the present application;

[0023] Figure 5 A flowchart of the emotion tag library module provided in the embodiment of the present application;

[0024] Figure 6 A flow chart of a secondary emotion extraction strategy provided in an embodiment of the present application;

[0025] Figure 7 A flow chart of an automated emotion inflection point monitoring strategy provided in an embodiment of the present application;

[0026] Figure 8Schematic diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners

[0027] Embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the embodiments of the present application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the embodiments of the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not used to limit the protection scope of the present application.

[0028] The characteristics of Weibo comment data are that the sentences are relatively short, Internet users prefer to use emoticons to express emotions, and the interactivity is relatively strong. The emotional changes vary greatly over time. In view of such data characteristics, a time-domain secondary sentiment analysis method based on text symbolization is provided. Figure 1 Flowchart of the time-domain secondary sentiment analysis method based on text symbolization provided by the embodiment of the present application, as Figure 1 shown, the method includes the following steps:

[0029] Step S102, obtain text data containing emoticons and preprocess the text data; obtain text data containing emoticons and preprocess it, including removing irrelevant characters, word segmentation, extracting emoticons, etc., to improve the quality and usability of the data.

[0030] The above text data containing emoticons includes but is not limited to Weibo comment data.

[0031] Step S104, construct an emoticon label library and a Chinese vocabulary library based on the preprocessed text data; wherein, the emoticon label library is used to correspond emoticons to specific emotion categories, and the Chinese vocabulary library is used to perform word segmentation and feature extraction on the preprocessed text data.

[0032] Step S106, perform encoding processing on the preprocessed text data, and perform sequence padding processing on the encoded data; perform encoding processing on the preprocessed text data, map each lexical unit to a unique digital code, and perform sequence padding processing on the encoded data to meet the input requirements of the deep learning model.

[0033] Step S108, using a deep learning model to analyze the data after sequence filling processing into emoticons, and combining the time domain secondary screening mechanism to mine potential emotions; wherein the deep learning model is trained based on the emoticon tag library and the Chinese vocabulary library; using a deep learning model (such as LSTM, BERT, etc.) to analyze the data after sequence filling processing into emoticons, and combining the time domain secondary screening mechanism to mine potential emotions. The model is trained based on the emoticon tag library and the Chinese vocabulary library, and can more accurately identify the emotional tendency in the text.

[0034] Step S110, track and judge the emotional inflection points according to the data source trigger mechanism, and label the analysis results and store them in the database. According to the data source trigger mechanism (such as monitoring time period, theme or topic, user account, etc.), collect the corresponding text data and perform emotional analysis. By analyzing the changes in emotional labels and their probability values, determine the emotional inflection points, judge the inflection points, and analyze the causes and trends of emotional changes. Finally, label the analysis results and store them in the database for subsequent query and tracing.

[0035] In the embodiments of the present application, by introducing secondary sentiment analysis, the present application can pay attention to the changes of primary and secondary emotions at the same time, provide a more comprehensive sentiment analysis perspective, and solve the problem of single sentiment analysis dimension in the prior art. Combined with the time domain secondary screening mechanism, the present application can track the changing trend of emotions in real time, identify the inflection points of emotions, realize dynamic monitoring and early warning of emotions, and solve the problem that the prior art lacks dynamic monitoring capabilities. By constructing an emoticon tag library and a Chinese vocabulary library, the present application makes full use of the association between emoticons and text emotions, improves the accuracy and depth of sentiment analysis, and solves the problem of insufficient utilization of emoticons in the prior art. Through the tracking and judgment of emotional inflection points, the present application can timely discover potential changes in emotions, provide early warning and decision-making support for public opinion monitoring and crisis management, and solve the problem of difficulty in identifying emotional inflection points in the prior art. The analysis results are labeled and stored in the database. The present application not only supports subsequent queries and tracing, but also has good scalability and can adapt to different application scenarios and data scales.

[0036] As an optional embodiment, the text data is preprocessed, including: performing primary screening of the text data through a filtering rule library to remove redundant characters, topic tags, and non-Chinese characters; performing deduplication processing on the text data that has passed the primary screening through a deduplication rule library; performing secondary processing on the deduplication processed text data through a symbol extraction library to locate and extract emoticons in the text; and performing tertiary adaptation on the secondary processed text data through a complex sample processing pool to convert one-to-many text into multiple sentences.

[0037] Optionally, construct a filtering rule library that contains regular expression rules for identifying and removing redundant characters, hashtags, and non-Chinese characters. Scan the original text data, apply the regular expressions in the filtering rule library, and remove the characters and tags that do not meet the requirements.

[0038] In the embodiments of the present application, removing irrelevant characters and tags reduces data noise and improves the efficiency and accuracy of subsequent processing. Retaining Chinese characters and necessary whitespace characters provides a cleaner text for subsequent word segmentation and semantic analysis.

[0039] Optionally, construct a deduplication rule library that defines a method for uniquely identifying text, such as based on the hash value of the text or the value of a specific field. Scan the text data that has undergone the first-level screening, calculate the hash value of each piece of text. Compare the hash values and remove duplicate text records, retaining only one.

[0040] In the embodiments of the present application, removing duplicate text avoids duplicate calculations in subsequent analysis and improves analysis efficiency. Ensuring the independence of each piece of text data avoids analysis biases caused by duplicate data.

[0041] Optionally, construct an emoji extraction library that contains rules for identifying and extracting emojis, such as regular expressions or symbol mapping tables. Scan the text data that has undergone deduplication processing, and use the rules in the emoji extraction library to locate emojis in the text. Extract the emojis and store them separately from the text content for subsequent analysis.

[0042] In the embodiments of the present application, separating emojis from the text content facilitates independent sentiment analysis of emojis. Emojis are an important part of emotional expression, and extracting the symbols can more accurately analyze the emotional tendency of the text.

[0043] Optionally, construct a complex sample processing pool that contains rules and algorithms for processing complex text structures, such as text splitting and recombination algorithms. Scan the text data that has undergone the second-level processing, and identify complex text samples that contain multiple emotional expressions. Split the complex text samples into multiple independent statements, with each statement corresponding to one emotional expression. For example, split "This is really too much fun and I'm addicted to it [tear][bad smile]" into "This is really too much fun and I'm addicted to it [tear]" and "This is really too much fun and I'm addicted to it [bad smile]".

[0044] In the embodiments of the present application, splitting complex text into multiple statements increases the number of samples in the corpus and improves the training effect of the model. By splitting complex text, it better meets the input requirements of the sentiment analysis model and improves the accuracy of sentiment analysis.

[0045] As an optional embodiment, the preprocessed text data is used to respectively construct an emoticon tag library and a Chinese vocabulary library, including: dividing the preprocessed text data into an emoticon set and a text data set; screening and counting the emoticon set to construct an emoticon tag library; analyzing and counting the text data set to construct a Chinese vocabulary library; and using a historical database to regularly update the Chinese vocabulary library by loading a training model.

[0046] Optionally, the preprocessed text data is scanned to separate the emoticons and plain text content therein. The emoticons are stored as an emoticon set, and the plain text content is stored as a text data set. For example, for the text "This is really fun [smirk]", "[smirk]" is extracted into the emoticon set, and "This is really fun" is stored into the text data set.

[0047] Perform statistical analysis on the emoji set and calculate the frequency of each emoji. According to the emotion classification theory and the daily usage habits of users, the emojis are divided into different emotion categories (such as "joy", "anger", "sadness", etc.). Build an emoji tag library to map each emoji with its corresponding emotion category. For example, map "[laughing and crying]" to "joy" and map "[anger]" to "anger". Update the emoji tag library regularly to adapt to new emojis and emotional expressions.

[0048] Perform word segmentation on the text dataset and split the text into independent vocabulary units. Count the frequency of each word, and filter out words related to sentiment analysis based on part-of-speech tagging and semantic analysis. Build a Chinese vocabulary library to store high-frequency words and their sentiment tendencies (such as "happy" corresponds to "positive emotion" and "angry" corresponds to "negative emotion"). Use natural language processing tools (such as Jieba word segmentation) for word segmentation and part-of-speech tagging.

[0049] Build a historical database to store text data and sentiment annotation results from past analyses. Extract data from the historical database regularly and load it into the training model for retraining. Use machine learning algorithms (such as Word2Vec, BERT, etc.) to update the Chinese vocabulary library and optimize the sentiment annotation and weight of the vocabulary. For example, through incremental learning, combine new data with the old model to update the sentiment weight of the vocabulary in the Chinese vocabulary library.

[0050] In the embodiments of the present application, by dividing the text data into an emoji set and a text data set, refined processing of the data is achieved, providing a clearer data basis for subsequent sentiment analysis. The mapping relationship between emojis and emotion categories makes sentiment analysis more accurate and enables better capture of the user's emotional tendency. Regularly updating the emoji label library can adapt to new emojis and emotional expression ways, ensuring the timeliness and accuracy of the system. Through word segmentation and sentiment annotation, the Chinese vocabulary library can better understand the emotional meaning of the text, improving the accuracy of sentiment analysis. Regularly updating the Chinese vocabulary library, combined with historical data and machine learning models, can continuously optimize the emotional weights of the vocabulary to adapt to the dynamic changes of the language. By constructing and optimizing the emoji label library and the Chinese vocabulary library, the sentiment analysis system can more accurately identify and understand the emotional information in the text.

[0051] As an alternative embodiment, the preprocessed text data is encoded, including: filtering and screening the preprocessed text data to remove irrelevant characters and redundant information; performing sentence word segmentation on the text data after filtering and screening to split the text into independent lexical units; mapping each lexical unit in the segmented short text data to a unique digital code through a preset coding library; organizing the mapped digital codes in a list form and representing the text data in a one-hot encoding form.

[0052] Optionally, use regular expressions or other text processing tools to further clean the text data and remove irrelevant characters (such as special symbols, HTML tags, URL links, etc.) that may interfere with the analysis. Remove duplicate words or phrases to avoid the impact of redundant information on subsequent analysis.

[0053] Use Chinese word segmentation tools (such as jieba segmentation, HanLP, etc.) to perform word segmentation on the text, splitting the sentence into independent lexical units. Clean the word segmentation results to remove stop words (such as "de", "shi", "he", etc.), which usually contribute less to sentiment analysis.

[0054] Construct a coding library that contains a vocabulary list and corresponding digital codes. The vocabulary list can be obtained through statistical analysis of a large amount of corpus. For each lexical unit after word segmentation, look up the corresponding digital code in the coding library. If the word is not in the coding library, assign a special "unknown" code. Replace each lexical unit with its corresponding digital code to form an encoded sequence.

[0055] Organize the encoded digital sequence into a list form, such as [1, 2]. Use one-hot encoding to convert each digital encoding into a vector of fixed length. Assuming the vocabulary size is 10, the number 1 is represented as [0, 1, 0, 0, 0, 0, 0, 0, 0, 0], and the number 2 is represented as [0, 0, 1, 0, 0, 0, 0, 0, 0, 0]. Represent the entire text data as a two-dimensional one-hot encoding matrix for subsequent processing by machine learning or deep learning models.

[0056] In the embodiments of the present application, by filtering and screening to remove irrelevant characters and redundant information, the interference of noise data on sentiment analysis can be reduced, and the data quality can be improved. Word segmentation processing splits the text into independent lexical units, which can better capture the semantic information of the text and provide a finer-grained input for subsequent sentiment analysis. Mapping the vocabulary to digital encoding can convert the text data into a numerical form that can be processed by machines, facilitating subsequent model training and prediction. By converting the text data into a vector of fixed length through one-hot encoding, the problem of inconsistent text lengths can be eliminated, and at the same time, the independence of the vocabulary can be retained, providing a standardized input for deep learning models. The encoded data can be directly input into machine learning or deep learning models, reducing the complexity of data preprocessing and improving the speed of model training.

[0057] As an alternative embodiment, the time-domain secondary screening mechanism includes: tracking the first weight value of the emotional expression in different time window periods in the time-domain space and recording the label; using a sliding time window to determine the secondary weight value of the next occurrence of the emotional expression for labeling, and recording the increase with the first value, screening out the secondary increased emotional expressions and making marks.

[0058] Optionally, according to the update frequency and application scenario of the data, set the size (such as 1 hour, 2 hours) and step size (such as 15 minutes, 30 minutes) of the sliding time window. For example, for Weibo comment data, 1 hour can be selected as the time window size and 30 minutes as the step size. Within each time window, perform sentiment analysis on the text data, calculate the probability value of each emotion label (such as "joy", "anger", "sadness"), and record it as the first weight value. Record each emotion label and its corresponding first weight value into the database to form time series data. For example, the emotion labels and their weight values within time window 1 are: {"joy": 0.8, "anger": 0.2, "sadness": 0.1}.

[0059] In the next time window, repeat the emotion weight calculation steps to get a new weight value (secondary weight value). For each emotion label, calculate the difference (increase) between the secondary weight value and the first weight value. For example, the emotion weight value in time window 2 is: {"joy":0.7,"anger":0.4,"sadness":0.15}, then the increase is: {"joy":-0.1,"anger":+0.2,"

[0060] Sadness":+0.05}. Filter out the emotion labels with the largest increase and mark them as "secondary emotions". For example, in the above example, the "angry" emotion has the largest increase (+0.2), so it is marked as a secondary emotion.

[0061] Add special tags to secondary emotions in the database to record their changing trajectories in the time window. Store the weight value, increase value and timestamp of secondary emotions in the database for subsequent analysis and tracking. Display the changing trajectories of secondary emotions in the form of charts to help analysts intuitively understand the dynamic changes of emotions.

[0062] In the embodiment of the present application, through the secondary screening mechanism in the time domain, the changing trend of emotions can be monitored in real time, and potential emotional fluctuations can be discovered in time. For example, in social media public opinion monitoring, secondary emotions that may trigger public opinion storms can be warned in advance, providing decision makers with more timely warning information. Traditional sentiment analysis usually only focuses on the main emotions, while the secondary emotion screening mechanism can capture the changes in secondary emotions and provide a more comprehensive perspective of sentiment analysis. For example, in complex social events, secondary emotions may reveal potential dissatisfaction or anxiety, which may not be obvious in the early stage, but may become dominant emotions over time. By sliding the time window and increasing the record, the changing trend of emotions can be captured more accurately, avoiding misjudgments caused by emotional judgments at a single time point. For example, when the change in the emotional weight value is not obvious, the increase record can reveal the potential changes in emotions and improve the accuracy of the analysis. By recording the emotional weight value and the increase value, the evolution trajectory of emotions can be tracked, providing historical data support for sentiment analysis. For example, when analyzing the emotional evolution of an event, the change trajectory of secondary emotions can be used to understand how emotions gradually evolve from secondary emotions to dominant emotions.

[0063] As an optional embodiment, the time domain secondary screening mechanism adopts a probability floating method in the time domain to predict potential emotions, including: calculating the probability value corresponding to each emotion label in the first time window; calculating the probability value corresponding to each emotion label in the second time window, and calculating the difference with the probability value in the first time window; screening out the emotion labels with increasing probability values and the largest increasing floating changes as secondary emotion labels to identify potential emotion trends.

[0064] Optionally, according to the dynamic characteristics of the data and the analysis requirements, a fixed time window length (such as 1 hour) is set as the initial analysis interval. Perform sentiment analysis on the text data within the current time window, and use deep learning models (such as LSTM, BERT, etc.) to calculate the probability values of each emotion label. Record each emotion label and its corresponding probability value to form an initial emotion probability distribution.

[0065] Slide the time window to the next interval (such as 1 hour later), and perform the same sentiment analysis on the text data within the new time window. Calculate the probability values of each emotion label again. Compare the probability values of the second time window with those of the first time window, and calculate the probability difference of each emotion label.

[0066] Filter out the emotion labels with increasing probability values from the difference results. Among the emotion labels with increasing probability, select the label with the largest increase in floating change as the secondary emotion label. In the above example, the probability of the "angry" label rises from 0.1 to 0.25, with a difference of +0.15, which is the emotion label with the largest increase in floating change, so it is identified as the secondary emotion. Record the secondary emotion label and its related information (such as timestamp, probability value, difference, etc.) in the database for subsequent analysis.

[0067] Analyze the potential direction of emotion evolution based on the change trend of the secondary emotion label. For example, if the "angry" emotion continues to rise in multiple consecutive time windows, it may indicate the accumulation of a certain negative emotion. Display the emotion probability values and the change trend of the secondary emotion in the form of a chart to help analysts intuitively understand the dynamic changes of emotions.

[0068] In the embodiment of this application, through the time-domain probability floating method, it is possible to capture the subtle changes of emotions in advance, especially those secondary emotions that have not yet become the dominant emotion but are rising rapidly. This helps to identify potential emotion risks in advance and provide early warnings for scenarios such as public opinion monitoring and crisis management. Traditional sentiment analysis usually only focuses on the static distribution of the current emotion, while the time-domain secondary screening mechanism can dynamically track the change trend of emotions and provide a richer analysis dimension. For example, in social media analysis, it is possible to more accurately capture the evolution process of emotions, rather than just the current emotion state. By calculating the probability difference, it is possible to more accurately identify the change trend of emotions and avoid misjudgment caused by the emotion fluctuations at a single time point. This method can more sensitively capture the potential changes of emotions and improve the overall accuracy of sentiment analysis.

[0069] As an alternative embodiment, sequence padding processing is performed on the encoded data, including: determining the maximum length of the data; padding the data with a length not exceeding the maximum length at the rear with padding elements being 0; truncating the data with a length exceeding the maximum length at the rear, and retaining the first maximum length of elements.

[0070] Optionally, in the preprocessed text data, the lengths of all text sequences (i.e., the number of lexical units) are counted. According to the statistical results, a suitable maximum length (max_length) is selected. For text sequences with a length less than max_length, 0 is padded at the end of the sequence until its length reaches max_length. For example, array operations in programming languages or dedicated machine learning libraries (such as TensorFlow or PyTorch) are used for padding. For text sequences with a length exceeding max_length, the extra parts are truncated from the end, and only the first max_length elements are retained.

[0071] In the embodiments of this application, through sequence padding and truncation, the lengths of all text sequences are unified to max_length, standardizing the format of the input data. This helps the model better process texts of different lengths and avoid computational problems caused by length differences. The standardized input data can significantly improve the speed and efficiency of model training. Especially when using deep learning models (such as RNN, LSTM, Transformer), the unified sequence length can reduce the waste of computational resources and avoid the negative impact of overly long sequences on the training process. Through padding and truncation, the model can learn more general text features instead of relying on texts of specific lengths.

[0072] As an alternative embodiment, sentiment inflection point tracking and judgment are performed according to the data source triggering mechanism, and the analysis results are tagged and stored in the database, including: determining the data source triggering mechanism, where the data source triggering mechanism includes setting the monitored time period, monitored theme or topic, and monitored user account; collecting corresponding text data according to the determined data source triggering mechanism; performing sentiment analysis on the collected text data to obtain sentiment labels and the probability values of the sentiment labels; determining the sentiment inflection point according to the sentiment labels and the probability values of the sentiment labels, where the sentiment inflection point is the point where the probability value of the sentiment label changes significantly; judging the determined sentiment inflection point, analyzing the reasons and trends of the sentiment change; performing tagging processing on the analysis results to generate tag data, and storing the tag data in the database.

[0073] Optionally, set the time range to be monitored, such as "the past 24 hours" or "from 9 am to 5 pm from Monday to Friday every week". Determine the hot topics or keywords to be concerned about, such as "the release of a new product of a certain brand" or "a certain public event". Specify the user accounts that need to be focused on, such as "opinion leader accounts" or "accounts of specific groups". According to the monitoring parameters, set the triggering conditions for data collection. For example, when content related to the specified topic appears within the monitored time period, trigger data collection.

[0074] According to the triggering mechanism, collect text data that matches the monitoring parameters from the specified data sources (such as social media platforms, news websites, forums, etc.). Store the collected text data in a database, and record metadata such as its source, release time, and user information.

[0075] Use a pre-trained sentiment analysis model (such as an emotion classification model based on LSTM or BERT) to analyze the text data. For each piece of text data, calculate its emotion label (such as "joy", "anger", "sadness") and its corresponding probability value. Store the emotion label and probability value in the database for subsequent analysis.

[0076] Calculate the change rate of the emotion label probability value, and detect significant change points (such as the probability value change exceeding the set threshold). Mark the points with significant probability value changes as emotion inflection points, and record their timestamps and emotion labels.

[0077] Combine information such as text content, release time, and user background to analyze the reasons for emotion changes. According to the change trend of the emotion inflection points, predict the possible future trend of emotions.

[0078] Convert the analysis results into a labeled data format for subsequent query and analysis. Store the labeled data in the database to support subsequent emotion monitoring and trend analysis.

[0079] In the embodiments of the present application, through the data source triggering mechanism, it is possible to accurately collect text data related to specific topics or users, avoid the interference of irrelevant information, and improve the pertinence of emotion monitoring. The tracking and research of emotion inflection points can capture the change trend of emotions in real time, provide dynamic emotion analysis, and help analysts understand the evolution process of emotions in a timely manner. Through the analysis of emotion inflection points, potential emotion risks or opportunities can be discovered in advance, providing early warnings and decision-making support for scenarios such as public opinion monitoring and crisis management. Combining the changes in emotion labels and probability values can more deeply analyze the reasons and trends of emotion changes, provide richer analysis dimensions, and support complex emotion analysis requirements. Labeling and storing the analysis results in the database is convenient for subsequent query and traceability, and at the same time supports the expansion and upgrade of the system to adapt to different application scenarios and data scales.

[0080] According to another aspect of an embodiment of the present application, a time-domain secondary sentiment analysis system based on text symbolization is provided. Figure 2 A schematic diagram of a time-domain secondary sentiment analysis system based on text symbolization provided in an embodiment of the present application, such as Figure 2 As shown, the time-domain secondary sentiment analysis system based on text symbolization includes: a data acquisition and preprocessing module 202, a database construction module 204, a data encoding and sequence filling module 206, a sentiment analysis module 208 and a sentiment turning point tracking and judging module 210. The time-domain secondary sentiment analysis system based on text symbolization is described in detail below.

[0081] The data acquisition and preprocessing module 202 is used to acquire text data containing emoticons and preprocess the text data;

[0082] A database construction module 204 is used to construct an expression tag library and a Chinese vocabulary library based on the preprocessed text data; wherein the expression tag library is used to correspond expression symbols to specific emotion categories, and the Chinese vocabulary library is used to perform word segmentation and feature extraction on the preprocessed text data;

[0083] The data encoding and sequence filling module 206 is used to encode the pre-processed text data and perform sequence filling on the encoded data;

[0084] The sentiment analysis module 208 is used to perform expression symbol analysis on the sequence-filled data using a deep learning model, and to mine potential emotions in combination with a time-domain secondary screening mechanism; wherein the deep learning model is trained based on an expression tag library and a Chinese vocabulary library;

[0085] The emotion turning point tracking and judging module 210 is used to track and judge the emotion turning point according to the data source trigger mechanism, and label the analysis results and store them in the database.

[0086] It should be noted here that the above-mentioned data acquisition and preprocessing module 202, database construction module 204, data encoding and sequence filling module 206, sentiment analysis module 208 and sentiment turning point tracking and judgment module 210 correspond to steps S102 to S110 in the method embodiment, and the examples and application scenarios implemented by the above-mentioned modules and corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned method embodiment.

[0087] The following uses Internet short text information as an example to explain the time-domain secondary sentiment analysis method of the present application in detail.

[0088] The embodiments of the present application provide a time-domain secondary sentiment analysis technical solution for massive data of short text information. By using a deep learning model to build a dedicated model library, establishing a text symbolization conversion channel, and introducing a brand-new "secondary emotion" concept, an automated emotion inflection point supervision strategy analysis method is constructed, which is specifically described as follows:

[0089] Figure 3 It is a flowchart of the time-domain secondary sentiment analysis method for Internet short text information provided by the embodiments of the present application. As Figure 3 shown, based on the HadoopSpark underlying cluster big data platform, the SpringCloud microservice end receives and reads the task data dispatched from upstream, uploads it to HDFS after basic data processing, listens to the HDFS distribution to the computing service of the Spark cluster periodically through Sparkstreaming, converts the text information into fine-grained emotion tags through the time-domain secondary sentiment analysis module, writes the analysis result in the form of key-value pairs into Redis, and synchronizes the task ID to Kafka, thereby realizing the process context of data in time-domain secondary sentiment analysis. Through the automated emotion inflection point supervision strategy analysis and judgment model, the input data is analyzed and judged, and the strategy content is displayed to the upstream data receiving end, achieving accurate analysis and timely judgment. Among them, the time-domain sentiment analysis method module mainly includes three parts: the information matrix module, the emotion tag library, and the time-domain secondary sentiment analysis module. The specific implementation methods are as follows:

[0090] Figure 4 It is a flowchart of the information matrix module provided by the embodiments of the present application. As Figure 4 shown, the information matrix module mainly preprocesses the data. The original information has certain redundancy and noise, and a series of methods are required for preprocessing to achieve better calculation results. It includes a filtering rule library, a deduplication rule library, a symbol extraction library, a complex sample processing pool, etc. After a series of processing, a unique language library that meets later matching is statistically formed. This language library is processed in a progressive manner. After the first-level screening, for example, redundant characters, topic tags, non-Chinese characters, etc. are filtered and deduplicated; the second-level processing is carried out. For example, the emoticons usually shown in the text are marked with "[]", and the emoticons are extracted by quickly locating this mark; finally, the third-level adaptation is carried out, that is, the one-to-many text is converted into multiple sentences. This not only enriches the content of the corpus but also is a means of amplifying the text data, enhancing the learning ability of the corpus. For example, "This is really too much fun and I'm addicted to it [tear][bad smile]" is converted into "This is really too much fun and I'm addicted to it [tear]" and "This is really too much fun and I'm addicted to it [bad smile]".

[0091] Figure 5 It is a flowchart of the emotion tag library module provided by the embodiments of the present application. As Figure 5As shown in the figure, the emotion label library is divided into two parts. One is to count the Weibo expressions extracted from a large amount of original data, and according to the emotion classification theory and users' daily usage habits, the expression system is divided into more than a hundred. The other is divided into 7 major categories according to the upstream output and display to express a fine-grained emotion system. Among them, a unique offline training mode is designed, that is, the expression library and the expression classification library are updated at fixed time intervals.

[0092] The time-domain secondary emotion analysis method is mainly a deep learning model for emoji analysis plus a time-domain secondary screening mechanism, which is an emotion analysis method that converts text information into the most appropriate and relevant emojis. The biggest difference between the training of the text emoji model and the previous ordinary text training model is that, first, the pure text prediction training library is transformed by self-built coding, that is, the preprocessed data is encoded. The encoding principle is to form a set after segmenting a large amount of corpus, and this set is encoded with non-repeating numbers. Then, this set dictionary is used as the original benchmark for feeding into the model, and each new sentence is encoded into a corpus set that conforms to the training according to the dictionary form. Second, a time-domain secondary emotion screening mechanism is added. In the past, only the emotional tendency expressed by the text of a sentence itself was simply analyzed, whether it was positive, negative, or neutral, etc. This kind of emotion analysis without context is of little significance. However, this mechanism extends the time length of the information, broadens the width of the data, fully integrates it into the situational context, and analyzes the potential emotions by calculating the emotional weight value each time to obtain the emotion emoji label with the largest fluctuation under a specific topic or time dimension. This is a method not adopted by other emotion analyses.

[0093] The training framework of the model part uses the TensorFlow learning framework, and the algorithms are word vectors, LSTM long short-term memory networks, attention mechanisms, etc. The specific steps are as follows:

[0094] The first step is data preprocessing

[0095] For training a model using Internet data to achieve the expected effect, the first step is always to perform data screening. In this stage, data cleaning, screening, and complex processing are carried out. For the short text data of Weibo comments, screening and processing are required, such as regular matching, symbol de-duplication, text redundancy, simplification of complex samples, symbol extraction, etc.

[0096] The second step is text feature extraction and training

[0097] First, divide the processed data into a dataset, including a training set, a validation set, and a test set. Encode the text into training data in the form of [[xx, xx, xx], a] according to the unique vocabulary and label libraries. After BatchNormalization, use a two-layer LSTM network for feature extraction, with outputs lstm_0_output and lstm_1_output respectively. Then, perform dimensional fusion with the original x input, calculate the average attention weights to obtain the weighted attention average value, and finally obtain each weight as the output through softmax. When the monitored value does not improve, the training function stops training and saves the current latest model.

[0098] Based on the emoji analysis model trained offline as a benchmark, a time-domain secondary screening mechanism is carried out during data prediction to obtain the emoji weight probabilities and corresponding emoji label markings under different time windows. This mechanism makes up for the pain point of simply using a deep model to obtain the emotional tendency at the time of emoji prediction without fully exploiting the internal information of the data. Combining the two can show the potential changes of the data over time according to the actual data scenario. The specific steps are as follows:

[0099] The first step is data encoding.

[0100] When new business data monitored by Spark needs to be analyzed and calculated, these raw data are uniformly processed. First, filter and screen them, then perform sentence tokenization, and finally convert the string data into a list through a unique encoding library and display it in the form of one-hot encoding.

[0101] The second step is padding sequence processing.

[0102] Since the lengths of the data after only one-hot encoding are different, in order to ensure the regularity and uniformity of the predicted data input, this step needs to pad the data. The padding parameter set here has maxlen as

[0103] max_length, and the padding position of padding is post (padding at the rear), and the truncating position of truncating the sequence is post (truncating at the rear of the sequence).

[0104] The third step is mechanism-based prediction.

[0105] After the data is prepared, prediction is carried out. Different from the previous conventional direct prediction, this solution uses a probability floating method in the time domain for output. Direct prediction can only see the expression emotion with the largest probability value, drowning out the expression emotions with an upward trend in emotion over time. For example, in the first time window, the calculated probability values corresponding to "[Surprise], [Anger], [Worry]..." are "0.95, 0.55, 0.35..." and in the second time window, the calculated probability values corresponding to "[Surprise], [Anger], [Worry]..." are "0.90, 0.76, 0.46..." (Note: The numbers here are for simplified illustration and not the actual calculated values). It can be seen that the expression symbol with the largest probability value in both calculations is '[Surprise]'. However, for the expression symbols ranked second and third, the probability values are increasing, and the one with the largest increase in the floating change is '[Anger]'. Thus, it is very likely to dominate the emotion as time goes by. From this, it can be seen that it is a potential expression symbol. For platforms like Weibo comments where the public opinion trend changes rapidly, grasping the potential emotional trend means seizing the opportunity of the trend. (Note: The specific method of the secondary emotion mechanism will be introduced in detail below). Set the number of predicted outputs according to business requirements. Here, the top 10 predicted output probability weights are preset. Each time after the analysis and calculation request initiated according to the time window, expression prediction is carried out, and the first weight, secondary weight, and the difference between the two weights are recorded and stored in sequence until the end of the time window.

[0106] Step 4, convert the output symbol image into text

[0107] The symbol representations and probability values predicted in the third step are matched to the corresponding labels according to the emoji labels in the emotion label library. After converting them into key-value pairs, the output result is the custom field dictionary value. Then, operations such as merging and sorting are performed according to the unique data identification code ID and written into Redis, and the task ID is synchronized to Kafka. Specifically, for example, after predicting and calculating a period of data, the top 10 symbol identifications and probability values of this group of data are obtained, in the form of: [{47, 0.07466985285282135, -0.002358945}, {26, 0.07886634767055511, 0.003625468},...]. The first number represents the symbol in the emoji library (the specific meaning is described in the "emotion label library"), the second group of numbers represents the weight of this emoji, and the third number represents the difference between the weight of this emoji and the previous time. The initial calculation can be set to null (the specific meaning is described in the "2. Secondary emotion extraction strategy"). Then, the corresponding emojis are traversed according to the form in the emotion label library: [{[angry scolding], 10}, {[bye], 11}, {[kneeling], 12}...]. Finally, the custom dictionary result after key-value conversion is in the form of: [{'result': [{'name': '[watching]', 'weight': 0.07466985285282135, 'categoryid': '6', 'diff': 0.002614352},...]}, where name represents the emoji, weight represents the weight, categoryid represents the specific category in the given emotion category library (the specific meaning is described in the "emotion label library"), and diff represents the difference between the calculated weight and the previous time. Finally, the unique identification code ID generated according to the time and data is merged with the above custom dictionary.

[0108] Secondary emotion extraction is based on the time domain space. According to different time periods of the time window, the first weight value of the emotion emoji in this time period is tracked and labeled. By sliding the time window, the secondary weight value of the next occurrence of the emotion emoji is determined and labeled, and the increase is recorded by comparing it with the first value. The secondary increased emotion emojis are screened out and marked, and so on within the set time window. Figure 6 The flowchart of the secondary emotion extraction strategy provided by the embodiment of this application is as Figure 6 shown, and the specific steps are as follows:

[0109] The first step, the time window is triggered

[0110] When the task data sender triggers the time window mechanism, the secondary emotion extraction strategy works. Considering the active time of short text comments, the data update collection frequency, and the actual application scenario, the time sliding window is set to an adjustable n hours, and recalculation is performed every half hour. The time information is encoded as part of the identification ID for subsequent uses such as storage record screening. The specific time setting can be controllably selected according to the actual application scenario and reasonably adjusted according to the data update frequency. Since this data collection system uses a half-hour frequency, that is, half an hour is the minimum step of the time window, the interception of the time window life cycle can be selected according to the actual situation, which can be a short-term time or a long-term time, and is reasonably selected according to the task.

[0111] Step 2, weight and difference recording

[0112] When the time window mechanism is triggered for the first time, the first weight data is obtained through calculation and analysis, marked as p1, p2, p3... After the second time the time mechanism is triggered and the calculation is completed, the secondary weight data and the difference in the numerical fluctuation from the previous time are recorded. The weights are marked as p11, p21, p31... and the differences are marked as ΔP1, ΔP2, Δp3... and so on until the calculation within the time window is completed. Since the results obtained each time are stored in the field, that is, ΔPmax is marked each time, and state={ΔP1, ΔP2,...ΔPm} (state is the status value label of the current analysis time and number of times), the largest value in the differences is found. In the data, we will focus on the data with the largest floating difference when the weight ranking remains unchanged, which represents a relatively large data fluctuation and is the core of the secondary emotion extraction method.

[0113] The automated emotion inflection point supervision strategy analyzes and sends specified task data to the cluster according to the trigger mechanism of different data sources, starts the analysis algorithm model to perform a series of analysis and calculations, stores the results in the database, and sends the analyzed data to the front end according to the analysis and judgment rules to obtain the emotional expression evolution trajectory displayed under different strategies and conduct the tracking and judgment of the emotion inflection point. Figure 7 The flowchart of the automated emotion inflection point supervision strategy provided by the embodiment of the present application is as Figure 7 shown, and the specific steps are as follows:

[0114] Step 1, task data source selection

[0115] Due to different task perspectives, the observation strategies are also different. Here, several different data sources are set up for better analysis and calculation. Vertically, it is divided into data under the same account within the same time period, especially data under key accounts, and data under the same topic within the same time period; horizontally, it is divided into data under different accounts within the same time period and data under different topics within the same time period. The data sources to be concerned about can be selected according to actual needs. After the data sources are selected, data analysis and calculation can be carried out.

[0116] Step 2: Judgment and result acquisition

[0117] After the task data sources are selected, it enters the analysis and judgment stage. The data is sent into the calculation model to obtain results. At this time, according to the time inflection points with secondary emotions in the previous results, and according to the data dimension strategy, it can be marked as which focus, whether it is the hidden emotion fluctuation under this key account or the hidden emotion under the topic is triggered at what time. According to actual needs, draw a fluctuation attention curve for convenient correlation judgment. For example, if the user plans to focus on the potential emotional tendency of the comments of netizens under a certain topic of a certain key account within a week, then a one-week monitoring time window life cycle can be selected according to the monitoring task. Next, the emotion fluctuation can be observed within the cycle. At this time, there will be two emotion curves represented by emoji that change with time, the main emotion and the secondary emotion. Extract the label with the larger weight difference value that extends with time in the secondary emotion mechanism. At this time, the user can view the secondary emotion fluctuation situation and comment data according to the marked time point status for judgment and process according to needs.

[0118] This application discloses a time-domain secondary emotion analysis method and system based on text symbolization. Relying on the big data distributed real-time analysis framework, through the emotion symbol dictionary constructed by corpus screening, using the LSTM algorithm model to build the transformation of text symbolization, and introducing a new concept of "secondary emotion", the emotional changes of emotion and secondary emotion are derived through the self-built threshold system in the time domain space. According to the automated emotion inflection point supervision strategy analysis, it has the characteristics of stronger potential risk mining, richer content, and easier prediction and judgment of risks.

[0119] (1) This application adopts a big data parallel streaming processing framework SparkStreaming, uses the deep learning algorithm model LSTM to build a text information emoji transformation matrix, constructs a self-built threshold system in the time domain space through information control in the time domain space, visualizes the text information emotion space, improves the sensitivity of text information risk mining, solves the existing simple positive / negative / neutral judgment phenomenon of statements, and enriches the perspective of internet information emotion control.

[0120] (2) This application proposes a brand-new "secondary emotion" theory. Generally, when people read a sentence or a passage of text, they mainly focus on the main emotion expressed, and do not pay extra attention to the secondary emotion. However, in today's complex and content-rich online environment, if the information shown by the secondary emotion can be discovered in a timely manner, it is actually possible to effectively infer some potential information risks that have not yet occurred. Therefore, the "secondary emotion" theory emerged on this basis, that is, by controlling the information in the time domain space to construct a self-built threshold system within the time domain space, evolving the emotional and secondary emotional changes, creating a very novel analysis point in the field of sentiment analysis, and improving the ability to mine potential risks.

[0121] (3) This application proposes an automated emotional inflection point supervision strategy analysis method. By obtaining data inputs in different dimensions according to different monitoring strategies, it can analyze the emotional evolution trajectory at the moment of the emotional inflection point in a more fine-grained manner, and label and store the monitored and analyzed information, enabling the trajectory to be traceable, the information to be tracked, and the emotion to be perceived, solving the pain points of simple judgment types and single content angles in ordinary sentiment analysis in the existing environment, and greatly improving the controllability and sensitivity of the monitored information.

[0122] The embodiment of this application also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The above-mentioned memory stores a computer program that can be executed by the above-mentioned at least one processor, and when the above-mentioned computer program is executed by the above-mentioned at least one processor, it is used to make the electronic device execute the time-domain secondary sentiment analysis method based on text symbolization in the embodiment of this application.

[0123] The embodiment of this application also provides a non-transitory machine-readable medium storing a computer program, wherein when the above-mentioned computer program is executed by a processor of a computer, it is used to make the above-mentioned computer execute the time-domain secondary sentiment analysis method based on text symbolization in the embodiment of this application.

[0124] The embodiment of this application also provides a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, it is used to make the computer execute the time-domain secondary sentiment analysis method based on text symbolization in the embodiment of this application.

[0125] Reference Figure 8, the structural block diagram of an electronic device that can be a server or a client in an embodiment of the present application will now be described. It is an example of a hardware device that can be applied to various aspects of the present application. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0126] As Figure 8 shown, the electronic device includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0127] Multiple components in the electronic device are connected to the I / O interface 805, including: an input unit 806, an output unit 807, a storage unit 808, and a communication unit 809. The input unit 806 can be any type of device that can input information into the electronic device. The input unit 806 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 807 can be any type of device that can present information, and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 808 can include, but is not limited to, magnetic disks, optical disks. The communication unit 809 allows the electronic device to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and can include, but is not limited to, a modem, a network card, an infrared communication device, and a wireless communication transceiver, such as a Bluetooth device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0128] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a CPU, a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing units, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above. For example, in some embodiments, the method embodiments of the present application can be implemented as a computer program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device via the ROM 802 and / or the communication unit 809. In some embodiments, the computing unit 801 can be configured to execute the above methods in any other suitable manner (e.g., by means of firmware).

[0129] The computer program for implementing the method of the embodiments of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the computer programs are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0130] In the context of the embodiments of the present application, the machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable signal medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared system, apparatus, or device, or any suitable combination of the above. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0131] It should be noted that the term "including" and its variants used in the embodiments of the present application are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The modifications of "one" and "a plurality" mentioned in the embodiments of the present application are illustrative rather than restrictive. Those skilled in the art should understand that unless clearly stated otherwise in the context, it should be understood as "one or more".

[0132] In the method embodiments provided by the present application, the steps recorded in the method embodiments can be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The protection scope of the present application is not limited in this regard.

[0133] The term "embodiment" in this specification means that the specific features, structures or characteristics described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various positions in the specification does not necessarily mean the same embodiment, nor does it mean being independent or alternative to other embodiments and mutually exclusive. The various embodiments in this specification are described in a related manner, and the same or similar parts between the various embodiments are cross-referred to. In particular, for the device, equipment, and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts refer to the partial description of the method embodiments.

[0134] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of patent protection. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A time-domain secondary sentiment analysis method based on text symbolization, characterized in that include: Acquire text data containing emoticons, and pre-process the text data; Constructing an expression tag library and a Chinese vocabulary library based on the preprocessed text data; wherein the expression tag library is used to correspond expression symbols to specific emotion categories, and the Chinese vocabulary library is used to perform word segmentation processing and feature extraction on the preprocessed text data; The pre-processed text data is encoded, and the encoded data is sequence filled; Using a deep learning model to perform expression symbol analysis on the data after sequence filling processing, and combining the time domain secondary screening mechanism to mine potential emotions; wherein the deep learning model is trained based on the expression tag library and the Chinese vocabulary library; Track and analyze emotional turning points based on the data source trigger mechanism, and label the analysis results and store them in the database.

2. The time-domain secondary sentiment analysis method based on text symbolization according to claim 1, wherein Preprocessing the text data includes: Performing a primary screening on the text data through a filtering rule library to remove redundant characters, topic tags, and non-Chinese characters; De-duplication is performed on the text data that has passed the first-level screening through the de-duplication rule library; Perform secondary processing on the deduplicated text data through the symbol extraction library to locate and extract emoticons in the text; The complex sample processing pool is used to perform tertiary adaptation on the text data that has undergone secondary processing, converting one-to-many text into multiple sentences.

3. The time-domain secondary sentiment analysis method based on text symbolization according to claim 1, wherein, The preprocessed text data is used to construct an expression tag library and a Chinese vocabulary library, respectively, including: Dividing the preprocessed text data into an emoticon set and a text data set; Screening and counting the emoticon set to construct the emoticon tag library; Analyzing and counting the text data set to construct the Chinese vocabulary database; The Chinese vocabulary library is regularly updated by loading the training model using the historical database.

4. The time-domain secondary sentiment analysis method based on text symbolization according to claim 1, wherein The preprocessed text data is encoded, including: Filtering the preprocessed text data to remove irrelevant characters and redundant information; Perform sentence segmentation on the filtered text data to split the text into independent vocabulary units; Each vocabulary unit in the short text data after word segmentation is mapped to a unique digital code through a preset coding library; The mapped numeric codes are organized in a list form and the text data is represented in the form of one-hot encoding.

5. The time-domain secondary sentiment analysis method based on text symbolization according to claim 1, wherein, The time domain secondary screening mechanism includes: Track the first weight values of emotional expressions in different window periods in the time domain space and record the labels; Use a sliding time window to determine the secondary weight value of the next occurrence of emotional expression, label it, and record the increase with the first value, filter out the secondary increase emotional expression and mark it.

6. The time-domain secondary sentiment analysis method based on text symbolization according to claim 5, characterized in that, The time domain secondary screening mechanism adopts a probability floating method in the time domain to predict potential emotions, including: The probability value corresponding to each emotion label is calculated in the first time window; The probability value corresponding to each emotion label is calculated in the second time window, and the difference between the probability value in the first time window is calculated; The sentiment tags with increasing probability values and the largest fluctuation in increasing fluctuations are selected as secondary sentiment tags to identify potential sentiment trends.

7. The time-domain secondary sentiment analysis method based on text symbolization according to claim 1, wherein Perform sequence filling processing on the encoded data, including: Determine the maximum length of the data; The data whose length does not exceed the maximum length is padded at the end, and the padded element is 0; Data that exceeds the maximum length will be truncated at the end, retaining the first maximum length elements.

8. The time-domain secondary sentiment analysis method based on text symbolization according to claim 1, wherein Track and analyze emotional turning points based on the data source trigger mechanism, and label the analysis results and store them in the database, including: Determine a data source trigger mechanism, wherein the data source trigger mechanism includes setting a monitoring time period, a monitoring subject or topic, and a monitoring user account; Collect corresponding text data according to the determined data source trigger mechanism; Performing sentiment analysis on the collected text data to obtain sentiment labels and probability values of the sentiment labels; Determining an emotion inflection point according to the emotion label and the probability value of the emotion label; wherein the emotion inflection point is a point at which the probability value of the emotion label changes significantly; Analyze the determined emotional turning points and the causes and trends of emotional changes; The analysis results are labeled to generate label data, and the label data is stored in a database.

9. A time-domain secondary sentiment analysis system based on text symbolization, characterized in that, include: A data acquisition and preprocessing module, used to acquire text data containing emoticons and preprocess the text data; A database construction module, used to construct an expression tag library and a Chinese vocabulary library based on the preprocessed text data; wherein the expression tag library is used to correspond expression symbols to specific emotion categories, and the Chinese vocabulary library is used to perform word segmentation and feature extraction on the preprocessed text data; A data encoding and sequence filling module, used for encoding the preprocessed text data and performing sequence filling on the encoded data; A sentiment analysis module, for performing expression symbol analysis on the sequence-filled data using a deep learning model, and mining potential emotions in combination with a time-domain secondary screening mechanism; wherein the deep learning model is trained based on the expression tag library and the Chinese vocabulary library; The emotion turning point tracking and analysis module is used to track and analyze emotion turning points according to the data source trigger mechanism, and label the analysis results and store them in the database.

10. An electronic device, comprising: A processor, and a memory for storing a program, characterized in that the program includes instructions, which, when executed by the processor, cause the processor to execute the time-domain secondary sentiment analysis method based on text symbolization as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Text generation auxiliary processing method and system based on machine learning

    CN121920388A