Network public opinion analysis system and method based on big data platform
By constructing a multi-source public opinion text set, reorganizing temporal features, and analyzing sentiment polarity, combined with information on the magnitude of sentiment changes and the scale of public opinion, the shortcomings of existing technologies in identifying public opinion risks have been addressed, and a more accurate and stable assessment of public opinion risks has been achieved.
Patent Information
- Application Number
- CN202610152033.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-03
- Publication Date
- 2026-05-15
AI Technical Summary
In existing technologies, methods for identifying public opinion risks fail to effectively characterize the magnitude of changes in emotions over adjacent time periods, making it difficult to identify risk signals implied by sudden changes in emotions. Furthermore, they do not incorporate scale information such as the number of public opinion texts, leading to small-scale, localized emotional fluctuations being misjudged as public opinion risks.
By constructing a multi-source public opinion text set, performing joint processing and temporal feature reorganization, combining sentiment polarity analysis and abnormal fluctuation identification, using sentiment change amplitude and public opinion scale information for risk verification, and employing multi-platform cross-validation to ensure the accuracy of identification.
It has achieved accurate identification of public opinion risks, reduced the risk of misjudgment, improved the accuracy and authenticity of identification, and enhanced the adaptability and stability of anomaly identification.
Smart Images

Figure CN122047253A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis technology, specifically to a network public opinion analysis system and method based on a big data platform. Background Technology
[0002] With the rapid development of internet and mobile communication technologies, online information dissemination is characterized by diversified platforms, fragmented content, and rapid spread. A large amount of public opinion, emotional expressions, and event discussions are continuously generated through social media, news platforms, forums, and other online channels, forming massive and complex online public opinion data. Analyzing online public opinion based on big data platforms has become an important technical means in fields such as government management, public safety, and corporate brand management.
[0003] Existing technologies have shortcomings in identifying public opinion risks: existing methods often focus on the level of sentiment values within a certain time window and judge risks through fixed thresholds or empirical rules. They fail to effectively characterize the magnitude of changes in sentiment over adjacent time periods, making it difficult to identify the risk signals implied by sudden changes in sentiment. Furthermore, when judging abnormal sentiment, they rely solely on sentiment values or their changes without considering scale information such as the number of public opinion texts, which may lead to small-scale, localized emotional fluctuations being misjudged as public opinion risks. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a network public opinion analysis system and method based on a big data platform to solve the problems mentioned in the background.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] In a first aspect, the present invention provides a network public opinion analysis system and method based on a big data platform, comprising the following steps:
[0007] S1. Constructing a multi-source public opinion text set to obtain a public opinion text set;
[0008] S2. Perform joint processing on the set of public opinion texts to obtain a set of basic semantic units;
[0009] S3. Reorganize the temporal features based on the basic semantic unit set to obtain the temporal feature set;
[0010] S4. Perform sentiment polarity analysis based on the time-series feature set to obtain the public opinion sentiment sequence;
[0011] S5. Identify abnormal fluctuations based on the public opinion sentiment sequence to obtain candidate intervals for public opinion risks;
[0012] S6. Verify the risk based on the candidate range of public opinion risk to obtain the public opinion risk judgment result.
[0013] To further optimize this technical solution, the construction of the multi-source public opinion text set in step S1 includes:
[0014] The target and time frame of public opinion are clearly defined, a semantic description is constructed, multi-source public opinion texts are screened, and text encoding and language standards are unified to construct a set of public opinion texts.
[0015] To further optimize this technical solution, the joint processing in step S2 includes:
[0016] Based on the obtained public opinion text set, a basic semantic unit set is constructed by combining lexical segmentation and semantic weight calculation, which has unified word form and standardized semantic expression.
[0017] To further optimize this technical solution, the temporal feature reassembly in step S3 includes:
[0018] Based on the obtained set of basic semantic units, a temporal feature expression form describing the evolution of public opinion semantics over time is constructed through time segmentation and semantic recombination, resulting in a temporal feature set.
[0019] To further optimize this technical solution, the time segmentation and semantic reorganization include:
[0020] Under a unified time reference, the timeline is divided into continuous time windows. Each basic semantic unit is assigned to the corresponding time window according to its time identifier. Semantic units within the same time window are summarized and reorganized, and semantic units with the same semantics in adjacent time windows are associated to form a continuous expression relationship of the same semantics in the time dimension.
[0021] To further optimize this technical solution, the sentiment polarity analysis in step S4 includes:
[0022] Based on the obtained temporal feature set, the sentiment polarity of semantic units is determined by the sentiment dictionary, and combined with the semantic weight of the semantic units, the sentiment value of the sentiment polarity results within the same time window is calculated to construct a public opinion sentiment sequence in chronological order.
[0023] To further optimize this technical solution, the comprehensive sentiment value calculation includes:
[0024]
[0025] in:
[0026] Overall sentiment score;
[0027] Semantic unit The value of the emotional polarity;
[0028] Semantic unit semantic weight;
[0029] The number of semantic units in the time window;
[0030] The overall sentiment value for that time window is obtained by weighting the sentiment polarity and semantic weight of all semantic units within the time window.
[0031] To further optimize this technical solution, the abnormal fluctuation identification in step S5 includes:
[0032] Based on the obtained public opinion sentiment sequence, the magnitude of sentiment change between adjacent time windows is calculated according to the time window sequence. Combined with the public opinion scale information of the time window, abnormal windows are jointly identified, and consecutive abnormal time windows are merged to obtain the public opinion risk candidate interval.
[0033] To further optimize this technical solution, the risk verification in step S6 includes:
[0034] Based on the obtained candidate intervals for public opinion risk, cross-validation across multiple platforms is used to determine the consistency of sentiment direction across different platforms. Intervals with consistent cross-platform characteristics are then selected to obtain the public opinion risk assessment results.
[0035] This technical solution has been further optimized, including the following functional modules:
[0036] The module includes a text semantic construction module, a semantic unit generation module, a temporal feature construction module, a public opinion sentiment analysis module, an abnormal fluctuation identification module, and a public opinion risk assessment module.
[0037] In a second aspect, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the computer program instructions, when executed by the processor, implement the steps of the network public opinion analysis system and method based on a big data platform as described in the first aspect of the present invention.
[0038] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program instructions are executed by a processor, they implement the steps of a network public opinion analysis system and method based on a big data platform as described in the first aspect of the present invention.
[0039] Compared with existing technologies, this invention provides a network public opinion analysis system and method based on a big data platform, which has the following beneficial effects:
[0040] This network public opinion analysis system and method based on a big data platform, through the identification of abnormal fluctuations, combines information on the magnitude of emotional changes and the scale of public opinion to achieve a joint judgment on public opinion risks, reducing the risk of misjudgment, improving the accuracy and authenticity of identification, and enhancing the adaptability and stability of abnormal identification. Attached Figure Description
[0041] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 This is a flowchart illustrating a network public opinion analysis method based on a big data platform proposed in this invention.
[0043] Figure 2 This is a schematic diagram of a network public opinion analysis system based on a big data platform proposed in this invention. Detailed Implementation
[0044] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0045] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0046] Secondly, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places throughout this specification does not necessarily refer to the same embodiment, nor is it a single embodiment or an embodiment selectively excluded from other embodiments.
[0047] Example 1:
[0048] Reference Figure 1 This is the first embodiment of the present invention, which provides a method for analyzing online public opinion based on a big data platform, including the following steps:
[0049] S1. Construct a multi-source public opinion text set to obtain a public opinion text set.
[0050] In this embodiment, the construction of the multi-source public opinion text set includes:
[0051] In the context of online public opinion analysis, public opinion information is inherently characterized by its dispersed sources, significant differences in expression, and unstable semantic orientation. Different online platforms exhibit significant differences in text length, expression style, publication time accuracy, and field structure, which directly makes it difficult to directly compare or aggregate texts of the same public opinion object on different platforms. Furthermore, a large number of online texts only overlap with the public opinion object in terms of surface keywords, but do not actually target the public opinion object in terms of semantics. If these texts are directly included in the analysis, they will significantly interfere with the subsequent feature calculation results.
[0052] By clearly defining the target and time frame of public opinion, constructing semantic descriptions, filtering public opinion texts from multiple sources, and unifying text encoding and language standards, a set of public opinion texts is constructed. This ensures the uniqueness of the target of public opinion, eliminates differences in text structure across multiple platforms, reduces the proportion of semantic noise, and provides deterministic input for subsequent analysis.
[0053] Methods for constructing multi-source public opinion text collections include:
[0054] Determine the target of public opinion and the time boundary: Clearly define the target of this public opinion analysis and take it as the sole analysis target. At the same time, set the time range of the public opinion analysis to limit the time validity of the text and avoid historical or premature information from interfering with the current analysis.
[0055] Constructing a semantic description of the public opinion object: In the actual network environment, relying solely on the object name for matching will introduce a large number of homonymous or semantically weakly related texts. For example, the object name may be used in metaphors, quotations, or other contexts, rather than as the subject of discussion. Therefore, it is necessary to construct a semantic description of the identified public opinion object, including the object name, common abbreviations, aliases, and core semantic expressions that are highly related to the object, as a basis for judging whether the text is related to the public opinion object.
[0056] Preliminary screening of multi-source texts: Within a preset time range, text content is obtained from multiple online information carrying platforms, and the texts are preliminarily screened based on the semantic description of the public opinion object. Only when the text meets both the condition of the appearance of the object's keyword and the condition of semantic association with the context is it included in the candidate text set. This avoids introducing texts that have no substantial connection with the public opinion object based on a single keyword match, thereby narrowing the text range and reducing the computational pressure.
[0057] Semantic consistency verification: Further semantic consistency verification is performed on the candidate text set to determine whether the text content revolves around the same public opinion object. Texts with obvious ambiguity or whose object keywords only appear in the title or local position but deviate from the content are removed to ensure that the texts entering the final set have clear semantic direction.
[0058] Unified text encoding and language standardization: Different network platforms have differences in text encoding, symbol usage and language standardization. If they are not processed uniformly, the same semantic meaning will be identified as different content at the computing level, thus affecting the feature statistics results. Therefore, unified encoding processing is performed on the text that passes the semantic verification, and the language expression is standardized, including symbol unification, abnormal character cleaning and language format standardization, to ensure that the text on different platforms has consistency at the character level.
[0059] Unified field structure mapping: Construct a unified data structure for each text, which includes at least three types of fields: time identifier, source identifier, and body content. And perform unified mapping processing for the inconsistent meaning of fields on different platforms, so that each public opinion text has a consistent data organization form in terms of structure, providing a structural foundation for parallel processing, time aggregation, and cross-platform comparison.
[0060] Generate a standardized set of public opinion texts: The texts that have undergone the above processing are uniformly summarized to form a standardized set of public opinion texts. Each text in the set is consistent in semantic reference, field structure and text specifications.
[0061] S2. Perform joint processing on the set of public opinion texts to obtain a set of basic semantic units.
[0062] In this embodiment, the joint processing includes:
[0063] Long texts contain a large amount of descriptive and decorative language. If they are used directly for statistics or comparison, it is difficult to distinguish the core content that truly carries public opinion information. Even when discussing the same public opinion object, different texts differ significantly in expression, word choice, and length, making direct horizontal comparison impossible. Therefore, it is necessary to transform the text from a natural language expression form to a basic semantic expression form to establish a unified semantic computing foundation for subsequent analysis.
[0064] Based on the obtained public opinion text set, a set of basic semantic units with unified word form and standardized semantic expression is constructed through joint processing of lexical segmentation and semantic weight calculation. This process decomposes long texts into stable basic semantic units, avoiding the influence of text length and expression style on subsequent analysis, distinguishing between core public opinion semantics and weak contribution semantics, improving the parallel processing efficiency of the big data platform, and providing input for subsequent steps.
[0065] The steps of the combined processing include:
[0066] Read the public opinion text set: Read the public opinion text set output by step S1. Each text in this set contains a uniform body content field, which provides stable input for subsequent lexical processing and ensures that all subsequent processing only applies to the text set that has passed the semantic consistency and structural specification constraints.
[0067] Lexical segmentation is performed on the main text of each public opinion text. Weighted conditional random field segmentation technology is used, which takes into account the context information during the segmentation process, so that the segmentation results remain stable in terms of semantic continuity and are more in line with actual language usage habits. It avoids semantic fragmentation caused by isolated segmentation. Through this process, continuous text is decomposed into several terms with clear boundaries, laying the foundation for subsequent semantic processing.
[0068] Constructing a term set and eliminating redundancy: Summarize the terms obtained from all texts to construct an initial term set, and remove informational terms that do not clearly carry public opinion semantics, such as connective words or expressions with very low semantic contribution, in order to avoid them interfering with subsequent weight calculations and make the remaining terms more focused on reflecting the public opinion content itself.
[0069] Semantic weight calculation: The term frequency-inverse document frequency (TF-IDF) feature weight calculation technique is performed on the retained terms. The frequency of occurrence of terms in the text and their distribution in the overall public opinion text set are considered to assign a clear semantic weight to each term and determine its importance in the current public opinion context. Through this weight calculation, frequent but indistinguishable terms can be distinguished from core terms with public opinion indicative significance.
[0070] Synonym merging: For terms after weight calculation, synonym merging is performed. Terms that are semantically equivalent or highly similar but different in word expression are uniformly merged into the same semantic expression form to avoid the same semantic being counted multiple times and to ensure that the semantic statistics results reflect the true public opinion content.
[0071] Perform lexical unification processing: Further perform lexical unification processing on the lexical items that have completed synonym merging, unify the expression of the same semantics under different grammatical forms into a standardized form, and ensure consistency and stability in subsequent statistics and comparisons;
[0072] Generate a set of basic semantic units: The semantic expressions that have been processed above are uniformly summarized to form a set of basic semantic units. Each semantic unit in the set represents the smallest semantic unit in the public opinion text that has a clear semantic direction, a unified expression form, consistent statistical standards, and can be independently calculated. The weight of a semantic unit is equal to the sum of the weights of all the original terms it contains.
[0073] Furthermore, the semantic weight calculation includes:
[0074]
[0075] in:
[0076] : term The semantic weight in the text is used to measure the importance of the semantic unit represented by the term in the overall public opinion. The value is greater than or equal to 0, and the value varies with the size of the public opinion text and the distribution of terms.
[0077] : term The word frequency weight in the text is used to reflect the local importance of the word in the current text, and ranges from 0 to 1;
[0078] : term The inverse document frequency is used to reflect the discriminative power of the term;
[0079] By combining the term frequency weight and the inverse document frequency of a term in the text, the semantic weight of a term in the text is obtained, so that terms that appear frequently in most texts are given lower weights, while terms that appear frequently in a small number of texts are given higher weights.
[0080] Furthermore, the word frequency weights include:
[0081]
[0082] in:
[0083] : term The frequency of occurrence in the text is obtained by statistical analysis of the lexical segmentation results;
[0084] The total number of occurrences of all terms in the text is used to normalize the word frequency, eliminate the impact of text length differences on weight calculation, and make the word frequency weights of different texts comparable.
[0085] The term frequency weight is obtained by calculating the percentage of times a term appears in the text.
[0086] Furthermore, the inverse document frequency includes:
[0087]
[0088] in:
[0089] The total number of public opinion texts is obtained by statistically analyzing the set of public opinion texts obtained from S1.
[0090] : Contains terms The number of texts containing the term is obtained by counting the number of texts containing that term.
[0091] The inverse document frequency is obtained by performing a logarithmic calculation based on the ratio of the total number of texts to the number of texts containing terms.
[0092] S3. Reorganize the temporal features based on the basic semantic unit set to obtain the temporal feature set.
[0093] In this embodiment, the temporal feature reconstruction includes:
[0094] In the context of online public opinion, the technical meaning expressed by the same semantics at different times can vary significantly. For example, the risk meaning of a negative semantics appearing briefly and repeatedly is completely different from that of a negative semantics appearing repeatedly. A semantics that suddenly intensifies in a short period of time usually means the trigger point of a public opinion event. A semantics that gradually diminishes reflects the natural decline of public attention. If the analysis is based solely on static semantic weights, it will be impossible to distinguish the above situations, which can easily lead to misjudgment of the public opinion situation.
[0095] Based on the obtained set of basic semantic units, a temporal feature expression form describing the evolution of public opinion semantics over time is constructed through time segmentation and semantic recombination, resulting in a set of temporal features. This distinguishes between sudden and continuous public opinion events, clarifies key time nodes in the evolution of public opinion, and transforms discrete text into continuous and analyzable temporal behavior, serving as the basis for subsequent steps. This reduces the complexity of analysis and improves technical stability and interpretability.
[0096] Furthermore, the time segmentation and semantic reorganization include:
[0097] Under a unified time base, the timeline is divided into continuous time windows. Each basic semantic unit is assigned to a corresponding time window according to its time identifier. Semantic units within the same time window are aggregated and reorganized, and semantic units with the same semantics in adjacent time windows are associated to form a continuous expression relationship of the same semantics in the time dimension. Specific implementation methods include:
[0098] Introducing time stamps and establishing a unified time benchmark: In the basic semantic unit set, each semantic unit comes from a specific public opinion text. The public opinion text itself naturally carries time information. Read the time stamp of the text corresponding to the semantic unit and unify it according to the preset time benchmark (hour, day or other fixed time granularity can be selected according to the needs of public opinion analysis). This enables all semantic units to be mapped to the same time axis, providing a unified reference scale for subsequent time series division.
[0099] Semantic units are segmented based on a time benchmark: Under a unified time benchmark, the time axis is divided into several consecutive time intervals, each time interval corresponding to a specific time window. Basic semantic units are assigned to the corresponding time windows according to their time identifiers, forming a correspondence between semantic units and time windows. This reduces the noise impact caused by time fluctuations in individual texts and enhances the stability and interpretability of subsequent time series analysis.
[0100] Reorganizing semantic units within a single time window: Within each time window, the semantic units assigned to that window are aggregated and processed so that the semantic expression within the time window can reflect the overall public opinion focus of that time period. Taking the semantic unit as the smallest unit of analysis, the occurrence of the semantic unit within the time window and the cumulative weight characteristics of the semantic unit within the time window are comprehensively considered to distinguish the importance of different semantics within the time window, thereby obtaining the semantic expression results under each time window. This avoids the bias caused by simple counting and makes the semantic expression within the time window closer to the real public opinion state.
[0101] Associating the same semantic unit in adjacent time windows: After summarizing the semantics within each time window, associating semantic units that point to the same semantics in different time windows, establishing the correspondence between semantic units in consecutive time windows according to the time order, so that the same semantics can form a continuous expression in the time dimension, thereby identifying the first appearance, continuous existence and disappearance time of semantics, and providing a direct basis for judging the evolution of public opinion.
[0102] Forming temporal features: After completing the semantic association across time windows, the continuous changes of each semantic on the time axis are organized into independent temporal features. Each temporal feature reflects the change behavior of a certain semantic in multiple time windows, forming a set of temporal features.
[0103] S4. Perform sentiment polarity analysis based on the time-series feature set to obtain the public opinion sentiment sequence.
[0104] In this embodiment, the emotion polarity analysis includes:
[0105] The risk, spread, and social impact of online public opinion often do not depend on whether the information itself appears, but rather on the direction and speed of change in emotional tendencies. For example, the spread of the same topic under positive emotions is usually normal discussion, but if negative emotions intensify in a short period of time, it may trigger public opinion risks. Analyzing only based on the semantic temporal characteristics of S3 can only observe changes in the quantity and frequency of semantics, and cannot directly determine the emotional attributes of the public opinion situation.
[0106] Based on the obtained temporal feature set, the sentiment polarity of semantic units is determined by the sentiment dictionary, and combined with the semantic weight of the semantic units, the sentiment value of the sentiment polarity results within the same time window is calculated. The public opinion sentiment sequence is constructed in chronological order, thereby quantitatively describing the changing trend of public opinion sentiment over time, identifying sentiment turning points and concentrated outbreak intervals, and providing direct input for subsequent public opinion situation judgment or risk assessment.
[0107] The steps involved in sentiment polarity analysis include:
[0108] Read the time-series features of public opinion: Read the time-series feature set output by step S3. This set has been organized according to the time window order and uses the existing time structure to ensure that the sentiment calculation results can correspond one-to-one with the original time-series features.
[0109] Extract semantic units within the time window: For each time window, extract all semantic units contained within that window as the basic analysis object for sentiment computing, thereby reducing the complexity of the sentiment computing stage, avoiding the reintroduction of text-level uncertainty, and ensuring consistency of computing standards between different time windows.
[0110] Determine the sentiment polarity of semantic units: For each semantic unit, a mature sentiment dictionary matching calculation technology (based on HowNet sentiment dictionary) is used to match the semantic unit with the entries in the sentiment dictionary to determine the sentiment tendency expressed by the semantic unit and determine its corresponding sentiment polarity category. If a semantic unit does not have an explicit sentiment label in the dictionary, it can be regarded as a sentiment-neutral semantic unit.
[0111] Calculate the overall sentiment value for the time window: After obtaining the sentiment polarity of the semantic unit, further combine the semantic weight of the semantic unit in the current time window to perform weighted processing on the sentiment contribution, and obtain an overall sentiment value in each time window to represent the overall sentiment state of public opinion in that time period. This makes important semantics occupy a higher proportion in the overall sentiment value, more realistically reflecting the dominant sentiment of public opinion in that time period, thereby avoiding low-importance semantics from having too much influence on the overall sentiment judgment or high-frequency but weak semantics from masking key signals.
[0112] Constructing a public opinion sentiment sequence in chronological order: After completing the sentiment calculation for all time windows, arrange the comprehensive sentiment values corresponding to each time window according to their chronological order to form a public opinion sentiment sequence.
[0113] Furthermore, the calculation of the comprehensive sentiment score includes:
[0114]
[0115] in:
[0116] The overall sentiment score represents the overall emotional state of public opinion within a time window. It serves as a basic element constituting the sentiment sequence of public opinion and ranges from -1 to 1. For example, a value of -0.6 indicates that negative semantic weight is relatively high, a value of 0.7 indicates that positive semantic weight is dominant, and a value close to 0 indicates that positive and negative sentiment are relatively balanced.
[0117] Semantic unit The sentiment polarity value is used to describe the emotional tendency expressed by the semantic unit. It is obtained by matching the semantic unit based on the HowNet sentiment dictionary. If the tendency is negative, the value is -1; if the tendency is positive, the value is 1; and if the tendency is neutral, the value is 0.
[0118] Semantic unit The semantic weight is used to measure the importance of the semantic unit in the overall public opinion, so that important semantics have a greater impact on the comprehensive sentiment value. It is obtained through step S2.
[0119] The number of semantic units in the time window;
[0120] The overall sentiment value for that time window is obtained by weighting the sentiment polarity and semantic weight of all semantic units within the time window.
[0121] S5. Identify abnormal fluctuations based on the public opinion sentiment sequence to obtain the candidate interval for public opinion risk.
[0122] In this embodiment, the abnormal fluctuation identification includes:
[0123] In real-world public opinion scenarios, risks are often not manifested as whether the sentiment value is negative, but rather as drastic changes in the sentiment value in a short period of time, a sudden shift in the sentiment trend from a stable state to a one-way amplification, or a significant increase in the magnitude of sentiment fluctuations compared to historical normal levels. Judging solely based on the sentiment value at a single moment can easily lead to misjudging long-term stable negative discussions as risks, or ignoring sudden but not yet extreme signs of risk.
[0124] Based on the obtained public opinion sentiment sequence, the magnitude of sentiment change between adjacent time windows is calculated according to the time window sequence. Combined with the public opinion scale information of the time window, abnormal windows are jointly identified, and continuous abnormal time windows are merged to obtain public opinion risk candidate intervals. This transforms the continuous sentiment sequence into a limited number of risk candidate intervals, narrowing the time range for subsequent detailed analysis or risk judgment, improving the accuracy and authenticity of identification, and enhancing the adaptability and stability of anomaly identification.
[0125] The specific steps for identifying abnormal fluctuations include:
[0126] Public sentiment sequence reading: Read the public sentiment sequence output in step S4, and determine its time order as the sole time reference for subsequent analysis to ensure that the sequence maintains a time arrangement consistent with the actual public sentiment evolution process before being input into the abnormal fluctuation identification process;
[0127] Calculation of the magnitude of sentiment change: In the actual evolution of public opinion, the risk often does not come from the level of sentiment value at a certain point in time, but from the drastic changes in sentiment in a short period of time. Therefore, while keeping the time sequence unchanged, the sentiment values corresponding to adjacent time windows in the sentiment sequence are compared to obtain the sentiment change magnitude sequence. This change magnitude is used to describe the degree of change of sentiment between consecutive time windows, thereby transforming the sentiment sequence into a change sequence that can highlight "mutation behavior".
[0128] Information on changes in the scale of public opinion: Public opinion risk is not only related to changes in emotions, but also closely related to the number of people participating in the discussion and the scale of information dissemination. Read the public opinion scale information that corresponds one-to-one with the emotional sequence time window. For example, obtain the number of public opinion texts in the time window based on the time identifier of the public opinion text in step S1, so as to obtain the change of the scale of public opinion participation over time. This is used to help judge whether the emotional change has a real basis for dissemination, and avoid misidentifying local and small sample emotional fluctuations as public opinion risk.
[0129] Anomaly Joint Judgment: The magnitude of emotional change and the scale of public opinion change are analyzed together. When the magnitude of emotional change is significantly higher than that of the adjacent interval and the scale of public opinion increases synchronously or changes in a concentrated manner within the same time window, the interval is considered to have abnormal fluctuation characteristics, thereby effectively filtering out the situation where the emotional change is obvious but the scale of dissemination is limited and the situation where the scale of dissemination changes but the emotional structure is stable.
[0130] Determination of public opinion risk candidate intervals: Continuous time windows that meet the abnormal fluctuation conditions are merged to form several continuous time intervals. Each time interval is marked as a public opinion risk candidate interval, which is used to indicate that there may be abnormalities in public opinion sentiment and dissemination behavior within that time range. All the identified intervals are summarized to form a set of public opinion risk candidate intervals.
[0131] Furthermore, the calculation of the magnitude of the emotional change includes:
[0132]
[0133] in:
[0134] The magnitude of emotional change, expressed within a time window. Time window The range of emotional changes between them is used to characterize the strength of changes in public opinion sentiment.
[0135] : No. The comprehensive sentiment value within a time window is obtained through step S4;
[0136] : No. The overall sentiment value within a time window, that is, the overall sentiment value corresponding to the previous time window adjacent to the current time window;
[0137] By calculating the absolute value of the combined sentiment value change between adjacent time windows, the influence of the direction of sentiment change is eliminated, and the magnitude of sentiment change, which reflects the intensity of the change, is obtained.
[0138] S6. Verify the risk based on the candidate range of public opinion risk to obtain the public opinion risk judgment result.
[0139] In this embodiment, the risk verification includes:
[0140] In the real-world network environment, different information publishing platforms may experience platform-specific abnormal fluctuations within a certain time period due to differences in user structure, dissemination mechanisms, and content preferences. Risk assessment on a single platform is easily affected by factors such as the homogeneity of the platform's user group structure, the concentrated dissemination of local events on a specific platform, and short-term anomalies caused by the platform's recommendation or display mechanisms. This can lead to risks identified on a single platform not having a wide social impact, resulting in misjudgments.
[0141] Based on the obtained candidate intervals of public opinion risk, cross-validation across multiple platforms is used to determine the consistency of sentiment direction across different platforms. Intervals with consistent cross-platform characteristics are selected to obtain the public opinion risk assessment results. This process eliminates local anomalies that only appear on a single platform, ensuring that the final assessed public opinion risk has cross-platform propagation characteristics and improving the stability, credibility, and interpretability of the public opinion risk conclusions.
[0142] The steps of the risk verification method include:
[0143] Obtain the set of candidate intervals for public opinion risks: Read the set of candidate intervals for public opinion risks output in step S5. Each candidate interval corresponds to a clear start time window and end time window, which are used to limit the time range of subsequent cross-platform analysis and ensure that subsequent cross-platform verification has clear time focus and calculation boundaries.
[0144] Platform public opinion data extraction: For each public opinion risk candidate interval, public opinion text data corresponding to that time interval is extracted from multiple predetermined information source platforms. The information source platforms are different online information release or dissemination platforms, and the data sources of each platform are independent of each other to ensure the objectivity of consistency verification.
[0145] Sentiment analysis of texts on each platform: While maintaining consistency in time intervals, sentiment analysis is performed on the public opinion texts within each platform, including calculating the comprehensive sentiment value and the magnitude of sentiment change, to obtain the overall sentiment performance of the platform within the candidate interval, which is used to describe the changes in public opinion sentiment on the platform within the time interval, such as overall positive, negative, or large magnitude of change.
[0146] Determine the consistency of emotional direction: Compare the emotional expressions obtained from each platform within the same candidate interval to determine whether their emotional direction is consistent, that is, whether the emotional polarity is the same or similar. When multiple platforms show the same or similar emotional direction changes within the same candidate interval, the candidate interval is considered to meet the cross-platform emotional consistency condition. It is not required that all platforms be completely consistent. It is sufficient if no less than a preset number (e.g., more than 60%) of the platforms meet the consistency condition.
[0147] Determine the final public opinion risk assessment result: For candidate intervals that meet the cross-platform sentiment consistency condition, they are confirmed as the final public opinion risk intervals and included in the final public opinion risk assessment result set. For candidate intervals that fail the consistency verification, they are not used as the final public opinion risk output, thereby avoiding misjudgment caused by single-platform anomalies.
[0148] Example 2:
[0149] Reference Figure 2 This is the second embodiment of the present invention, which provides a network public opinion analysis system based on a big data platform, including the following functional modules:
[0150] Text Semantic Construction Module: This module is used to perform unified semantic processing on public opinion texts from different information source platforms. By aligning the semantics of multi-source texts and unifying their expressions, it constructs a semantically consistent set of public opinion texts, thereby eliminating the impact of platform differences, expression differences, and language form differences on subsequent analysis and providing standardized input for subsequent semantic analysis.
[0151] Semantic Unit Generation Module: This module performs lexical and semantic joint processing on a standardized set of public opinion texts, identifies the core semantic content in the text, constructs a set of basic semantic units, unifies the different expressions scattered in the text into stable semantic units, and assigns a corresponding semantic weight to each semantic unit to characterize its importance in public opinion.
[0152] The temporal feature construction module is used to reorganize the basic set of public opinion semantic units in terms of time dimension, aggregate the semantic units according to the preset time window, form the temporal features of public opinion, and provide a time basis for subsequent sentiment analysis and trend analysis.
[0153] The public opinion sentiment analysis module is used to calculate the sentiment polarity of public opinion based on the time-series characteristics of public opinion. It combines semantic weights with sentiment polarity information to generate a public opinion sentiment sequence that changes over time. The resulting sentiment sequence not only reflects the direction of emotions, but also reflects the degree of influence of different semantics in public opinion.
[0154] Abnormal fluctuation identification module: used to analyze the public opinion sentiment sequence, identify abnormal fluctuations in sentiment over time, and determine the candidate range of public opinion risk in combination with changes in the scale of public opinion;
[0155] The public opinion risk assessment module is used to perform cross-platform consistency verification of public opinion risk candidate intervals. By comparing the sentiment performance of different platforms within the same time interval, the final public opinion risk assessment result is determined, thereby effectively eliminating local anomalies on a single platform and improving the reliability and stability of the public opinion risk assessment result.
[0156] Example 3:
[0157] In practical applications, this invention can be applied to the monitoring of online public opinion risks during public emergencies. It is used to continuously analyze and identify risks in the context of multiple online information platforms. The following description uses the monitoring of online public opinion during sudden public transportation failures in urban areas as a typical scenario.
[0158] In this application scenario, the system first accesses multiple online information source platforms within a preset time frame, including news platforms, social media platforms, and local forum platforms, to uniformly process online texts related to urban public transportation operations. Addressing the differences in expression, descriptive perspectives, and language styles across different platforms, the system performs semantic unification processing on the acquired text content, merging different expressions pointing to the same event or issue into a semantically consistent set of public opinion texts, thereby avoiding information fragmentation caused by platform differences.
[0159] After semantic unification, a joint lexical and semantic analysis is performed on the key semantic content in the public opinion text. Content reflecting public concerns is extracted as basic semantic units, and each semantic unit is assigned a specific semantic weight based on its occurrence and distribution characteristics in the text set. This process can objectively reflect the importance of different public opinion concerns in the overall discussion.
[0160] The basic semantic units of public opinion are reorganized according to continuous time windows to form temporal characteristics reflecting the changes in public opinion content over time. Based on this, the sentiment polarity of public opinion content within each time window is calculated by combining semantic weights, resulting in a public opinion sentiment sequence reflecting the process of changes in public sentiment. This sentiment sequence can continuously depict the evolution of public sentiment from a stable state to a state of intense change.
[0161] After obtaining the public opinion sentiment sequence, the magnitude of sentiment changes between adjacent time windows is analyzed, and the changes in the number of public opinion texts within each time window are simultaneously incorporated. When the magnitude of sentiment changes significantly amplifies in a short period of time, and the scale of public opinion participation increases synchronously within the corresponding time period, this continuous time range is identified as a candidate interval for public opinion risk, used to characterize the possible clustering of public opinion risks within this time period.
[0162] For the identified candidate intervals of public opinion risk, a comparative analysis of public opinion sentiment performance within the same time range is further conducted on different online information platforms. When multiple platforms show a consistent direction of sentiment change within the same time interval, this time interval is confirmed as the actual public opinion risk interval, and it is output as the public opinion risk judgment result. This cross-platform consistency verification method can effectively avoid misjudgments caused by local anomalies on a single platform.
[0163] Through the above specific implementation, the present invention can accurately identify the time range of abnormal changes in public sentiment in real-world events such as sudden malfunctions of urban public transportation, providing reliable technical support for relevant departments to grasp the public opinion situation in a timely manner and formulate response measures.
[0164] Example 4:
[0165] This embodiment also provides a computer device applicable to a network public opinion analysis system and method based on a big data platform, including a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to realize the network public opinion analysis system and method based on a big data platform as proposed in the above embodiment.
[0166] This embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements a network public opinion analysis system and method based on a big data platform as proposed in the above embodiments.
[0167] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0168] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0169] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-including system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0170] More specific examples (a non-exhaustive list) of computer-readable media include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0171] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0172] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for analyzing online public opinion based on a big data platform, characterized in that, Includes the following steps: S1. Constructing a multi-source public opinion text set to obtain a public opinion text set; S2. Perform joint processing on the set of public opinion texts to obtain a set of basic semantic units; S3. Reorganize the temporal features based on the basic semantic unit set to obtain the temporal feature set; S4. Perform sentiment polarity analysis based on the time-series feature set to obtain the public opinion sentiment sequence; S5. Identify abnormal fluctuations based on the public opinion sentiment sequence to obtain candidate intervals for public opinion risks; S6. Verify the risk based on the candidate range of public opinion risk to obtain the public opinion risk judgment result.
2. The method for analyzing online public opinion based on a big data platform according to claim 1, characterized in that, The construction of the multi-source public opinion text set in step S1 includes: The target and time frame of public opinion are clearly defined, a semantic description is constructed, multi-source public opinion texts are screened, and text encoding and language standards are unified to construct a set of public opinion texts.
3. The method for analyzing online public opinion based on a big data platform according to claim 1, characterized in that, The joint processing in step S2 includes: Based on the obtained public opinion text set, through the joint processing of lexical segmentation and semantic weight calculation, basic semantic units with unified word form and standardized semantic expression are constructed, forming a basic semantic unit set.
4. The method for analyzing online public opinion based on a big data platform according to claim 1, characterized in that, The temporal feature reconstruction in step S3 includes: Based on the obtained set of basic semantic units, a temporal feature expression form describing the evolution of public opinion semantics over time is constructed through time segmentation and semantic recombination, resulting in a temporal feature set.
5. The online public opinion analysis method based on a big data platform according to claim 4, characterized in that, The time segmentation and semantic reorganization include: Under a unified time reference, the timeline is divided into continuous time windows. Each basic semantic unit is assigned to the corresponding time window according to its time identifier. Semantic units within the same time window are summarized and reorganized, and semantic units with the same semantics in adjacent time windows are associated to form a continuous expression relationship of the same semantics in the time dimension.
6. The method for analyzing online public opinion based on a big data platform according to claim 1, characterized in that, The sentiment polarity analysis in step S4 includes: Based on the obtained temporal feature set, the sentiment polarity of semantic units is determined by the sentiment dictionary, and combined with the semantic weight of the semantic units, the sentiment value of the sentiment polarity results within the same time window is calculated to construct a public opinion sentiment sequence in chronological order.
7. The method for analyzing online public opinion based on a big data platform according to claim 6, characterized in that, The calculation of the comprehensive sentiment score includes: in: Overall sentiment score; Semantic unit The value of the emotional polarity; Semantic unit semantic weight; The number of semantic units in the time window; The overall sentiment value for that time window is obtained by weighting the sentiment polarity and semantic weight of all semantic units within the time window.
8. The method for analyzing online public opinion based on a big data platform according to claim 1, characterized in that, The abnormal fluctuation identification in step S5 includes: Based on the obtained public opinion sentiment sequence, the magnitude of sentiment change between adjacent time windows is calculated according to the time window sequence. Combined with the public opinion scale information of the time window, abnormal windows are jointly identified, and consecutive abnormal time windows are merged to obtain the public opinion risk candidate interval.
9. The method for analyzing online public opinion based on a big data platform according to claim 1, characterized in that, The risk verification in step S6 includes: Based on the obtained candidate intervals for public opinion risk, cross-validation across multiple platforms is used to determine the consistency of sentiment direction across different platforms. Intervals with consistent cross-platform characteristics are then selected to obtain the public opinion risk assessment results.
10. A network public opinion analysis system based on a big data platform, constructed based on the network public opinion analysis method based on a big data platform as described in any one of claims 1-9, characterized in that, Includes the following functional modules: The module includes a text semantic construction module, a semantic unit generation module, a temporal feature construction module, a public opinion sentiment analysis module, an abnormal fluctuation identification module, and a public opinion risk assessment module.