Public opinion analysis method and device, equipment and storage medium

By constructing the initial public opinion map and generating a multi-layer public opinion relationship network and public opinion hotspot map, the problem in the existing technology is difficult to track and accurately analyze the evolution process of public opinion in real time, and the accurate insight into public opinion and strong timeliness analysis results are achieved.

CN120216704APending Publication Date: 2025-06-27CHONGQING JINKANG NEW ENERGY VEHICLE CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510376275.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-27

Smart Images

  • Figure CN120216704A_ABST
    Figure CN120216704A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a public opinion analysis method and device, equipment and a storage medium. The public opinion analysis method comprises the steps of obtaining original public opinion data; constructing an initial public opinion graph according to the original public opinion data; the initial public opinion graph comprises a connection relationship between nodes, the nodes correspond to named entities in the original public opinion data, and the connection relationship between the nodes corresponds to an entity relationship between the named entities; generating a multilayer public opinion relationship network according to the initial public opinion graph; generating a public opinion hotspot map according to the initial public opinion map; and obtaining a public opinion analysis result according to the multilayer public opinion relationship network and the public opinion hotspot map. According to the method, the public opinions can be accurately insighted, and the public opinion hotspots are tracked and accurately captured in real time, so that a public opinion analysis result with high timeliness and high accuracy is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of public opinion analysis, and particularly to a public opinion analysis method and apparatus, an electronic device, and a storage medium. Background Art

[0002] In the current era of information explosion, the importance of monitoring and analyzing public opinion has become increasingly prominent. In related technologies, public opinion analysis mainly relies on the following methods to achieve, namely: keyword search-based, using simple statistical analysis tools, and applying traditional natural language processing (NLP) techniques. Among them, the keyword search-based method is to preset keywords related to a specific topic, and then search for content containing these keywords in a large amount of text data, such as news reports and social media posts; the method of using simple statistical analysis tools is to perform word frequency analysis on text data using simple statistical analysis tools; the method of applying traditional NLP techniques, such as part-of-speech tagging and named entity recognition, is to perform preliminary processing on text data to identify entities and the part-of-speech information of words.

[0003] However, the above several public opinion analysis methods still have problems such as being unable to deeply excavate the deep meaning of text and being difficult to track the evolution process of public opinion in real time and accurately, resulting in being unable to provide an in-time and accurate public opinion analysis report for enterprises as a decision-making basis.

[0004] Therefore, how to accurately insight into public opinion during the public opinion monitoring process, track and accurately capture the hotspots of public opinion in real time, so as to obtain a public opinion analysis result with strong timeliness and high accuracy has become an urgent problem to be solved. Summary of the Invention

[0005] Embodiments of the present application provide a public opinion analysis method to solve the problem that it is difficult to track the evolution process of public opinion in real time and accurately, resulting in being unable to provide an in-time and accurate public opinion analysis report for enterprises as a decision-making basis.

[0006] Correspondingly, embodiments of the present application also provide a public opinion analysis apparatus, an electronic device, and a storage medium to ensure the implementation and application of the above method.

[0007] To solve the above problems, embodiments of the present application disclose a public opinion analysis method, the method comprising:

[0008] Obtain original public opinion data;

[0009] Construct an initial public opinion map based on the original public opinion data; the initial public opinion map includes the connection relationships between nodes, the nodes correspond to the named entities in the original public opinion data, and the connection relationships between the nodes correspond to the entity relationships between the named entities;

[0010] Generate a multi-layer public opinion relationship network based on the initial public opinion map;

[0011] Generate a public opinion hot spot map based on the initial public opinion map;

[0012] Obtain a public opinion analysis result based on the multi-layer public opinion relationship network and the public opinion hot spot map.

[0013] Optionally, the generating a multi-layer public opinion relationship network based on the initial public opinion map includes:

[0014] Calculate the node importance scores corresponding to each node in the initial public opinion map;

[0015] Determine the key nodes in the initial public opinion map according to the node importance scores;

[0016] Obtain the connection relationships corresponding to the key nodes;

[0017] Divide the initial public opinion map into multiple topic community sub-networks according to the connection relationships between the nodes in the initial public opinion map;

[0018] Traverse the nodes and the connection relationships between the nodes in the initial public opinion map to identify the connected components and triangle structures in the initial public opinion map; the triangle structure corresponds to the structure in which three nodes in the initial public opinion map are connected pairwise;

[0019] Generate the multi-layer public opinion relationship network according to the key nodes, the connection relationships corresponding to the key nodes, the topic community sub-networks, the connected components and the triangle structures.

[0020] Optionally, the generating a multi-layer public opinion relationship network according to the key nodes, the connection relationships corresponding to the key nodes, the topic community sub-networks, the connected components and the triangle structures includes:

[0021] Extract the key nodes and the connection relationships corresponding to the key nodes in the initial public opinion map to generate a top-level public opinion relationship network; the top-level public opinion relationship network includes core nodes, edge nodes and connection relationships between nodes, the core nodes correspond to the key nodes, and the edge nodes are directly or indirectly connected to the core nodes through the connection relationships between the nodes;

[0022] Extract the topic community sub-networks in the initial public opinion map to generate a middle-level public opinion relationship network;

[0023] Extract the connected components and the triangular structures in the initial public opinion graph to generate an underlying public opinion relationship network;

[0024] The top-level public opinion relationship network, the middle-level public opinion relationship network, and the underlying public opinion relationship network constitute the multi-level public opinion relationship network.

[0025] Optionally, generating a public opinion hot spot graph based on the initial public opinion graph includes:

[0026] Extract the connected subgraphs in the initial public opinion graph;

[0027] Calculate the comprehensive score corresponding to the connected subgraph;

[0028] Determine potential hot spots based on the comprehensive score;

[0029] Generate the public opinion hot spot graph based on the potential hot spots.

[0030] Optionally, the initial public opinion graph includes node attributes and connection relationship attributes. Calculating the comprehensive score corresponding to the connected subgraph includes:

[0031] Determine the sentiment tendency of the connection relationship between the nodes in the connected subgraph; the sentiment tendency is determined by performing sentiment analysis on the original public opinion data corresponding to the connection relationship between the nodes;

[0032] Calculate the node importance score in the connected subgraph; the node importance score is determined according to the degree of the node, the node weight, and the sentiment tendency of the connection relationship corresponding to the node; the node attributes include the degree of the node and the node weight, and the connection relationship attributes include the sentiment tendency of the connection relationship corresponding to the node;

[0033] Calculate the comprehensive score of the connected subgraph according to the node importance score and the sentiment tendency.

[0034] Optionally, generating the public opinion hot spot graph based on the potential hot spots includes:

[0035] Monitor the potential hot spots to obtain potential hot spot data of the potential hot spots; the potential hot spot data at least includes the duration, the heat change value, and the hot spot propagation path;

[0036] Generate a public opinion hot spot distribution map according to the potential hot spots and the hot spot propagation path;

[0037] Generate a hot spot evolution curve graph according to the duration and the heat change value of the potential hot spots; the public opinion hot spot graph includes the public opinion hot spot distribution map and the hot spot evolution curve graph.

[0038] Optionally, obtaining the public opinion analysis result based on the multi-layer public opinion relationship network and the public opinion hotspot map includes:

[0039] Generating a comprehensive public opinion dissemination analysis report based on the multi-layer public opinion relationship network;

[0040] Generating a hotspot analysis report based on the public opinion hotspot map;

[0041] Obtaining the public opinion analysis result based on the comprehensive public opinion dissemination analysis report and the hotspot analysis report.

[0042] An embodiment of the present application also discloses a public opinion analysis device, and the device includes:

[0043] An acquisition module, configured to acquire original public opinion data;

[0044] An initial public opinion map construction module, configured to construct an initial public opinion map according to the original public opinion data; the initial public opinion map includes connection relationships between nodes, the nodes correspond to named entities in the original public opinion data, and the connection relationships between the nodes correspond to entity relationships between the named entities;

[0045] A public opinion relationship network construction module, configured to generate a multi-layer public opinion relationship network according to the initial public opinion map;

[0046] A public opinion hotspot map construction module, configured to generate a public opinion hotspot map according to the initial public opinion map;

[0047] An analysis module, configured to obtain a public opinion analysis result based on the multi-layer public opinion relationship network and the public opinion hotspot map.

[0048] An embodiment of the present application also discloses an electronic device, including: a processor; and a memory, on which executable code is stored, and when the executable code is executed, the processor is caused to execute one or more of the public opinion analysis methods in the embodiments of the present application.

[0049] An embodiment of the present application also discloses one or more machine-readable media, on which executable code is stored, and when the executable code is executed, a processor is caused to execute one or more of the public opinion analysis methods in the embodiments of the present application.

[0050] Compared with the prior art, the embodiments of the present application include the following advantages:

[0051] In the embodiments of the present application, original public opinion data is obtained. An initial public opinion map can be constructed based on the original public opinion data. The initial public opinion map includes connection relationships between nodes, where the nodes correspond to named entities in the original public opinion data, and the connection relationships between nodes correspond to entity relationships between named entities. A multi-layer public opinion relationship network is generated based on the initial public opinion map. Based on the generated multi-layer public opinion relationship network, complex and abstract public opinion information can be presented in an intuitive and understandable visual form, and further, multi-level and multi-angle public opinion analysis can be performed on the public opinion data. A public opinion hot spot map is generated based on the initial public opinion map, which can achieve precise tracking of the entire life cycle of public opinion hot spots and clearly display the evolution trend and dissemination situation of hot spots. Strongly time-sensitive and highly accurate public opinion analysis results can be obtained based on the multi-layer public opinion relationship network and the public opinion hot spot map, providing real-time and accurate decision-making support for enterprises. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0053] Figure 1 is a flowchart of the steps of an embodiment of a public opinion analysis method of the present application;

[0054] Figure 2 is a public opinion dissemination architecture diagram of an embodiment of a public opinion analysis method of the present application;

[0055] Figure 3 is a structural block diagram of an embodiment of a public opinion analysis device of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.

[0057] In today's era of information explosion, the importance of monitoring and analyzing public opinion and public sentiment has become increasingly prominent. In related technologies, public opinion analysis methods mainly rely on the following several technologies:

[0058] Traditional public opinion analysis systems mostly use keyword search technology. Such systems will preset keywords related to specific topics and then search for content containing these keywords in a vast amount of text data. For example, for public opinion monitoring in the automotive industry, keywords such as "automobile", "automobile manufacturing", and "sedan" can be set. The advantage of this method is its simplicity and directness, which can quickly locate the text containing the keywords. However, its limitations are also very obvious. Due to the complexity of language, relying solely on keyword search is very likely to miss a large amount of important information that is semantically related but does not contain specific keywords. For example, some emerging internet buzzwords or metaphorical expressions may not be covered in the preset keywords.

[0059] Some public opinion analysis tools can perform simple statistical analysis on text data, such as word frequency statistics. The popularity of a certain topic is measured by counting the frequency of occurrence of specific words in the text. For example, by counting the frequency of the word "environmental protection" in news reports within a specific time period, if the frequency is high, it is considered that the topic of environmental protection has a high degree of popularity. However, this method lacks a deep understanding of the semantics of the text, cannot distinguish the meaning differences of the same word in different contexts, and is easily interfered by noise data. For example, the word "apple" may refer to a fruit or a technology company, and simple word frequency statistics cannot accurately judge the public opinion popularity related to which concept.

[0060] Some public opinion analyses adopt traditional NLP techniques, such as part-of-speech tagging and named entity recognition. These techniques can perform preliminary processing on the text, identifying entities such as people, places, and organizations in the text as well as the part-of-speech information of words. However, traditional NLP techniques have problems with poor efficiency when dealing with large-scale and high-real-time public opinion data.

[0061] However, in the face of complex and diverse public opinion data, it is difficult for related technologies to accurately understand the true emotions and potential semantic relationships therein. With the development of the Internet, public opinion data has grown exponentially, and related technologies are also difficult to meet the urgent needs of real-time public opinion monitoring and analysis and hot topic tracking analysis. Moreover, it is difficult to build a perfect public opinion analysis model by using these technologies alone, and it is impossible to accurately track the evolution trend of hot topics.

[0062] To address the above problems, including the problem that it is impossible to deeply explore the deep meaning of the text and difficult to conduct public opinion analysis in real time and accurately track the evolution process of public opinion, this application proposes a public opinion analysis method that can accurately understand public opinion, track and capture public opinion hotspots in real time, so as to obtain public opinion analysis results with strong timeliness and high accuracy.

[0063] Refer to Figure 1 , which is the step flowchart of an embodiment of the public opinion analysis method of this application, including the following steps:

[0064] Step 101, obtain the original public opinion data;

[0065] Among them, the original public opinion data can be relevant public opinion data collected from multiple data sources, such as news websites, social media platforms, and automotive forums.

[0066] Step 102: Construct an initial public opinion graph based on the original public opinion data; the initial public opinion graph includes the connection relationships between nodes, the nodes correspond to the named entities in the original public opinion data, and the connection relationships between the nodes correspond to the entity relationships between the named entities.

[0067] Among them, the initial public opinion graph can be a full-scale public opinion graph constructed based on the original public opinion data. The nodes in the initial public opinion graph correspond to all the named entities in the original public opinion data. The connection relationships between the nodes, that is, the edges of the initial public opinion graph, correspond to the entity relationships between the named entities in the graph. The two endpoints of the edge correspond to two named entities with an entity relationship.

[0068] Step 103: Generate a multi-layer public opinion relationship network based on the initial public opinion graph;

[0069] Among them, the multi-layer public opinion relationship network is a multi-level public opinion dissemination relationship network extracted by analyzing the initial public opinion graph, which can meet the public opinion analysis needs at different levels and can also present the characteristics and laws of public opinion dissemination from multiple perspectives.

[0070] Step 104: Generate a public opinion hot spot graph based on the initial public opinion graph;

[0071] Among them, the public opinion hot spot graph is extracted by analyzing the initial public opinion graph and can reflect hot public opinion events.

[0072] Step 105: Obtain a public opinion analysis result based on the multi-layer public opinion relationship network and the public opinion hot spot graph.

[0073] In an embodiment provided by the present application, original public opinion data from multiple sources can be collected, and then an initial public opinion graph can be constructed based on the collected original public opinion data. The initial public opinion graph includes the connection relationships between nodes. The nodes in the graph can correspond to the named entities in the original public opinion data, and the connection relationships between the nodes in the graph can correspond to the entity relationships between the named entities. A multi-layer public opinion relationship network and a public opinion hot spot graph can be generated based on the initial public opinion graph, and a public opinion analysis result can be obtained based on the generated multi-layer public opinion relationship network and public opinion hot spot graph. According to the above implementation process, public opinion can be accurately insighted, the real-time tracking and accurate capture of public opinion hot spots can be carried out, so as to obtain a public opinion analysis result with strong timeliness and high accuracy.

[0074] Exemplarily, when obtaining the original public opinion data, web crawler technology can be used to capture articles on major news websites according to the preset keywords related to automobile brands, and extract information such as news content containing keywords, release time, and news source. Data can be obtained by leveraging the Application Programming Interface (API) provided by social media platforms. For example, relevant information such as social platform content, user information, number of likes, number of forwards, number of comments, and release time can be obtained through the social platform API. Crawler technology can also be used for well-known automobile forums to collect the titles, contents, authors, release times, reply contents, etc. of forum posts as public opinion data. After collecting the original public opinion data, the Graph.fromElements() method of Flink Gelly (graph processing library) can be used to construct an initial public opinion graph in the stream processing environment of Flink. Among them, Graph.fromElements() is a method for dynamically constructing a graph object from a set of edges or a set of edges and vertices.

[0075] In an embodiment provided by the present application, after obtaining the original public opinion data, data cleaning and preprocessing can also be performed on the original public opinion data. Exemplarily, methods such as regular expressions can be used to remove HTML (Hyper Text Markup Language) tags and special characters from the data collected from web pages, and the data format can be unified into a plain text format. At the same time, the data formats from different sources are unified. For example, the time format is unified into "YYYY-MM-DD HH:MM:SS", and the text encoding format is uniformly converted to UTF-8 (8-bit, Universal Character Set / Unicode Transformation Format) encoding, etc. After completing data cleaning and preprocessing, each piece of the collected original public opinion data can be used as a message and sent to the Kafka cluster using the client API provided by the distributed publish-subscribe messaging system Kafka according to a certain theme, such as the "automobile public opinion data" theme. Exemplarily, in Java, the KafkaProducer class can be used to send messages. After configuring parameters such as the address, port of the Kafka cluster, and the theme to be sent, the processed data can be sent to Kafka one by one.

[0076] In an embodiment provided by the present application, the original public opinion data in Kafka can be consumed in the stream processing environment of Flink. According to the consumed original public opinion data, the nodes of the initial public opinion graph are determined, and the edges of the initial public opinion graph are constructed. The named entity recognition technology can be used to identify various entities from the text data, such as automobile brand names, vehicle models, people, organizations, locations, etc., as the nodes of the initial public opinion graph. Node attributes can also be added to each identified node. Among them, the node attributes can include node type, occurrence frequency, first occurrence time, last occurrence time, etc. And the node attributes in the initial public opinion graph can be calculated and updated in real time according to the statistics and analysis of the data during the Flink stream processing. The occurrence frequency of a node can be determined by counting the occurrences of the named entity corresponding to the node in each piece of data. The first occurrence time and the last occurrence time of a node can be determined according to the timestamps in the original public opinion data containing the node. Exemplarily, entity recognition can be implemented by integrating relevant NLP libraries in the Flink job or calling external NLP services. For example, in Java, some open-source NLP libraries such as Stanford NLP can be used or the NLP API based on cloud services can be called to perform entity recognition operations. For example, from the message "The endurance of the T brand M model is really good, the driving experience is great, but the price is a bit high.", the two named entities "T brand" and "M model" can be identified as nodes. Among them, the node type of "T brand" is the automobile brand type, and the node type of "M model" is the vehicle model type. The identified named entities can also be encapsulated into a data stream composed of NamedEntity objects.

[0077] Based on the entity relationships between the named entities identified from the original public opinion data, the connection relationships between the corresponding nodes in the initial public opinion graph can be determined, and the edges in the initial public opinion graph can be correspondingly generated. The two endpoints of an edge correspond to two nodes with a connection relationship. Among them, the entity relationships can include mention relationships, belonging relationships, association relationships, etc. Connection relationship attributes can also be added to the connection relationships between nodes, and the edge attributes of the correspondingly generated edges can be generated. Among them, the connection relationship attributes can include the weight of the connection relationship, sentiment tendency, first appearance time, and most recent appearance time, etc. The weight of the connection relationship can be determined by counting the frequency of occurrence of the corresponding two entities in the same text segment and normalizing the frequency of occurrence. The sentiment tendency can be obtained by performing sentiment analysis on the sentences containing the corresponding entities. The connection relationship attributes can also be updated and calculated in real time during the Flink stream processing. Exemplary, when two entities appear in a text segment of the original public opinion data, an edge can be established between the nodes corresponding to these two entities, and the entity relationship between these two entities corresponds to the type of the edge. More than two entities can also appear in a text segment of the original public opinion data. At this time, NLP can be used to judge the specific relationship between entities by learning the common patterns of entity relationships in a large number of texts, and NLP can also judge the entity relationship by understanding the text semantics. NLP's relationship analysis, semantic understanding, and context analysis technologies can also be used to comprehensively judge the entity relationship between entities. Exemplary, for the sentence "Zhang San bought a mobile phone in Company A" in the original public opinion data, entities such as "Zhang San", "Company A", and "mobile phone" can be identified through the named entity recognition technology in NLP. Through NLP analysis, it can be obtained that the relationship between "Company A" and "mobile phone" is a "production" relationship, and an edge of "belonging relationship" can be established between them. And the relationship between "Zhang San" and "Company A" is a "consumer behavior occurs in" relationship, and an edge of the corresponding type can be established accordingly. NLP can also judge entity relevance by understanding text semantics. For example, "Li Si praised Wang Wu's design plan", NLP can judge that there is a "mention relationship" between "Li Si" and "Wang Wu" from the semantics. Also, the analysis of the text context is very important. For example, "In the office, the computer used by Zhao Liu is L, and L Company is a well-known technology enterprise", NLP can know through context analysis that the relationship between "Zhao Liu" and "computer" is a "usage relationship", the relationship between "computer" and "L" is a "belonging relationship", and "Zhao Liu" and "L" are indirectly associated through "computer" and are not directly connected. According to the above implementation process, blindly connecting all the identified entities in pairs can be avoided, thereby constructing a more reasonable and effective public opinion graph.

[0078] In an embodiment of the present application, after named entity recognition of the original public opinion data, lexical analysis can also be performed on the original public opinion data, including part-of-speech tagging and stemming. Exemplarily, the part-of-speech tagging function of a library such as Stanford NLP can be used to perform part-of-speech tagging on the text corresponding to each recognized named entity, and the result obtained is a data stream composed of CoreMap objects containing sentence part-of-speech tagging information. Exemplarily, for the text "The battery life of the T brand and M model is really good, the driving experience is great, but the price is a bit high.", "T brand" can be tagged as a noun (automobile brand name), "M model" can be tagged as a noun (model name), "of" can be tagged as a particle, "battery life" can be tagged as a noun, "ability" can be tagged as a noun, "really" can be tagged as an adverb, "very good" can be tagged as an adjective phrase, etc. For the original public opinion data that contains English words, a stemming algorithm can also be used to restore the words to their stem forms, which can be achieved by calling the relevant stemming algorithm library in Flink. Exemplarily, after stemming, "driving" can be restored to "drive", "experiences" can be restored to "experience", etc. Finally, a text data stream after stemming is obtained, where each text is the form of the original sentence after stemming. Through lexical analysis, text data after part-of-speech tagging and stemming can be obtained, and these processed results provide a more standardized and easily processed basic data form for subsequent syntactic analysis, semantic analysis, sentiment tendency recognition, etc.

[0079] In an embodiment of the present application, after named entity recognition of the original public opinion data, syntactic analysis can also be performed on the original public opinion data, including constructing a syntactic structure tree. Exemplarily, syntactic analysis can be performed on each text after stemming, and a syntactic structure tree of the sentence can be constructed in the stream processing environment of Flink using a syntactic analyzer based on probabilistic context-free grammar. For the sentence "The battery life of the T brand and M model is really good, the driving experience is great, but the price is a bit high.", its syntactic structure tree may show hierarchical structure relationships such as "T brand and M model" being part of the subject, "battery life" being the head, and "really good" being the predicate part, etc. By analyzing the syntactic structure tree, the structure and semantics of the sentence can be understood more deeply, providing a more accurate basis for subsequent semantic analysis and sentiment tendency recognition, etc.

[0080] In an embodiment of the present application, after named entity recognition of the original public opinion data, semantic analysis can also be performed on the original public opinion data, including calculating the semantic similarity between words and mining the semantic relationships between entities. Through semantic analysis, the semantic relationships between words in the text can be better understood, providing a basis for subsequent mining of the semantic relationships between entities. Exemplarily, pre-trained word vector models such as Word2Vec (word to vector) or BERT (Bidirectional Encoder Representations from Transformers) can be used to calculate the semantic similarity between words in the stream processing environment of Flink. Combining the named entity recognition results and the semantic similarity calculation results, the semantic relationships between entities can be mined, and information such as this relationship, related entities, and semantic similarity can be encapsulated into a data stream composed of SemanticRelation objects. By mining the semantic relationships between entities, the semantic information of the edges in the graph structure can also be enriched, better reflecting the real relationships between entities in the public opinion text. The mined semantic relationships between entities are applied to the entity relationships that build the connection relationships between nodes in the initial public opinion graph.

[0081] By deeply integrating one or more of various advanced NLP technologies, including named entity recognition, lexical analysis, syntactic analysis, and semantic analysis, comprehensive and in-depth processing of the original public opinion data text can be achieved, deeply understanding the semantic connotation of the original public opinion data text, thereby accurately insighting into public opinion and improving the accuracy of public opinion analysis results.

[0082] Exemplarily, as Flink continuously consumes new original public opinion data in Kafka, the incremental graph update algorithm of FlinkGelly can be used to update the graph structure. When new nodes and edges are recognized, the graph is updated in the stream processing environment of Flink through the graph.addVertex() and graph.addEdge() methods.

[0083] In an embodiment of the present application, generating a multi-layer public opinion relationship network according to the initial public opinion graph includes:

[0084] Calculating the node importance scores corresponding to each node in the initial public opinion graph;

[0085] Determining the key nodes in the initial public opinion graph according to the node importance scores;

[0086] Among them, the importance score of a node can be used as an importance indicator of the node in the initial public opinion graph. Nodes with high importance scores can be used as key nodes for the spread of public opinion. The importance score of a node can be calculated based on the sum of the out-degree and in-degree of the node and the quality of the node.

[0087] Obtain the connection relationships corresponding to the key nodes;

[0088] Among them, the connection relationships corresponding to the key nodes are expanded around the key nodes according to the actual association relationships between the nodes.

[0089] According to the connection relationships between the nodes in the initial public opinion graph, divide the initial public opinion graph into multiple topic community sub-networks;

[0090] Among them, a topic community sub-network is a local network in the initial public opinion graph. Each topic community sub-network represents a relatively independent and internally closely related public opinion topic, which is the key to reflecting topic aggregation.

[0091] Traverse the nodes and the connection relationships between the nodes in the initial public opinion graph, and identify the connected components and triangle structures in the initial public opinion graph; the triangle structure corresponds to the structure in the initial public opinion graph where three nodes are pairwise connected;

[0092] Among them, a connected component, also called a maximal connected subgraph. Any two nodes in a connected component can be directly or indirectly connected. An initial public opinion graph can include one or more connected components, and each connected component represents a different topic and the corresponding propagation path.

[0093] Generate the multi-layer public opinion relationship network based on the key nodes, the connection relationships corresponding to the key nodes, the topic community sub-networks, the connected components, and the triangle structures.

[0094] After constructing the initial public opinion graph, the node importance scores corresponding to each node in the initial public opinion graph can be calculated. Nodes with high importance scores are determined as key nodes in the initial public opinion graph, and the connection relationships corresponding to the key nodes are obtained. Exemplarily, the PageRank (web page ranking) algorithm can be used to evaluate the importance of each node in the initial public opinion graph. In the public opinion graph, an edge represents the relationship of mutual mention, association, etc. between nodes. The PageRank algorithm can calculate its importance score based on the sum of the out-degree and in-degree of the node and the quality of the node. When implemented in Flink Gelly, it can also be iterated multiple times to update the PageRank value of the node according to the importance of the source nodes of the incoming edges of each node and the weights of the edges. For example, in the initial public opinion graph in the automotive field, the PageRank value of the automotive brand nodes often mentioned by authoritative automotive evaluation media is relatively high.

[0095] According to the connection relationships between nodes in the initial public opinion graph, the initial public opinion graph can be divided into multiple topic community sub-networks. In the automotive brand public opinion data, there are various topic communities, such as automotive technology discussions, exterior design comments, after-sales service feedback, etc. Exemplarily, after constructing the initial public opinion graph, a community discovery algorithm can be applied to divide multiple topic community sub-networks from the initial public opinion graph. Using the community discovery algorithm in Flink Gelly, such as the Louvain algorithm, the graph can be divided according to the connection relationships between nodes. Through iterative optimization, the internal connections of the community sub-networks are made tight, and the connections between community sub-networks are relatively sparse.

[0096] By traversing all the connection relationships between nodes in the initial public opinion graph, the connected components and triangle structures in the initial public opinion graph can be identified. There can be one or more connected components in the initial public opinion graph. If all the nodes in the initial public opinion graph can be directly or indirectly connected, then the initial public opinion graph has only one connected component, which is the initial public opinion graph itself. If there is a part of the nodes in the initial public opinion graph that have no connection relationship with the rest of the nodes, then there are multiple connected components in the initial public opinion graph. The existence of multiple connected components can reflect to a certain extent the information dissemination limitations or topic independence in the process of public opinion dissemination. Exemplarily, the connected component algorithm of Flink Gelly can be used to traverse the nodes and connection relationships of the initial public opinion graph, and the mutually reachable nodes can be divided into one connected component. The connected component algorithm can be implemented based on depth-first search or breadth-first search. The triangle enumerator algorithm can be used to traverse all combinations of nodes and connection relationships in the initial public opinion graph to find all triangle structures in the graph.

[0097] A multi-layer public opinion relationship network can be generated based on the key nodes, the connection relationships corresponding to the key nodes, the topic community sub-networks, the connected components, and the triangle structures obtained from the initial public opinion graph.

[0098] In an embodiment of the present application, the generating of the multi-layer public opinion relationship network according to the key nodes, the connection relationships corresponding to the key nodes, the topic community sub-networks, the connected components, and the triangle structures includes:

[0099] Extracting the key nodes and the connection relationships corresponding to the key nodes in the initial public opinion graph to generate a top-layer public opinion relationship network; the top-layer public opinion relationship network includes core nodes, edge nodes, and connection relationships between nodes, the core nodes correspond to the key nodes, and the edge nodes are directly or indirectly connected to the core nodes through the connection relationships between nodes;

[0100] Extracting the topic community sub-networks in the initial public opinion graph to generate a middle-layer public opinion relationship network;

[0101] Extract the connected components and the triangular structures in the initial public opinion graph to generate an underlying public opinion relationship network;

[0102] The top-level public opinion relationship network, the middle-level public opinion relationship network, and the underlying public opinion relationship network constitute the multi-layer public opinion relationship network.

[0103] The key nodes and the connection relationships corresponding to the key nodes in the initial public opinion graph can be extracted. Taking the key nodes as the core nodes and the other nodes associated with the key nodes as the edge nodes, and expanding around the key nodes according to the connection relationships between the nodes, a top-level public opinion relationship network centered on the key nodes is formed. In addition, the node importance scores of each node can be reflected through the visual features of the nodes in the top-level public opinion relationship network, such as by updating the size and color depth of each node. Exemplarily, the key nodes with high importance extracted from the initial public opinion graph can be well-known automobile brands or authoritative evaluation institutions, etc. The top-level public opinion relationship network generated according to the key nodes and the connection relationships corresponding to the key nodes can intuitively present the core-edge structure of public opinion dissemination.

[0104] The topic community sub-networks divided in the initial public opinion graph can be extracted to generate a middle-level public opinion relationship network. Exemplarily, in the initial public opinion graph of automobile brands, a topic community sub-network centered on "the battery life problem of new energy vehicles" can be divided, and its nodes can include relevant automobile brands, models, consumers and experts concerned about battery life, battery life technology R & D institutions, etc. Its connection relationship, that is, the edge in the topic community sub-network, can correspond to the discussion, research, and feedback relationships of each node around the battery life problem. The middle-level public opinion relationship network generated according to the topic community sub-network can clearly present the classification and aggregation of public opinion topics, which is conducive to analyzing the characteristics and development trends of public opinion in different topics.

[0105] All connected components and triangular structures identified in the initial public opinion map can be extracted to generate an underlying public opinion relationship network. Each connected component represents a different sub-topic of public opinion or an independent information dissemination group. Exemplarily, connected components representing antique car restoration and modern electric vehicle technology discussions can be extracted from the initial public opinion map, and the two connected components are completely independent of each other. Connected components are the key to reflecting topic independence and dividing public opinion modules in the public opinion analysis map, which helps classify complex public opinion data and conduct targeted analysis. In public opinion analysis, the triangular structure reveals closer relationships between entities and is the key to analyzing the tightness of relationships and stable public opinion groups, which helps analyze the tightness of relationships between entities and potential group structures in the public opinion network. Exemplarily, a triangular structure composed of a car brand, a car club, and a well-known car modifier can be extracted from the initial public opinion map, and this triangular structure can indicate closer interactions among the car brand, the car club, and the well-known car modifier in public opinion dissemination and topic development. Small groups with little influence in public opinion dissemination can also be discovered based on the tight triangular relationships among the extracted car brands, industry associations, and opinion leaders. The underlying public opinion relationship network generated based on the extracted connected components and triangular structures enriches the levels and diversity of the relationship network, reflects the independence and segmentation of public opinion topics, and helps discover small groups or stable relationships with important influence in public opinion dissemination and topic development.

[0106] The multi-layer public opinion relationship network can be composed of a top-layer public opinion relationship network, a middle-layer public opinion relationship network, and an underlying public opinion relationship network. According to the above implementation process, a multi-level public opinion dissemination relationship network can be formed to meet the public opinion analysis needs at different levels.

[0107] Exemplarily, when new key nodes, connection relationships corresponding to the key nodes, topic community sub-networks, connected components, and triangular structures are extracted from the initial public opinion map, the multi-layer public opinion relationship network can be incrementally updated. With the help of Flink's powerful stream processing framework, the problem of insufficient data dynamic change processing ability in related technologies can be effectively solved. It can consume a large amount of public opinion data from Kafka at an extremely high speed and quickly update the initial public opinion map and the multi-layer public opinion relationship network in a real-time stream processing environment, so as to achieve public opinion analysis with strong timeliness and high accuracy.

[0108] Exemplarily, during the process of constructing a multi-layer public opinion relationship network, the dissemination scope and speed of specific information sources can also be concerned, and the single-source shortest path algorithm is used to obtain the shortest path lengths from the information source nodes to other nodes, and the key paths in public opinion dissemination are correspondingly obtained. Among them, the specific information source can be the evaluation of a car brand by a car blogger. The nodes in the initial public opinion graph cover various public opinion-related entities, including bloggers, car brands, models, related topics, etc., and the connection relationships represent the associations between entities. The connection relationship attributes also include the weights of the connection relationships, which can reflect the information dissemination probability or association intensity. The specific information source node can be set as the source point, and the single-source shortest path algorithm of Flink Gelly, such as the Dijkstra algorithm, is used to traverse the connection relationships between the nodes in the initial public opinion graph, and the path lengths are updated according to the weights of the connection relationships, so as to calculate the shortest paths from this source point to all other nodes. The weights of the connection relationships correspond to the lengths of the edges. Calculating the paths from the source point to all other nodes is to calculate the sum of the weights of the edges passed from the source point to each node. Obtaining the shortest path lengths from the information source to other entities helps to analyze the dissemination efficiency and influence scope of key information in the public opinion network. Exemplarily, based on the source point of "the official release of a new car model by a car brand", by obtaining its corresponding shortest path length, the corresponding key path can be determined as the fastest channel for the message to spread to the consumer group, and the nodes and connection relationships on the key path can also constitute the key part related to the information dissemination efficiency in the public opinion relationship network. Applying the single-source shortest path algorithm to different information sources can determine the dissemination paths of information from each source. Multiple key paths are intertwined in the multi-layer public opinion relationship network, showing the flow trajectories of information between different nodes, constituting the information dissemination path part of the public opinion relationship network, and can be used for subsequent public opinion analysis.

[0109] Exemplarily, the nodes in the initial public opinion graph can be classified and labeled. Among them, the classification labels can include positive, negative, neutral, etc. The LabelPropagation algorithm can be used to start from the initially labeled nodes and update the classification labels of the unlabeled nodes according to the label conditions of adjacent nodes. The initial nodes can be determined by comprehensively considering the degree of the nodes, the importance of the nodes, the node attributes, and the connection relationship attributes of the corresponding connections of the nodes. In Flink Gelly, the label information is iteratively transmitted so that each node has a label. For example, the initial label of the evaluation node of a certain automobile brand by some consumers is positive, and the consumer nodes discussing the same model around may be affected and labeled as positive. By classifying and labeling the nodes, the public opinion tendency and classification can be quickly identified, and the overall situation of the public opinion can be reflected according to the distribution and connection conditions of the nodes with different labels. For example, during a certain period, a large number of nodes related to a certain automobile brand are labeled as negative and closely related, indicating that the brand may be facing a public opinion crisis. Exemplarily, the obtained classification labels can also be represented by different colors or symbols in the multi-layer public opinion relationship network. For example, green is used to represent positive nodes, red is used to represent negative nodes, and gray is used to represent neutral evaluation nodes. In this way, the multi-layer public opinion relationship network can not only display the node connections and information dissemination paths, but also intuitively present the public opinion tendency and emotional atmosphere, and quickly grasp the public opinion status in different regions during the public opinion analysis.

[0110] Referring to Figure 2 As shown in the public opinion dissemination architecture diagram of an embodiment of the public opinion analysis method of the present application, web crawlers can be used to obtain the original public opinion data from multiple information sources such as automobile content websites, social platforms, and automobile trading websites, and then the obtained original public opinion data is sent to the distributed publish-subscribe message system Kafka. The original public opinion data in the Kafka cluster is processed by a graph processing library, such as Flink Gelly, and natural language processing NLP. Then, the processed public opinion data can be input into a distributed graph database, such as Nebula Graph, and a comprehensive multi-layer public opinion relationship network is constructed using the distributed graph database. Taking a certain automobile brand (assumed to be "X Automobile") as an example, the constructed multi-layer public opinion relationship network can include the following content:

[0111] 1. Top layer - Core driving layer: Through the PageRank algorithm, it can be calculated that the "X Automobile" brand itself is in a key position in the entire public opinion network and is frequently mentioned and discussed by various media and consumers. Therefore, "X Automobile" is regarded as the most important node, and its size is the most prominent in the graph. The senior leaders or official spokespersons of the automobile brand, such as the spokesperson of X Automobile, the CEO of X Automobile, and the product manager of X Automobile, their remarks and decisions have a significant impact on the brand's public opinion and are also nodes with relatively high PageRank values. For example, the remarks of the CEO of X Automobile and other content can become the focus of public opinion and are closely connected to the brand node. Authoritative automobile evaluation agencies, such as professional automobile magazines and well-known automobile evaluation bloggers, their evaluations and rankings of X Automobile will affect consumers' views and have a strong association with the brand. Therefore, they can be presented as relatively large and important nodes in the graph. There are strong association edges between the brand, the leaders, and the evaluation agencies. For example, when a new brand product is launched, the evaluation agency is invited to participate, and the leader appears in the interview of the evaluation agency. The edges formed by these interactions reflect the flow of information among these core roles and also reflect the influence dissemination path of the brand in the industry. The thickness of the edges can represent the degree of closeness of the association. For example, a well-known automobile evaluation blogger has 9 million fans, and a well-known automobile evaluation blogger with whom the brand has a long-term cooperation has 5 million fans. However, the association between the brand and the well-known automobile evaluation blogger with whom it has a long-term cooperation is closer, and the corresponding edge is thicker. There will also be associations between different automobile evaluation bloggers, such as between a well-known automobile evaluation blogger and Automobile Evaluation Blogger A.

[0112] 2. Middle Layer - Topic Community Layer: The topic community layer includes a technology discussion community, a design review community, and a after-sales service feedback community. Nodes in the technology discussion community can include R & D engineers of X cars, automotive technology experts, consumers who are concerned about the technological innovation of X cars, etc. Edges are formed through discussions in technology forums and social media technology groups among the nodes. These edges represent the communication on the technological topics of X cars. For example, R & D engineers disclose some stories behind technology R & D to consumer A, and technology experts answer consumer B's questions about the new power system, forming an information dissemination network centered around technological topics. Nodes in the design review community can include automotive designers, consumers who are concerned about the exterior and interior design of cars, design-related media, etc. Edges are formed through exchanges on design-related platforms among the nodes. For example, design media XX Design Magazine publishes an article interpreting the design of X cars. After external designers read the article published by XX Design Magazine, they post improvement suggestions on the design details in the comment section. Consumer C discusses the improvement suggestions posted by external designers in the comment section. Consumer D posts his own opinions in the comment section after reading the article published by XX Design Magazine, constituting a community network themed on design reviews. Nodes in the after-sales service feedback community can include employees of the after-sales service center of X cars, car owners who encounter after-sales problems, consumer rights protection organizations, etc. The feedback and handling process of after-sales problems form edges. For example, the communication chain where the problem car owner complains to the after-sales staff of the after-sales service center and the after-sales staff responds and solves the problem. Or the attention and intervention of consumer rights protection organizations in after-sales problems can also form a sub-network of the topic community related to after-sales service.

[0113] 3. Bottom layer - Micro - detail layer: The micro - detail layer can include connected components and triangular structures. Exemplarily, a connected component can be related to a specific vehicle model. The nodes in the connected component can include the owners of a popular model of a certain vehicle of X Automobile, such as the owners of Model X1, and can also include potential purchasers and automotive parts suppliers of this model. The communication between owner A and owner B about the driving experience of this model, owner B's description of the vehicle to potential purchasers, the communication between the parts supplier and owner A regarding parts quality feedback, and the potential purchaser's understanding of parts from the parts supplier, etc. form edges. This connected component is independent of discussions related to other vehicle models, forming a relatively closed information - spreading group that focuses on the public opinion of Model X1 of X Automobile. Exemplarily, the triangular structure can be a triangular structure with nodes being the official of X Automobile, the X Automobile Club, and the X Automobile modifier. The cooperation between the official of X Automobile and the X Automobile Club in organizing activities and promoting brand culture forms an edge; the X Automobile modifier sharing modification experience with members of the X Automobile Club forms an edge; the X Automobile modifier and the official may form an edge in terms of cooperation in modification events or the official's recognition of the modifier's works. The triangular structure composed of these three nodes reflects the stable relationship in the fields of automotive culture dissemination and personalized modification, and the interaction between the nodes plays an important role in the development of public opinion of X Automobile in specific fields.

[0114] Exemplarily, assume that X Automobile releases a new model. Starting from the source of the official release of the news, the information may be spread through the first - layer dissemination nodes, such as the official website and the official social media accounts, and then to the second - layer dissemination nodes, such as automotive media and automotive bloggers, and then to the third - layer dissemination nodes, such as ordinary consumers and potential consumers. Through the single - source shortest - path algorithm, it can be determined which edges are the key paths for the information to reach consumers fastest during this dissemination process. For example, if it is found that the reposting by well - known long - term cooperative automotive review bloggers on social platforms plays a key role in the rapid spread of information, then the associated edges between the official and these bloggers and the edges between the bloggers and their fans are the key paths for the spread of public opinion during the new product release, and can be highlighted in the graph in special colors or bold fonts, etc.

[0115] Exemplarily, in the entire public opinion dissemination architecture diagram, nodes can be marked according to the LabelPropagation algorithm. For example, nodes with more positive comments on X cars, such as nodes related to consumers' praise for vehicle comfort and cost-effectiveness, are marked green, indicating positive public opinion; nodes with more complaints about vehicle quality problems, such as nodes related to consumers' reports of engine failures and other problems on the complaint platform, are marked red, indicating negative public opinion; for some neutral evaluations, such as nodes related to evaluations that are not satisfied with a certain vehicle configuration but are generally acceptable, are marked gray. Through this color marking, the public opinion tendency in different regions of the entire public opinion network can be intuitively seen, and the brand's word-of-mouth in different aspects can be quickly understood.

[0116] In an embodiment of the present application, generating the public opinion hotspot map according to the initial public opinion map includes:

[0117] Extracting the connected subgraphs in the initial public opinion map;

[0118] Among them, the connected subgraph is a part of the initial public opinion map, including the connection relationships between some nodes in the initial public opinion map. Any two nodes in the connected subgraph can be directly or indirectly connected.

[0119] Calculating the comprehensive score corresponding to the connected subgraph;

[0120] Determining potential hotspots according to the comprehensive score;

[0121] Generating the public opinion hotspot map based on the potential hotspots.

[0122] Mark a node in the initial public opinion graph as the starting node. Starting from the starting node, along the connection relationships of this node, mark all the nodes passed until there are no new nodes to traverse. In this way, the connection relationships between the marked nodes form a connected subgraph. Then, select one of the unmarked nodes as the new starting node and repeat the above operations until all nodes in the initial public opinion graph are marked. In this way, all connected subgraphs in the initial public opinion graph can be extracted. The connected component is the largest connected subgraph. After extracting the connected subgraphs, the comprehensive scores of each connected subgraph can be calculated. Select the connected subgraphs whose comprehensive scores exceed the preset threshold, whose comprehensive scores continue to rise within the preset time, and whose emotional tendencies of the connection relationships within the subgraph show strong consistency. Determine the topics represented by such connected subgraphs as potential hotspots. The emotional tendencies of the connection relationships within the subgraph showing strong consistency can be that most of the emotional tendencies of the connection relationships in this connected subgraph are positive or most are negative. Based on the determined potential hotspots, a corresponding public opinion hotspot graph can be generated. Exemplarily, the hotspot judgment and tracking operations can be implemented by writing corresponding logical codes in the stream processing environment of Flink. According to the above implementation process, the subtle dynamic changes of public opinion hotspots can be captured in a timely manner, the accurate tracking of the full life cycle of public opinion hotspots of automobile brands can be realized, real-time and accurate public opinion analysis can be carried out, decision-making support can be provided for enterprises, and enterprises can be helped to make rapid responses.

[0123] In an embodiment of the present application, the initial public opinion graph includes node attributes and connection relationship attributes. Calculating the comprehensive score corresponding to the connected subgraph includes:

[0124] Determine the emotional tendency of the connection relationships between the nodes in the connected subgraph; the emotional tendency is determined by performing sentiment analysis on the original public opinion data corresponding to the connection relationships between the nodes;

[0125] Calculate the node importance score in the connected subgraph; the node importance score is determined according to the degree of the node, the node weight, and the emotional tendency of the connection relationship corresponding to the node; the node attributes include the degree of the node and the node weight, and the connection relationship attributes include the emotional tendency of the connection relationship corresponding to the node;

[0126] Calculate the comprehensive score of the connected subgraph according to the node importance score and the emotional tendency.

[0127] After performing named entity recognition on the original public opinion data, sentiment tendency recognition can also be carried out on the original public opinion data. By performing sentiment analysis on the original public opinion data corresponding to the connection relationship between nodes, the sentiment tendency corresponding to the connection relationship is determined. And the determined sentiment tendency is added to the connection attribute of the corresponding connection relationship. When calculating the comprehensive score of a node, the corresponding sentiment tendency can be obtained from the connection attribute of the connection relationship corresponding to the node. Exemplarily, it can be implemented by integrating a relevant sentiment classification model library or calling an external sentiment classification model service in the Flink stream processing environment, and this application does not limit this. A sentiment dictionary can be constructed based on lexical and context analysis, including positive words, negative words, etc. By constructing the sentiment dictionary, sentiment analysis can be performed on the text corresponding to each identified named entity, the number of positive words and negative words is counted, and the sentiment tendency of the sentence is judged in combination with the grammatical structure and context of the text. For example, when performing sentiment tendency recognition on the text "The battery life of the M model of T brand is really good, and the driving experience is great, but the price is a bit high.", it can be judged that the corresponding sentiment tendency is "positive" because the number of occurrences of positive words such as "really good" and "great" is more than the number of occurrences of negative words such as "a bit high". It is also possible to use a sentiment classification model based on deep learning, such as a model based on a convolutional neural network or a recurrent neural network. In the Flink stream processing environment, the sentence is input into the model, and the sentiment tendency category is output.

[0128] The formula for calculating the node importance score is I = α * d + β * w + γ * s. Here, I is the node importance score, d is the degree of the node, α is the weight coefficient of the degree, w is the weight of the node, β is the weight coefficient of the weight, s is the sentiment tendency value corresponding to the connection relationship of the node. The sentiment tendency value can be set correspondingly as positive +1, negative -1, and neutral 0, and γ is the weight coefficient of the sentiment tendency. If the degree of the node is relatively high, the occurrence frequency is high, and the sentiment tendencies of the connected edges are mostly positive, the importance score is relatively high. As new data is added and the structure of the initial public opinion graph is updated, the degree, weight, and sentiment tendency of the connection relationship of the node will change. Therefore, the node importance scores of each node can be updated in real time according to the changes in the initial public opinion graph. The weight coefficients of the degree, weight, and sentiment tendency of the node can be adaptively set according to the specific characteristics of the public opinion data or business requirements, and this application does not limit this. Exemplarily, if it is desired that the weight and degree of the node have a greater impact on the score and the sentiment tendency has a smaller impact, the weight coefficients of each parameter can be adjusted correspondingly when calculating the node importance score. For example, set α to 0.3, β to 0.5, and γ to 0.2. The initial public opinion graph includes node attributes and connection attributes. The node attributes include the degree of the node and the node weight, and the connection relationship attributes include the sentiment tendency of the connection relationship corresponding to the node. When calculating the node importance scores of each node in the connected subgraph, the degree and node weight of the node can be obtained from the node attributes of the corresponding node, and the corresponding sentiment tendency can be obtained from the connection attributes of the connection relationship corresponding to the node, and then the corresponding sentiment tendency value can be obtained according to the preset value.

[0129] According to the node importance scores of each node in the connected subgraph and the sentiment tendency of the corresponding connection relationship, the comprehensive score of each connected subgraph can be calculated. The formula for the comprehensive score is:

[0130]

[0131] where S represents the comprehensive score of the connected subgraph, w i represents the node importance score of the i-th node in the connected subgraph, n represents the number of nodes in the connected subgraph, γ represents the weight coefficient of the sentiment tendency of the connection relationship in the connected subgraph, e j represents the sentiment tendency score of the j-th connection relationship in the connected subgraph, and m represents the number of connection relationships in the connected subgraph.

[0132] Exemplarily, by performing real-time analysis and statistics on the original public opinion data consumed from Kafka, the original public opinion graph can be updated in real time, and the information such as the node importance scores of the nodes in the obtained connected subgraph and the sentiment tendency of the connection relationship can be updated in real time, thereby updating the comprehensive score of the connected subgraph in real time.

[0133] In an embodiment of this application, generating the public opinion hotspot graph based on the potential hotspot includes:

[0134] Monitoring the potential hotspots to obtain potential hotspot data of the potential hotspots; the potential hotspot data at least includes duration, heat change value, and hotspot propagation path;

[0135] Generating a public opinion hotspot distribution map according to the potential hotspots and the hotspot propagation path;

[0136] Generating a hotspot evolution curve graph according to the duration and the heat change value of the potential hotspots; the public opinion hotspot map includes the public opinion hotspot distribution map and the hotspot evolution curve graph.

[0137] By continuously monitoring the development of potential hotspots, potential hotspot data of the potential hotspots can be obtained, which at least includes information such as the duration of the hotspots, heat change values, and hotspot propagation paths. Among them, the heat change value of the hotspots can be determined by the comprehensive score of the connected subgraph corresponding to the hotspots and the change in the importance score of the nodes in the connected subgraph. The hotspot propagation path can be used to analyze the actual positions and relationships of newly added nodes and connection relationships in the connected subgraph corresponding to the hotspots. Exemplarily, the hotspot evolution tracking operation can be implemented by integrating relevant tracking record libraries or writing corresponding logical codes in Flink jobs.

[0138] A public opinion hotspot distribution map can be generated according to potential hotspots and the monitored hotspot propagation path. Exemplarily, after determining the potential hotspots, relevant data of the potential hotspots can be collected from various public opinion monitoring channels, such as information such as content, comments, likes, and reposts posted by users related to the potential hotspots. Then, through natural language processing technology, the collected data is analyzed to identify keywords and topics related to the hotspots, so as to determine the specific content and scope of the hotspots. The hotspot propagation path can also be obtained by analyzing the propagation method of hotspot information in the network, including the propagation source, propagation nodes, propagation channels, and the relationships between propagation nodes. Then, according to the potential hotspots and the hotspot propagation path, a public opinion hotspot distribution map is obtained. The potential hotspots can be one or more.

[0139] Exemplarily, when generating a public opinion hot spot distribution map, nodes of different shapes and colors can be used to represent different types of entities. For example, circles represent car brands, rectangles represent car models, triangles represent people, etc. The size of the nodes can be proportional to the node importance score. For example, car brand nodes with a high market share and high discussion heat have a high node importance score and can be displayed as large circles in the graph. The colors of the nodes can be distinguished according to the sentiment tendency. For example, green represents positive sentiment, red represents negative sentiment, and gray represents neutral sentiment. For example, when showing the public opinion hot spot map of brand T, the brand node of "brand T" is represented by a larger circle. If its sentiment tendency is positive, it is displayed in green; if it is negative, it is red. Different colors and thicknesses of edges can also be used to represent the types and weights of connection relationships. For example, solid lines can be used to represent "ownership relationships" edges, and dashed lines can be used to represent "association relationships" edges; the thickness of the edges can be proportional to the weight of the connection relationship, indicating the degree of closeness between two named entities. The color of the edges can be set according to the sentiment tendency, with green representing positive, red representing negative, and gray representing neutral). For example, the edge connecting the brand node of "brand T" and the car model node of "model M" represents an ownership relationship and can be drawn with a solid line. If the sentiment tendency of this edge is positive, it is displayed in green, and the thickness of the edge is determined according to the weight. The nodes and edges are reasonably arranged in a two-dimensional plane to make the graph structure clear and intuitive, showing the relationships and hot spot distributions among various entities in the public opinion of brand T. According to the above steps, a public opinion hot spot distribution map can be generated and visually displayed, intuitively showing various entities and their relationships in the hot public opinion.

[0140] Based on the duration of potential hot spots and the heat change values, a hot spot evolution curve graph can be generated. Exemplarily, according to the real-time monitored public opinion information related to potential hot spots, such as the release volume, interaction volume (comments, likes, forwards, etc.) and dissemination scope related to potential hot spots, etc., the connection relationships, node attributes, and connection attributes among the nodes in the connected subgraph corresponding to the potential hot spots can be updated, further updating the comprehensive score of the connected subgraph corresponding to the potential hot spots and the node importance scores of each node in the graph. The heat value of each potential hot spot is obtained by synthesizing the comprehensive score of the connected subgraph corresponding to the hot spot and the importance scores of the nodes in the connected subgraph, and the heat change value is obtained according to the heat values at different time points.

[0141] Exemplarily, when generating a hot topic evolution curve graph, time can be used as the horizontal axis, and public opinion data can be divided according to a certain time interval. Among them, the time interval can be hours or days, and the present application does not limit this. The comprehensive score of the connected subgraph corresponding to the potential topic or the node importance score of the key nodes in the connected subgraph can be used as the vertical axis index to draw a curve that changes over time. Exemplarily, taking days as the unit, the collected public opinion data of brand T can be divided according to the date, and a curve showing the change of the comprehensive score of the hot topic of "price reduction of model M of brand T" over time can be drawn, or a curve showing the change of the brand node importance score of "brand T" over time can be drawn. According to the above steps, a hot topic evolution curve graph can be generated and visually displayed. The hot topic change curve in the graph can intuitively present each stage of the generation, development, climax and decline of the hot topic.

[0142] In an embodiment of the present application, the obtaining of the public opinion analysis result according to the multi-layer public opinion relationship network and the public opinion hot topic graph includes:

[0143] Generating a comprehensive public opinion dissemination analysis report according to the multi-layer public opinion relationship network;

[0144] Generating a hot topic analysis report according to the public opinion hot topic graph;

[0145] Obtaining the public opinion analysis result according to the comprehensive public opinion dissemination analysis report and the hot topic analysis report.

[0146] When conducting public opinion analysis, based on the top-level public opinion relationship network, the key driving forces and corresponding important dissemination paths in public opinion dissemination can be analyzed. Based on the middle-level public opinion relationship network, the public opinion dissemination structure and sentiment tendency under different topics can be analyzed, and the dissemination laws of each topic can be further analyzed. Based on the bottom-level public opinion relationship network, the public opinion details within a local range and the close relationship of entity keys can be analyzed, and the microscopic public opinion dynamics can be further explored. According to the multi-layer public opinion relationship network, multi-level analysis results can be obtained from the macroscopic industry control, specific topics to public opinion small groups. In addition, information such as information dissemination efficiency, topic discovery, influence distribution, topic independence, relationship tightness and public opinion tendency can also be analyzed to analyze the characteristics and laws of public opinion dissemination from multiple angles. Combining the above multi-level and multi-angle analysis contents, a comprehensive public opinion dissemination analysis report can be generated.

[0147] When conducting public opinion analysis, based on the public opinion hot spot distribution map, it is possible to quickly identify entity roles, importance, and sentiment tendencies, gain insights into the tightness of entity associations, and grasp the overall public opinion situation from a macro perspective. Based on the hot spot evolution curve graph, it is possible to grasp the dynamic change trend of hot spot public opinion in real time, analyze the development process of hot spots according to the generation, development, climax, and decline stages of the hot spots presented in the graph, predict the public opinion trend, and compare different hot spots or nodes. By comprehensively extracting and analyzing the information obtained from the public opinion hot spot map, a hot spot analysis report can be generated.

[0148] According to the comprehensive public opinion dissemination analysis report and the hot spot analysis report, the final public opinion analysis result can be obtained, providing strong support for public opinion monitoring, public opinion response, and public opinion decision-making.

[0149] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present application are not limited by the described action sequence, because according to the embodiments of the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present application.

[0150] Based on the above embodiments, the present embodiment further provides a public opinion analysis device, which is applied to electronic devices such as terminal devices and servers.

[0151] Referring to Figure 3 , a structural block diagram of an embodiment of a public opinion analysis device of the present application is shown, which specifically may include the following modules:

[0152] The acquisition module 301 is used to acquire the original public opinion data;

[0153] The initial public opinion map construction module 302 is used to construct an initial public opinion map according to the original public opinion data; the initial public opinion map includes the connection relationships between nodes, the nodes correspond to the named entities in the original public opinion data, and the connection relationships between the nodes correspond to the entity relationships between the named entities;

[0154] The public opinion relationship network construction module 303 is used to generate a multi-layer public opinion relationship network according to the initial public opinion map;

[0155] The public opinion hot spot map construction module 304 is used to generate a public opinion hot spot map according to the initial public opinion map;

[0156] The analysis module 305 is used to obtain a public opinion analysis result according to the multi-layer public opinion relationship network and the public opinion hot spot map.

[0157] The public opinion relationship network construction module further includes:

[0158] A node importance score calculation sub-module, which is used to calculate the node importance scores corresponding to each node in the initial public opinion graph;

[0159] A key node determination sub-module, which is used to determine the key nodes in the initial public opinion graph according to the node importance scores;

[0160] A connection relationship acquisition sub-module corresponding to the key nodes, which is used to acquire the connection relationships corresponding to the key nodes;

[0161] A topic community sub-network division sub-module, which is used to divide the initial public opinion graph into multiple topic community sub-networks according to the connection relationships between the nodes in the initial public opinion graph;

[0162] A connected component and triangle structure recognition sub-module, which is used to traverse the nodes and the connection relationships between the nodes in the initial public opinion graph to identify the connected components and triangle structures in the initial public opinion graph; the triangle structure corresponds to the structure in which three nodes in the initial public opinion graph are connected pairwise;

[0163] A multi-layer public opinion relationship network generation sub-module, which is used to generate the multi-layer public opinion relationship network according to the key nodes, the connection relationships corresponding to the key nodes, the topic community sub-networks, the connected components and the triangle structures.

[0164] The multi-layer public opinion relationship network generation sub-module further includes:

[0165] A top-layer public opinion relationship network generation unit, which is used to extract the key nodes and the connection relationships corresponding to the key nodes in the initial public opinion graph to generate a top-layer public opinion relationship network; the top-layer public opinion relationship network includes core nodes, edge nodes and the connection relationships between the nodes, the core nodes correspond to the key nodes, and the edge nodes are directly or indirectly connected to the core nodes through the connection relationships between the nodes;

[0166] A middle-layer public opinion relationship network generation unit, which is used to extract the topic community sub-networks in the initial public opinion graph to generate a middle-layer public opinion relationship network;

[0167] A bottom-layer public opinion relationship network generation unit, which is used to extract the connected components and the triangular structures in the initial public opinion graph to generate a bottom-layer public opinion relationship network;

[0168] The top-layer public opinion relationship network, the middle-layer public opinion relationship network and the bottom-layer public opinion relationship network constitute the multi-layer public opinion relationship network.

[0169] The public opinion hot spot graph construction module includes:

[0170] A connected sub-graph extraction sub-module, which is used to extract the connected sub-graphs in the initial public opinion graph;

[0171] A comprehensive score calculation sub-module for calculating the comprehensive score corresponding to the connected sub-graph;

[0172] A potential hot spot determination sub-module for determining potential hot spots according to the comprehensive score;

[0173] An opinion hot spot map generation sub-module for generating the opinion hot spot map based on the potential hot spots.

[0174] The initial opinion map includes node attributes and connection relationship attributes, and the comprehensive score calculation sub-module is further used for:

[0175] Determining the sentiment tendency of the connection relationship between the nodes in the connected sub-graph; the sentiment tendency is determined by performing sentiment analysis on the original opinion data corresponding to the connection relationship between the nodes;

[0176] Calculating the node importance score in the connected sub-graph; the node importance score is determined according to the degree of the node, the node weight, and the sentiment tendency of the connection relationship corresponding to the node; the node attributes include the degree of the node and the node weight, and the connection relationship attributes include the sentiment tendency of the connection relationship corresponding to the node;

[0177] Calculating the comprehensive score of the connected sub-graph according to the node importance score and the sentiment tendency.

[0178] The opinion hot spot map generation sub-module further includes:

[0179] A monitoring unit for monitoring the potential hot spots to obtain potential hot spot data of the potential hot spots; the potential hot spot data at least includes the duration, the heat change value, and the hot spot propagation path;

[0180] An opinion hot spot distribution map generation unit for generating an opinion hot spot distribution map according to the potential hot spots and the hot spot propagation path;

[0181] A hot spot evolution curve graph generation unit for generating a hot spot evolution curve graph according to the duration and the heat change value of the potential hot spots; the opinion hot spot map includes the opinion hot spot distribution map and the hot spot evolution curve graph.

[0182] The analysis module is further used for:

[0183] Generating an opinion comprehensive communication analysis report according to the multi-layer opinion relationship network;

[0184] Generating a hot spot analysis report according to the opinion hot spot map;

[0185] Obtaining the opinion analysis result according to the opinion comprehensive communication analysis report and the hot spot analysis report.

[0186] An opinion analysis device provided by an embodiment of the present application can collect raw opinion data from multiple sources, and then construct an initial opinion map based on the collected raw opinion data. The initial opinion map includes connection relationships between nodes, where the nodes in the map can correspond to named entities in the raw opinion data, and the connection relationships between the nodes in the map can correspond to entity relationships between the named entities. A multi-layer opinion relationship network and an opinion hot spot map can be generated based on the initial opinion map, and an opinion analysis result can be obtained based on the generated multi-layer opinion relationship network and opinion hot spot map. According to the above implementation process, the opinion can be accurately insighted, the real-time tracking and precise capture of opinion hot spots can be carried out, so as to obtain an opinion analysis result with strong timeliness and high accuracy.

[0187] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment.

[0188] An embodiment of the present application also provides a non-volatile readable storage medium, in which one or more modules (programs) are stored. When the one or more modules are applied to a device, the device can be enabled to execute instructions (instructions) for each method step in the embodiment of the present application.

[0189] An embodiment of the present application provides one or more machine-readable media, on which instructions are stored. When executed by one or more processors, an electronic device is enabled to execute one or more of the methods as described in the above embodiments. In the embodiment of the present application, the electronic device includes various types of devices such as terminal devices, servers (clusters), etc.

[0190] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other.

[0191] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable opinion analysis terminal devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable opinion analysis terminal devices generate a device for implementing the functions specified in one Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0192] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable public opinion analysis terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 specified in one block or multiple blocks.

[0193] These computer program instructions can also be loaded onto a computer or other programmable public opinion analysis terminal device, such that a series of operation steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 specified in one block or multiple blocks.

[0194] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the embodiments of the present application.

[0195] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or terminal device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or terminal device comprising the said element.

[0196] The above has introduced in detail a public opinion analysis method and apparatus, an electronic device and a storage medium provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A method for analyzing public opinion, characterized in that: The method comprises: Obtain original public opinion data; Constructing an initial public opinion graph according to the original public opinion data; the initial public opinion graph includes connection relationships between nodes, the nodes correspond to named entities in the original public opinion data, and the connection relationships between the nodes correspond to entity relationships between the named entities; Generate a multi-layer public opinion relationship network according to the initial public opinion graph; Generate a public opinion hotspot map according to the initial public opinion map; The public opinion analysis result is obtained according to the multi-layer public opinion relationship network and the public opinion hotspot map.

2. The method according to claim 1, characterized in that Generating a multi-layer public opinion relationship network according to the initial public opinion graph includes: Calculate the node importance score corresponding to each node in the initial public opinion graph; Determine the key nodes in the initial public opinion graph according to the node importance scores; Obtaining the connection relationship corresponding to the key node; According to the connection relationship between the nodes in the initial public opinion graph, the initial public opinion graph is divided into a plurality of topic community sub-networks; Traversing the nodes in the initial public opinion graph and the connection relationships between the nodes, identifying connected components and triangular structures in the initial public opinion graph; the triangular structure corresponds to a structure in which three nodes in the initial public opinion graph are connected in pairs; The multi-layer public opinion relationship network is generated according to the key nodes, the connection relationships corresponding to the key nodes, the topic community sub-network, the connected components and the triangular structure.

3. The method according to claim 2, characterized in that The generating a multi-layer public opinion relationship network according to the key nodes, the connection relationships corresponding to the key nodes, the topic community sub-network, the connected components and the triangular structure includes: Extracting the key nodes in the initial public opinion graph and the connection relationships corresponding to the key nodes to generate a top-level public opinion relationship network; the top-level public opinion relationship network includes core nodes, edge nodes and connection relationships between nodes, the core nodes correspond to the key nodes, and the edge nodes are directly or indirectly connected to the core nodes through the connection relationships between the nodes; Extracting the topic community subnetwork in the initial public opinion graph to generate a middle-level public opinion relationship network; Extracting the connected components and the triangular structure in the initial public opinion graph to generate an underlying public opinion relationship network; The top-level public opinion relationship network, the middle-level public opinion relationship network and the bottom-level public opinion relationship network constitute the multi-layer public opinion relationship network.

4. The method according to claim 1, characterized in that: Generating a public opinion hotspot map according to the initial public opinion map includes: Extracting a connected subgraph from the initial public opinion graph; Calculating a comprehensive score corresponding to the connected subgraph; determining potential hot spots based on the composite score; The public opinion heat map is generated based on the potential hot spots.

5. The method according to claim 4, characterized in that The initial public opinion graph includes node attributes and connection relationship attributes, and the calculation of the comprehensive score corresponding to the connected subgraph includes: Determine the sentiment tendency of the connection relationship between the nodes in the connected subgraph; the sentiment tendency is determined by performing sentiment analysis on the original public opinion data corresponding to the connection relationship between the nodes; Calculating the node importance score in the connected subgraph; the node importance score is determined according to the node degree, node weight and the sentiment tendency of the node corresponding connection relationship; the node attribute includes the node degree and the node weight, and the connection relationship attribute includes the sentiment tendency of the node corresponding connection relationship; The comprehensive score of the connected subgraph is calculated according to the node importance score and the sentiment tendency.

6. The method according to claim 4, characterized in that The generating the public opinion hotspot map based on the potential hotspots includes: Monitoring the potential hotspot to obtain potential hotspot data of the potential hotspot; the potential hotspot data at least includes duration, heat change value and hotspot propagation path; Generate a public opinion hotspot distribution map according to the potential hotspots and the hotspot propagation paths; A hotspot evolution curve graph is generated according to the duration and the heat change value of the potential hotspot; the public opinion hotspot map includes the public opinion hotspot distribution map and the hotspot evolution curve graph.

7. The method according to claim 1, characterized in that The obtaining of the public opinion analysis result according to the multi-layer public opinion relationship network and the public opinion hotspot map includes: Generate a comprehensive public opinion communication analysis report based on the multi-layer public opinion relationship network; Generate a hotspot analysis report based on the public opinion hotspot map; The public opinion analysis result is obtained based on the comprehensive public opinion communication analysis report and the hot spot analysis report.

8. A public opinion analysis device, characterized in that: The device comprises: Acquisition module, used to obtain original public opinion data; An initial public opinion graph construction module is used to construct an initial public opinion graph according to the original public opinion data; the initial public opinion graph includes connection relationships between nodes, the nodes correspond to named entities in the original public opinion data, and the connection relationships between the nodes correspond to entity relationships between the named entities; A public opinion relationship network construction module is used to generate a multi-layer public opinion relationship network according to the initial public opinion graph; A public opinion heat map construction module, used to generate a public opinion heat map according to the initial public opinion map; The analysis module is used to obtain the public opinion analysis results based on the multi-layer public opinion relationship network and the public opinion hotspot map.

9. An electronic device, characterized in that: include: processor; and A memory having executable codes stored thereon, which, when executed, enables the processor to execute the public opinion analysis method as described in one or more of claims 1-7.

10. A computer-readable storage medium having executable code stored thereon, which, when executed, enables a processor to execute the public opinion analysis method as described in one or more of claims 1-7.

Citation Information

Cited By

  • Public opinion tracing method based on multi-source heterogeneous data

    CN121302289A

  • Opinion traceability method based on multi-source heterogeneous data

    CN121302289B