GPT-based directional label public opinion analysis method
By using a GPT-based targeted tag sentiment analysis method, the first sentiment data under the data tags of target users is obtained, generating high-quality sentiment topics and sentiment categories. This solves the problem of inaccurate results in sentiment analysis and achieves more accurate sentiment analysis and a more comprehensive assessment of user influence.
Patent Information
- Application Number
- CN202411633115.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-14
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2044-11-14
AI Technical Summary
Existing technologies have inaccurate results in public opinion analysis, especially in identifying and extracting keywords or key phrases, which cannot selectively filter information, leading to the omission of important information and interference from irrelevant data.
We employ a GPT-based targeted tagging-based public opinion analysis method. By acquiring the first public opinion data under the data tags of target users, we use GPT to generate high-quality public opinion topics and sentiment categories, and combine them with user information to generate public opinion analysis results, thereby reducing interference from irrelevant data.
It improves the accuracy and efficiency of public opinion analysis, enabling more precise targeting of specific topics or areas of interest to users, capturing complex and implicit emotional changes, and generating more accurate public opinion analysis results.
Smart Images

Figure CN119740137B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and more specifically, to a targeted tagging-based public opinion analysis method based on GPT. Background Technology
[0002] Public opinion analysis refers to the collection and analysis of publicly available data from social networks, news platforms, forums, etc., to identify public attitudes and sentiments towards a particular topic, field, or event. However, when conducting public opinion analysis, related technologies generally rely on keyword filtering techniques to identify and extract keywords or key phrases from the collected public opinion data, and then generate public opinion reports or analysis results based on the extracted keywords or key phrases. This approach can lead to inaccurate public opinion analysis results. Summary of the Invention
[0003] The purpose of this disclosure is to provide a GPT-based targeted tagging sentiment analysis method to solve the aforementioned technical problems.
[0004] To achieve the above objectives, the first aspect of this disclosure provides a targeted tagging-based public opinion analysis method based on GPT, the method comprising:
[0005] In response to the data tags configured for the target user, the system obtains the first public opinion data of the target user for the data tags, wherein the target user is a domain expert in the target field, the data tags are data tags used to identify data categories in the target field, and the target user is determined based on the user influence value and domain influence value of the user in the target field, wherein the user influence value is used to characterize the degree of influence of the user on other users, and the domain influence value is used to characterize the degree of influence of the user on the target field;
[0006] Based on GPT and the first public opinion data, the public opinion theme and sentiment category for the first public opinion data are obtained, wherein GPT is used to output the corresponding public opinion theme and sentiment category according to the input first public opinion data;
[0007] Based on the public opinion topic, the sentiment category, and the user information of the target user, a public opinion analysis result is generated.
[0008] Optionally, the target users include multiple groups, and correspondingly, the public opinion topics and sentiment categories include multiple groups. The step of generating public opinion analysis results based on the public opinion topics, sentiment categories, and user information of the target users includes:
[0009] Using the public opinion theme as a classification dimension, the sentiment category corresponding to the first public opinion data is divided into the public opinion theme corresponding to the first public opinion data, thus obtaining the sentiment category under each public opinion theme;
[0010] For each of the aforementioned public opinion topics, the sentiment categories under the aforementioned public opinion topics are aggregated to obtain the sentiment distribution information of the aforementioned public opinion topics;
[0011] Based on the sentiment distribution information of each of the aforementioned public opinion topics, the user information of multiple public opinion topics and multiple target users, a public opinion analysis result is generated.
[0012] Optionally, the target user is obtained in the following way:
[0013] The second public opinion data of users on the target network platform targeting preset data tags is obtained through the application programming interface of the target network platform, wherein the preset data tags are data tags in the target domain;
[0014] Based on the second public opinion data, determine the public opinion impact value of the user corresponding to the second public opinion data in the target field;
[0015] The target user is determined based on the public opinion influence value of the user corresponding to the second public opinion data in the target field.
[0016] Optionally, the second public opinion data includes content data published by the user on the target network platform, user data of the user on the target network platform, and interaction data of the user on the target network platform. Determining the public opinion influence value of the user corresponding to the second public opinion data in the target domain based on the second public opinion data includes:
[0017] Based on the user data, a first influence value of the user on other users in the target network platform is determined, and the first influence value is determined as the user's influence value.
[0018] Based on the content data and the interaction data, a second influence value of the content published by the user on the target network platform in the target domain is determined, and the second influence value is determined as the domain influence value.
[0019] Based on the user influence value and the domain influence value, the public opinion influence value of the user corresponding to the second public opinion data in the target domain is determined.
[0020] Optionally, the user data includes first user data, second user data, and third user data. The first user data is used to indicate the credibility of the content posted by the user on the target network platform. The second user data includes the user's fan information on the target network platform. The third user data includes the user's follower information on the target network platform. Determining the first influence value of the user on other users on the target network platform based on the user data includes:
[0021] Based on the preset correspondence between the trust score and the credibility, and the first user data, a third influence value is determined;
[0022] Based on the fan information and the follower information, the social network of the user in the target network platform is constructed, and the degree centrality, tightness centrality, betweenness centrality and eigenvector centrality of the user in the social network are determined according to the social network. A fourth influence value is determined according to the degree centrality, tightness centrality, betweenness centrality and eigenvector centrality.
[0023] Using the target constant as the base and the number of followers in the follower information as the exponent, the target value is obtained, and the target value is normalized to obtain the fifth influence value.
[0024] Based on the third influence value, the fourth influence value, and the fifth influence value, the first influence value of the user on other users in the target network platform is determined.
[0025] Optionally, the interaction data includes first interaction data, second interaction data, third interaction data, and fourth interaction data. The first interaction data includes the number of likes, the second interaction data includes the number of comments, the third interaction data includes the number of shares, and the fourth interaction data includes the number of replies. Determining the second influence value of the content published by the user on the target network platform in the target domain based on the content data and the interaction data includes:
[0026] The system acquires trending topics in the target domain and target historical content data published by the user on the target network platform. The trending topics are those with a network popularity value greater than a preset network popularity value, and the target historical content data are the content data published by the user within a preset time interval.
[0027] Based on the semantic recognition model, the historical content data, and the trending topics, the semantic similarity between the historical content data and the trending topics is determined.
[0028] When the semantic similarity is greater than a preset semantic similarity threshold, the sum of the number of likes, comments, reposts, and replies is calculated, and the ratio of the sum to the duration corresponding to the preset time interval is calculated to obtain the content interaction value.
[0029] Based on the semantic similarity and the content interaction value, a second influence value of the content published by the user on the target network platform in the target domain is determined.
[0030] A second aspect of this disclosure provides a GPT-based targeted tagging sentiment analysis device, the device comprising:
[0031] The first acquisition module is used to acquire first public opinion data of the target user in response to the data tag configured for the target user, wherein the target user is a domain expert in the target field, the data tag is a data tag used to identify data category in the target field, and the target user is determined based on the user influence value and domain influence value of the user in the target field, wherein the user influence value is used to characterize the degree of influence of the user on other users, and the domain influence value is used to characterize the degree of influence of the user on the target field;
[0032] The first processing module is used to obtain the public opinion theme and sentiment category for the first public opinion data based on GPT and the first public opinion data, wherein GPT is used to output the corresponding public opinion theme and sentiment category according to the input first public opinion data;
[0033] The second processing module is used to generate public opinion analysis results based on the public opinion topic, the sentiment category, and the user information of the target user.
[0034] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in any of the first aspects.
[0035] Fourthly, this disclosure provides an electronic device, comprising:
[0036] A storage device on which computer programs are stored;
[0037] A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of the first aspects.
[0038] Fifthly, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in any one of the first aspects.
[0039] The above technical solution allows for the configuration of data tags for target users, enabling the targeted acquisition of primary public opinion data for those users under those tags. This data can then be used to generate corresponding public opinion analysis results. Therefore, events within a target domain can be monitored specifically based on user-defined data tags, allowing for targeted acquisition and analysis of primary public opinion data. This reduces the processing volume of primary public opinion data and improves analysis efficiency. Furthermore, because the primary public opinion data is tagged, the analysis results generated from it are free of irrelevant data, improving accuracy and preventing interference from irrelevant information.
[0040] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0041] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the following detailed description to explain the present disclosure, but do not constitute a limitation thereof. In the drawings:
[0042] Figure 1 This is a flowchart illustrating a GPT-based targeted tagging sentiment analysis method according to an exemplary embodiment of this disclosure;
[0043] Figure 2 This is a flowchart illustrating another GPT-based targeted tag sentiment analysis method according to an exemplary embodiment of this disclosure;
[0044] Figure 3 This is a schematic diagram illustrating a portion of the results of public opinion analysis in a visual manner, according to an exemplary embodiment of this disclosure;
[0045] Figure 4 This is a structural block diagram of a GPT-based targeted tagging public opinion analysis device according to an exemplary embodiment of the present disclosure;
[0046] Figure 5 This is a block diagram illustrating an electronic device according to an exemplary embodiment;
[0047] Figure 6 This is a block diagram illustrating another electronic device according to an exemplary embodiment. Detailed Implementation
[0048] The specific embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit this disclosure.
[0049] As mentioned in the background section, related technologies for public opinion analysis generally rely on keyword filtering techniques to identify and extract keywords or key phrases from collected public opinion data. These extracted keywords or key phrases are then used to generate public opinion reports or analysis results. However, because this method is a general approach, it cannot differentiate between the specific needs of different fields or selectively filter information. On the one hand, this results may contain a large amount of data irrelevant to the topics users are interested in, making it difficult for users to quickly find valuable information. On the other hand, this general method may fail to effectively identify public opinion data specific to certain fields or tags, leading to the omission of important information. Consequently, the analysis results may not accurately reflect the public's attitudes and / or emotions towards a particular topic, field, or event.
[0050] In view of this, the present disclosure provides a GPT-based targeted tag sentiment analysis method to solve the above-mentioned technical problems.
[0051] The embodiments of this disclosure will be further explained below with reference to the accompanying drawings.
[0052] Figure 1 and Figure 2 This is a flowchart illustrating a GPT-based targeted tagging sentiment analysis method according to an exemplary embodiment of this disclosure, with reference to... Figure 1 and Figure 2 This public opinion analysis method may include the following steps:
[0053] S101: In response to the data tag configured for the target user, obtain the first public opinion data of the target user for the data tag, wherein the target user is a domain expert in the target field, the data tag is a data tag used to identify data category in the target field, the target user is determined based on the user influence value and domain influence value of the user in the target field, the user influence value is used to characterize the degree of influence of the user on other users, and the domain influence value is used to characterize the degree of influence of the user on the target field.
[0054] In this embodiment, there can be one or more target users. When there are multiple target users, the same data tags can be configured for multiple target users to identify areas of consensus or disagreement among them. Alternatively, different data tags can be configured for multiple target users to identify the core interests of each target user.
[0055] In this embodiment, the data tags and the first public opinion data can be set according to actual conditions, and this disclosure embodiment does not impose any restrictions on them. For example, in the field of virtual assets, data tags may include "blockchain regulation" or "cryptocurrency," etc. In the field of image processing, data tags may include "image enhancement," "feature extraction," or "semantic segmentation," etc. For example, the first public opinion data may include content posted by users on the target network platform, comments posted by users on the target network platform, likes posted by users on the target network platform, and user information on the target network platform, etc.
[0056] In this embodiment, obtaining the first public opinion data of the target user regarding the data tag can be achieved by automatically collecting all content related to the data tag posted by users on the social network through a social network API (such as the API of an information publishing platform), and then filtering out the content posted by the target user from the collected content based on the target user's user identifier, thereby obtaining the first public opinion data. Of course, the first public opinion data of the target user regarding the data tag can also be obtained through other methods, and this embodiment does not impose any restrictions on this.
[0057] S102: Based on GPT and the first public opinion data, obtain the public opinion theme and sentiment category for the first public opinion data. GPT is used to output the corresponding public opinion theme and sentiment category according to the input first public opinion data.
[0058] It should be understood that related technologies generally generate public opinion topics based on probabilistic models (such as Latent Dirichlet Allocation (LDA) and Non-negative Matrix Factorization (NMF)) and models that consider temporal dynamics and topic relevance (such as Dynamic Topic Model (DTM) and Correlated Topic Model (CTM)). However, these models often generate topics of poor quality when processing short texts or texts with complex contexts, thus affecting the accuracy of the analysis results. GPT (Generative Pre-trained Transformer) can better understand textual context. Therefore, this embodiment inputs the first public opinion data into GPT to utilize TopicGPT technology for high-quality topic generation, thereby improving the accuracy of public opinion analysis results.
[0059] Furthermore, in related technologies, methods generally rely on dictionary-based approaches (i.e., using a predefined sentiment dictionary to assess the sentiment tendency of text), machine learning-based approaches (i.e., predicting sentiment by training classifiers such as support vector machines or random forests), or deep learning-based approaches (i.e., using recurrent neural networks, convolutional neural networks, or Transformer models to process sequential data and capture long-term dependencies in the text to obtain the sentiment tendency of the text) to determine the sentiment of input text. While these methods have achieved some success in handling explicit sentiment expressions, they often struggle to accurately identify complex and implicit sentiment expressions, such as irony, humor, or metaphor. Therefore, this embodiment inputs the first public opinion data into GPT to utilize GPT's THOR framework to capture complex sentiment changes in the first public opinion data and identify implicit sentiments within it, thereby improving the accuracy of sentiment categories and further enhancing the accuracy of public opinion analysis results.
[0060] It is worth noting that the method of generating public opinion topics and sentiment categories based on GPT and the first public opinion data in this embodiment is only illustrative. In possible embodiments, GPT can include two GPTs: one GPT is used to output the corresponding public opinion topic based on the input first public opinion data, and the other GPT is used to output the corresponding sentiment category based on the input first public opinion data. In other possible embodiments, the first public opinion data can also be input into BERT (Bidirectional Encoder Representations from Transformers), RoBERTa (Robustly optimized BERT approach), T5 (Text-to-Text Transfer Transformer), or XLNet (eXtreme Language Model Pre-training), which have excellent text understanding and semantic analysis capabilities, to obtain the public opinion topic and sentiment category corresponding to the first public opinion data.
[0061] S103: Generate public opinion analysis results based on the public opinion theme, sentiment category, and user information of the target users.
[0062] For example, the public opinion topic, sentiment category, and target user information can be input into a large language model, such as a T5 model, to generate corresponding public opinion analysis results. The public opinion analysis results can be determined based on actual circumstances, and this embodiment of the disclosure does not impose any limitations on this. For example, the public opinion analysis results can be used to indicate the public opinion topics that target users are interested in within a target domain, as well as the sentiment distribution of content under those topics.
[0063] After obtaining the public opinion analysis results, to facilitate users' intuitive understanding of the current public opinion situation, the results can be displayed in various visualization methods such as charts, word clouds, and bar charts, for example... Figure 3 As shown. Furthermore, to ensure users can understand the current public opinion situation in a timely and accurate manner, the public opinion analysis results can be updated when the first public opinion data under the data tag changes.
[0064] The above technical solution allows for the configuration of data tags for target users, enabling the targeted acquisition of primary public opinion data for those users under those tags. This data can then be used to generate corresponding public opinion analysis results. Therefore, events within a target domain can be monitored specifically based on user-defined data tags, allowing for targeted acquisition and analysis of primary public opinion data. This reduces the processing volume of primary public opinion data and improves analysis efficiency. Furthermore, because the primary public opinion data is tagged, the analysis results generated from it are free of irrelevant data, improving accuracy and preventing interference from irrelevant information.
[0065] To facilitate understanding of the GPT-based targeted tag sentiment analysis method provided in this disclosure, the possible implementation methods of this solution are described below.
[0066] In possible approaches, the target users can include multiple users, and correspondingly, the public opinion topics and sentiment categories can also include multiple categories. Accordingly, based on the public opinion topics, sentiment categories, and user information of the target users, the generated public opinion analysis results can include:
[0067] Using public opinion topics as a classification dimension, the sentiment category corresponding to the first public opinion data is assigned to the corresponding public opinion topic, resulting in the sentiment category under each public opinion topic. For each public opinion topic, the sentiment categories under the topic are aggregated to obtain the sentiment distribution information of the public opinion topic. Based on the sentiment distribution information of each public opinion topic, multiple public opinion topics, and user information of multiple target users, public opinion analysis results are generated.
[0068] For example, public opinion topics include Topic 1, Topic 2, and Topic 3, and sentiment categories include Sentiment 1, Sentiment 2, Sentiment 3, Sentiment 4, and Sentiment 5. Specifically, sentiment categories for Topic 1 include Sentiment 1, Sentiment 4, and Sentiment 5; sentiment categories for Topic 2 include Sentiment 2 and Sentiment 4; and sentiment categories for Topic 3 include Sentiment 2, Sentiment 3, and Sentiment 5. Therefore, after assigning sentiment categories to their corresponding public opinion topics, by aggregating and statistically analyzing the quantity of each sentiment category within each public opinion topic, we can obtain the sentiment distribution information under each public opinion topic, i.e., the quantity or proportion of each sentiment category. For example, the sentiment distribution information can be represented as:
[0069] Theme 1: Emotion 1 (30%), Emotion 4 (40%), Emotion 5 (30%);
[0070] Theme 2: Emotion 2 (60%), Emotion 4 (40%);
[0071] Theme 3: Emotion 2 (20%), Emotion 3 (50%), Emotion 5 (30%).
[0072] Then, by combining the sentiment distribution information of multiple public opinion topics with the user information of target users, a comprehensive public opinion analysis result can be generated. For example, the generated public opinion analysis result may include: the overall sentiment tendency of each public opinion topic, the distribution of each sentiment category in different topics, and / or predictions of public opinion hotspots and trends, etc.
[0073] In this embodiment, by performing thematic classification and sentiment analysis on public opinion data, it is possible to more accurately pinpoint specific topics or areas of interest to users, thereby improving the results of public opinion analysis. Furthermore, by aggregating sentiment categories and generating sentiment distribution information, it is helpful to gain a deeper understanding of the public's emotional attitudes towards specific topics, enriching the content of the public opinion analysis results.
[0074] It should be understood that the content in the above public opinion analysis results is for illustrative purposes only and does not constitute a limitation on the proposed solution. In other possible approaches, implicit content and / or viewpoints regarding the first public opinion data can be obtained based on GPT and the first public opinion data, thereby generating public opinion analysis results. Furthermore, the implicit content and / or viewpoints can be combined to generate more accurate public opinion analysis results.
[0075] Among the possible methods, target users can be obtained in the following ways:
[0076] By using the application programming interface of the target network platform, second public opinion data of users on the target network platform targeting preset data tags are obtained, wherein the preset data tags are data tags in the target domain; based on the second public opinion data, the public opinion influence value of the user corresponding to the second public opinion data in the target domain is determined; based on the public opinion influence value of the user corresponding to the second public opinion data in the target domain, the target user is determined.
[0077] In this embodiment, the target network platform can be a social networking platform, a news platform, a video sharing platform, or any other network platform; this disclosure does not impose any limitations on this. There can be one or more target network platforms. When there is only one target network platform, the second public opinion data corresponding to each user on the target network platform can be obtained through the platform's API. Then, the second public opinion data corresponding to each user is analyzed to obtain the public opinion impact value for each user. When there are multiple target network platforms, users with accounts on multiple target network platforms can be first selected. Then, for each selected user, the second public opinion data for that user on each target network platform can be obtained sequentially through the platform's application programming interface (API). Finally, the second public opinion data for that user on each target network platform is analyzed to obtain the user's public opinion impact value.
[0078] In this embodiment, the target user can be determined based on the public opinion influence value of the user corresponding to the second public opinion data in the target domain. Target users can be selected based on the magnitude of each user's public opinion influence value. For example, the user with the highest public opinion influence value can be selected as the target user, or the top N users with the highest public opinion influence values, where N is a positive integer. Alternatively, a public opinion influence threshold can be pre-set, and users with public opinion influence values greater than or equal to the threshold can be selected as target users.
[0079] In this embodiment, the second public opinion data can be determined according to the actual situation, and this disclosure does not impose any restrictions on it. For example, the second public opinion data may include content posted by users on the target network platform, comments posted by users on the target network platform, likes posted by users on the target network platform, and user information on the target network platform, etc.
[0080] In some possible ways, the second public opinion data may include content data posted by users on the target online platform, user data on the target online platform, and user interaction data on the target online platform. Accordingly, based on the second public opinion data, determining the public opinion impact value of the user corresponding to the second public opinion data in the target domain may include:
[0081] Based on user data, determine the first influence value of the user on other users on the target network platform, and define the first influence value as the user influence value; based on content data and interaction data, determine the second influence value of the content published by the user on the target network platform in the target domain, and define the second influence value as the domain influence value; based on the user influence value and the domain influence value, determine the public opinion influence value of the user corresponding to the second public opinion data in the target domain.
[0082] In this embodiment, content data, user data, and interaction data can all be determined according to actual circumstances, and this disclosure does not impose any limitations on them. For example, content data may include content published by a user on the target network platform, the time, frequency, and quantity of such content. User data may include the number of followers, user status (e.g., unverified user, verified user, active user, official user), and number of followers. Interaction data may include likes, reposts, shares, or access data.
[0083] In this embodiment, the first influence value of a user on other users in the target network platform can be determined based on three factors: user credibility, network centrality, and number of followers. User credibility can be determined based on user status; for example, the credibility of an unverified user is lower than that of an active user, and the credibility of an active user is lower than that of an official user. Therefore, user credibility can be determined based on user status; for example, the credibility of an unverified user can be set to 0.8, the credibility of an active user to 0.9, and the credibility of an official user to 1.0. Thus, after obtaining user data, user credibility can be obtained based on user status.
[0084] Network centrality is an indicator of the importance / influence of nodes in a network, and it can include four aspects: degree centrality, compactness centrality, betweenness centrality, and eigenvector centrality. Degree centrality refers to the number of connections a node has in the network, i.e., the node's degree. Higher degree centrality indicates more connections a node has in the network, and thus a greater influence. Compactness centrality is the distance between a node and other nodes, i.e., the reciprocal of the shortest path length to other nodes. Higher compactness centrality indicates stronger connections between nodes, and faster information transmission. Betweenness centrality refers to the degree of betweenness of a node in the network, i.e., the number of times a node appears on the shortest path between other nodes. High betweenness centrality indicates that the node plays an important role in information transmission and connections within the network. Eigenvector centrality indicates the importance of a node in the network, which is related to the importance of its neighboring nodes. Higher eigenvector centrality indicates that the nodes surrounding the node are more important, and the node's influence is greater. Therefore, network centrality can be derived by constructing a social network using user follower and fan information, and then obtaining degree centrality, tightness centrality, betweenness centrality, and eigenvector centrality through analysis of the social network. For example, a weight can be assigned to each of the degree centrality, tightness centrality, betweenness centrality, and eigenvector centrality. Thus, after obtaining the degree centrality, tightness centrality, betweenness centrality, and eigenvector centrality, network centrality can be obtained based on these values and their corresponding weights.
[0085] The number of followers is the most direct way to reflect a user's influence. However, since the number of followers varies greatly, a portion of a user's influence can be obtained by taking the logarithm of the number of followers and normalizing it.
[0086] In other words, in possible ways, user data may include first user data, second user data, and third user data. The first user data indicates the credibility of content posted by the user on the target online platform. The second user data includes the user's follower information on the target online platform. The third user data includes the user's follower information on the target online platform. Accordingly, based on the user data, determining the user's first influence value on other users on the target online platform may include:
[0087] Based on the pre-defined correspondence between trust scores and credibility, and the first user data, a third influence value is determined. Based on fan and follower information, a social network for the user on the target network platform is constructed. According to this social network, the user's degree centrality, tightness centrality, betweenness centrality, and eigenvector centrality within the social network are determined. Based on these factors, a fourth influence value is determined. Using a target constant as the base and the number of followers in the follower information as the exponent, a target value is obtained. This target value is then normalized to obtain a fifth influence value. Based on the third, fourth, and fifth influence values, the user's first influence value on other users on the target network platform is determined.
[0088] After obtaining the third influence value (i.e., user credibility), the fourth influence value (i.e., network centrality), and the fifth influence value (i.e., number of followers) through the above methods, the first influence value of the user on other users in the target network platform is obtained by substituting the third, fourth, and fifth influence values into the following formula:
[0089]
[0090] in, Indicates user The first impact value on other users in the target network platform. This indicates the third influence value. This indicates the fourth influence value. This represents the fifth influence value. , and This represents a coefficient, the magnitude of which can be determined according to actual circumstances, and this disclosure does not impose any limitations on this. For example, .
[0091] It should be understood that related technologies generally assess user influence based on simple metrics such as the number of followers, the following relationships between users, and interactions (e.g., likes, comments, retweets). For example, methods like FansRank, PageRank, and TwitterRank are used to evaluate user influence. FansRank assesses influence based on the number of followers, assuming more followers equal greater influence, but it ignores interaction quality and the existence of fake followers, easily overestimating a user's influence. PageRank assesses a user's importance in the network by analyzing the strength of connections between users, focusing on the structure of the social network but neglecting content quality and domain relevance. TwitterRank, building on PageRank, introduces topic relevance to assess a user's influence on specific topics. It is suitable for identifying key users in specific domains but lacks sufficient analysis of user interaction quality. Therefore, while the above algorithms each have their own emphasis, they all simply use one or two metrics to measure user influence, lacking comprehensive consideration, resulting in an incomplete and inaccurate assessment of user influence. When evaluating user influence based on the method in this embodiment, not only is the user's social network structure considered, but also the interaction with the user's published content and its domain relevance. This allows for a more comprehensive evaluation from multiple dimensions, thereby overcoming the limitations of single-dimensional analysis in related technologies and improving the accuracy of user influence evaluation.
[0092] In possible embodiments, interaction data may include first interaction data, second interaction data, third interaction data, and fourth interaction data. First interaction data may include likes, second interaction data may include comments, third interaction data may include shares, and fourth interaction data may include replies. Accordingly, based on content data and interaction data, determining the second impact value of content posted by a user on a target network platform within a target domain may include:
[0093] The system acquires trending topics in the target domain and historical content data posted by users on the target online platform. Trending topics are those with a network popularity value greater than a preset network popularity value, and historical content data refers to content data posted by users within a preset time interval. Based on a semantic recognition model, historical content data, and trending topics, the system determines the semantic similarity between historical content data and trending topics. When the semantic similarity is greater than a preset semantic similarity threshold, the system calculates the sum of likes, comments, shares, and replies, and calculates the ratio of this sum to the corresponding duration of the preset time interval to obtain the content interaction value. Based on the semantic similarity and content interaction value, the system determines the second influence value of the content posted by users on the target online platform in the target domain.
[0094] In this embodiment, the preset network popularity value, preset time interval, semantic recognition model, and semantic similarity threshold can all be determined according to the actual situation, and this disclosure embodiment does not impose any restrictions on them.
[0095] For example, the preset network popularity value can be set to 10,000, the preset time interval can be set to 15 days, the semantic recognition model can be set to the SBERT (Sentence-BERT) model, and the semantic similarity threshold can be set to 0.3. Therefore, when the network popularity value of a topic in the target domain is greater than 10,000, it can be considered a trending topic. Then, all content data posted by the target user within the past 15 days is obtained. For each piece of content data, the content data and the trending topic are input into the SBERT model to obtain the semantic similarity between the content data and the trending topic. If the semantic similarity is less than 0.3, it means that the content data is not related to the target domain. The semantic similarity between the next piece of content data and the trending network is then calculated. If the semantic similarity is greater than or equal to 0.3, it means that the content data is related to the target domain. The number of likes, comments, shares, and replies for the content data is then obtained, and these numbers are substituted into the following formula to obtain the content interaction value:
[0096]
[0097] in, Indicates the first The content interaction value of each historical content data item. Indicates the first The number of likes on each piece of historical content data. Indicates the first The number of comments on each piece of historical content data. Indicates the first The number of times historical content was forwarded. Indicates the first The number of replies to historical content data. Indicates the preset time interval.
[0098] After obtaining the content interaction value and semantic similarity of each historical content data point, the second influence value is obtained by substituting these values into the following formula:
[0099]
[0100] in, Indicates user Second impact value, Indicates the first Semantic similarity between historical content data and trending networks. Indicates user Historical content data.
[0101] After obtaining the first and second impact values, the user's public opinion impact value in the target domain is obtained by substituting the first and second impact values into the following formula:
[0102]
[0103] in, Indicates user Public opinion impact value and This represents a coefficient, the magnitude of which can be determined according to actual circumstances, and this disclosure does not impose any limitations on this. For example, .
[0104] In summary, through the above technical solutions, this approach can, on the one hand, conduct public opinion analysis based on user-defined targeted tags, enabling targeted monitoring of specific areas or events and improving the accuracy of data processing and analysis. On the other hand, it can combine analysis with dimensions such as user credibility, network centrality, number of followers, interactivity of user-posted content, and its relevance to the relevant domain, thereby achieving a more comprehensive and accurate assessment of a user's public opinion influence under specific tags, further enhancing the accuracy of public opinion analysis results. Furthermore, this solution also utilizes GPT's THOR framework for sentiment analysis, enabling the capture of more subtle sentiment changes and further improving the accuracy of public opinion analysis results.
[0105] Based on the same concept, this disclosure also provides a GPT-based targeted tagging sentiment analysis device, such as... Figure 4 As shown, the public opinion analysis device 400 may include:
[0106] The first acquisition module 401 is used to acquire first public opinion data of the target user in response to the data tag configured for the target user, wherein the target user is a domain expert in the target field, the data tag is a data tag used to identify data category in the target field, and the target user is determined based on the user influence value and domain influence value of the user in the target field, wherein the user influence value is used to characterize the degree of influence of the user on other users, and the domain influence value is used to characterize the degree of influence of the user on the target field;
[0107] The first processing module 402 is used to obtain the public opinion theme and sentiment category for the first public opinion data based on GPT and the first public opinion data, wherein the GPT is used to output the corresponding public opinion theme and sentiment category according to the input first public opinion data;
[0108] The second processing module 403 is used to generate public opinion analysis results based on the public opinion topic, the sentiment category, and the user information of the target user.
[0109] The aforementioned public opinion analysis device 400 allows for the configuration of data tags for target users, enabling the targeted acquisition of primary public opinion data for those users under those tags. This data can then be used to generate corresponding public opinion analysis results. Therefore, events within a target domain can be monitored specifically based on user-defined data tags, allowing for targeted acquisition and analysis of primary public opinion data. This reduces the processing volume of primary public opinion data and improves analysis efficiency. Furthermore, since the primary public opinion data is tagged, the generated analysis results are free from irrelevant data, improving accuracy and preventing interference from irrelevant information.
[0110] In possible embodiments, the target users may include multiple groups, and correspondingly, the public opinion topics and sentiment categories may include multiple groups, and accordingly, the second processing module 403 may include:
[0111] The first processing unit is used to use the public opinion theme as a classification dimension to classify the sentiment category corresponding to the first public opinion data into the public opinion theme corresponding to the first public opinion data, so as to obtain the sentiment category under each public opinion theme;
[0112] The second processing unit is used to aggregate the sentiment categories under each of the public opinion topics to obtain the sentiment distribution information of the public opinion topics.
[0113] The third processing unit is used to generate public opinion analysis results based on the sentiment distribution information of each of the public opinion topics, the user information of multiple public opinion topics and multiple target users.
[0114] In some possible configurations, the public opinion analysis device 400 may also include:
[0115] The second acquisition module is used to acquire second public opinion data of users in the target network platform for preset data tags through the application programming interface of the target network platform, wherein the preset data tags are data tags in the target domain;
[0116] The first determining module is used to determine the public opinion influence value of the user corresponding to the second public opinion data in the target field based on the second public opinion data;
[0117] The second determining module is used to determine the target user based on the public opinion influence value of the user corresponding to the second public opinion data in the target field.
[0118] In some possible embodiments, the second public opinion data includes content data posted by the user on the target network platform, user data of the user on the target network platform, and interaction data of the user on the target network platform. Accordingly, the first determining module may include:
[0119] The first determining unit is configured to determine, based on the user data, a first influence value of the user on other users in the target network platform, and to determine the first influence value as the user's influence value.
[0120] The second determining unit is configured to determine, based on the content data and the interaction data, a second influence value of the content published by the user on the target network platform in the target domain, and to determine the second influence value as the domain influence value;
[0121] The third determining unit is used to determine the public opinion influence value of the user corresponding to the second public opinion data in the target field based on the user influence value and the field influence value.
[0122] In a possible manner, the user data includes first user data, second user data, and third user data. The first user data is used to indicate the credibility of the content posted by the user on the target network platform. The second user data includes the user's follower information on the target network platform. The third user data includes the user's follower information on the target network platform. Accordingly, the first determining unit may include:
[0123] The first determining subunit is used to determine the third influence value based on the preset correspondence between the trust score and the credibility and the first user data;
[0124] The second determining subunit is used to construct the user's social network on the target network platform based on the fan information and the follower information, and to determine the user's degree centrality, tightness centrality, betweenness centrality and eigenvector centrality in the social network based on the social network, and to determine a fourth influence value based on the degree centrality, tightness centrality, betweenness centrality and eigenvector centrality.
[0125] The third determining subunit is used to obtain the target value by using the target constant as the base and the number of followers in the follower information as the exponent, and to normalize the target value to obtain the fifth influence value.
[0126] The fourth determining subunit is used to determine the first influence value of the user on other users in the target network platform based on the third influence value, the fourth influence value, and the fifth influence value.
[0127] In a possible manner, the interaction data includes first interaction data, second interaction data, third interaction data, and fourth interaction data, wherein the first interaction data includes the number of likes, the second interaction data includes the number of comments, the third interaction data includes the number of shares, and the fourth interaction data includes the number of replies. Accordingly, the second determining unit may include:
[0128] The acquisition subunit is used to acquire hot topics in the target domain and target historical content data published by the user on the target network platform. The hot topics are topics with a network popularity value greater than a preset network popularity value, and the target historical content data are content data published by the user within a preset time interval.
[0129] The fifth determining subunit is used to determine the semantic similarity between the historical content data and the hot topics based on the semantic recognition model, the historical content data, and the hot topics;
[0130] The sixth determining subunit is used to calculate the sum of the number of likes, comments, reposts and replies when the semantic similarity is greater than a preset semantic similarity threshold, and to calculate the ratio of the sum to the duration corresponding to the preset time interval to obtain the content interaction value;
[0131] The seventh determining subunit is used to determine the second influence value of the content published by the user on the target network platform in the target domain based on the semantic similarity and the content interaction value.
[0132] Regarding the public opinion analysis device 400 in the above embodiments, the specific methods by which each module performs its operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0133] Based on the same concept, embodiments of this disclosure also provide a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of any of the above-described GPT-based targeted tag sentiment analysis methods.
[0134] Based on the same concept, this disclosure also provides an electronic device that may include:
[0135] A storage device on which computer programs are stored;
[0136] A processing device for executing a computer program stored in a storage device to implement the steps of any of the above-described GPT-based targeted tag sentiment analysis methods.
[0137] Based on the same concept, this disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described GPT-based targeted tag sentiment analysis methods.
[0138] Figure 5 This is a block diagram illustrating an electronic device 500 according to an exemplary embodiment. For example... Figure 5 As shown, the electronic device 500 may include a processor 501 and a memory 502. The electronic device 500 may also include one or more of a multimedia component 503, an input / output (I / O) interface 504, and a communication component 505.
[0139] The processor 501 controls the overall operation of the electronic device 500 to complete all or part of the steps in the GPT-based targeted tag sentiment analysis method described above. The memory 502 stores various types of data to support the operation of the electronic device 500. This data may include, for example, instructions for any application or method operating on the electronic device 500, and application-related data such as contact data, sent and received messages, images, audio, video, etc. The memory 502 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. Multimedia component 503 may include a screen and an audio component. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in memory 502 or transmitted via communication component 505. The audio component also includes at least one speaker for outputting audio signals. I / O interface 504 provides an interface between processor 501 and other interface modules, such as a keyboard, mouse, buttons, etc. These buttons may be virtual or physical buttons. Communication component 505 is used for wired or wireless communication between the electronic device 500 and other devices. Wireless communication may include Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination of these. Therefore, the corresponding communication component 505 may include a Wi-Fi module, a Bluetooth module, or an NFC module.
[0140] In an exemplary embodiment, the electronic device 500 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the aforementioned GPT-based targeted tag sentiment analysis method.
[0141] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the GPT-based targeted tagging sentiment analysis method described above. For example, the computer-readable storage medium may be the memory 502 including the program instructions, which may be executed by the processor 501 of the electronic device 500 to complete the GPT-based targeted tagging sentiment analysis method described above.
[0142] In another exemplary embodiment, a computer program product is also provided, which includes a computer program executable by a processor, which, when executed by the processor, implements the steps of the GPT-based targeted tag sentiment analysis method described above.
[0143] Figure 6 This is a block diagram illustrating an electronic device 600 according to an exemplary embodiment. For example, the electronic device 600 may be provided as a server. (Refer to...) Figure 6 The electronic device 600 includes a processor 622, which may be one or more, and a memory 632 for storing computer programs executable by the processor 622. The computer programs stored in the memory 632 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processor 622 may be configured to execute the computer program to perform the aforementioned GPT-based targeted tagging sentiment analysis method.
[0144] Additionally, the electronic device 600 may also include a power supply component 626 and a communication component 650. The power supply component 626 can be configured to perform power management of the electronic device 600, and the communication component 650 can be configured to enable communication of the electronic device 600, such as wired or wireless communication. Furthermore, the electronic device 600 may also include an input / output (I / O) interface 658. The electronic device 600 can operate on an operating system stored in memory 632, such as Windows Server™, Mac OSX™, Unix™, Linux™, etc.
[0145] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the GPT-based targeted tagging sentiment analysis method described above. For example, the computer-readable storage medium may be the memory 632 including the program instructions, which may be executed by the processor 622 of the electronic device 600 to complete the GPT-based targeted tagging sentiment analysis method described above.
[0146] In another exemplary embodiment, a computer program product is also provided, which includes a computer program executable by a processor, which, when executed by the processor, implements the steps of the GPT-based targeted tag sentiment analysis method described above.
[0147] The preferred embodiments of this disclosure have been described in detail above with reference to the accompanying drawings. However, this disclosure is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this disclosure, various simple modifications can be made to the technical solutions of this disclosure, and these simple modifications all fall within the protection scope of this disclosure.
[0148] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, this disclosure will not describe the various possible combinations separately.
[0149] Furthermore, various different embodiments of this disclosure can be combined in any way, as long as they do not violate the spirit of this disclosure, they should also be regarded as the content disclosed in this disclosure.
Claims
1. A targeted tagging-based public opinion analysis method based on GPT, characterized in that, The method includes: In response to the data tags configured for the target user, the system obtains the first public opinion data of the target user for the data tags, wherein the target user is a domain expert in the target field, the data tags are data tags used to identify data categories in the target field, and the target user is determined based on the user influence value and domain influence value of the user in the target field, wherein the user influence value is used to characterize the degree of influence of the user on other users, and the domain influence value is used to characterize the degree of influence of the user on the target field; Based on GPT and the first public opinion data, the public opinion theme and sentiment category for the first public opinion data are obtained, wherein GPT is used to output the corresponding public opinion theme and sentiment category according to the input first public opinion data; Based on the public opinion topic, the sentiment category, and the user information of the target user, a public opinion analysis result is generated.
2. The targeted tagging sentiment analysis method based on GPT according to claim 1, characterized in that, The target users include multiple groups, and correspondingly, the public opinion topics and sentiment categories include multiple groups. The step of generating public opinion analysis results based on the public opinion topics, sentiment categories, and user information of the target users includes: Using the public opinion theme as a classification dimension, the sentiment category corresponding to the first public opinion data is divided into the public opinion theme corresponding to the first public opinion data, thus obtaining the sentiment category under each public opinion theme; For each of the aforementioned public opinion topics, the sentiment categories under the aforementioned public opinion topics are aggregated to obtain the sentiment distribution information of the aforementioned public opinion topics; Based on the sentiment distribution information of each of the aforementioned public opinion topics, the user information of multiple public opinion topics and multiple target users, a public opinion analysis result is generated.
3. The GPT-based targeted tagging public opinion analysis method according to claim 1 or 2, characterized in that, The target users are obtained through the following methods: The second public opinion data of users on the target network platform targeting preset data tags is obtained through the application programming interface of the target network platform, wherein the preset data tags are data tags in the target domain; Based on the second public opinion data, determine the public opinion impact value of the user corresponding to the second public opinion data in the target field; The target user is determined based on the public opinion influence value of the user corresponding to the second public opinion data in the target field.
4. The GPT-based targeted tagging sentiment analysis method according to claim 3, characterized in that, The second public opinion data includes content data published by the user on the target network platform, user data of the user on the target network platform, and interaction data of the user on the target network platform. Determining the public opinion influence value of the user corresponding to the second public opinion data in the target domain based on the second public opinion data includes: Based on the user data, a first influence value of the user on other users in the target network platform is determined, and the first influence value is determined as the user's influence value. Based on the content data and the interaction data, a second influence value of the content published by the user on the target network platform in the target domain is determined, and the second influence value is determined as the domain influence value. Based on the user influence value and the domain influence value, the public opinion influence value of the user corresponding to the second public opinion data in the target domain is determined.
5. The GPT-based targeted tagging sentiment analysis method according to claim 4, characterized in that, The user data includes first user data, second user data, and third user data. The first user data is used to indicate the credibility of the content posted by the user on the target network platform. The second user data includes the user's fan information on the target network platform. The third user data includes the user's follower information on the target network platform. Determining the user's first influence value on other users on the target network platform based on the user data includes: Based on the preset correspondence between the trust score and the credibility, and the first user data, a third influence value is determined; Based on the fan information and the follower information, the social network of the user in the target network platform is constructed, and the degree centrality, tightness centrality, betweenness centrality and eigenvector centrality of the user in the social network are determined according to the social network. A fourth influence value is determined according to the degree centrality, tightness centrality, betweenness centrality and eigenvector centrality. Using the target constant as the base and the number of followers in the follower information as the exponent, the target value is obtained, and the target value is normalized to obtain the fifth influence value. Based on the third influence value, the fourth influence value, and the fifth influence value, the first influence value of the user on other users in the target network platform is determined.
6. The GPT-based targeted tagging sentiment analysis method according to claim 4, characterized in that, The interactive data includes first interactive data, second interactive data, third interactive data, and fourth interactive data. The first interactive data includes the number of likes, the second interactive data includes the number of comments, the third interactive data includes the number of shares, and the fourth interactive data includes the number of replies. Determining the second influence value of the content published by the user on the target network platform in the target domain based on the content data and the interactive data includes: The system acquires trending topics in the target domain and target historical content data published by the user on the target network platform. The trending topics are those with a network popularity value greater than a preset network popularity value, and the target historical content data are the content data published by the user within a preset time interval. Based on the semantic recognition model, the historical content data, and the trending topics, the semantic similarity between the historical content data and the trending topics is determined. When the semantic similarity is greater than a preset semantic similarity threshold, the sum of the number of likes, comments, reposts, and replies is calculated, and the ratio of the sum to the duration corresponding to the preset time interval is calculated to obtain the content interaction value. Based on the semantic similarity and the content interaction value, a second influence value of the content published by the user on the target network platform in the target domain is determined.
7. A targeted tagging-based public opinion analysis device based on GPT, characterized in that, The device includes: The first acquisition module is used to acquire first public opinion data of the target user in response to the data tag configured for the target user, wherein the target user is a domain expert in the target field, the data tag is a data tag used to identify data category in the target field, and the target user is determined based on the user influence value and domain influence value of the user in the target field, wherein the user influence value is used to characterize the degree of influence of the user on other users, and the domain influence value is used to characterize the degree of influence of the user on the target field; The first processing module is used to obtain the public opinion theme and sentiment category for the first public opinion data based on GPT and the first public opinion data, wherein GPT is used to output the corresponding public opinion theme and sentiment category according to the input first public opinion data; The second processing module is used to generate public opinion analysis results based on the public opinion topic, the sentiment category, and the user information of the target user.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method described in any one of claims 1-6.
9. An electronic device, characterized in that, include: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method according to any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method described in any one of claims 1-6.
Citation Information
Patent Citations
Public opinion text sentiment analysis method based on semantic dependency relationship fusion features
CN115098634A
Public opinion analysis and arrangement system based on AIGC
CN115982473A