A method for determining opinion leaders

By obtaining the network evaluation set, using the LDA model and similarity matrix to identify opinion leaders, it solves the problem of difficult to quickly and accurately identify opinion leaders in the existing technology, and achieves fast and accurate opinion leaders screening and future prediction.

CN119003766BActive Publication Date: 2025-07-22QINGDAO UNIV OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410989512.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2025-07-22
Estimated Expiration
2044-07-23

AI Technical Summary

Technical Problem

It is difficult to quickly and accurately identify opinion leaders in online public opinion events, especially in a large number of user messages.

Method used

By obtaining the network evaluation set, using the LDA model to determine the optimal number of topics, dividing the comment text into multiple topic sets, combining the text similarity and emotional similarity matrix, calculating the topic leader group and opinion leaders, using PageRank values to determine influence, and predicting future opinion leaders.

Benefits of technology

It has achieved rapid and accurate identification of opinion leaders in online public opinion events, can predict future opinion leaders, and improve the accuracy and prediction capabilities of public opinion monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119003766B_ABST
    Figure CN119003766B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of public opinion monitoring, and particularly relates to a method for determining opinion leaders. This application provides a method for determining opinion leaders, including: obtaining a network evaluation set of a target event, where the network evaluation set includes multiple network review texts; determining the optimal number of topics of the target event according to the network evaluation set; dividing the network evaluation set into multiple topic sets according to the optimal number of topics; and determining the current opinion leaders of each topic set according to each network review text in the topic set. Through the analysis of a large number of message texts, it is possible to quickly and accurately identify opinion leaders.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of public opinion monitoring, and specifically relates to a method for determining opinion leaders. Background Art

[0002] With the assistance of modern information technology, the Internet has become the main front for Internet users to express their views and opinion tendencies. Based on the current situation, the influence of opinion leaders has become increasingly prominent. An opinion leader refers to a person who is active in the interpersonal network and has high representativeness and prestige. The influence of opinion leaders can affect social public opinion and public opinions in a more direct and rapid manner through various communication channels. Therefore, it is very necessary to accurately identify opinion leaders in sudden public opinion events. Since there are a large number of user messages under network events and the views of different users are different, it is difficult for the existing technology to quickly and accurately find opinion leaders among numerous users. Summary of the Invention

[0003] In view of the above technical problems, the invention provides a method for determining opinion leaders. This application provides a method for determining opinion leaders, including: obtaining a network evaluation set of a target event, where the network evaluation set includes multiple network evaluation texts; determining the optimal number of topics of the target event according to the network evaluation set; dividing the network evaluation set into multiple topic sets according to the optimal number of topics; determining the current opinion leaders of each topic set according to each network evaluation text in the topic set. By analyzing a large number of message texts, it is possible to quickly and accurately identify opinion leaders.

[0004] In a first aspect, this application provides a method for determining opinion leaders, including: obtaining a network evaluation set of a target event, where the network evaluation set includes multiple network evaluation texts; determining the optimal number of topics of the target event according to the network evaluation set; dividing the network evaluation set into multiple topic sets according to the optimal number of topics; determining the current opinion leaders of each topic set according to each network evaluation text in the topic set.

[0005] In some embodiments, the determining the number of topics of the target event according to the network evaluation set includes: determining the perplexity index under different numbers of topics according to the network evaluation set; determining the optimal number of topics according to the perplexity index.

[0006] In some embodiments, dividing the network evaluation set into multiple topic sets according to the optimal number of topics includes: using the LDA model to determine the topic probabilities of each network evaluation text in the network evaluation set belonging to each topic according to the optimal number of topics; determining the topic to which each network evaluation text belongs according to the topic probabilities; and dividing the network evaluation set into multiple topic sets according to the topics of each network evaluation text.

[0007] In some embodiments, determining the current opinion leader of each topic set according to the topic set includes: performing the following operations for any topic set: determining a text similarity matrix between each first evaluation text in the topic set according to the first evaluation text in the topic set; determining a topic leader group according to the text similarity matrix; determining an emotional similarity matrix between each second evaluation text according to the second evaluation text corresponding to the topic leader group; and determining the current opinion leader in the topic leader group according to the emotional similarity matrix.

[0008] In some embodiments, determining the topic leader group according to the text similarity matrix includes: obtaining a preset first percentage threshold and a first PageRank value threshold; converting the text similarity matrix into a one-dimensional text similarity vector, where the one-dimensional text similarity vector includes multiple text similarity factors; arranging the text similarity factors from largest to smallest to form a text similarity factor sequence; determining a text similarity factor threshold according to the text similarity factor sequence and the first percentage threshold; determining the text similarity factors greater than or equal to the text similarity factor threshold in the one-dimensional text similarity vector as important text factors; establishing an undirected text graph according to the important text factors; calculating the text PageRank value of each node in the undirected text graph; and determining the topic leader group of the topic according to the text PageRank value of each node and the first PageRank value threshold.

[0009] In some embodiments, determining the current opinion leader in the topic leader group according to the sentiment similarity matrix includes: obtaining a preset second percentage threshold and a second PageRank value threshold; converting the sentiment similarity matrix into a one-dimensional sentiment similarity vector, where the one-dimensional sentiment similarity vector includes a plurality of sentiment similarity factors; arranging the sentiment similarity factors from largest to smallest to form a sentiment similarity factor sequence; determining a sentiment similarity factor threshold according to the sentiment similarity factor sequence and the second percentage threshold; determining the sentiment similarity factors greater than or equal to the sentiment similarity factor threshold in the one-dimensional sentiment similarity vector as important sentiment factors; establishing a sentiment undirected graph according to the important sentiment factors; calculating the sentiment PageRank values of each node of the sentiment undirected graph; and determining the current opinion leader of the topic according to the sentiment PageRank values of each node and the second PageRank value threshold.

[0010] In some embodiments, determining the text similarity matrix between each first review text in the topic set according to the first review text in the topic set includes: calculating the word frequency of each word in each first review text; calculating the inverse document frequency of each word in the topic set; determining the weight of each word in each first review text according to the word frequency and the inverse document frequency; and determining the text similarity between each first review text according to the weights of each word, thereby obtaining the text similarity matrix.

[0011] In some embodiments, determining the sentiment similarity matrix between each second review text according to the second review text corresponding to the topic leader group includes: obtaining a preset sentiment dictionary; determining the sentiment words in each second review text and the sentiment polarity corresponding to the sentiment words according to the sentiment dictionary; determining the sentiment tendency score of each second review text according to the sentiment words and the sentiment polarity; and determining the sentiment similarity between each second review text according to the sentiment tendency score, thereby obtaining the sentiment similarity matrix.

[0012] In some embodiments, the method further includes: obtaining the user activity information of each current opinion leader; determining a leader feature importance sequence according to the user activity information; determining the weights of each feature in the user activity information according to the leader feature importance sequence; and predicting future opinion leaders according to the weights of each feature.

[0013] In some embodiments, predicting future opinion leaders according to the weights of each feature includes: obtaining the target activity information of a target user; determining the leader score of the target user according to the target activity information and the weights of each feature; and predicting whether the target user can become an opinion leader in the future according to the leader score.

[0014] Advantages of the present invention: The present application provides a method for determining opinion leaders, including: obtaining a set of online evaluations of a target event, where the set of online evaluations includes multiple online evaluation texts; determining the optimal number of topics of the target event according to the set of online evaluations; dividing the set of online evaluations into multiple topic sets according to the optimal number of topics; and determining the current opinion leaders of each topic set according to each online evaluation text in the topic set. By analyzing a large number of message texts, opinion leaders can be quickly and accurately identified. Brief Description of the Drawings

[0015] The scope of the present disclosure can be better understood by reading the following detailed description of exemplary embodiments in conjunction with the accompanying drawings. The accompanying drawings included are:

[0016] Figure 1 It is the overall flowchart of a method for determining opinion leaders provided by an embodiment of the present application;

[0017] Figure 2 It is the overall flowchart of a method for predicting opinion leaders provided by an embodiment of the present application. Detailed Description of Specific Embodiments

[0018] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limitations of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0019] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0020] If similar descriptions such as "first / second / third" appear in the application documents, the following explanation is added. In the following description, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0022] With the assistance of modern information technology, the Internet has become the main front for netizens to express their views and opinions. Based on the current situation, the influence of opinion leaders has become increasingly prominent. Opinion leaders refer to those who are active in the interpersonal network and have high representativeness and prestige. The influence of opinion leaders can affect public opinion and public views in a more direct and rapid manner through various communication channels. Therefore, it is very necessary to accurately identify opinion leaders in sudden public opinion events. Since there are a large number of user messages under online events and the views of different users are different, it is difficult for existing technologies to quickly and accurately find opinion leaders among many users.

[0023] In view of the problems existing in the prior art, such as Figure 1 As shown, the present application provides a method for determining an opinion leader. The method is applied to an electronic device, and the electronic device can be a server, a mobile terminal, a computer, a cloud platform, etc. The functions realized by the device data processing in the embodiments of the present application can be implemented by a processor of the electronic device calling program code. Among them, the program code can be stored in a computer storage medium. The method for determining an opinion leader includes:

[0024] Step S1: Obtain a network evaluation set of a target event, where the network evaluation set includes multiple network review texts.

[0025] In the stage when the online public opinion of a certain public event starts to ferment and heat up, use the web crawler software python to collect in real time the microblog posts, comments and other review text contents of the topic to be monitored and the relevant data of the users from the Weibo platform. Use the Chinese word segmentation tool jieba to process the review text, remove the meaningless words in the text according to the stop word list, filter out the irrelevant content, and thus obtain each network review text in the network evaluation dataset.

[0026] Step S2: Determine the optimal number of topics of the target event according to the network evaluation set.

[0027] In some embodiments, step S2 "determine the optimal number of topics of the target event according to the network evaluation set" includes:

[0028] Step S21: Determine the perplexity index under different numbers of topics according to the network evaluation set.

[0029] Step S22: Determine the optimal number of topics according to the perplexity index.

[0030] Step S3: Divide the network evaluation set into multiple topic sets according to the optimal number of topics.

[0031] In some embodiments, step S3, "divide the network evaluation set into multiple topic sets according to the optimal number of topics", includes:

[0032] Step S31: Use the LDA model to determine the topic probabilities of each network evaluation text in the network evaluation set belonging to each topic according to the optimal number of topics.

[0033] Step S32: Determine the topic to which each network evaluation text belongs according to the topic probabilities.

[0034] Step S33: Divide the network evaluation set into multiple topic sets according to the topics of each network evaluation text.

[0035] During the process of public opinion diffusion, various true and false views emerge in an endless stream, and facts and rumors are dazzling. Previous opinion leader identification methods did not conduct classification and selection, and the identification results sometimes could not hit the key point, but were intertwined with positive and negative remarks. Since various viewpoints emerge in an endless stream in various comments, the discussion of a certain event can be divided into multiple viewpoints, that is, the topics mentioned in this application. However, sometimes the topics of some comments are difficult to clarify, or new topics will be generated, but not all topics are worthy of being detected.

[0036] Therefore, in this application, before classifying each evaluation text according to the topic, it is also necessary to determine the optimal number of topics. Perplexity is an index to measure the ability of the model to predict new documents. The lower the perplexity, the better the prediction ability of the model. Therefore, in this application, the optimal number of topics is determined through the perplexity index, where the number of topics corresponding to when the perplexity decreases and tends to be flat or at the inflection point is the optimal number of topics.

[0037] When determining the optimal number of topics, first train the LDA model with different numbers of topics, then use the network evaluation set to calculate the perplexity of the LDA model with different numbers of topics, and establish a relationship table between perplexity and the number of models. Find the point where the perplexity decrease rate slows down or the inflection point in the relationship table. At this time, the number of topics corresponding to such a point is the optimal number of topics. Since the interpretability of each topic also needs to be considered, the largest number of topics cannot be selected as the optimal number of topics, which will lead to the topics being overly subdivided and difficult to interpret each topic.

[0038] After determining the optimal number of topics, the network evaluation texts in the network evaluation set can be classified using the LDA model according to the optimal number of topics, so as to form multiple data sets with different topics, that is, topic sets.

[0039] The LDA model is a method for clustering and extracting the topic information of text data, and gives the article topics in the form of probability distributions, which can quickly identify the topic types from large-scale text data, and thus become an important basis for judging the similarity degree of text topics.

[0040] Therefore, in this application, after clustering topics by LDA and obtaining the topic probabilities of the text, the cosine similarity method is used to calculate the similarity between topics to verify whether the number of topics and the division between each topic are reasonable.

[0041] Step S4: Determine the current opinion leaders of each of the topic sets according to the network review texts in the topic set.

[0042] In the discussion of an event, there will be corresponding topic leaders and opinion leaders for each topic. Therefore, for each topic set, the following steps can be performed to determine the current opinion leaders corresponding to each topic.

[0043] In the past, the spread of public opinion was usually studied from the perspective of social networks, with the interaction behavior between users as a reference. However, the mutual influence between comments is also important in the process of information dissemination and opinion formation. Therefore, this application realizes the identification of opinion leaders through the calculation of text similarity and sentiment similarity, expanding the traditional method of evaluating influence or identifying user activity based on interaction behavior characteristics.

[0044] In some embodiments, step S4, "Determine the current opinion leaders of each of the topic sets according to the network review texts in the topic set", includes:

[0045] Step S41: Determine the text similarity matrix between each of the first review texts in the topic set according to the first review texts in the topic set.

[0046] In some embodiments, step S41, "Determine the text similarity matrix between each of the first review texts in the topic set according to the first review texts in the topic set", includes:

[0047] Step S411: Calculate the word frequency of each word in each of the first review texts.

[0048] Step S412: Calculate the inverse document frequency of each word in the topic set.

[0049] Step S413: Determine the weight of each word in each of the first review texts according to the word frequency and the inverse document frequency.

[0050] Step S414: Determine the text similarity between each of the first review texts according to the weights of each word, and then obtain the text similarity matrix.

[0051] Although multiple network rating texts are classified into multiple topic sets according to clustering in this application, everyone has their own ideas, so it is difficult to ensure that the content expressed by each first review text in each topic set is exactly the same. It can only be said that the first review texts are consistent to a certain extent. Moreover, due to the mutual influence between comments, if the similarity between certain comments is high enough, it means that there are relatively consistent ideas between these comments. Then, the first review text with a more extensive influence can be determined through text similarity, and then the topic leader can be determined.

[0052] Since text similarity is a relationship between two texts, in this application, the text similarity of each first review document in a topic set is represented in the form of a matrix.

[0053] The text similarity between two comments is mainly determined by the expressions of various words. Therefore, when calculating the text similarity between two first review texts, the weights of each word in each first review text can be determined first, and then the text similarity can be calculated according to the weights of each word.

[0054] Therefore, in this application, TF-IDF can be used to measure the weight of each word in each first review text, and then the cosine similarity between the weights of each word in two first review texts can be calculated to obtain the text similarity between the two first review texts.

[0055] Among them, TF-IDF is a method for measuring the importance of words in a text. Its basic principle is to assign a weight to each word in the text and consider the frequency of the word in the entire text collection. TF (term frequency) represents the frequency of a word in the text, that is: IDF (inverse document frequency) reflects the importance of a word in the entire text collection, that is:

[0056]

[0057] TF-IDF is the product of TF and IDF, which is used to measure the importance of a word in a text and considers the frequency in the entire text collection: .

[0058] Once the TF-IDF weights of each word in each first review text are calculated, these weights can be used to calculate the similarity between texts. A common method is cosine similarity:

[0059] .

[0060] Step S42: Determine the topic leader group according to the text similarity matrix.

[0061] In some embodiments, step S42, "determine the topic leader group according to the text similarity matrix", includes:

[0062] Step S421: Obtain a preset first percentage threshold and a first PageRank value threshold.

[0063] Step S422: Convert the text similarity matrix into a one-dimensional text similarity vector, where the one-dimensional text similarity vector includes multiple text similarity factors.

[0064] Step S423: Arrange the text similarity factors from largest to smallest to form a text similarity factor sequence.

[0065] Step S424: Determine a text similarity factor threshold according to the text similarity factor sequence and the first percentage threshold.

[0066] Step S425: Determine the important text factors as the text similarity factors in the one-dimensional text similarity vector that are greater than or equal to the text similarity factor threshold.

[0067] Step S426: Establish a text undirected graph according to the important text factors.

[0068] Step S427: Calculate the text PageRank values of each node in the text undirected graph.

[0069] Step S428: Determine the topic leader group of the topic according to the text PageRank values of each node and the first PageRank value threshold.

[0070] Text similarity can express the similarity relationship between each first review text, but it is impossible to directly find the topic leader among many first review texts through the similarity relationship. Therefore, in this application, it is also necessary to process the text similarity matrix.

[0071] Since the text similarity matrix can only reflect the relationship between two first review texts and cannot reflect the relationship between the entire group, this application converts the text similarity matrix into a one-dimensional vector. Here, in order to distinguish it from the calculation of the current opinion leader below, this one-dimensional vector is called a one-dimensional text similarity vector, and the one-dimensional vector is composed of multiple factors. Similarly, in order to distinguish, each factor in the one-dimensional vector is called a text similarity factor.

[0072] Because the ideas between texts with higher similarity are more identical, and it can also be considered that the mutual influence between the two texts is relatively serious. Therefore, in this application, each first review text is filtered by the first percentage threshold, and the texts with higher text similarity are retained. In the one-dimensional vector, it is to retain the text similarity factor with a larger value, that is, the important text factor. Since there is a corresponding relationship between the similarity factor, the text similarity, and the first review text, the text undirected graph can also be constructed through the important text factor. In the text undirected graph, each node is each first review text, and the edge of the undirected graph represents the text similarity between two first review texts. Only the text similarity corresponding to the important text factor can be called an edge. After obtaining the undirected graph, calculate the text PageRank value of each node in the text undirected graph according to the text undirected graph. Determine the influence of each first review text in the entire topic through the text PageRank value, and determine the first review text whose influence or text PageRank value meets the threshold as the topic leader. And since in a topic, there is more than one text with higher influence, and there is more than one user, there is more than one topic leader, thus forming a topic leader group.

[0073] Step S43: Determine the emotional similarity matrix between each of the second review texts according to the second review texts corresponding to the topic leader group.

[0074] In some embodiments, step S43, "Determine the emotional similarity matrix between each of the second review texts according to the second review texts corresponding to the topic leader group", includes:

[0075] Step S431: Obtain a preset emotional dictionary.

[0076] Step S432: Determine the emotional words in each of the second review texts and the emotional polarities corresponding to the emotional words according to the emotional dictionary.

[0077] Step S433: Determine the emotional tendency score of each of the second review texts according to the emotional words and the emotional polarities.

[0078] Step S434: Determine the emotional similarity between each of the second review texts according to the emotional tendency scores, and then obtain the emotional similarity matrix.

[0079] Emotional similarity can indicate the stance among various review texts. To find the current opinion leader, it is necessary to analyze the emotional similarity among the second review texts in the topic leader group and clarify the emotional influence among the second review texts. Since each review contains multiple emotional words, and these emotional words have positive and negative emotions, in this application, before determining the emotional similarity among the second review texts, it is necessary to first obtain a preset emotional dictionary, which is determined based on the emotional word classification generally recognized by the group, such as the Chinese Appraisal Lexicon by Li Jun of Tsinghua University, NTUSD, and HowNet Emotional Dictionary. Then, text processing is performed on each second review text, including steps such as word segmentation and stop word removal. The second review text after text processing is matched with the preset emotional dictionary, and various emotional words and the emotional tendencies of each emotional word in each second review text are counted. Then, according to the emotional words and emotional tendencies, the emotional tendency score of each second review text is determined. Finally, the similarity of the emotional tendency scores of two second review texts is calculated through cosine similarity, and then the emotional similarity matrix between the two second review texts is obtained.

[0080] Among them, for any second review text , different emotional scores are assigned to different emotional words in the emotional dictionary.

[0081] Step S44: Determine the current opinion leader in the topic leader group according to the emotional similarity matrix.

[0082] In some embodiments, step S44 "Determine the current opinion leader in the topic leader group according to the emotional similarity matrix" includes:

[0083] Step S441: Obtain a preset second percentage threshold and a second PageRank value threshold.

[0084] Step S442: Convert the emotional similarity matrix into a one-dimensional emotional similarity vector, and the one-dimensional emotional similarity vector includes multiple emotional similarity factors.

[0085] Step S443: Arrange the emotional similarity factors from large to small to form an emotional similarity factor sequence.

[0086] Step S444: Determine the emotional similarity factor threshold according to the emotional similarity factor sequence and the second percentage threshold.

[0087] Step S445: Determine the important emotional factors as the emotional similarity factors in the one-dimensional emotional similarity vector that are greater than or equal to the emotional similarity factor threshold.

[0088] Step S446: Establish an emotional undirected graph according to the important emotional factors.

[0089] Step S447: Calculate the sentiment PageRank value of each node in the sentiment undirected graph.

[0090] Step S448: Determine the current opinion leader of the topic according to the sentiment PageRank value of each node and the second PageRank value threshold.

[0091] The sentiment similarity can express the similarity relationship between each second review text, but it is impossible to directly find the opinion leader among many second review texts through the similarity relationship. Therefore, in this application, it is also necessary to process the sentiment similarity matrix.

[0092] Since the sentiment similarity matrix can only reflect the relationship between two second review texts and cannot reflect the relationship between the entire group, in this application, the sentiment similarity matrix is first converted into a one-dimensional vector using cosine similarity. Here, this one-dimensional vector is called the one-dimensional sentiment similarity vector, and the one-dimensional vector is composed of multiple factors. Similarly, for the sake of distinction, each factor in the one-dimensional vector is called the sentiment similarity factor.

[0093] Because the more similar the texts are, the more the same their sentiment tendencies are. At the same time, it can also be considered that the mutual sentiment influence between the two texts is more serious. Therefore, in this application, each second review text is filtered by the second percentage threshold, and the texts with higher sentiment similarity are retained. In the one-dimensional vector, it is reflected by retaining the sentiment similarity factors with larger values, that is, the important sentiment factors. Since there is a corresponding relationship between the similarity factor, the sentiment similarity, and the second review text, the sentiment undirected graph can also be constructed through the important sentiment factors. In the sentiment undirected graph, each node is each second review text, and the edge of the undirected graph represents the sentiment similarity between two second review texts. Only the sentiment similarity corresponding to the important sentiment factor can be called an edge. After obtaining the undirected graph, calculate the sentiment PageRank value of each node in the sentiment undirected graph. Determine the influence of each second review text in the entire topic through the sentiment PageRank value, and determine the second review text whose influence or sentiment PageRank value meets the threshold as the opinion leader.

[0094] In addition, since existing technologies mostly start from the characteristics of Weibo users themselves, such as the number of followers, reposts, likes, comments, and the number of follows, they ignore the importance of the text features of the posts or comments themselves, such as the text length, post time, comment time, etc.; at the same time, since the spread of public opinion shows certain spatial characteristics, the IP addresses of Weibo users cannot be ignored in the research, which has often been overlooked in previous studies; in addition, in the identification algorithm of opinion leaders, different weights are not given to different features, resulting in the inaccuracy of the currently determined opinion leaders and the inability to predict future opinion leaders.

[0095] Regarding the above-mentioned existing problems, the present application can better determine the important feature sequence for becoming an opinion leader through the currently determined opinion leaders in the above steps. According to the important feature sequence, the future opinion leaders of this event can be predicted, and then their published content can be verified with emphasis to prevent problems before they occur and detect and supervise in time before the large-scale outbreak of public opinion; in addition, news official can imitate the important feature sequence of opinion leaders to conduct traffic confrontation with opinion leaders spreading false information and give full play to the role of official opinion leaders.

[0096] Therefore, in some embodiments, the method further includes:

[0097] Step S5: Obtain the user activity information of each of the currently determined opinion leaders.

[0098] Step S6: Determine the importance sequence of leader features according to the user activity information.

[0099] Step S7: Determine the weights of each feature in the user activity information according to the importance sequence of leader features.

[0100] Step S8: Predict future opinion leaders according to the weights of each feature.

[0101] In some embodiments, step S8, "Predict future opinion leaders according to the weights of each feature", includes:

[0102] Step S81: Obtain the target activity information of the target user.

[0103] Step S82: Determine the leader score of the target user according to the target activity information and the weights of each feature.

[0104] Step S83: Predict whether the target user can become an opinion leader in the future according to the leader score.

[0105] Since the second review text corresponding to the current opinion leader has been determined, we can determine the user activity information of the current opinion leader based on the second review text. This user activity information includes: the user's IP address, the publication time of the comment, the text length, and related attribute data such as the number of the user's fans, the degree of attention received, and the degree of attention to others. These attribute data of all opinion leaders are formed into a dataset to be analyzed together. The trained neural network model for feature extraction is used to determine important features from these datasets to be analyzed, and at the same time, the importance of these important features in the current opinion leader sample group is determined, thereby obtaining a feature importance sequence.

[0106] After obtaining the feature importance sequence, we can determine the weights of each important feature according to the feature importance sequence, clarify the important features as opinion leaders, and the weights of each important feature. At this time, we can use the weights of these important features to predict future opinion leaders among the existing users under this event. At the same time, it also enables the official to borrow these important features for imitation, so that the official account can become a future opinion leader and conduct traffic confrontation with opinion leaders who spread false information.

[0107] Among them, when predicting future opinion leaders, obtain the target activity information of the target user, and calculate the leader score of the target user according to the attributes in the target activity information and the weights of each feature. Then obtain a preset score threshold. When the leader score of the target user is greater than or equal to the score threshold, it is determined that the target user may become an opinion leader in the future.

[0108] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0109] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures, or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the "in one embodiment" or "in an embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the magnitude of the serial numbers of the above processes does not mean the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.

[0110] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0111] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments. The foregoing storage medium includes: various media such as removable storage devices, read-only memory (ROM, Read Only Memory), magnetic disks, or optical discs that can store program codes.

[0112] Alternatively, if the above integrated unit of the present application is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application essentially or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a controller to execute all or part of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as removable storage devices, ROM, magnetic disks, or optical discs that can store program codes.

[0113] As described above, it is only the implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the said claims.

Claims

1. A method for determining an opinion leader, characterized in that, including; obtaining a network evaluation set of a target event, where the network evaluation set includes a plurality of network review texts; determining an optimal number of topics of the target event according to the network evaluation set; dividing the network evaluation set into a plurality of topic sets according to the optimal number of topics; determining current opinion leaders of each of the topic sets according to each network review text in the topic sets; The determining current opinion leaders of each of the topic sets according to the topic sets includes: performing the following operations for each topic set: determining a text similarity matrix between each first review text in the topic set according to the first review text in the topic set; determining a topic leader group according to the text similarity matrix; determining an emotional similarity matrix between each second review text corresponding to the topic leader group according to the second review text corresponding to the topic leader group; determining current opinion leaders in the topic leader group according to the emotional similarity matrix.

2. The method according to claim 1, wherein The determining the number of topics of the target event according to the network evaluation set includes: determining perplexity indicators under different numbers of topics according to the network evaluation set; The determining perplexity indicators under different numbers of topics according to the network evaluation set includes: training an LDA model with different numbers of topics; calculating the perplexity of the LDA model with different numbers of topics using the network evaluation set, and establishing a relationship table between perplexity and the number of models; determining the optimal number of topics according to the perplexity indicators, where an inflection point where the perplexity decline rate slows down is found in the relationship table, and the number of topics corresponding to the inflection point is the optimal number of topics.

3. The method according to claim 1, characterized in that The dividing the network evaluation set into a plurality of topic sets according to the optimal number of topics includes: using the LDA model to determine the topic probabilities of each network review text in the network evaluation set belonging to each topic according to the optimal number of topics; determining the topics to which each network review text belongs according to the topic probabilities; dividing the network evaluation set into a plurality of topic sets according to the topics to which each network review text belongs.

4. The method according to claim 1, characterized in that, The determining the topic leader group according to the text similarity matrix includes: obtaining a preset first percentage threshold and a first PageRank value threshold; converting the text similarity matrix into a one-dimensional text similarity vector, where the one-dimensional text similarity vector includes a plurality of text similarity factors; arranging the text similarity factors from largest to smallest to form a text similarity factor sequence; determining a text similarity factor threshold according to the text similarity factor sequence and the first percentage threshold; determining text similarity factors greater than or equal to the text similarity factor threshold in the one-dimensional text similarity vector as important text factors; establishing a text undirected graph according to the important text factors; calculating the text PageRank values of each node in the text undirected graph; determining the topic leader group of the topic according to the text PageRank values of each node and the first PageRank value threshold.

5. The method according to claim 1, wherein The determining current opinion leaders in the topic leader group according to the emotional similarity matrix includes: obtaining a preset second percentage threshold and a second PageRank value threshold; Convert the emotional similarity matrix into a one-dimensional emotional similarity vector, where the one-dimensional emotional similarity vector includes multiple emotional similarity factors; Arrange the emotional similarity factors from largest to smallest to form an emotional similarity factor sequence; Determine an emotional similarity factor threshold according to the emotional similarity factor sequence and the second percentage threshold; Determine the important emotional factors from the emotional similarity factors in the one-dimensional emotional similarity vector that are greater than or equal to the emotional similarity factor threshold; Establish an emotional undirected graph based on the important emotional factors; Calculate the emotional PageRank values of each node in the emotional undirected graph; Determine the current opinion leader of the topic according to the emotional PageRank values of each node and the second PageRank value threshold.

6. The method according to claim 1, characterized in that, The determining the text similarity matrix between each first review text in the topic set according to the first review text in the topic set includes: Calculate the word frequency of each word in each first review text; Calculate the inverse document frequency of each word in the topic set; Determine the weight of each word in each first review text according to the word frequency and the inverse document frequency; Determine the text similarity between each first review text according to the weights of each word, and then obtain the text similarity matrix.

7. The method according to claim 1, characterized in that The determining the emotional similarity matrix between each second review text according to the second review text corresponding to the topic leader group includes: Obtain a preset emotional dictionary; Determine the emotional words in each second review text and the emotional polarities corresponding to the emotional words according to the emotional dictionary; Determine the emotional tendency score of each second review text according to the emotional words and the emotional polarities; Determine the emotional similarity between each second review text according to the emotional tendency scores, and then obtain the emotional similarity matrix.

8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: Obtain the user activity information of each current opinion leader; Determine a leader feature importance sequence according to the user activity information; Determine the weights of each feature in the user activity information according to the leader feature importance sequence; Predict future opinion leaders according to the weights of each feature.

9. The method according to claim 8, wherein The predicting future opinion leaders according to the weights of each feature includes: Obtain the target activity information of the target user; Determine the leader score of the target user according to the target activity information and the weights of each feature; Predict whether the target user can become an opinion leader in the future according to the leader score.

Citation Information

Patent Citations

  • Method and device for identifying opinion leader

    CN111460317A

  • Social network opinion leader mining method based on whole and part of graph structure

    CN114492455A