Network public opinion hot topic identification method, device and equipment

Through Chinese pre-trained language model and manifold learning dimensionality reduction technology, social media text data is embedded and dimensionality reduction, combined with density clustering technology, hot topics on online public opinion are identified, solving the problems of poor timeliness and low accuracy in the existing technology, and achieving efficient and accurate recognition of public opinion topics.

CN120235148APending Publication Date: 2025-07-01TIANJIN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510194534.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing technology has problems of poor timeliness and low accuracy when identifying hot topics on online public opinion. It is difficult to adapt to the rapidly changing network environment and cannot capture the deep meaning behind the topic.

Method used

The social media text data is embedded through a Chinese pre-trained language model, and the dimensionality reduction algorithm based on manifold learning is used to reduce the dimensionality of the high-dimensional vector representation. Then, the text data after the dimensionality reduction is density clustered, and keywords are extracted as the recognized topic.

Benefits of technology

It can accurately identify and summarize public opinion topics with high attention and representativeness from massive social media data, improve the automation level of public opinion hotspot recognition, and provide more accurate and timely information support for public opinion analysis and decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235148A_ABST
    Figure CN120235148A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing, in particular to a network public opinion hot topic recognition method, device and equipment, and the network public opinion hot topic recognition method comprises the steps: embedding social media text data through a Chinese pre-training language model to obtain corresponding high-dimensional vector representation; performing dimension reduction on the high-dimensional vector representation by adopting a dimension reduction algorithm based on manifold learning to obtain dimension-reduced text data; and performing density clustering on the text data subjected to dimension reduction, and extracting keywords from a clustering result as recognized topics. According to the method, public opinion topics with high attention and representativeness can be accurately recognized and summarized from massive social media data, the automation level of public opinion hotspot recognition is improved, and more accurate and timely information support is provided for public opinion analysis and decision making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of natural language processing, and in particular, to a method, apparatus, and device for identifying hot topics in online public opinion. Background Art

[0002] With the rapid development of social media, platforms such as Weibo and Twitter have become important channels for people to obtain information, express opinions, and participate in public discussions. These platforms generate a huge amount of information content every day, covering multiple fields such as society, economy, and culture. However, the sharp increase in the amount of information has also significantly increased the cost for people to quickly identify and obtain true and useful information from this information. At the same time, as an important form reflecting public attitudes, emotions, and opinions, online public opinion, various hot topics and events have rapidly emerged and spread widely on social media platforms, having a profound impact on the formation and development of social public opinion, and thus attracting much attention from all sectors of society. This phenomenon further highlights the necessity of efficient analysis and management of online public opinion. How to efficiently identify hot topics and extract key content in such a complex information environment has become an urgent problem to be solved.

[0003] Hot topic detection is a core task in the research of online public opinion, and its purpose is to automatically identify the currently most popular or most concerned topics from a large amount of social media content. Traditional hot topic detection methods mainly rely on keyword matching, frequency statistics, and simple machine learning algorithms. Although these methods can discover hot topics to a certain extent, they often have problems such as poor timeliness and low accuracy, and are difficult to adapt to the rapidly changing network environment. In addition, traditional methods have relatively limited understanding of semantics and cannot capture the deep meaning behind topics. Summary of the Invention

[0004] This application provides a method, apparatus, and device for identifying hot topics in online public opinion to solve the problems in the above background art.

[0005] In a first aspect, this application provides a method for identifying hot topics in online public opinion, including:

[0006] Embedding the social media text data through a Chinese pre-trained language model to obtain a corresponding high-dimensional vector representation;

[0007] Using a dimensionality reduction algorithm based on manifold learning to reduce the dimensionality of the high-dimensional vector representation to obtain the dimensionality-reduced text data;

[0008] Performing density clustering on the dimensionality-reduced text data, and extracting keywords from the clustering results as the identified topics.

[0009] Further, before embedding the social media text data through the Chinese pre-trained language model to obtain the corresponding high-dimensional vector representation, it further includes:

[0010] Collect social media text data and perform preprocessing, where the preprocessing includes removing noise data, stop word filtering, and text segmentation.

[0011] Further, embedding the social media text data through the Chinese pre-trained language model to obtain the corresponding high-dimensional vector representation includes:

[0012] Input the preprocessed social media text data into the Chinese-BERT-WWM model and output the corresponding high-dimensional vector representation.

[0013] Further, using the dimensionality reduction algorithm based on manifold learning to reduce the dimensionality of the high-dimensional vector representation to obtain the dimensionality-reduced text data includes:

[0014] Use the UMAP algorithm to construct a local adjacency graph, a global manifold structure, and a low-dimensional mapping for the high-dimensional vector representation to obtain the dimensionality-reduced text data.

[0015] Further, performing density clustering on the dimensionality-reduced text data and extracting keywords from the clustering results as the identified topics includes:

[0016] Use the HDBSCAN algorithm to calculate the core distance, construct a similarity graph, generate a density tree, and cut the density tree for the dimensionality-reduced text data to generate a clustering result;

[0017] Generate a bag-of-words representation for each cluster through the bag-of-words model, and use the c-TF-IDF weighting algorithm to identify the characteristic words of each cluster as the topic corresponding to each cluster.

[0018] Further, it further includes:

[0019] According to the topic and the representative text corresponding to the topic, generate an abstract corresponding to the topic through a large language model.

[0020] Further, generating an abstract corresponding to the topic through a large language model according to the topic and the representative text corresponding to the topic includes:

[0021] Construct a prompt according to the topic and the representative text corresponding to the topic;

[0022] Input the topic, the representative text corresponding to the topic, and the prompt into the Atom-7B-Chat model to generate an abstract corresponding to the topic.

[0023] In a second aspect, the present application provides a device for identifying hot topics in online public opinion, including:

[0024] A data embedding module, configured to embed social media text data through a Chinese pre-trained language model to obtain a corresponding high-dimensional vector representation;

[0025] A vector dimensionality reduction module, configured to perform dimensionality reduction on the high-dimensional vector representation by using a dimensionality reduction algorithm based on manifold learning to obtain the dimensionality-reduced text data;

[0026] A topic extraction module, configured to perform density clustering on the dimensionality-reduced text data, and extract keywords from the clustering results as the identified topics.

[0027] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned method for identifying hot topics in online public opinion is implemented.

[0028] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above-mentioned method for identifying hot topics in online public opinion is implemented.

[0029] The above technical solutions of the present application have the following advantages:

[0030] The method for identifying hot topics in online public opinion provided in the first aspect of the present application embeds social media text data through a Chinese pre-trained language model to obtain a corresponding high-dimensional vector representation, performs dimensionality reduction on the high-dimensional vector representation by using a dimensionality reduction algorithm based on manifold learning to obtain the dimensionality-reduced text data, performs density clustering on the dimensionality-reduced text data, and extracts keywords from the clustering results as the identified topics. It can accurately identify and summarize public opinion topics with high attention and representativeness from a large amount of social media data, improve the automation level of public opinion hot topic identification, and provide more accurate and timely information support for public opinion analysis and decision-making.

[0031] It can be understood that the beneficial effects of the above second aspect, third aspect, and fourth aspect can refer to the relevant descriptions in the above first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0033] Figure 1 Flow chart of the method for identifying hot topics of online public opinion provided by this application;

[0034] Figure 2 Basic work flow chart of hot topic detection provided by this application;

[0035] Figure 3 Basic work flow chart of topic summary generation provided by this application;

[0036] Figure 4 Structure diagram of the device for identifying hot topics of online public opinion provided by this application;

[0037] Figure 5 Structure schematic diagram of the electronic device provided by this application. Detailed implementation manners

[0038] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures, technologies, etc. are presented to thoroughly understand the embodiments of this application. However, those skilled in the art should clearly understand that this application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of this application.

[0039] It should be understood that when used in the specification of this application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0040] In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0041] The reference to "one embodiment" or "some embodiments" etc. in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways. "Multiple" means "two or more".

[0042] The following will further describe in detail the specific implementation manners of the present application in combination with the accompanying drawings and embodiments. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.

[0043] As Figure 1 shown, the embodiment of the present application provides a method for identifying hot topics of online public opinion, which specifically includes the following steps: embedding the social media text data through a Chinese pre-trained language model to obtain a corresponding high-dimensional vector representation; using a dimensionality reduction algorithm based on manifold learning to reduce the dimensionality of the high-dimensional vector representation to obtain the dimensionality-reduced text data; performing density clustering on the dimensionality-reduced text data, and extracting keywords from the clustering results as the identified topics.

[0044] In some embodiments, before embedding the social media text data through a Chinese pre-trained language model to obtain a corresponding high-dimensional vector representation, it further includes: collecting social media text data and performing preprocessing, where the preprocessing includes removing noise data, stop word filtering, and text segmentation.

[0045] In some embodiments, embedding the social media text data through a Chinese pre-trained language model to obtain a corresponding high-dimensional vector representation includes: inputting the preprocessed social media text data into the Chinese-BERT-WWM model and outputting the corresponding high-dimensional vector representation.

[0046] In some embodiments, using a dimensionality reduction algorithm based on manifold learning to reduce the dimensionality of the high-dimensional vector representation to obtain the dimensionality-reduced text data includes: using the UMAP algorithm to construct a local adjacency graph, a global manifold structure, and a low-dimensional mapping for the high-dimensional vector representation to obtain the dimensionality-reduced text data.

[0047] In some embodiments, performing density clustering on the dimensionality-reduced text data and extracting keywords from the clustering results as the identified topics includes: using the HDBSCAN algorithm to calculate the core distance, construct a similarity graph, generate a density tree, and cut the density tree for the dimensionality-reduced text data to generate a clustering result; generating a bag-of-words representation for each cluster through a bag-of-words model, and using the c-TF-IDF weighting algorithm to identify the characteristic words of each cluster as the topics corresponding to each cluster.

[0048] In some embodiments, it further includes: generating an abstract corresponding to the topic according to the topic and the representative text corresponding to the topic through a large language model.

[0049] In some embodiments, generating the abstract corresponding to the topic through a large language model according to the topic and the representative text corresponding to the topic includes: constructing a prompt word according to the topic and the representative text corresponding to the topic; inputting the topic, the representative text corresponding to the topic, and the prompt word into the Atom-7B-Chat model to generate the abstract corresponding to the topic.

[0050] Hot topic detection uses a technical solution based on text clustering to process social media data to automatically identify the most representative and attention-grabbing hot topics. First, preprocess the text data in the social media platform, including steps such as removing stop words, word segmentation, and stemming, to ensure the effectiveness and accuracy of the data. Subsequently, use unsupervised learning methods to perform topic modeling and clustering on the processed text data.

[0051] To efficiently process Chinese text and obtain an accurate semantic representation, this application uses the Chinese-BERT-WWM model for embedding. Chinese-BERT-WWM is a Chinese pre-trained language model based on BERT that adopts a more refined vocabulary masking strategy (Whole Word Masking, WWM). Compared with the traditional character-level masking method, WWM can retain complete Chinese vocabulary information during the training process, thus effectively solving the word embedding problem in Chinese text processing. Specifically, the Chinese text to be processed, including single sentences, paragraphs, or entire articles, is first input into the model. To ensure consistency with the model training process, special marker symbols [CLS] and [SEP] need to be added at the beginning and end of the text respectively. Among them, the [CLS] marker is used to indicate the starting position of the sentence and serves as the output representation in the classification task; the [SEP] marker is used to distinguish different sentences or paragraphs. After being processed by this model, the input text will be converted into a high-dimensional vector representation, facilitating subsequent operations such as clustering analysis and topic extraction.

[0052] After embedding the document, each input text will obtain a corresponding high-dimensional vector representation. This vector is usually a fixed-length array (such as 768-dimensional, 1024-dimensional, etc.), which contains rich semantic information of the text. Before clustering, the high-dimensional representation of the text data usually needs to be dimensionally reduced to reduce the computational complexity and improve the clustering effect. This application uses UMAP (Uniform Manifold Approximation and Projection) technology for dimensional reduction. UMAP is a dimensional reduction method based on manifold learning, with strong non-linear dimensional reduction ability, which can preserve the local and global structure of the text vector. Compared with traditional linear dimensional reduction methods such as PCA, UMAP can better capture the complex relationships in the text data, especially suitable for dimensional reduction and visualization of text embedding vectors, making the clustering results more representative and operable. The core idea of UMAP is to construct an adjacency graph of the data and use the adjacency information to capture the manifold structure of the data. The algorithm is divided into two stages: constructing a local adjacency graph and performing low-dimensional embedding through an optimization process.

[0053] UMAP first determines a set of adjacent points for each data point in the high-dimensional space. To represent the local structure, UMAP uses the k-nearest neighbor algorithm (K-NN) to construct the relationship between data points and represents the weights of these relationships through a Gaussian distribution. Specifically, for the data point x i , calculate its distance from other points x j , select the k nearest points as neighbors, and use a distance-weighted Gaussian distribution to determine the similarity:

[0054]

[0055] where d(x i , x j ) is the Euclidean distance between the point x i and x j , and σ i is the standard deviation of the Gaussian distribution adjusted by the local scale parameter.

[0056] UMAP not only considers the local neighbor relationship but also tries to capture the global structure of the dataset. By using a random adjacency graph (usually a connected graph) to represent the global manifold of the data, UMAP adopts an optimization method related to graph theory to map the data from the high-dimensional space to the low-dimensional space. In the low-dimensional space, UMAP uses the negative sampling method to learn the low-dimensional representation, and the goal is to minimize the difference in similarity between adjacent points in the high-dimensional and low-dimensional spaces. The optimization goal is usually represented by the cross-entropy loss function:

[0057]

[0058] where P(xi |x j ) and Q(y i |y j ) represent the similarity between data points in high-dimensional and low-dimensional spaces respectively. By minimizing this objective, UMAP can effectively capture the global and local structures of the data in the low-dimensional space.

[0059] To identify topics with different densities and shapes, this application uses HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise) as the clustering algorithm. HDBSCAN is a density clustering algorithm. Compared with the traditional K-means algorithm, it does not require a preset number of clusters and can handle clusters with different densities. HDBSCAN can automatically determine the number and shape of clusters according to the local density changes of the data, which is very suitable for dealing with the diversity and complexity of topics in social media. The core idea of HDBSCAN is to perform clustering based on the density distribution of data points. Its workflow can be divided into the following steps:

[0060] HDBSCAN first calculates the core distance of each data point. The core distance refers to the distance from a data point to its k-th nearest neighbor. The core distance for each point x i can be expressed as:

[0061]

[0062] where d(x i , x j ) is the distance between data points x i and x j , and N k (x i ) is the k-th nearest neighbor of point x i . HDBSCAN constructs a similarity graph by calculating the reachable distance between each pair of data points. The reachable distance is defined based on the core distance and the distance between points and is used to measure the density relationship between points:

[0063] reachability(x i , x j ) = max(core_distance(x i ), d(x i , x j ))

[0064] This distance metric represents the distance from point x i to point x j"Reachability". If the reachable distance between two points is small, it indicates that they belong to the same density region.

[0065] HDBSCAN identifies the hierarchical structure of data by constructing a density tree (also known as a hierarchical structure diagram). In this tree, each node represents a density level. The root of the tree represents the global density of the dataset, while the leaf nodes represent the smallest clustering units. To identify the final clusters, HDBSCAN cuts the density tree. By selecting an appropriate density threshold, the algorithm cuts the tree into multiple clusters. Generally speaking, lower density indicates noise or outliers, while higher density indicates the main clusters of the data. Finally, HDBSCAN generates the clustering results based on the hierarchical structure and the cut density tree, and labels the low-density points as noise.

[0066] The advantage of HDBSCAN is that it does not rely on the assumption of centroids, can adaptively identify regions with higher density and group them into the same class, and at the same time has good robustness to noise and outliers. Through this clustering step, different public opinion hot topics can be effectively identified from a large amount of text.

[0067] When performing clustering analysis, traditional centroid-based topic representation methods (such as K-means) assume that each cluster has a common centroid, but for clusters with different densities and shapes, this method is often not applicable. Therefore, this application uses the Bag-of-Words technology to represent the characteristics of each cluster. Specifically, first, all the texts in each cluster are merged into a "super document", and then the frequency of each word in this document is calculated to generate the bag-of-words representation of the cluster. The bag-of-words model is a statistical method based on all the documents in the cluster, which can effectively avoid the assumption of the cluster structure. In this way, a word frequency distribution for each cluster can be obtained, and further analysis can be carried out on which words are more representative in a specific cluster. This method avoids the assumption of the cluster shape and ensures an accurate grasp of the cluster topic. To take into account the differences in cluster sizes, L1 normalization is performed during the calculation of the bag-of-words representation to ensure that the word frequencies in larger clusters do not overly affect the final representation result.

[0068] To extract the most representative topic words from the generated clusters, this application adopts the c-TF-IDF (class-based Term Frequency-Inverse Document Frequency) method, which is an improvement on the traditional TF-IDF (Term Frequency-Inverse Document Frequency) technology and is specifically used for topic representation at the cluster level. The core idea of c-TF-IDF is to use clusters (i.e., topics) as the calculation unit instead of individual documents, so as to identify the characteristic words in each cluster. The specific implementation steps are as follows:

[0069] Calculation of Term Frequency (TF): Merge all the texts in each cluster to form a "super document" at the cluster level. Count the frequency of each word in the cluster and calculate its relative term frequency. The formula is:

[0070]

[0071] This formula measures the relative importance of word w in cluster c.

[0072] Calculation of Inverse Document Frequency (IDF): Calculate the distribution of each word in all clusters to measure the rarity of the word among different clusters. If a word is widely distributed in multiple clusters, its contribution to distinguishing a specific cluster is low; conversely, if a word is mainly concentrated in a certain cluster, its discrimination ability is strong. The formula is:

[0073]

[0074] Among them, N represents the total number of clusters, and DF(w) represents the number of clusters containing word w.

[0075] Calculation of c-TF-IDF weighting: Combine the term frequency and the inverse document frequency to calculate the weighted value of each word in a specific cluster. The formula is:

[0076] c-TF-IDF(w,c) = TF(w,c) · IDF(w)

[0077] This weighted value comprehensively considers the frequency of the word in the cluster and its rarity across clusters, and can effectively highlight the characteristic vocabulary of each cluster.

[0078] Through the c-TF-IDF weighting scheme, discriminative keywords can be extracted from each cluster to form high-quality topic representations. Compared with traditional TF-IDF, c-TF-IDF takes the cluster as the core calculation unit, can better adapt to the topic representation requirements in unstructured text data, and at the same time avoids excessive assumptions about the cluster structure. This method is particularly suitable for processing diverse content in social media data, providing a solid foundation for subsequent public opinion analysis and topic-level summary generation.

[0079] The topic summary generation adopts the Atom-7B-Chat model, which is a language model pre-trained in Chinese based on the Llama2 (Large Language Model Meta AI 2) architecture. Llama2 has obtained powerful semantic understanding and generation capabilities through pre-training on a large amount of text data. Compared with the original Llama2 model, Atom-7B-Chat has been optimized in the Chinese context and can better handle the grammar and semantic characteristics of Chinese texts. This model can generate natural language texts in a generative manner and is suitable for the task of topic summary generation.

[0080] According to the obtained cluster label words and representative documents, combined with the Atom-7B-Chat model, topic summaries with strong readability and rich information are generated. Based on the Chinese pre-trained model of Llama2 and combined with Prompt engineering techniques, appropriate Prompts are established to generate accurate and human-language-habit-compliant summary content. The design of the prompt words aims to enable the model to understand the summary content to be generated. According to the label words and representative documents of each cluster, an input format with context information is constructed. The prompt words consist of three parts: system prompt words, example prompt words, and main prompt words:

[0081] The system prompt, as the global prompt for all conversations, clarifies the role positioning and main tasks of the model. Through a simple and clear description, the model is informed that its identity is an assistant focused on topic annotation. The main purpose of the system prompt is to set the behavior benchmark of the model, ensure that the model maintains consistency throughout the task, follows the set rules, and avoids generating content that deviates from the task goal. The example prompt conveys a clear output demonstration to the model by showing a specific input-output instance. By providing an example of a topic document and keywords, the model is demonstrated how to convert the input (document and keywords) into accurate and concise topic labels. The main prompt is the template for actually generating labels, and its design goal is to combine the information (documents and keywords) extracted by the model from the cluster to generate short labels related to the topic. This prompt is the core guidance for the model to generate topic labels in the actual task.

[0082] Input the constructed prompt words into the Atom-7B-Chat model. Since Atom-7B-Chat is a Chinese pre-trained model optimized based on the Llama2 architecture, it can understand and process the details of Chinese text and generate summaries with fluent language and clear structure. By combining the Atom-7B-Chat model with the Prompt project, high-quality topic summaries can be automatically generated from the generated clustered label words and representative documents. This solution is highly customizable and adaptable, and can quickly identify and generate public opinion summaries with generalization and information content in diverse social media data, providing strong support for subsequent public opinion monitoring and analysis.

[0083] The following is an explanation through specific embodiments.

[0084] Example

[0085] This embodiment is mainly used in social media public opinion analysis scenarios, by automatically detecting hot topics from a large amount of unstructured text data on platforms such as Weibo and Twitter, and generating topic summaries with high readability and accuracy. Figure 2 The present invention is a flow chart of a method for detecting hot topics of online public opinion, which may include the following steps:

[0086] Step 1: Data collection and storage. Collect a large amount of public text data from social media platforms such as Weibo and Twitter, and store it in csv format. The fields include text content and release time. For example:

[0087]

#Cultural Tourism and Creative Industry in Luoyang#

[0088] Step 2: Data preprocessing. Clean and standardize the collected data, including removing noise data and filtering stop words. Then, use Chinese word segmentation technology to segment the text and generate basic units for subsequent analysis. For example:

[0089] Looking at Luoyang's Cultural and Tourism Cultural and Creative Industries The Henan Cultural and Tourism Cultural and Creative Industries Development Conference The main aspects of the project signing arrangements for this conference are as follows: First, signing of cultural and tourism industry projects. So far, key cultural and tourism projects have been sorted out with a total investment of 52.56 billion yuan. Major projects have been selected for on-site signing with a total investment of 36.58 billion yuan. The projects include digital development of cultural relics, construction of cultural and creative industrial parks, cultural projects covering construction of tourist resorts, building of tourist hotels and homestays, traditional format projects such as tourist scenic area development and construction of cultural and tourism complexes, cosmic base immersive performance format projects, fully reflecting the characteristics and trends of the cultural and tourism development in our province. Second, signing of projects to attract tourists into Henan. Mainly, the cultural and tourism departments of our province, cultural and tourism enterprises, leading travel and well-known OTA platforms, and cultural and tourism departments in key tourist source areas have signed agreements to attract tourists into Henan to continuously expand the tourist source market outside the province.

[0090] Step 3: Topic clustering. Use the Chinese pre-trained model Chinese-BERT-WWM to convert text data into high-dimensional vector embeddings to capture the semantic features of the text. Use UMAP to perform dimensionality reduction on the high-dimensional embeddings, retaining the main semantic structure features while reducing the data dimension, thereby improving the computational efficiency. The dimensionality-reduced embeddings are subject to topic clustering by HDBSCAN (hierarchical density clustering). HDBSCAN is suitable for clusters with unbalanced densities, can identify topic clusters of different shapes and densities, and generate topic labels. Calculate the weight of each word in the cluster by the c-TF-IDF method to generate the label words for each topic. This method can effectively express the differences between clusters and highlight the representative words of each topic. Select the top 5 documents most relevant to the topic in each cluster as the representative documents for the topic, providing a basis for subsequent abstract generation.

[0091] Figure 3 It is a flowchart of a topic abstract generation method based on a large model, and this method can include the following steps:

[0092] Step 1: Construct a suitable Prompt, including a system prompt, an example prompt, and a main prompt, to enhance the generation effect of the model. Refer to Table 1, the system prompt defines the global task role of the model, such as "a useful assistant for marking topics"; refer to Table 2, the example prompt conveys a clear output demonstration to the model by showing a specific input-output instance; refer to Table 3, the main prompt contains specific input text examples and topic label descriptions to guide the model to generate accurate abstracts. Among them, the [INST] tag is used to identify the beginning and end of the prompt; [DOCUMENTS] contains the top 5 documents most relevant to the topic; [KEYWORDS] contains the top 10 keywords most relevant to the topic generated by c-TF-IDF. Integrate the output results of the topic detection module (including topic label words and representative documents) into the Prompt template to form the input for abstract generation.

[0093] Step 2: Abstract Generation. Table 4 shows an example of abstract generation. The Prompt constructed in Step 1 (including system prompt, example prompt, main prompt, and the labeled words and representative documents output by the topic detection module) is passed as input to Atom-7B-Chat. Based on the provided input, the Atom-7B-Chat model will understand the core content of the topic and generate a concise and representative abstract by combining the context information of the documents and keywords. The generated abstract will accurately summarize the topic, provide clear key information, and ensure high readability.

[0094]

[0095]

[0096]

[0097]

[0098] This embodiment can efficiently identify hot topics in social media and generate high-quality topic summaries, thereby helping users quickly grasp the public opinion dynamics and improve the information extraction and decision-making efficiency.

[0099] The method for identifying hot topics of network public opinion provided by the embodiment of the present application aims to automatically identify and summarize highly concerned and representative public opinion topics from a large amount of social media data. This method mainly includes hot topic detection and topic summary generation. Hot topic detection performs text embedding through the Chinese pre-trained language model Chinese-BERT-WWM, uses UMAP for dimensionality reduction, and combines the HDBSCAN algorithm for density clustering to automatically identify hot topics of public opinion. Subsequently, the c-TF-IDF weighting method is used to further extract keywords in the clustering to ensure the accuracy and representativeness of topic expression. Topic summary generation is based on the Atom-7B-Chat model of the Llama2 architecture, combined with prompt engineering technology, to automatically generate a smooth and information-rich topic summary. This method not only improves the automation level of hot topic identification of public opinion, but also provides more accurate and timely information support for public opinion analysis and decision-making, and has broad application prospects.

[0100] Corresponding to the method for identifying hot topics of network public opinion described in the above embodiment, as Figure 4 shown, the embodiment of the present application also provides a device for identifying hot topics of network public opinion, and the device for identifying hot topics of network public opinion includes:

[0101] A data embedding module, configured to perform embedding on social media text data through a Chinese pre-trained language model to obtain a corresponding high-dimensional vector representation;

[0102] A vector dimensionality reduction module, which is used to reduce the dimensionality of the high-dimensional vector representation by using a dimensionality reduction algorithm based on manifold learning to obtain the text data after dimensionality reduction;

[0103] A topic extraction module, which is used to perform density clustering on the text data after dimensionality reduction and extract keywords from the clustering results as the identified topics.

[0104] It should be noted that for the information interaction, execution process, etc. between the above modules / units, since they are based on the same concept as the method embodiment of the present application, their specific functions and the technical effects brought can be specifically referred to in the method embodiment part, and will not be elaborated here.

[0105] Those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated here.

[0106] The embodiment of the present application also provides an electronic device, as Figure 5 shown, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the network public opinion hot topic recognition method provided in the first aspect.

[0107] In applications, the electronic device may include, but is not limited to, a processor and a memory, Figure 5 which are only examples of the electronic device and do not constitute a limitation to the electronic device. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, input / output devices, network access devices, etc. The input / output devices may include cameras, audio acquisition / playback devices, displays, etc. The network access device may include a network module for performing wireless network with external devices.

[0108] In an application, the processor may be a Central Processing Unit (CPU), and the processor may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0109] In an application, in some embodiments, the memory may be an internal storage unit of an electronic device, such as a hard disk or memory of the electronic device. In other embodiments, the memory may also be an external storage device of the electronic device, for example, a plug-in hard disk equipped on the electronic device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. The memory may also include both an internal storage unit and an external storage device of the electronic device. The memory is used to store an operating system, application programs, a BootLoader, data, and other programs, such as program codes of computer programs. The memory may also be used to temporarily store data that has been output or will be output.

[0110] The embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments can be implemented.

[0111] To implement all or part of the processes in the above method embodiments of the present application, it can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps in the above method embodiments can be implemented. Among them, the computer program includes computer program codes, and the computer program codes can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program codes to an electronic device, a recording medium, a computer memory, a Read-Only Memory (ROM), a Random Access Memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc.

[0112] Those of ordinary skill in the art will appreciate that the devices and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented using electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled artisans may use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0113] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. Additionally, the couplings or direct couplings or communication connections shown or discussed among each other can be through some interfaces, and the devices are indirectly coupled or communication-connected, which can be in electrical, mechanical, or other forms.

[0114] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for identifying hot topics in online public opinion, characterized in that: include: Embed social media text data through the Chinese pre-trained language model to obtain the corresponding high-dimensional vector representation; Using a dimensionality reduction algorithm based on manifold learning to reduce the dimension of the high-dimensional vector representation to obtain text data after dimensionality reduction; Density clustering is performed on the text data after dimension reduction, and keywords are extracted from the clustering results as identified topics.

2. The method for identifying hot topics of online public opinion according to claim 1, characterized in that: Before embedding the social media text data through the Chinese pre-trained language model to obtain the corresponding high-dimensional vector representation, the method further includes: The social media text data is collected and preprocessed, and the preprocessing includes removing noise data, filtering stop words and segmenting text.

3. The method for identifying hot topics of online public opinion according to claim 1, characterized in that: The social media text data is embedded by the Chinese pre-trained language model to obtain the corresponding high-dimensional vector representation, including: The preprocessed social media text data is input into the Chinese-BERT-WWM model and the corresponding high-dimensional vector representation is output.

4. The method for identifying hot topics of online public opinion according to claim 1, characterized in that: The dimensionality reduction algorithm based on manifold learning is used to reduce the dimension of the high-dimensional vector representation to obtain text data after dimensionality reduction, including: The UMAP algorithm is used to construct a local adjacency graph, a global manifold structure and a low-dimensional mapping for the high-dimensional vector representation to obtain text data after dimensionality reduction.

5. The method for identifying hot topics of online public opinion according to claim 1, characterized in that: The step of performing density clustering on the dimensionally reduced text data and extracting keywords from the clustering results as identified topics includes: The HDBSCAN algorithm is used to calculate the core distance of the text data after dimension reduction, construct a similarity graph, generate a density tree, cut the density tree, and generate a clustering result; The bag-of-words model is used to generate the bag-of-words representation of each cluster, and the c-TF-IDF weighted algorithm is used to identify the characteristic words of each cluster as the topic corresponding to each cluster.

6. The method for identifying hot topics of online public opinion according to claim 1, characterized in that: Also includes: According to the topic and representative texts corresponding to the topic, a summary corresponding to the topic is generated by a large language model.

7. The method for identifying hot topics of online public opinion according to claim 6, characterized in that: The step of generating a summary corresponding to the topic by using a large language model according to the topic and the representative text corresponding to the topic includes: Constructing prompt words according to the topic and representative texts corresponding to the topic; The topic, the representative text corresponding to the topic and the prompt word are input into the Atom-7B-Chat model to generate a summary corresponding to the topic.

8. A device for identifying hot topics in online public opinion, characterized in that: include: The data embedding module is used to embed social media text data through the Chinese pre-trained language model to obtain the corresponding high-dimensional vector representation; A vector dimensionality reduction module, used to reduce the dimension of the high-dimensional vector representation by using a dimensionality reduction algorithm based on manifold learning to obtain text data after dimensionality reduction; The topic extraction module is used to perform density clustering on the text data after dimension reduction, and extract keywords from the clustering results as identified topics.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for identifying hot topics in online public opinion as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for identifying hot topics in online public opinion as described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Hot topic monitoring method and system based on improved HDBSCAN clustering algorithm

    CN120429486A

  • Open domain Chinese event pattern induction method and system based on multi-dimensional feature fusion

    CN121835691A