A method for analyzing polar activity based on multi-domain public opinion big data

Through the polar activity analysis method based on multi-field public opinion big data, the problem that the existing technology cannot effectively capture the cross-domain correlation of low-frequency key information and models is solved, and in-depth mining and real-time monitoring of polar public opinion is achieved, and accurate event prediction and decision-making support is provided.

CN119989236BActive Publication Date: 2025-06-13POLAR RES INST OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510436066.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-06-13
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

Existing polar public opinion data analysis technology cannot effectively capture low-frequency key information, it is difficult to model cross-domain correlation, and traditional static models cannot adapt to the rapid changes in data distribution, resulting in insufficient analysis accuracy and real-time.

Method used

The polar activity analysis method based on multi-field public opinion big data is adopted to capture low-frequency information through screening and analysis of polar sparse characteristics, and the potential connections between the characteristics are mined through correlation analysis to predict the development trend of events. Specific steps include data cleaning, sparse feature screening, feature vector construction, spatial density clustering processing, correlation feature analysis and dynamic adjustment.

Benefits of technology

It significantly improves the depth and breadth of polar public opinion information, can capture low-frequency but critical information, enhance the monitoring of data, timely predict event development trends, and provide accurate decision-making basis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989236B_ABST
    Figure CN119989236B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for analyzing polar activity based on multi-domain public opinion big data, which relates to the technical field of data processing. The specific steps of this method are as follows: S100 Screen the sparse features of the poles regarding public opinion; S200 Mine the implicit associations between different sparse features of the poles; S300 Mark the abnormally active points to generate a heat zone map; S400 Statistically analyze the frequently occurring feature pairs in the active points and dynamically adjust them; S500 Obtain the analysis result of polar activity. The method for analyzing polar activity based on multi-domain public opinion big data proposed by the present invention can significantly improve the depth and breadth of mining public opinion information about the poles. Through the effective screening and analysis of the sparse features of the poles, it is no longer limited to relying on high-frequency features, but can capture those low-frequency but key information, and through association analysis, mine the potential connections between them and other features, enhance the monitoring intensity of the data, and ensure that the public opinion dynamics are grasped in the first time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a polar activity analysis method based on multi-field public opinion big data. Background Art

[0002] Polar activity analysis is a core means of monitoring environmental changes, resource development, and geopolitical dynamics in polar regions. Its core task is to identify the spatiotemporal evolution patterns and cross-domain correlations of polar events by integrating multi-source heterogeneous data. With the increasing complexity of polar activities, existing public opinion analysis tools and technologies often cannot meet these needs, especially when faced with low-frequency keywords and implicit associations.

[0003] At present, traditional analysis methods face three challenges:

[0004] First, polar data has natural sparseness and long-tail distribution characteristics, which makes key signals easily submerged by noise;

[0005] Second, the cross-domain correlation of polar events is difficult to model using single-modality data;

[0006] Third, the dynamic evolution of polar activity requires real-time analysis capabilities, while traditional static models cannot adapt to the rapid drift of data distribution.

[0007] In summary, the existing polar public opinion data analysis technology has many shortcomings and cannot meet the needs of comprehensive and in-depth analysis of the activity in the polar regions. Therefore, a new analysis method is needed to overcome these shortcomings and improve the accuracy and effectiveness of polar activity analysis. Summary of the invention

[0008] The purpose of the present invention is to make up for the shortcomings of the prior art and to provide a polar activity analysis method based on multi-field public opinion big data. It can effectively screen and analyze polar sparse features, no longer be limited to relying on high-frequency features, but can capture low-frequency but critical information, and through association analysis, explore the potential connection between it and other features, and predict the development trend of events.

[0009] In order to solve the above technical problems, the present invention provides the following technical solutions: a polar activity analysis method based on multi-field public opinion big data, the specific steps of the method are:

[0010] S100. Collect multi-domain polar public opinion big data, clean the collected data, remove duplicate, erroneous and irrelevant information, and filter out polar sparse features about public opinion from the processed public opinion data set, and output a sparse feature list and original geographic coordinates ;

[0011] S200. Map the polar sparse features into feature vectors, mine the implicit associations between different polar sparse features, and obtain the associated feature set T;

[0012] S300. Perform data sampling on the remaining public opinion data of the screened polar sparse features, perform spatial density clustering processing after stratifying the data features, calculate the similarity between data points, compare the similarity with the average similarity, and divide it into three types of regions, including the dense region, the transition region, and the sparse region, and mark the abnormally active points, and generate a heat map and a list of cross-region active points as ;

[0013] S400. Perform association analysis on the associated feature set obtained in S200 and the list of cross-region active points in S300, count the feature pairs that frequently appear in the active points, perform dynamic adjustment, and output the updated feature association relationship table;

[0014] S500. Map the high-frequency associated feature pairs in the updated association feature table to the abnormally active points. For each abnormally active point, find the high-frequency associated feature pairs corresponding to it, analyze the relationship between the public opinion information represented by these feature pairs and the active point, and obtain the polar activity analysis result.

[0015] Furthermore, the S100 calculates the occurrence frequency of each feature , and this occurrence frequency is used to measure the frequency of occurrence of this feature in the entire public opinion dataset.

[0016] By calculating the ratio of the number of occurrences of each feature to the total number of features in the dataset, its occurrence frequency is obtained , that is , where represents the number of occurrences of feature , T represents the total number of features in the dataset, combined with the occurrence frequency threshold , retain those that satisfy .

[0017] Even further, the S100 calculates the distribution entropy of the feature . The distribution entropy is used to measure the degree of dispersion of the feature in different data sources. When the distribution of the feature in each data source is uniform, that is, when the distribution entropy is high, retain the features whose distribution entropy is greater than the set threshold . The distribution entropy is where is the distribution probability in the th data source, and M is the total number of data sources. Retain the features that satisfy to obtain the sparse feature set , while recording the geographical coordinates corresponding to each sparse feature to form a geographical coordinate set .

[0018] Furthermore, the S200 converts each sparse feature in the sparse feature set processed by S100 into a feature vector. The process is as follows: for the public opinion data feature set containing different words, each sparse feature corresponding vector each dimension in represents the frequency of the word in , through the word set , for the sparse feature , its feature vector the th dimension is the number of occurrences of the word in , that is , where represents the number of occurrences of the word in the sparse feature .

[0019] Furthermore, the S200 calculates the similarity between different sparse feature vectors to measure the implicit association degree between features. By calculating the cosine value of the included angle between two vectors, the similarity degree of sparse features is judged, so as to judge the implicit association strength between two sparse features. That is, for two sparse features and corresponding feature vectors and the similarity between is: , where is the dot product of vectors, and are the norms of vectors and respectively. When the similarity > β, it is considered that there is an implicit association between features and , record these two features and their corresponding cosine similarity values , and construct an associated feature set T.

[0020] Furthermore, S300 performs data sampling on the remaining data after S100 extracts polar sparse features. This data sampling includes stratified sampling and random sampling. Among them, the stratified sampling stratifies according to the source, time, and topic category of the remaining public opinion data, and retains the representative data of each layer. For the random sampling, for each subset after stratification, data points are extracted at a sampling ratio of 20% to form a preliminary sampling data set. Then, the preliminary sampling data set is randomly sampled again, and 80% of the data is extracted from the preliminary sampling data set to form the final sampling data set. The final sampling data set is vectorized using the construction of the sparse feature vector in S200, and the feature vector is normalized. The sampled data points will be used as the input for spatial density clustering processing;

[0021] The spatial density clustering calculates the similarity between data points. Each data point is represented by its feature vector. The cosine similarity between vectors is calculated to measure the similarity between data points. According to the similarity and neighborhood relationship between data points, the data points are divided into core points, boundary points, and noise points. Using the distribution of data points, the region is divided into a dense area, a transition area, and a sparse area. That is, when the number of core points in a region is large and the similarity is greater than the similarity mean, it is a dense area. When core points and boundary points are mixed and the similarity is equal to the similarity mean, it is a transition area. When the number of noise points is large and the similarity is less than the similarity mean, it is a sparse area;

[0022] Statistical abnormally active points are generated to form a list of cross-region active points, which is combined with the original geographical coordinates and the specific locations are displayed in the heat zone map.

[0023] Furthermore, the specific steps of S400 are as follows:

[0024] S401. Compare each feature pair in the feature association set obtained in S200 with the public opinion content involved in each active point in the list of cross-region active points generated by S300. For each cross-region active point, check whether its public opinion text contains the feature pair in the feature association set. That is, for the feature pair and the active point , define the matching function , where the active point is one of the active points in the list of cross-region active points ;

[0025] S402. Traverse all cross-region active points and count the number of times each feature pair appears in the active points and , is the number of cross-region active points. The more times it appears, it means that this feature pair appears frequently in the active points and has a strong correlation with the abnormal situation of polar activity;

[0026] S403. Determine the high-frequency association threshold and the low-efficiency association threshold , and , for determining whether a feature pair is a high-frequency associated feature pair for determining whether a feature pair is a low-efficiency associated feature pair;

[0027] S404. Dynamically adjust the feature pairs, including enhancing the high-frequency associations and eliminating the low-efficiency associations;

[0028] S405. After dynamic adjustment, obtain an updated feature association relationship table.

[0029] Furthermore, the S403 is based on the high-frequency association threshold and the low-efficiency association threshold . When the occurrence times of the feature pair , enhance the association strength of the feature pair in the association feature table. When , eliminate the feature pair from the association feature table. The updated association feature table , where is the enhancement coefficient, is the feature pair in the association feature set of the association strength.

[0030] Compared with the prior art, the method for analyzing polar activity based on multi-domain public opinion big data has the following beneficial effects:

[0031] First, the method for analyzing polar activity based on multi-domain public opinion big data proposed by the present invention can significantly improve the depth and breadth of mining polar public opinion information. Through effective screening and analysis of polar sparse features, it is no longer limited to the dependence on high-frequency features, but can capture those low-frequency but key information, and through association analysis, mine the potential connections between it and other features, enhance the monitoring intensity of data, and ensure timely grasp of public opinion dynamics.

[0032] Second, in terms of the processing of public opinion field data, by accurately dividing the dense area, transition area and sparse area of the data and marking the abnormally active points, this enables us to clearly understand the overall situation and local changes of the public opinion field, timely discover potential hot events and risks, and predict the development trend of events by analyzing the distribution and association features of the abnormally active points, providing timely and accurate basis for decision-making.

[0033] Other advantages, objects, and features of the present invention will be set forth in part in the following description, and in part will be obvious to those skilled in the art upon examination of the following, or may be learned from the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0035] Figure 1 It is a flowchart of a method for analyzing polar activity based on multi-domain public opinion big data;

[0036] Figure 2 It is a schematic diagram of a dynamic self-optimizing closed loop for polar activity analysis in Embodiment 2. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following will, in conjunction with the accompanying drawings and preferred embodiments, describe in detail the specific embodiments, structures, features, and effects of the present invention as follows.

[0038] Embodiment 1

[0039] As Figure 1 shown, this embodiment elaborates in detail the specific application process of a method for analyzing polar activity based on multi-domain public opinion big data. This method collects multi-domain polar public opinion big data, realizes the screening and analysis of sparse polar features, the spatial density clustering processing of public opinion field data, and then mines the implicit associations between features, and finally obtains the polar activity analysis result.

[0040] First, enter the data collection and preprocessing stage (S100). Collect polar public opinion big data from multiple channels. These data cover various forms such as text, pictures, and videos, and have a wide range of sources and multi-domain characteristics. There is a large amount of duplicate, incorrect, and irrelevant information in the collected data. To ensure the accuracy of subsequent analysis, it is necessary to clean it. Remove duplicate data from each piece of data and ensure that the data format is unified. For different data types, remove data outside the range. Through keyword matching and semantic analysis techniques, remove data irrelevant to polar public opinion. For example, use a predefined polar-related keyword library to check whether the data contains these keywords. If it does not contain and is also irrelevant to the polar topic after semantic analysis, it will be removed. For the cleaned data, calculate each feature Occurrence frequency , in order to identify those features with relatively low occurrence frequencies in the entire public opinion dataset, because these low-frequency features contain important information that is not captured by conventional analysis methods. According to , where represents the number of occurrences of feature , which is statistically obtained by traversing the entire dataset,[ represents the total number of features in the dataset, which is obtained by counting all features. Combining the occurrence frequency threshold , retain the features that satisfy . Here, the threshold is used to distinguish low-frequency features from high-frequency features. Then, calculate the distribution entropy of feature . The distribution entropy is used to measure the dispersion degree of a feature across different data sources. When a feature is evenly distributed across various data sources, it indicates that it is involved in multiple data sources and has a more extensive representativeness. According to , where is 's distribution probability in the th data source, which is obtained by calculating the proportion of the number of occurrences of in each data source to the total number of occurrences,[ is the total number of data sources, that is, the number of different channels for collecting data. Retain the features with a distribution entropy greater than the set threshold . The threshold is used to filter out features with a relatively high distribution entropy. After filtering, a sparse feature set is obtained, and at the same time, record the geographical coordinates corresponding to each sparse feature to form a geographical coordinate set . This geographical coordinate information will be used in subsequent analyses to deeply explore polar public opinion information in combination with the geographical location dimension.[

[0041] Then, enter the sparse feature vector construction and correlation analysis stage (S200), and further process the sparse feature set obtained in S100. Convert each sparse feature into a feature vector for subsequent correlation analysis, that is, for the public opinion data feature set with different words. For each sparse feature , construct its feature vector . Each dimension in vector represents the occurrence frequency of the corresponding word in . Through the vocabulary set , for the sparse feature , the th dimension of its feature vector is the word The number of occurrences of in , that is represents the vocabulary in the sparse feature . It is obtained by counting the vocabulary in . In this way, each sparse feature is transformed into a feature vector with dimensions. These vectors will serve as the basic data results for subsequent correlation analysis. Calculate the similarity between different sparse feature vectors to measure the implicit correlation degree between features. Use the method of calculating the cosine value of the included angle between two vectors to judge the similarity degree of sparse features, and then judge the implicit correlation strength between two sparse features. That is, for two sparse features and , the corresponding feature vectors and , their similarity is calculated by . Among them, is the dot product of vectors, which is obtained by multiplying the corresponding dimension elements and then summing. and are the norms of vectors and respectively, calculated according to the calculation of the vector norm . When the similarity ≥β, it is considered that there is an implicit correlation between features and . The threshold here is used to judge whether there is a significant implicit correlation between features, record these two features and their corresponding cosine similarity values , and construct the associated feature set . The associated feature set records all feature pairs that meet the implicit correlation conditions and their similarities, providing key association information for subsequent analysis.

[0042] Subsequently, it enters the stage of density clustering processing of public opinion field data space (S300). The remaining data after extracting polar sparse features by S100 is used for data sampling. This data sampling includes stratified sampling and random sampling. Among them, the stratified sampling stratifies according to the source, time, and theme category of the remaining public opinion data, and retains the representative data of each layer. For the random sampling, for each subset after stratification, data points are extracted according to a sampling ratio of 20% to form a preliminary sampling data set. The preliminary sampling data set is randomly sampled again, and 80% of the data is extracted from the preliminary sampling data set to form the final sampling data set. The final sampling data set is subjected to feature vectorization processing using the construction of sparse feature vectors in S200, and the feature vectors are normalized. The sampled data points will be used as the input for density clustering processing in the space to analyze the distribution of public opinion field data. Each data point is represented by its feature vector, and the similarity between data points is also measured by calculating the cosine similarity between vectors. For data points and the corresponding feature vectors are respectively and , and their similarity . The calculation method is the same as the calculation method of sparse feature vector similarity. In this way, a similarity matrix between data points is obtained for subsequent clustering analysis. According to the similarity and neighborhood relationship between data points, the data points are divided into core points, boundary points, and noise points. Set the minimum neighborhood point number threshold and the neighborhood radius threshold for the data point , and its neighborhood . If , then is a core point. If is not a core point, but there exists a core point such that , then is a boundary point, otherwise is a noise point. Using the distribution of data points, the region is divided into a dense area, a transition area, and a sparse area. When the number of core points in a region is large and the similarity is greater than the similarity mean, it is a dense area, which means that the data points in this region are closely clustered and the public opinion information is relatively concentrated. When core points and boundary points are mixed and the similarity is equal to the similarity mean, it is a transition area, and this region is in a transitional state between the dense area and the sparse area. When there are many noise points and the similarity is less than the similarity mean, it is a sparse area, indicating that the data points in this region are relatively scattered and the public opinion information is relatively less. Set the activity threshold , for the data point , when its activity , then it is marked as an abnormally active point. At the same time, record the abnormally active points across regions and generate a list of cross-region active points These abnormally active points and cross-regional active points represent hot events or important developments in polar public opinion.

[0043] Next, we enter the stage of correlation analysis and dynamic adjustment of correlation features and active points (S400), where we conduct correlation analysis on the correlation feature set obtained in S200 and the cross-region active point list in S300, further explore feature pairs related to abnormal polar activity, and dynamically adjust the correlation features. We compare the feature pairs in the feature correlation set obtained in S200 with the public opinion content involved in each active point in the cross-region active point list generated in S300 one by one. and active points , define the matching function , by traversing all cross-region active points, counting each feature pair Number of times it appears in active points ,according to ,in is the number of cross-region active points. The more times it appears, the more frequently the feature pair appears in the active points, and the stronger the correlation with the abnormal polar activity. Then, the high-frequency correlation threshold is determined. and inefficient association threshold ,and , Used to determine whether a feature pair is a high-frequency associated feature pair. Used to determine whether a feature pair is an inefficiently associated feature pair. When , the association strength of the feature pair in the association feature table is enhanced; when When , the feature pair is removed from the associated feature table, and the updated associated feature table ,in It is an enhancement coefficient used to strengthen the correlation strength of high-frequency correlation feature pairs, making them more influential in subsequent analysis. is the feature pair in the associated feature set The correlation strength is dynamically adjusted to obtain an updated feature correlation table. This updated table removes the feature pairs that are weakly correlated with the polar activity anomalies, strengthens the role of the strongly correlated feature pairs, and more accurately reflects the feature correlation related to polar activity, providing more reliable data support for subsequent polar activity analysis.

[0044] Finally, enter the polar activity analysis stage (S500). Map the high-frequency associated feature pairs in the updated associated feature table to the abnormal activity points, analyze the relationship between the public opinion information represented by these feature pairs and the abnormal activity points, and obtain the polar activity analysis results. That is, for each abnormal activity point, find the corresponding high-frequency associated feature pairs, analyze the public opinion information represented by these high-frequency associated feature pairs, and combine the relevant attributes of the abnormal activity points (such as geographical location, activity value, etc.) to explore the internal connection between them. Considering the mapping relationship between all abnormal activity points and high-frequency associated feature pairs comprehensively, evaluate the polar activity from multiple dimensions.

[0045] In summary, this embodiment details the complete implementation process of a polar activity analysis method based on multi-domain public opinion big data. Starting from data collection and preprocessing, by screening polar sparse features, constructing feature vectors and analyzing their associations, then performing spatial density clustering on the data in the public opinion field, as well as the association analysis and dynamic adjustment of associated features and active points, finally, the accurate analysis of polar activity is achieved. In this process, the method can effectively mine the potential information in polar public opinion data, capture low-frequency but key features, clearly show the situation and changes in the public opinion field, and provide comprehensive and accurate data basis.

[0046] Embodiment 2

[0047] Based on Embodiment 1, this embodiment provides the specific steps of a polar activity analysis method based on multi-domain public opinion big data in the public opinion analysis of polar resource development projects, as Figure 2 shown, a dynamic self-optimizing closed loop for polar activity analysis is constructed through the two-way drive of feature mapping, thermal zoning and abnormal activity points.

[0048] In specific implementation, the specific steps of a polar activity analysis method based on multi-domain public opinion big data in the public opinion analysis of polar resource development projects are as follows:

[0049] Multi-platform collection: Determine that this analysis focuses on the oil resource development project in the polar region, and collect public opinion data related to the oil resource development project from channels such as social media platforms, energy information websites, and forums;

[0050] Keyword setting: Set keywords such as "polar oil development", "project name", "region name", etc. to ensure that the collected data is closely related to the target project;

[0051] Duplicate removal processing: Perform duplicate removal on the collected data, check whether there are exactly the same contents, and remove duplicate data entries to avoid repeated analysis;

[0052] Invalid data elimination: Use keyword matching and semantic analysis to eliminate data unrelated to polar oil development projects;

[0053] Initial frequency screening: Identify features with relatively low frequencies;

[0054] Distribution evaluation: Examine the distribution of low-frequency features across different data sources. When a low-frequency feature appears in multiple data sources, it indicates that although its frequency is low, it has a certain degree of universality, and it is retained as a polar sparse feature;

[0055] Vector transformation: For each selected polar sparse feature, construct a feature vector based on its occurrence in different data;

[0056] Association mining: By comparing the similarity of different feature vectors, explore the potential relationships between features. If two feature vectors are similar, it means that the corresponding sparse features have an implicit association, thus constructing an association feature set;

[0057] Vector representation of data points: For the remaining data after sparse feature screening, take each piece of data as a unit and transform it into a feature vector based on the keywords, semantic themes, etc. it contains;

[0058] Similarity calculation and clustering: Calculate the similarity between the feature vectors of these data points and cluster the data points with high similarity together;

[0059] Dynamic density heat map zoning: Based on the clustering of data points, divide into dense areas, transition areas, and sparse areas;

[0060] Marking of active points: Set a measure of activity, mark the abnormally active points by comparing with the measure, and record the abnormally active points across regions to generate a list of cross-region active points;

[0061] Abnormal association mining: Compare the constructed association feature set with the list of cross-region active points one by one to check whether the content of each cross-region active point contains the feature pairs in the association feature set;

[0062] Frequency statistics: Count the number of times each feature pair appears in the cross-region active points to measure its frequency of occurrence;

[0063] Threshold setting and adjustment: According to the statistical results, set a high-frequency association threshold and a low-efficiency association threshold. For feature pairs with the number of occurrences higher than the high-frequency association threshold, enhance their importance in association analysis. For feature pairs with the number of occurrences lower than the low-efficiency association threshold, remove them from the association feature set;

[0064] Update the association feature table: After dynamic adjustment, obtain an association feature relationship table that more accurately reflects the public opinion of polar oil development projects;

[0065] Feature pair and active point mapping: Map and analyze the high-frequency associated feature pairs in the updated associated feature relationship table with the abnormal active points, and feedback the remaining data to vector conversion and re-optimize it;

[0066] Activity evaluation: Based on the mapping relationships between all abnormal active points and high-frequency associated feature pairs, evaluate the activity of the polar oil development project from multiple dimensions, and finally obtain the activity analysis result of the polar region in this resource development project, providing a reference basis for project decision-making.

[0067] The above is only a preferred embodiment of the present invention, and does not impose any form of limitation on the present invention. Although the present invention has been disclosed above with a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments by using the above-disclosed technical content without departing from the technical solution of the present invention. However, as long as it does not depart from the technical solution content of the present invention, any brief modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. A polar activity analysis method based on multi-field public opinion big data, characterized in that: The specific steps of this method are: S100. Collect multi-domain polar public opinion big data, clean the collected data, remove duplicate, erroneous and irrelevant information, and filter out polar sparse features about public opinion from the processed public opinion data set, and output a sparse feature list and original geographic coordinates ; S200, mapping the polar sparse features into feature vectors, mining implicit associations between different polar sparse features, and obtaining an associated feature set T; S300, sampling the remaining public opinion data with extremely sparse features after screening, stratifying the data features and performing spatial density clustering, calculating the similarity between data points, and comparing the similarity with the similarity mean to divide the areas into three categories, including dense areas, transition areas, and sparse areas, and marking abnormally active points, generating a thermal partition map and a list of cross-region active points. ; S400, performing correlation analysis on the correlation feature set obtained in S200 and the cross-region active point list in S300, counting the feature pairs that frequently appear in the active points, performing dynamic adjustment, and outputting an updated feature correlation relationship table; S500, mapping the high-frequency correlation feature pairs in the updated correlation feature table with the abnormally active points, finding the high-frequency correlation feature pairs corresponding to each abnormally active point, analyzing the connection between the public opinion information represented by these feature pairs and the active point, and obtaining the polar activity analysis results.

2. According to claim 1, a polar activity analysis method based on multi-field public opinion big data is characterized in that: The S100 calculates each feature The frequency of occurrence of this feature is used to measure The frequency of occurrence in the entire public opinion data set is obtained by calculating the ratio of the number of occurrences of each feature to the total number of features in the data set. ,Right now ,in, Representation characteristics The number of occurrences, T represents the total number of features in the data set, combined with the frequency threshold , keep satisfied .

3. According to claim 2, a polar activity analysis method based on multi-field public opinion big data is characterized in that: The S100 computing features The distribution entropy of , the distribution entropy is used to measure the degree of dispersion of features in different data sources. When the distribution of features in each data source is uniform, that is, the distribution entropy is high, the distribution entropy greater than the set threshold is retained. The characteristic of the distribution entropy ,in for In the The distribution probability of data sources, M is the total number of data sources, and the features, and obtain a sparse feature set , and record the geographic coordinates corresponding to each sparse feature to form a geographic coordinate set .

4. According to claim 1, a polar activity analysis method based on multi-field public opinion big data is characterized in that: S200 collects the sparse features processed by S100 Each sparse feature in is converted into a feature vector, and the process is as follows: public opinion data feature sets with different vocabulary, each sparse feature The corresponding vector Each dimension in represents the word The frequency of occurrence in the vocabulary set , for sparse features , Its eigenvector No. Dimensions For vocabulary exist The number of occurrences in ,in Words to express In sparse features The number of occurrences in .

5. According to claim 4, a polar activity analysis method based on multi-field public opinion big data is characterized in that: The S200 calculates the similarity between different sparse feature vectors to measure the implicit association degree between features. The similarity of sparse features is determined by calculating the cosine value of the angle between two vectors, thereby determining the implicit association strength between two sparse features. and The corresponding eigenvector and Similarity between for: ,in is the vector dot product, and They are vectors and The modulus of >β, the characteristic and There is an implicit association between the two features and their corresponding cosine similarity values. , and construct the associated feature set T.

6. According to claim 1, a polar activity analysis method based on multi-field public opinion big data is characterized in that: The S300 uses the remaining data after S100 extracts the polar sparse features to perform data sampling, and the data sampling includes stratified sampling and random sampling, wherein the stratified sampling is stratified according to the source, time and subject category of the remaining public opinion data, and the representative data of each level is retained. For each stratified subset, the random sampling extracts data points according to a sampling ratio of 20% to form a preliminary sampling data set, and the preliminary sampling data set is randomly sampled again, and 80% of the data is extracted from the preliminary sampling data set to form a final sampling data set. The final sampling data set is subjected to feature vectorization processing using the construction of the sparse feature vector in S200, and the feature vector is normalized. The sampled data points will be used as input for spatial density clustering processing; The spatial density clustering calculates the similarity between data points, each data point is represented by its feature vector, and the similarity between data points is measured by calculating the cosine similarity between vectors. According to the similarity and neighborhood relationship between data points, the data points are divided into core points, boundary points and noise points. The distribution of data points is used to divide the area into dense area, transition area and sparse area. That is, when there are many core points in an area and the similarity is greater than the similarity mean, it is a dense area. When the core points and boundary points are mixed and the similarity is equal to the similarity mean, it is a transition area. When there are many noise points and the similarity is less than the similarity mean, it is a sparse area. Count the abnormally active points, generate a list of cross-region active points, and combine them with the original geographic coordinates to display the specific locations in the thermal zoning map.

7. According to claim 1, a polar activity analysis method based on multi-field public opinion big data is characterized in that: The specific steps of S400 are: S401, compare the feature pairs in the feature association set obtained in S200 with the public opinion content involved in each active point in the cross-region active point list generated in S300 one by one, and for each cross-region active point, check whether its public opinion text contains the feature pair in the feature association set, that is, for the feature pair and active points , define the matching function , where active points It is a list of cross-region active points S402, traverse all cross-region active points, and count each feature pair Number of times it appears in active points and , is the number of cross-region active points. The more times it appears, the more frequent this feature pair appears in the active points, which is strongly correlated with the anomalies of polar activity; S403: Determine high frequency correlation threshold and inefficient association threshold ,and , Used to determine whether a feature pair is a high-frequency associated feature pair. Used to determine whether a feature pair is an inefficiently associated feature pair; S404, dynamically adjusting the feature pairs, including enhancing high-frequency associations and eliminating low-efficiency associations; S405: After dynamic adjustment, an updated feature association relationship table is obtained.

8. According to claim 7, a polar activity analysis method based on multi-field public opinion big data is characterized in that: S403 is based on the high frequency correlation threshold and inefficient association threshold , when the number of occurrences of feature pairs When , the association strength of the feature pair in the association feature table is enhanced. When , the feature pair is removed from the associated feature table, and the updated associated feature table ,in is the enhancement factor, is the feature pair in the associated feature set The strength of association.

Citation Information

Patent Citations

  • Online public opinion monitoring method for enterprise crisis public gateway

    CN112632218A

  • Traffic safety public opinion analysis method based on SQ-LDA topic model

    CN115757776A