Polar region activeness analysis method based on multi-field public opinion big data

Through the polar activity analysis method based on multi-field public opinion big data, the problem that the existing technology cannot effectively capture the cross-domain correlation of low-frequency key information and models is solved, and in-depth mining and dynamic monitoring of polar public opinion is achieved, and accurate event prediction and decision-making support is provided.

CN119989236AActive Publication Date: 2025-05-13POLAR RES INST OF CHINA

Patent Information

Application Number
CN202510436066.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-13
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

Existing polar public opinion data analysis technology cannot effectively capture low-frequency key information, it is difficult to model cross-domain correlation, and traditional static models cannot adapt to the needs of rapid data changes.

Method used

The polar activity analysis method based on multi-field public opinion big data is adopted to capture low-frequency information through screening and analysis of polar sparse characteristics, and the potential connections between the characteristics are mined through correlation analysis to predict the development trend of events. Specific steps include data cleaning, feature mapping, spatial density clustering, correlation feature analysis and dynamic adjustment.

Benefits of technology

It significantly improves the depth and breadth of polar public opinion information, can capture low-frequency but critical information, clearly understand the situation in the public opinion field, timely discover potential hot spots and risks, predict event development trends, and provide a timely and accurate basis for decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989236A_ABST
    Figure CN119989236A_ABST
Patent Text Reader

Abstract

The invention discloses a polar region activeness analysis method based on multi-field public opinion big data, and relates to the technical field of data processing, and the method comprises the following specific steps: S100, screening polar region sparse features about public opinions, S200, mining implicit association among different polar region sparse features, S300, marking abnormal active points to generate a thermodynamic partition map, and S400, extracting the thermodynamic partition map; according to the polar region activeness analysis method based on the multi-field public opinion big data, the mining depth and breadth of polar region public opinion information can be remarkably improved, the polar region activeness can be effectively analyzed by effectively screening and analyzing polar region sparse features, and the polar region activeness can be effectively analyzed. According to the method, the dependence on high-frequency features is not limited any more, low-frequency but key information can be captured, potential relations between the low-frequency but key information and other features are mined through correlation analysis, the monitoring strength on data is enhanced, and it is ensured that the dynamic state of public opinions is mastered at the first time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a polar activity analysis method based on multi-field public opinion big data. Background Art

[0002] Polar activity analysis is a core means of monitoring environmental changes, resource development, and geopolitical dynamics in polar regions. Its core task is to identify the spatiotemporal evolution patterns and cross-domain correlations of polar events by integrating multi-source heterogeneous data. With the increasing complexity of polar activities, existing public opinion analysis tools and technologies often cannot meet these needs, especially when faced with low-frequency keywords and implicit associations.

[0003] At present, traditional analysis methods face three challenges: First, polar data has natural sparseness and long-tail distribution characteristics, which makes key signals easily submerged by noise; Second, the cross-domain correlation of polar events is difficult to model using single-modality data; Third, the dynamic evolution of polar activity requires real-time analysis capabilities, while traditional static models cannot adapt to the rapid drift of data distribution.

[0004] In summary, the existing polar public opinion data analysis technology has many shortcomings and cannot meet the needs of comprehensive and in-depth analysis of the activity in the polar regions. Therefore, a new analysis method is needed to overcome these shortcomings and improve the accuracy and effectiveness of polar activity analysis. Summary of the invention

[0005] The purpose of the present invention is to make up for the shortcomings of the prior art and to provide a polar activity analysis method based on multi-field public opinion big data. It can effectively screen and analyze polar sparse features, no longer be limited to relying on high-frequency features, but can capture low-frequency but critical information, and through association analysis, explore the potential connection between it and other features, and predict the development trend of events.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: a polar activity analysis method based on multi-field public opinion big data, the specific steps of the method are: S100. Collect multi-domain polar public opinion big data, clean the collected data, remove duplicate, erroneous and irrelevant information, and filter out polar sparse features about public opinion from the processed public opinion data set, and output a sparse feature list and original geographic coordinates ; S200, mapping the polar sparse features into feature vectors, mining implicit associations between different polar sparse features, and obtaining an associated feature set T; S300, sampling the remaining public opinion data with extremely sparse features after screening, stratifying the data features and performing spatial density clustering, calculating the similarity between data points, and comparing the similarity with the similarity mean to divide the areas into three categories, including dense areas, transition areas, and sparse areas, and marking abnormally active points, generating a thermal partition map and a list of cross-region active points. ; S400, performing correlation analysis on the correlation feature set obtained in S200 and the cross-region active point list in S300, counting the feature pairs that frequently appear in the active points, performing dynamic adjustment, and outputting an updated feature correlation relationship table; S500, mapping the high-frequency correlation feature pairs in the updated correlation feature table with the abnormally active points, finding the high-frequency correlation feature pairs corresponding to each abnormally active point, analyzing the connection between the public opinion information represented by these feature pairs and the active point, and obtaining the polar activity analysis results.

[0007] Furthermore, the S100 calculates each feature The frequency of occurrence of this feature is used to measure The frequency of occurrence in the entire public opinion data set, The frequency of occurrence of each feature is obtained by calculating the ratio of the number of occurrences of each feature to the total number of features in the data set. ,Right now ,in, Representation characteristics The number of occurrences, T represents the total number of features in the data set, combined with the frequency threshold , keep satisfied .

[0008] Furthermore, the S100 calculates the feature The distribution entropy of , distribution entropy is used to measure the degree of dispersion of features in different data sources. When the distribution of features in various data sources is uniform, that is, the distribution entropy is high, the distribution entropy greater than the set threshold is retained. The characteristic of the distribution entropy ,in for In the The distribution probability of data sources, M is the total number of data sources, and the features, and obtain a sparse feature set , and record the geographic coordinates corresponding to each sparse feature to form a geographic coordinate set .

[0009] Furthermore, the S200 sets the sparse feature set processed by S100 Each sparse feature in is converted into a feature vector, and the process is as follows: public opinion data feature sets with different vocabulary, each sparse feature The corresponding vector Each dimension in represents the word The frequency of occurrence in the vocabulary set , for sparse features , whose eigenvector No. Dimensions For vocabulary exist The number of occurrences in ,in Words to express In sparse features The number of occurrences in .

[0010] Furthermore, the S200 calculates the similarity between different sparse feature vectors to measure the implicit association degree between the features, and determines the similarity of the sparse features by calculating the cosine value of the angle between the two vectors, thereby determining the implicit association strength between the two sparse features. and The corresponding eigenvector and Similarity between for: ,in is the vector dot product, and They are vectors and The modulus of >β, the characteristic and There is an implicit association between the two features and their corresponding cosine similarity values. , and construct the associated feature set T.

[0011] Furthermore, the S300 performs data sampling using the remaining data after the polar sparse features are extracted in S100. The data sampling includes stratified sampling and random sampling. The stratified sampling is stratified according to the source, time and subject category of the remaining public opinion data, and representative data of each level is retained. The random sampling extracts data points according to a sampling ratio of 20% for each stratified subset to form a preliminary sampling data set. The preliminary sampling data set is randomly sampled again, and 80% of the data is extracted from the preliminary sampling data set to form a final sampling data set. The final sampling data set is subjected to feature vectorization processing using the construction of the sparse feature vector in S200, and the feature vector is normalized. The sampled data points will be used as input for spatial density clustering processing. The spatial density clustering calculates the similarity between data points, each data point is represented by its feature vector, and the similarity between data points is measured by calculating the cosine similarity between vectors. According to the similarity and neighborhood relationship between data points, the data points are divided into core points, boundary points and noise points. The distribution of data points is used to divide the area into dense area, transition area and sparse area. That is, when there are many core points in an area and the similarity is greater than the similarity mean, it is a dense area. When the core points and boundary points are mixed and the similarity is equal to the similarity mean, it is a transition area. When there are many noise points and the similarity is less than the similarity mean, it is a sparse area. Count the abnormally active points, generate a list of cross-region active points, and combine them with the original geographic coordinates to display the specific locations in the thermal zoning map.

[0012] Furthermore, the specific steps of S400 are: S401, compare the feature pairs in the feature association set obtained in S200 with the public opinion content involved in each active point in the cross-region active point list generated in S300 one by one, and for each cross-region active point, check whether its public opinion text contains the feature pair in the feature association set, that is, for the feature pair and active points , define the matching function , where active points It is a list of cross-region active points One of the active points in S402, traverse all cross-region active points, and count each feature pair Number of times it appears in active points and , is the number of cross-region active points. The more times it appears, the more frequent this feature pair appears in the active points, which is strongly correlated with the anomalies of polar activity; S403: Determine high frequency correlation threshold and inefficient association threshold ,and , Used to determine whether a feature pair is a high-frequency associated feature pair. Used to determine whether a feature pair is an inefficiently associated feature pair; S404, dynamically adjusting the feature pairs, including enhancing high-frequency associations and eliminating low-efficiency associations; S405: After dynamic adjustment, an updated feature association relationship table is obtained.

[0013] Furthermore, the step S403 is based on the high frequency correlation threshold and inefficient association threshold , when the number of occurrences of feature pairs When , the association strength of the feature pair in the association feature table is enhanced. When , the feature pair is removed from the associated feature table, and the updated associated feature table ,in is the enhancement factor, is the feature pair in the associated feature set The strength of association.

[0014] Compared with the existing technology, this polar activity analysis method based on multi-field public opinion big data has the following beneficial effects: 1. The polar activity analysis method based on multi-field public opinion big data proposed in the present invention can significantly improve the depth and breadth of mining polar public opinion information. Through the effective screening and analysis of polar sparse features, it is no longer limited to the reliance on high-frequency features, but can capture those low-frequency but critical information, and through correlation analysis, it can explore the potential connection between it and other features, enhance the monitoring of data, and ensure that the public opinion dynamics are grasped at the first time.

[0015] 2. In terms of processing public opinion data, the present invention accurately divides the data into dense areas, transition areas and sparse areas, and marks abnormal active points. This enables us to clearly understand the overall situation and local changes in the public opinion field, timely discover potential hot events and risks, and predict the development trend of events by analyzing the distribution and correlation characteristics of abnormal active points, providing timely and accurate basis for decision-making.

[0016] Other advantages, objectives and features of the present invention will be set forth in part in the following description and, in part, will be apparent to those skilled in the art based on an examination of the following or may be taught from the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0018] Figure 1 This is a flow chart of a polar activity analysis method based on multi-domain public opinion big data; Figure 2 Schematic diagram of the dynamic self-optimization closed loop of polar activity analysis in Example 2. DETAILED DESCRIPTION

[0019] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the specific implementation mode, structure, characteristics and effects of the present invention are described in detail below in combination with the accompanying drawings and preferred embodiments.

[0020] Embodiment 1 like Figure 1 As shown, this embodiment elaborates on the specific application process of a polar activity analysis method based on multi-field public opinion big data. The method collects multi-field polar public opinion big data, realizes the screening and analysis of polar sparse features, and the spatial density clustering processing of public opinion field data, and then mines the implicit correlation between features, and finally obtains the polar activity analysis results.

[0021] First, enter the data collection preprocessing stage (S100), collect polar public opinion big data from multiple channels. These data cover multiple forms such as text, pictures, videos, etc., and are from a wide range of sources and have multi-field characteristics. There are a lot of duplicate, erroneous and irrelevant information in the collected data. In order to ensure the accuracy of subsequent analysis, it is necessary to clean it, remove duplicate data from each data, and ensure that the data format is unified. For different data types, remove data that exceeds the range, and remove data that is not related to polar public opinion through keyword matching and semantic analysis technology. For example, use the predefined polar-related keyword library to check whether the data contains these keywords. If not and the data is not related to polar topics after semantic analysis, it will be removed. For the cleaned data, calculate each feature Frequency of occurrence In order to find out the features that appear less frequently in the entire public opinion dataset, because these low-frequency features contain important information that is not paid attention to by conventional analysis methods, according to ,in Representation characteristics The number of occurrences of is counted by traversing the entire data set. Represents the total number of features in the dataset, obtained by counting all features and combining them with the frequency threshold , keep satisfied The feature, here the threshold Used to distinguish low-frequency features from high-frequency features, and then calculate the features The distribution entropy of ,Distribution entropy is used to measure the degree of dispersion of features in different data sources. When a feature is evenly distributed in various data sources, it means that it is involved in multiple data sources and has a wider range of representativeness. ,in for In the The distribution probability of the data source is calculated by The proportion of the number of occurrences in each data source to the total number of occurrences is obtained. is the total number of data sources, that is, the number of different channels for collecting data, and the distribution entropy is retained to be greater than the set threshold The feature, threshold Used to filter out features with higher distribution entropy. After filtering, a sparse feature set is obtained. , and record the geographic coordinates corresponding to each sparse feature to form a geographic coordinate set ,These geographic coordinate information will be used in subsequent analysis to combine the geographic location dimension to deeply explore polar public opinion information.

[0022] Then, the sparse feature vector construction and association analysis phase (S200) is entered to construct the sparse feature vector obtained in S100. Further processing is performed to convert each sparse feature into a feature vector for subsequent association analysis, that is, public opinion data feature sets with different vocabulary, for each sparse feature , construct its feature vector ,vector Each dimension in represents the corresponding word in The frequency of occurrence in the vocabulary set , for sparse features , whose eigenvector No. Dimensions For vocabulary exist The number of occurrences in ,in Words to express In sparse features The number of occurrences in The vocabulary is counted in , so that each sparse feature is converted into a The vectors are used as the basic data results of the subsequent association analysis to calculate the similarity between different sparse feature vectors to measure the implicit association between features. The method of calculating the cosine value of the angle between two vectors is used to judge the similarity of sparse features, and then to judge the implicit association strength between two sparse features. and The corresponding eigenvector and , the similarity between them Depend on Calculated, among which is the vector dot product, obtained by multiplying the elements of the corresponding dimensions and then summing them. and They are vectors and The modulus of the vector is calculated based on Calculate, when the similarity β, it is considered that the feature and There is an implicit association between Used to determine whether there is a significant implicit association between features, and record the two features and their corresponding cosine similarity values , and build a related collection , associated feature set All feature pairs that meet the implicit association conditions and their similarities are recorded, providing key association information for subsequent analysis.

[0023] Subsequently, the spatial density clustering processing stage (S300) of the public opinion data is entered, and the remaining data after the polar sparse features are extracted by S100 is used for data sampling. The data sampling includes stratified sampling and random sampling, wherein the stratified sampling is stratified according to the source, time and subject category of the remaining public opinion data, and the representative data of each level is retained. For each stratified subset, the random sampling extracts data points according to the sampling ratio of 20% to form a preliminary sampling data set, and the preliminary sampling data set is randomly sampled again, and 80% of the data are extracted from the preliminary sampling data set to form a final sampling data set. The final sampling data set is feature vectorized using the construction of the sparse feature vector in S200, and the feature vector is normalized. The sampled data points will be used as the input of the spatial density clustering processing to analyze the distribution of the public opinion data, and each data point will be represented by its feature vector. The similarity between the data points is also measured by calculating the cosine similarity between the vectors. For the data points and The corresponding eigenvectors are and , their similarity The calculation method is consistent with the sparse feature vector similarity calculation method. In this way, the similarity matrix between data points is obtained for subsequent clustering analysis. According to the similarity and neighborhood relationship between data points, the data points are divided into core points, boundary points and noise points, and the minimum neighborhood point threshold of the core point is set. and neighborhood radius threshold , for the data point , its neighborhood .like ,but As the core point, if Not a core point, but there is a core point Make ,but is a boundary point, otherwise The distribution of data points is used to divide the area into dense area, transition area and sparse area. When there are many core points in an area and the similarity is greater than the mean similarity, it is a dense area, which means that the data points in the area are closely clustered and the public opinion information is relatively concentrated; when the core points and boundary points are mixed and the similarity is equal to the mean similarity, it is a transition area, and the area is in a transition state between dense area and sparse area; when there are many noise points and the similarity is less than the mean similarity, it is a sparse area, which means that the data points in the area are relatively scattered and the public opinion information is relatively small. Set the activity threshold , for the data point , when its activity , it will be marked as an abnormally active point. At the same time, the cross-region abnormally active points are recorded to generate a cross-region active point list These abnormally active points and cross-regional active points represent hot events or important developments in polar public opinion.

[0024] Next, we enter the stage of correlation analysis and dynamic adjustment of correlation features and active points (S400), where we conduct correlation analysis on the correlation feature set obtained in S200 and the cross-region active point list in S300, further explore feature pairs related to abnormal polar activity, and dynamically adjust the correlation features. We compare the feature pairs in the feature correlation set obtained in S200 with the public opinion content involved in each active point in the cross-region active point list generated in S300 one by one. and active points , define the matching function , by traversing all cross-region active points, counting each feature pair Number of times it appears in active points ,according to ,in is the number of cross-region active points. The more times it appears, the more frequently the feature pair appears in the active points, and the stronger the correlation with the abnormal polar activity. Then, the high-frequency correlation threshold is determined. and inefficient association threshold ,and , Used to determine whether a feature pair is a high-frequency associated feature pair. Used to determine whether a feature pair is an inefficiently associated feature pair. When , the association strength of the feature pair in the association feature table is enhanced; when When , the feature pair is removed from the associated feature table, and the updated associated feature table ,in It is an enhancement coefficient used to strengthen the correlation strength of high-frequency correlation feature pairs, making them more influential in subsequent analysis. is the feature pair in the associated feature set The correlation strength is dynamically adjusted to obtain an updated feature correlation table. This updated table removes the feature pairs that are weakly correlated with the polar activity anomalies, strengthens the role of the strongly correlated feature pairs, and more accurately reflects the feature correlation related to polar activity, providing more reliable data support for subsequent polar activity analysis.

[0025] Finally, the polar activity analysis stage (S500) is entered, and the high-frequency correlation feature pairs in the updated correlation feature table are mapped to the abnormally active points, and the relationship between the public opinion information represented by these feature pairs and the abnormally active points is analyzed to obtain the polar activity analysis results, that is, for each abnormally active point, the corresponding high-frequency correlation feature pairs are found, and the public opinion information represented by these high-frequency correlation feature pairs is analyzed. Combined with the relevant attributes of the abnormally active points (such as geographical location, activity value, etc.), the intrinsic connection between them is excavated, and the mapping relationship between all abnormally active points and high-frequency correlation feature pairs is comprehensively considered to evaluate the polar activity from multiple dimensions.

[0026] In summary, this embodiment demonstrates in detail the complete implementation process of a polar activity analysis method based on multi-field public opinion big data, starting from data collection and preprocessing, by screening polar sparse features, constructing feature vectors and analyzing their correlation, and then performing spatial density clustering processing on public opinion field data, as well as correlation analysis and dynamic adjustment of correlation features and active points, and finally achieving accurate analysis of polar activity. In this process, the method can effectively mine potential information in polar public opinion data, capture low-frequency but key features, clearly show the situation and changes of the public opinion field, and provide a comprehensive and accurate data basis.

[0027] Embodiment 2 On the basis of Example 1, this example provides a method for analyzing polar activity based on multi-field public opinion big data in specific steps of public opinion analysis of polar resource development projects, such as Figure 2 As shown in the figure, a dynamic self-optimizing closed loop of polar activity analysis is constructed through feature mapping, thermal zoning and bidirectional driving of abnormally active points.

[0028] In the specific implementation, a polar activity analysis method based on multi-field public opinion big data has the following specific steps in the public opinion analysis of polar resource development projects: Multi-platform collection: This analysis focuses on oil resource development projects in the polar regions, and collects public opinion data related to oil resource development projects from social media platforms, energy information websites, forums and other channels; Keyword setting: Set keywords, such as "polar oil development", "project name", "region name", etc., to ensure that the collected data is closely related to the target project; Deduplication: Deduplication is performed on the collected data to check whether there is exactly the same content, remove duplicate data entries, and avoid repeated analysis; Elimination of invalid data: Using keyword matching and semantic analysis, eliminate data irrelevant to the polar oil development project; Frequency screening: find out the features with low frequency; Distribution evaluation: Inspect the distribution of low-frequency features in different data sources. When a low-frequency feature appears in multiple data sources, it means that although its frequency is low, it has a certain degree of universality. It is retained as a very sparse feature. Vector transformation: For each selected polar sparse feature, construct a feature vector based on its occurrence in different data; Association mining: By comparing the similarities of different feature vectors, the potential connections between features are mined. If two feature vectors are similar, it means that the corresponding sparse features have implicit associations, thereby constructing an association feature set. Vector representation of data points: The remaining data after sparse feature screening is converted into feature vectors based on the keywords, semantic topics, etc. contained in each data item; Similarity calculation clustering: Calculate the similarity between the feature vectors of these data points and cluster the data points with high similarity together; Dynamic density thermal zoning: based on the aggregation of data points, it divides the dense area, transition area and sparse area; Active point marking: Set activity measurement standards, mark abnormal active points by comparing with the measurement standards, record abnormal active points across regions, and generate a list of cross-region active points; Abnormal association mining: Compare the constructed association feature set with the cross-region active point list one by one to check whether the content of each cross-region active point contains the feature pairs in the association feature set; Frequency statistics: Count the number of times each feature pair appears in cross-region active points to measure its frequency of occurrence; Threshold setting and adjustment: According to the statistical results, set the high-frequency association threshold and the low-efficiency association threshold. For feature pairs with a frequency higher than the high-frequency association threshold, enhance their importance in the association analysis. For feature pairs with a frequency lower than the low-efficiency association threshold, remove them from the association feature set. Update the correlation feature table: After dynamic adjustment, a correlation feature relationship table that more accurately reflects the public opinion of the polar oil development project is obtained; Mapping feature pairs with active points: Map and analyze the high-frequency associated feature pairs in the updated associated feature relationship table with abnormal active points, and feed the remaining data back to vector conversion and re-optimize; Activity assessment: By comprehensively mapping the relationship between all abnormally active points and high-frequency associated feature pairs, the activity of polar oil development projects is assessed from multiple dimensions, and finally the activity analysis results of the polar regions in the resource development projects are obtained, providing a reference for project decision-making.

[0029] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technical personnel in this field can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.

Claims

1. A polar activity analysis method based on multi-field public opinion big data, characterized in that: The specific steps of this method are: S100. Collect multi-domain polar public opinion big data, clean the collected data, remove duplicate, erroneous and irrelevant information, and filter out polar sparse features about public opinion from the processed public opinion data set, and output a sparse feature list and original geographic coordinates ; S200, mapping the polar sparse features into feature vectors, mining implicit associations between different polar sparse features, and obtaining an associated feature set T; S300, sampling the remaining public opinion data with extremely sparse features after screening, stratifying the data features and performing spatial density clustering, calculating the similarity between data points, and comparing the similarity with the similarity mean to divide the areas into three categories, including dense areas, transition areas, and sparse areas, and marking abnormally active points, generating a thermal partition map and a list of cross-region active points. ; S400, performing correlation analysis on the correlation feature set obtained in S200 and the cross-region active point list in S300, counting the feature pairs that frequently appear in the active points, performing dynamic adjustment, and outputting an updated feature correlation relationship table; S500, mapping the high-frequency correlation feature pairs in the updated correlation feature table with the abnormally active points, finding the high-frequency correlation feature pairs corresponding to each abnormally active point, analyzing the connection between the public opinion information represented by these feature pairs and the active point, and obtaining the polar activity analysis results.

2. According to claim 1, a polar activity analysis method based on multi-field public opinion big data is characterized in that: The S100 calculates each feature The frequency of occurrence of this feature is used to measure The frequency of occurrence in the entire public opinion data set is obtained by calculating the ratio of the number of occurrences of each feature to the total number of features in the data set. ,Right now ,in, Representation characteristics The number of occurrences, T represents the total number of features in the data set, combined with the frequency threshold , keep satisfied .

3. According to claim 2, a polar activity analysis method based on multi-field public opinion big data is characterized in that: The S100 computing features The distribution entropy of , the distribution entropy is used to measure the degree of dispersion of features in different data sources. When the distribution of features in each data source is uniform, that is, the distribution entropy is high, the distribution entropy greater than the set threshold is retained. The characteristic of the distribution entropy ,in for In the The distribution probability of data sources, M is the total number of data sources, and the features, and obtain a sparse feature set , and record the geographic coordinates corresponding to each sparse feature to form a geographic coordinate set .

4. According to claim 1, a polar activity analysis method based on multi-field public opinion big data is characterized in that: S200 collects the sparse features processed by S100 Each sparse feature in is converted into a feature vector, and the process is as follows: public opinion data feature sets with different vocabulary, each sparse feature The corresponding vector Each dimension in represents the word The frequency of occurrence in the vocabulary set , for sparse features , Its eigenvector No. Dimensions For vocabulary exist The number of occurrences in ,in Words to express In sparse features The number of occurrences in .

5. According to claim 4, a polar activity analysis method based on multi-field public opinion big data is characterized in that: The S200 calculates the similarity between different sparse feature vectors to measure the implicit association degree between features. The similarity of sparse features is determined by calculating the cosine value of the angle between two vectors, thereby determining the implicit association strength between two sparse features. and The corresponding eigenvector and Similarity between for: ,in is the vector dot product, and They are vectors and The modulus of >β, the characteristic and There is an implicit association between the two features and their corresponding cosine similarity values. , and construct the associated feature set T.

6. According to claim 1, a polar activity analysis method based on multi-field public opinion big data is characterized in that: The S300 uses the remaining data after S100 extracts the polar sparse features to perform data sampling, and the data sampling includes stratified sampling and random sampling, wherein the stratified sampling is stratified according to the source, time and subject category of the remaining public opinion data, and the representative data of each level is retained. For each stratified subset, the random sampling extracts data points according to a sampling ratio of 20% to form a preliminary sampling data set, and the preliminary sampling data set is randomly sampled again, and 80% of the data is extracted from the preliminary sampling data set to form a final sampling data set. The final sampling data set is subjected to feature vectorization processing using the construction of the sparse feature vector in S200, and the feature vector is normalized. The sampled data points will be used as input for spatial density clustering processing; The spatial density clustering calculates the similarity between data points, each data point is represented by its feature vector, and the similarity between data points is measured by calculating the cosine similarity between vectors. According to the similarity and neighborhood relationship between data points, the data points are divided into core points, boundary points and noise points. The distribution of data points is used to divide the area into dense area, transition area and sparse area. That is, when there are many core points in an area and the similarity is greater than the similarity mean, it is a dense area. When the core points and boundary points are mixed and the similarity is equal to the similarity mean, it is a transition area. When there are many noise points and the similarity is less than the similarity mean, it is a sparse area. Count the abnormally active points, generate a list of cross-region active points, and combine them with the original geographic coordinates to display the specific locations in the thermal zoning map.

7. According to claim 1, a polar activity analysis method based on multi-field public opinion big data is characterized in that: The specific steps of S400 are: S401, compare the feature pairs in the feature association set obtained in S200 with the public opinion content involved in each active point in the cross-region active point list generated in S300 one by one, and for each cross-region active point, check whether its public opinion text contains the feature pair in the feature association set, that is, for the feature pair and active points , define the matching function , where active points It is a list of cross-region active points S402, traverse all cross-region active points, and count each feature pair Number of times it appears in active points and , is the number of cross-region active points. The more times it appears, the more frequent this feature pair appears in the active points, which is strongly correlated with the anomalies of polar activity; S403: Determine high frequency correlation threshold and inefficient association threshold ,and , Used to determine whether a feature pair is a high-frequency associated feature pair. Used to determine whether a feature pair is an inefficiently associated feature pair; S404, dynamically adjusting the feature pairs, including enhancing high-frequency associations and eliminating low-efficiency associations; S405: After dynamic adjustment, an updated feature association relationship table is obtained.

8. According to claim 7, a polar activity analysis method based on multi-field public opinion big data is characterized in that: S403 is based on the high frequency correlation threshold and inefficient association threshold , when the number of occurrences of feature pairs When , the association strength of the feature pair in the association feature table is enhanced. When , the feature pair is removed from the associated feature table, and the updated associated feature table ,in is the enhancement factor, is the feature pair in the associated feature set The strength of association.

Citation Information

Patent Citations

  • Data analysis method and apparatus

    CN106557558A

  • Modular public opinion monitoring method and system aiming at network public opinion events

    CN109446394A

  • Online public opinion monitoring method for enterprise crisis public gateway

    CN112632218A

  • Traffic safety public opinion analysis method based on SQ-LDA topic model

    CN115757776A

  • Network public opinion analysis method and apparatus, and computer-readable storage medium

    WO2019227710A1

Cited By

  • Method, device and equipment for identifying ground side intranet structure of low earth orbit satellite network

    CN120979955A