Network public opinion situation analysis method and system

By obtaining public opinion data from the target data source, calculating and reordering dark blocks of different matrices and dynamically adjusting the sliding window distance, the problem of insufficient prediction ability of public opinion evolution trends and difficulty in determining the number of clusters in traditional methods is solved, and a more in-depth and forward-looking public opinion situation analysis is achieved.

CN120067720AInactive Publication Date: 2025-05-30NANJING CHENGYI FUTURE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510197946.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional methods of online public opinion situation analysis are difficult to deeply explore the evolution trend of public opinion, lack the ability to predict the evolution trend of public opinion, and lack effective mechanisms to determine the number of clusters.

Method used

By obtaining public opinion data from the target data source, cleaning and feature extraction, a preset sliding distance sliding window is used to obtain a subset of data, dark blocks reordering different matrices, dynamically adjusting the sliding window distance to automatically identify the number of clusters and process outliers.

Benefits of technology

It realizes better automatic identification of cluster cluster numbers, effectively identify and process outliers, reduces their impact on clustering results, and improves the forward-looking and accurate public opinion situation analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067720A_ABST
    Figure CN120067720A_ABST
Patent Text Reader

Abstract

The invention relates to a network public opinion situation analysis method and system, and the method comprises the steps: obtaining public opinion data from a target data source, carrying out the cleaning of the public opinion data, extracting the features of the public opinion data, and obtaining a subset of the public opinion data in a mode of sliding a window at a preset sliding distance; obtaining the minimum value of the cluster; calculating a reordering dissimilar matrix of the current subset, extracting a dark block of the reordering dissimilar matrix of the current subset based on the minimum value, calculating the similarity between the reordering dissimilar matrix of the current subset and the reordering dissimilar matrix of the previous subset, and if the similarity meets a preset condition, extracting the dark block of the reordering dissimilar matrix of the current subset. If yes, increasing the sliding distance according to the dark blocks of the reordering dissimilar matrixes of the current subset and the previous subset to obtain the current subset again, and otherwise, obtaining the marks of the dark blocks according to the public opinion data set corresponding to the dark blocks in the reordering dissimilar matrixes of the current subset; and calculating a change curve of the public opinion data sets of the same labeled dark blocks of different subsets, and taking all labeled change curve graphs as analysis results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data analysis, and particularly to a method and system for analyzing the situation of online public opinion. Background Art

[0002] The analysis of the situation of online public opinion is a method for monitoring, analyzing, and predicting online public opinion. For enterprises, the analysis of the situation of online public opinion can help them understand consumers' evaluations of their products or services, timely discover the risks of damaged brand images, and take corresponding measures for crisis public relations. By analyzing the online public opinion of competitors, enterprises can also understand market dynamics and provide references for product R & D and marketing. In addition, for individuals, the analysis of the situation of online public opinion can also help them understand social hotspots, grasp the trend of public opinion, and improve their information literacy.

[0003] Traditional analysis methods often can only capture the surface information of public opinion, but cannot deeply explore the evolution trend of public opinion, lack the ability to predict the evolution trend of public opinion, and cannot provide decision-makers with forward-looking information. The clustering trend analysis method can cluster similar public opinion views and topics together by clustering the online public opinion data, forming different public opinion groups and topic clusters. Then, by analyzing the evolution trends of these public opinion groups and topic clusters, the internal structure and evolution law of public opinion can be deeply explored, providing more comprehensive, in-depth, and forward-looking information for decision-makers. However, there is no effective mechanism for determining the number of clusters. If the number of clusters is small, some new and small public opinions and other types of public opinions will be grouped into one cluster. If the number of clusters is large, the clusters will be over-divided, and then some public opinions belonging to the same cluster will be separated. How to determine the number of clusters is crucial for the analysis of the situation of public opinion. Summary of the Invention

[0004] In view of the above problems, in the first aspect of the present invention, a method for analyzing the situation of online public opinion is provided. The method includes: Obtaining public opinion data from a target data source, cleaning the public opinion data, extracting the features of the public opinion data, sorting the public opinion data by time, and obtaining a subset of the public opinion data by using a preset sliding distance to slide a window; obtaining the minimum value for forming a cluster; Calculating the re-ordered dissimilarity matrix of the current subset, extracting the dark blocks of the re-ordered dissimilarity matrix of the current subset based on the minimum value, binarizing the re-ordered dissimilarity matrices of the current subset and the previous subset by using the same threshold, calculating the similarity between the binarized re-ordered dissimilarity matrices of the current subset and the previous subset. If the similarity meets the preset condition, re-obtain the current subset by increasing the sliding distance according to the dark blocks of the re-ordered dissimilarity matrices of the current subset and the previous subset. Otherwise, obtain the annotation of the dark block according to the public opinion data set corresponding to the dark block in the re-ordered dissimilarity matrix of the current subset. Calculate the change curve of the public opinion data set of the dark blocks with the same annotation in different subsets, and use the change curve graphs of all annotations as the analysis results.

[0005] Preferably, the extraction of the dark blocks of the re-ordered dissimilarity matrix of the current subset based on the minimum value is specifically as follows: Continuously move along the main diagonal of the re-ordered dissimilarity matrix of the current subset. If the average value of the square block covering the minimum number of elements centered on the current point on the main diagonal is less than the preset value, then take the current point as a candidate point; Determine the dark blocks of the candidate points, and merge the dark blocks to obtain the dark blocks of the re-ordered dissimilarity matrix of the current subset.

[0006] Preferably, the determination of the dark blocks of the candidate points is specifically as follows: Obtain the square block of the candidate point, increase the length and width of the square block by 2 to get a new square block, and the new square block is still centered on the candidate point. Judge the change situation of the average value of the new square block relative to the average value of the original square block. If the change situation meets the conditions, then take the new square block as the square block of the candidate point, and repeat continuously until the change situation does not meet the conditions; Take the square block of the candidate point as the dark block of the candidate point.

[0007] Preferably, the merging of the dark blocks to obtain the dark blocks of the re-ordered dissimilarity matrix of the current subset is specifically as follows: Calculate the intersection-over-union ratio of the dark blocks. If the intersection-over-union ratio is greater than the set value, then merge the two dark blocks.

[0008] Preferably, the method for obtaining the threshold is as follows: Calculate the average value of the re-ordered dissimilarity matrix of the current subset; Calculate the average value of the re-ordered dissimilarity matrix of the previous subset; Take the minimum value or the maximum value or the average value of the average values of the current subset and the previous subset as the threshold.

[0009] Preferably, the re-acquisition of the current subset by increasing the sliding distance according to the dark blocks of the re-ordered dissimilarity matrices of the current subset and the previous subset is specifically as follows: Calculate the comprehensive change rate of the dark block area and the dark block average value of the re-ordered dissimilarity matrices of the current subset and the previous subset; Determine the increased sliding distance according to the comprehensive change rate and the preset distance, and slide the current sliding window to the right by the increased sliding distance to obtain the current subset.

[0010] In a second aspect of the present invention, a network public opinion situation analysis system is provided, and the system includes: A subset acquisition unit, which is used to acquire public opinion data from a target data source, extract the features of the public opinion data after cleaning the public opinion data, sort the public opinion data by time, and acquire a subset of the public opinion data in a manner of sliding a window with a preset sliding distance; and acquire the minimum value that becomes a cluster. An identification unit, which is used to calculate the re - sorted dissimilarity matrix of the current subset, extract the dark blocks of the re - sorted dissimilarity matrix of the current subset based on the minimum value, binarize the re - sorted dissimilarity matrices of the current subset and the previous subset using the same threshold, calculate the similarity between the binarized re - sorted dissimilarity matrices of the current subset and the previous subset. If the similarity meets the preset conditions, re - acquire the current subset by increasing the sliding distance according to the dark blocks of the re - sorted dissimilarity matrices of the current subset and the previous subset; otherwise, obtain the annotation of the dark block according to the public opinion data set corresponding to the dark block in the re - sorted dissimilarity matrix of the current subset. An analysis unit, which is used to calculate the change curve of the public opinion data set of the dark blocks with the same annotation in different subsets, and take all the change curve graphs of the annotations as the analysis results.

[0011] Preferably, the extracting the dark blocks of the re - sorted dissimilarity matrix of the current subset based on the minimum value is specifically: Continuously move along the main diagonal of the re - sorted dissimilarity matrix of the current subset. If the average value of the square covering the minimum number of elements centered on the current point on the main diagonal is less than a preset value, then take the current point as a candidate point. Determine the dark block of the candidate point, and merge the dark blocks to obtain the dark blocks of the re - sorted dissimilarity matrix of the current subset.

[0012] Preferably, the determining the dark block of the candidate point is specifically: Obtain the square of the candidate point, increase the length and width of the square by 2 to get a new square, and the new square is still centered on the candidate point. Judge the change situation of the average value of the new square relative to the average value of the original square. If the change situation meets the conditions, take the new square as the square of the candidate point, and repeat continuously until the change situation does not meet the conditions. Take the square of the candidate point as the dark block of the candidate point.

[0013] Preferably, the merging the dark blocks to obtain the dark blocks of the re - sorted dissimilarity matrix of the current subset is specifically: Calculate the intersection - over - union ratio of the dark blocks. If the intersection - over - union ratio is greater than a set value, then merge the two dark blocks.

[0014] Preferably, the method for obtaining the threshold is: Calculate the average value of the re - sorted dissimilarity matrix of the current subset; Calculate the average value of the re - sorted dissimilarity matrix of the previous subset; Use the minimum value, maximum value, or average value of the current subset and the average value of the previous subset as the threshold.

[0015] Preferably, the re-obtaining of the current subset by increasing the sliding distance of the dark blocks in the re-ordered dissimilarity matrix according to the current subset and the previous subset is specifically as follows: Calculate the comprehensive change rate of the dark block area and the average value of the dark blocks in the re-ordered dissimilarity matrix of the current subset and the previous subset; Determine the increased sliding distance according to the comprehensive change rate and the preset distance, and slide the current sliding window to the right by the increased sliding distance to obtain the current subset.

[0016] According to the minimum value that becomes a cluster, the present invention extracts the dark blocks of the re-ordered dissimilarity matrix of the current subset from the re-ordered dissimilarity matrix of the public opinion data calculated for the current window, and uses the dark blocks to complete the clustering analysis of the public opinion data, which can better automatically identify the number of clusters of the clustering, effectively identify and process outliers, and reduce their impact on the clustering result. Moreover, according to the binarization results of the re-ordered dissimilarity matrices of the current window and the previous window, the distance of the sliding window is dynamically adjusted. When the public opinion changes little, the sliding distance is increased to reduce redundant calculations. Description of the Drawings

[0017] Figure 1 Is the flowchart of Embodiment 1; Figure 2 Is the schematic diagram of the sliding window; Figure 3 Is the schematic diagram of the dissimilarity matrix; Figure 4 Is the schematic diagram of the re-ordered dissimilarity matrix; Figure 5 Is the public opinion heat trend chart. Detailed Embodiments

[0018] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0019] The meaning of the term "at least one" in the present application is one or more, and the meaning of the term "a plurality" in the present application is two or more. For example, a plurality of second messages refers to two or more second messages. The terms "system" and "network" are often used interchangeably herein.

[0020] It should be understood that the terms used in the description of the various examples herein are for the purpose of describing specific examples only and are not intended to be limiting. As used in the description of the various examples and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0021] It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. The term "and / or" describes an associative relationship between associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the character " / " in this application generally indicates that the associated objects before and after are in an "or" relationship.

[0022] It should further be understood that in the various embodiments of the present application, the magnitudes of the serial numbers of the various processes do not imply the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0023] It should be understood that determining B based on A does not mean determining B solely based on A, and B can also be determined based on A and / or other information.

[0024] Embodiment 1, as Figure 1 shown, a method for analyzing the situation of network public opinion, the method comprising: Step 1, obtaining public opinion data from a target data source, cleaning the public opinion data and then extracting the characteristics of the public opinion data, sorting the public opinion data according to time, and obtaining subsets of the public opinion data by sliding a window with a preset sliding distance; obtaining the minimum value for forming clusters; Scraping text information related to a specific topic from pre-set platforms such as websites, social media, and forums, or obtaining public opinion data from a database, wherein the scraping of text information is carried out within the scope permitted by law and / or with user authorization. For example, to analyze the public opinion of new energy vehicles, posts, comments, articles containing relevant keywords are scraped from Weibo, car forums, news websites, etc. Removing noise in the data, such as invalid characters, duplicate information, advertisements, etc., and converting the text data into numerical values or vectors that can be processed by a computer, for example, extracting keywords, sentiment tendencies, author information, etc. Sorting the data after cleaning and feature extraction in chronological order. In one embodiment, it is sorted in the order from left to right, and the closer to the right, the more recent the public opinion data. Using the sliding window method to divide the time series data into continuous and overlapping subsets, as Figure 2As shown. The sliding window size is based on the number of pieces of public opinion data. For example, 1000 pieces of public opinion data are used as a window, and the sliding distance is 100 or 200, that is, each time 100 or 200 pieces of public opinion data are slid. In the present invention, the sliding window and the subset corresponding to the sliding window are in one-to-one correspondence. The current subset and the current window, and the previous subset and the previous window, represent the same meaning and can be replaced with each other.

[0025] In public opinion analysis, if the number of pieces of public opinion data in a cluster is too small, there is no need to pay attention to these public opinions, that is, these very small clusters do not form public opinions. Here, the minimum value of the cluster is set for the following extraction of dark blocks. Preferably, the minimum value is the square value of an odd number, such as 5×5, 15×15, 17×17, etc.

[0026] Step 2, calculate the re-ordered dissimilarity matrix of the current subset, extract the dark blocks of the re-ordered dissimilarity matrix of the current subset based on the minimum value, binarize the re-ordered dissimilarity matrices of the current subset and the previous subset using the same threshold, calculate the similarity between the binarized re-ordered dissimilarity matrices of the current subset and the previous subset. If the similarity meets the preset conditions, re-obtain the current subset by increasing the sliding distance according to the dark blocks of the re-ordered dissimilarity matrices of the current subset and the previous subset; otherwise, obtain the annotation of the dark block according to the public opinion data set corresponding to the dark block in the re-ordered dissimilarity matrix of the current subset. For the subset of public opinion data within the current sliding window extracted by the sliding window, calculate the dissimilarity matrix. Each element of the dissimilarity matrix represents the degree of dissimilarity or distance between any two pieces of public opinion data in the subset. Figure 3 Shows the dissimilarity matrix of 10 data. Re-order the dissimilarity matrix by means of the minimum spanning tree, etc., so that similar data points are clustered together in the matrix, and dissimilar data points are scattered, so as to present obvious dark blocks, as Figure 4 As shown. In one embodiment, VAT (Visual Assessment of cluster Tendency) or iVAT (improved VAT) is used to calculate the re-ordered dissimilarity matrix of the current subset.

[0027] In the re-ordered dissimilarity matrix, the dark blocks represent clusters of public opinion data with similar characteristics. Identifying the dark blocks is equivalent to identifying public opinion events or topics. Using the minimum value, slide in the matrix to find squares that meet the conditions, such as the average value being less than the preset value, and then expand and merge the squares, etc., so as to extract the dark blocks.

[0028] In a more specific embodiment, the extraction of the dark blocks of the re-ordered dissimilarity matrix of the current subset based on the minimum value is specifically: Continuously move along the main diagonal of the reordered dissimilarity matrix of the current subset. If the average value of the square covering the minimum number of elements centered at the current point on the main diagonal is less than the preset value, then take the current point as a candidate point; Determine the dark blocks of the candidate points, and merge the dark blocks to obtain the dark blocks of the reordered dissimilarity matrix of the current subset.

[0029] In the reordered dissimilarity matrix, the value of the main diagonal is at least 0. Starting from the main diagonal of the reordered dissimilarity matrix, move along the main diagonal. The current position on the main diagonal determines a square, and the size of the square, that is, the number of elements contained in the square, is the minimum value. For example, if the minimum value is 5×5, then the size of the square is also 5×5, where the current point or the current position on the main diagonal is located at the center of the square.

[0030] Calculate the average value of all elements within this square of the reordered dissimilarity matrix. Take this square as a mask and calculate the average value of the reordered dissimilarity matrix within this mask. If this average value is lower than the preset value, the public opinion data in this area is relatively concentrated and is a candidate cluster.

[0031] For each candidate point, further determine the specific range of the dark block it represents. In one embodiment, it is achieved by iteratively expanding the size of the square until the average value within the square no longer meets the condition or the preset value. Since different candidate points may represent overlapping dark blocks, these dark blocks need to be merged to obtain the final set of dark blocks.

[0032] In another embodiment, the determination of the dark blocks of the candidate points is specifically as follows: Obtain the square of the candidate point, increase the length and width of the square by 2 to get a new square, and the new square is still centered at the candidate point. Judge the change of the average value of the new square relative to the average value of the original square. If the change meets the condition, then take the new square as the square of the candidate point and repeat continuously until the change does not meet the condition; Take the square of the candidate point as the dark block of the candidate point.

[0033] For a candidate point, obtain one of the squares centered at this point. Then, increase the length and width of this square by 2 pixels to obtain a new and larger square, still centered at this candidate point. For example, if the original square is 5×5, it becomes 7×7 after expansion. Calculate the average value of all elements within the new square and compare it with the average value of the original square to determine the change between them. If the change in the average value of the new square relative to the average value of the original square is within a preset range, for example, the average value of the new square is still less than the preset value or the difference between the average value of the new square and the average value of the original square is less than a certain threshold, then use the new square as the square of the candidate point and repeat the above steps to continue expanding the size of the square. If the change does not meet the conditions, for example, the average value of the new square is significantly higher than that of the original square or exceeds the preset range, then stop expanding the size of the square. By iteratively expanding the size of the square, the boundary of the dark block can be gradually determined while avoiding over-expansion.

[0034] In an alternative embodiment, the dark blocks for merging to obtain the reordered dissimilarity matrix of the current subset are specifically as follows: Calculate the intersection over union (IoU) of the dark blocks. If the IoU is greater than a set value, then merge the two dark blocks. Among them, the IoU is the ratio of the intersection and union of the public opinion data corresponding to the two dark blocks in the reordered dissimilarity matrix. The value range of the IoU is from 0 to 1, and the larger the value, the higher the overlap degree of the two dark blocks. If the IoU is greater than the set value, such as 0.2 or 0.3, then merge the two dark blocks. The reordered dissimilarity matrix is an n×n matrix, and each element (i, j) is the distance or dissimilarity between the i-th public opinion data and the j-th public opinion data. The public opinion data corresponding to the dark block can be determined by the position of the dark block in the reordered dissimilarity matrix.

[0035] When the sliding window slides, it will cause changes in the reordered dissimilarity matrix. Both the reordered dissimilarity matrix and the distances in the extracted dark blocks will change, but the clustering result may remain unchanged. Binarize the reordered dissimilarity matrices of the current subset and the previous subset using the same threshold, and calculate the similarity between the binarized reordered dissimilarity matrices of the current subset and the previous subset. Even if there are some changes in the internal values of the dark block, as long as its overall structure remains unchanged, the binarized matrices will remain similar. Moreover, binarization simplifies the matrix into two states, 0 and 1, highlighting the main structure in the matrix, making the calculation of similarity more stable and robust, and reducing the influence of minor numerical changes on the result.

[0036] Binarize both the reordered dissimilarity matrix of the public opinion data subset of the current window, i.e., the current subset, and the reordered dissimilarity matrix of the public opinion data subset of the previous window, i.e., the previous subset. The binarization means converting the values in the matrix into two states, 0 and 1. If it is greater than the threshold, it is converted to 1; otherwise, it is converted to 0. Calculate the similarity degree between the two binarized matrices. The calculation methods of similarity include but are not limited to mean squared error, structural similarity index, cosine similarity, etc. If the calculated similarity meets the preset conditions, in one embodiment, the preset condition is greater than or equal to the similarity threshold, then the two matrices are similar enough, that is, the public opinion situation has not changed significantly. At this time, to improve the analysis efficiency, increase the sliding distance of the sliding window, that is, reduce the overlapping part of the time subsets, thereby reducing redundant calculations. Then re-obtain the current subset. If the similarity is less than the similarity threshold, it is considered that the public opinion situation has changed significantly, and it is necessary to label the dark blocks of the current subset, that is, analyze the public opinion events or topics represented by the dark blocks.

[0037] When binarizing the reordered dissimilarity matrix, too large or too small a threshold will affect the final similarity judgment. In one embodiment, the method for obtaining the threshold is: calculate the average value of the reordered dissimilarity matrix of the current subset; calculate the average value of the reordered dissimilarity matrix of the previous subset; use the minimum value or maximum value or average value of the average values of the current subset and the previous subset as the threshold.

[0038] Calculate the average values of the reordered dissimilarity matrices of the current subset and the previous subset respectively, and then use the minimum value, maximum value or average value of these two average values as the binarization threshold. Using the average value of the matrix as the threshold can reflect the overall numerical distribution of the matrix, making the threshold more representative. At the same time, considering the factors of data changes in adjacent time periods, using the minimum value, maximum value or average value of the average values of the two subsets can make the threshold more adaptable.

[0039] If the public opinion situations in adjacent time periods are very similar, then there is no need to conduct repeated detailed analysis. At this time, the sliding distance of the sliding window can be increased to reduce redundant calculations and improve the analysis efficiency. In one embodiment, the method of re-obtaining the current subset by increasing the sliding distance according to the dark blocks of the reordered dissimilarity matrices of the current subset and the previous subset is specifically as follows: Calculate the comprehensive change rate of the dark block area and the dark block average value of the reordered dissimilarity matrices of the current subset and the previous subset; Determine the increased sliding distance according to the comprehensive change rate and the preset distance, and slide the current sliding window to the right by the increased sliding distance to obtain the current subset.

[0040] By calculating the comprehensive change rate of the dark block area and the average value of the dark block in the re - sorted dissimilarity matrix between the current subset and the previous subset, the change degree of the public opinion situation is quantified. The larger the comprehensive change rate is, the more significant the change of the public opinion situation is, and vice versa. Then, according to the comprehensive change rate and the preset sliding distance, the increased sliding distance is determined. If the comprehensive change rate is very small, it indicates that the public opinion situation has not changed significantly. At this time, the sliding distance can be greatly increased to reduce redundant calculations. If the comprehensive change rate is relatively large, it indicates that the public opinion situation has changed significantly. At this time, the sliding distance should be maintained at the preset sliding distance to capture the details of the public opinion change. In one embodiment, calculating the comprehensive change rate of the dark block area and the average value of the dark block in the re - sorted dissimilarity matrix between the current subset and the previous subset is specifically to calculate the change rate of the dark block area between the current subset and the previous subset, and the change rate of the average value of the dark block, and then sum or perform weighted summation on these two change rates to obtain the comprehensive change rate. In one embodiment, according to the comprehensive change rate and the preset distance to determine the increased sliding distance, and sliding the current sliding window to the right by the increased sliding distance to obtain the current subset is specifically to calculate the increased value in the way of d(1 - r), and take the sum of the preset sliding distance and the increased value as the increased sliding distance, where d is the preset sliding distance and r is the comprehensive change rate. If it is a decimal, it is rounded. For example, if the preset sliding distance is 100 items and the increased value is 30, then the current window slides 30 more to the right.

[0041] In one embodiment, the operation of increasing the sliding distance in step 2 is only executed once. That is, no matter whether the similarity between the re - sorted dissimilarity matrix of the current subset and the previous subset after binarization meets the preset conditions, the annotation of the dark block is obtained according to the public opinion data set corresponding to the dark block in the re - sorted dissimilarity matrix of the current subset. In another embodiment, the operation of increasing the sliding distance in step 2 is executed multiple times until the similarity between the re - sorted dissimilarity matrix of the current subset and the previous subset after binarization does not meet the preset conditions, and then the annotation of the dark block is obtained according to the public opinion data set corresponding to the dark block in the re - sorted dissimilarity matrix of the current subset.

[0042] After obtaining the dark blocks of the current subset, the annotation of the dark block is obtained according to the public opinion data set corresponding to the dark block. For example, taking the high - frequency words as the annotation, or performing sentiment analysis on the public opinion data set corresponding to the dark block and taking the sentiment analysis result as the annotation, etc.

[0043] Step 3, calculate the change curve of the public opinion data set of the dark blocks with the same annotation in different subsets, and take all the change curve graphs of the annotations as the analysis result.

[0044] Extract the dark blocks with the same annotation in different time subsets, that is, the public opinion data of the same public opinion event or theme, and calculate their changes over time. If it is a new dark block annotation, use this annotation as the starting point of the change curve of this annotation. For example, a line chart can be used to represent the change in the discussion heat of a certain public opinion event over time. In one embodiment, the line chart will mark the heat, the annotation of the dark block, etc. By observing these change curves, the changes in the public opinion situation can be comprehensively understood, and hot events, trends, and anomalies can be identified. Figure 5 Shows the regional heat trend chart generated using the solution of the present invention.

[0045] Embodiment 2 provides a network public opinion situation analysis system, and the system includes: A subset acquisition unit, configured to acquire public opinion data from a target data source, extract the characteristics of the public opinion data after cleaning the public opinion data, sort the public opinion data according to time, and acquire subsets of the public opinion data in a manner of sliding a window with a preset sliding distance; acquire the minimum value that becomes a cluster; An identification unit, configured to calculate the re-ordered dissimilarity matrix of the current subset, extract the dark blocks of the re-ordered dissimilarity matrix of the current subset based on the minimum value, binarize the re-ordered dissimilarity matrices of the current subset and the previous subset using the same threshold, calculate the similarity between the binarized re-ordered dissimilarity matrices of the current subset and the previous subset, and if the similarity meets the preset condition, re-acquire the current subset by increasing the sliding distance according to the dark blocks of the re-ordered dissimilarity matrices of the current subset and the previous subset, otherwise, obtain the annotation of the dark block according to the public opinion data set corresponding to the dark block in the re-ordered dissimilarity matrix of the current subset; An analysis unit, configured to calculate the change curves of the public opinion data sets of the dark blocks with the same annotation in different subsets, and use the change curves of all annotations as the analysis result.

[0046] Preferably, the extracting the dark blocks of the re-ordered dissimilarity matrix of the current subset based on the minimum value is specifically: Continuously move along the main diagonal of the re-ordered dissimilarity matrix of the current subset. If the average value of the square covering the minimum number of elements centered on the current point of the main diagonal is less than the preset value, use the current point as a candidate point; Determine the dark blocks of the candidate points, and merge the dark blocks to obtain the dark blocks of the re-ordered dissimilarity matrix of the current subset.

[0047] Preferably, the determining the dark blocks of the candidate points is specifically: Obtain the square of the candidate point, increase the length and width of the square by 2 to obtain a new square, and the new square is still centered on the candidate point. Judge the change situation of the average value of the new square relative to the average value of the original square. If the change situation meets the condition, use the new square as the square of the candidate point, and repeat continuously until the change situation does not meet the condition; Take the block of the candidate point as the dark block of the candidate point.

[0048] Preferably, the dark blocks obtained by merging the dark blocks to obtain the re - sorted dissimilarity matrix of the current subset are specifically as follows: Calculate the intersection - over - union ratio of the dark blocks. If the intersection - over - union ratio is greater than the set value, then merge the two dark blocks.

[0049] Preferably, the method for obtaining the threshold is as follows: Calculate the average value of the re - sorted dissimilarity matrix of the current subset; Calculate the average value of the re - sorted dissimilarity matrix of the previous subset; Take the minimum value or the maximum value or the average value of the average values of the current subset and the previous subset as the threshold.

[0050] Preferably, the method for re - obtaining the current subset by increasing the sliding distance according to the dark blocks of the re - sorted dissimilarity matrix of the current subset and the previous subset is specifically as follows: Calculate the comprehensive change rate of the area and the average value of the dark blocks of the re - sorted dissimilarity matrix of the current subset and the previous subset; Determine the increased sliding distance according to the comprehensive change rate and the preset distance, and slide the current sliding window to the right by the increased sliding distance to obtain the current subset.

[0051] Through the description of the above - mentioned embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general - purpose hardware. Of course, it can also be implemented by dedicated hardware including application - specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits or dedicated circuits, etc. However, for this application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of this application.

[0052] In the above - mentioned embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0053] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium, etc.

[0054] The above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. However, such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for analyzing network public opinion situation, characterized in that: The method comprises: Obtain public opinion data from the target data source, clean the public opinion data and extract its features, sort the public opinion data by time, and obtain a subset of the public opinion data using a sliding window with a preset sliding distance; obtain the minimum value that becomes a cluster; Calculate the reordered dissimilarity matrix of the current subset, extract the dark block of the reordered dissimilarity matrix of the current subset based on the minimum value, binarize the reordered dissimilarity matrix of the current subset and the previous subset using the same threshold, calculate the similarity of the binarized reordered dissimilarity matrix of the current subset and the previous subset, and if the similarity meets the preset conditions, increase the sliding distance according to the dark blocks of the reordered dissimilarity matrix of the current subset and the previous subset to re-acquire the current subset, otherwise, obtain the annotation of the dark block according to the public opinion data set corresponding to the dark block in the reordered dissimilarity matrix of the current subset; Calculate the change curves of the public opinion data set of dark blocks with the same annotations in different subsets, and use the change curve graphs of all annotations as the analysis results.

2. The method according to claim 1, characterized in that The step of extracting the dark block of the reordered dissimilar matrix of the current subset based on the minimum value is specifically as follows: Continuously move along the main diagonal of the reordered dissimilarity matrix of the current subset, and if the average value of the block covering the minimum value elements with the current point on the main diagonal as the center is less than a preset value, the current point is taken as a candidate point; Determine the dark blocks of the candidate points, and merge the dark blocks to obtain the dark blocks of the reordered dissimilarity matrix of the current subset.

3. The method according to claim 2, characterized in that The dark block of the candidate point is determined as follows: Get the square of the candidate point, increase the length and width of the square by 2 to get a new square, the new square is still centered on the candidate point, and determine the change of the average value of the new square relative to the average value of the original square. If the change meets the conditions, the new square is used as the square of the candidate point, and the process is repeated until the change does not meet the conditions; The square of the candidate point is regarded as the dark block of the candidate point.

4. The method according to claim 2, characterized in that The dark blocks are merged to obtain the dark blocks of the reordered dissimilar matrix of the current subset, specifically: Calculate the intersection-and-union ratio of the dark blocks. If the intersection-and-union ratio is greater than the set value, merge the two dark blocks.

5. The method according to claim 1, characterized in that The method for obtaining the threshold is: Calculate the average of the reordered dissimilarity matrix of the current subset; Calculate the average of the reordered dissimilarity matrix of the previous subset; The minimum value, maximum value or average value of the average values ​​of the current subset and the previous subset is used as the threshold.

6. The method according to claim 1, characterized in that The dark blocks of the reordered dissimilar matrices of the current subset and the previous subset are increased in sliding distance to reacquire the current subset, specifically: Calculate the comprehensive change rate of the dark block area and the dark block average value of the reordered dissimilarity matrix of the current subset and the previous subset; An increased sliding distance is determined according to the comprehensive change rate and the preset distance, and the current sliding window is slid rightward by the increased sliding distance to obtain the current subset.

7. A network public opinion situation analysis system, characterized in that: The system comprises: The subset acquisition unit is used to acquire public opinion data from the target data source, extract the features of the public opinion data after cleaning the public opinion data, sort the public opinion data according to time, and acquire a subset of the public opinion data by using a sliding window with a preset sliding distance; and acquire the minimum value of the cluster; an identification unit, used to calculate the reordered dissimilarity matrix of the current subset, extract the dark blocks of the reordered dissimilarity matrix of the current subset based on the minimum value, binarize the reordered dissimilarity matrix of the current subset and the previous subset using the same threshold, calculate the similarity of the binarized reordered dissimilarity matrix of the current subset and the previous subset, and if the similarity meets the preset condition, reacquire the current subset by increasing the sliding distance according to the dark blocks of the reordered dissimilarity matrix of the current subset and the previous subset; otherwise, obtain the annotation of the dark blocks according to the public opinion data set corresponding to the dark blocks in the reordered dissimilarity matrix of the current subset; The analysis unit is used to calculate the change curve of the public opinion data set of dark blocks with the same annotations in different subsets, and use the change curve graphs of all annotations as the analysis results.

8. The system according to claim 7, characterized in that The step of extracting the dark block of the reordered dissimilar matrix of the current subset based on the minimum value is specifically as follows: Continuously move along the main diagonal of the reordered dissimilarity matrix of the current subset, and if the average value of the block covering the minimum value elements with the current point on the main diagonal as the center is less than a preset value, the current point is taken as a candidate point; Determine the dark blocks of the candidate points, and merge the dark blocks to obtain the dark blocks of the reordered dissimilarity matrix of the current subset.

9. The system according to claim 8, characterized in that The dark block of the candidate point is determined as follows: Get the square of the candidate point, increase the length and width of the square by 2 to get a new square, the new square is still centered on the candidate point, and determine the change of the average value of the new square relative to the average value of the original square. If the change meets the conditions, the new square is used as the square of the candidate point, and the process is repeated until the change does not meet the conditions; The square of the candidate point is regarded as the dark block of the candidate point.

10. The system according to claim 7, characterized in that The dark blocks of the reordered dissimilar matrices of the current subset and the previous subset are increased in sliding distance to reacquire the current subset, specifically: Calculate the comprehensive change rate of the dark block area and the dark block average value of the reordered dissimilarity matrix of the current subset and the previous subset; An increased sliding distance is determined according to the comprehensive change rate and the preset distance, and the current sliding window is slid rightward by the increased sliding distance to obtain the current subset.