Network information processing method, device and equipment
Through grouping, weight determination and random sampling of network information, combined with sentence vector dimensionality reduction and clustering processing, the problem of low sampling accuracy of massive multi-source network information is solved, and efficient network information analysis and utilization is achieved.
Patent Information
- Application Number
- CN202510531471.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, the sampling accuracy of massive multi-source network information is not high, resulting in low subsequent analysis and utilization efficiency.
By obtaining multi-source network information related to preset topics, preprocessing and grouping, and determining the number of samples based on the weights and preset parameters of each packet, randomly selecting sample network information, combining sentence vector dimensionality reduction and clustering processing, calculating the similarity and distance of the cluster, and determining the weight of the analysis type.
实现了对海量多源网络信息的高精准抽样和高效聚合,提升了对相关主题的认知深度及决策支持力度。
Smart Images

Figure CN120281669A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of Internet data processing, and in particular, to a processing technology for network information. Background Art
[0002] With the rapid development and popularization of Internet technology, users' online activities are becoming increasingly active, and network information shows an explosive growth trend, with the characteristics of being massive, multi-source, heterogeneous, and dynamically updated. In the existing Internet environment, the analysis and utilization of network information resources face challenges such as large data scale, diverse sources, strong timeliness, and complex information semantics.
[0003] Existing methods for analyzing network information usually adopt sampling techniques to extract network information, and then aggregate it for analysis and utilization to improve processing efficiency. Due to the unique characteristics of network information such as dynamicity and real-time nature, the accuracy of sampling is insufficient, resulting in low efficiency of the subsequent aggregated analysis and processing of the sampled network information.
[0004] Therefore, how to improve the sampling accuracy of massive multi-source network information to facilitate subsequent efficient analysis and utilization has become an urgent problem to be solved. Summary of the Invention
[0005] The purpose of this application is to provide a method, device, and equipment for processing network information to at least partially solve the technical problem of low sampling accuracy of massive multi-source network information in the prior art.
[0006] According to one aspect of this application, a method for processing network information is provided, where the method includes:
[0007] Obtain multi-source network information of a preset statistical period related to a preset theme, and perform preprocessing to obtain a number of network information, and determine the source category of each network information;
[0008] Group all the network information, and determine the weight of each group according to the number of all the network information and the number of network information included in each group;
[0009] Based on preset parameters, determine the number of network information to be extracted from all the network information as the first sample quantity, and determine the number of network information to be extracted from each group as the second sample quantity corresponding to the group according to the weight of each group and the first sample quantity, where the sum of the second sample quantities corresponding to each group is equal to the first sample quantity;
[0010] Randomly extract the number of network information corresponding to the second sample quantity of each group from each group as sample network information.
[0011] Optionally, wherein the grouping of all network information includes:
[0012] Determine the stage of public opinion development of the preset theme to which each network information belongs;
[0013] Group all network information according to the stage of public opinion development.
[0014] Optionally, wherein the grouping of all network information includes:
[0015] Determine the key event nodes of the preset theme to which each network information belongs;
[0016] Group all network information according to the key event nodes.
[0017] Optionally, the method for processing network information further includes:
[0018] Preprocess each sample network information, convert each preprocessed sample network information into a sentence vector, and perform dimensionality reduction processing on the sentence vector to obtain the sample data corresponding to the sample network information;
[0019] Perform clustering processing on all sample data to obtain a number of clusters, and calculate the internal similarity index and inter-cluster distance of each cluster, where each cluster includes one or more data points, one of which is the core point, and each data point corresponds to a sample data.
[0020] 5. The method according to claim 4, wherein the method for processing network information further includes:
[0021] Determine the source weight according to the source category of each sample network information, and determine the weight of the sample network information according to its grouping weight and source weight;
[0022] Determine the analysis type corresponding to the cluster according to the sample network information corresponding to the core point of each cluster, and determine the weight and weight ratio of the analysis type corresponding to the cluster according to the weights of the sample network information corresponding to each data point in each cluster.
[0023] According to another aspect of the present application, there is provided a device for processing network information, wherein the device includes:
[0024] The first module is used to obtain multi-source network information of a preset statistical period related to a preset theme, perform preprocessing to obtain a number of network information, and determine the source category of each network information;
[0025] The second module is used to group all network information and determine the weight of each group according to the number of all network information and the number of network information included in each group;
[0026] A third module, configured to determine, based on preset parameters, the number of network information extracted from all network information as the first sample quantity, and determine, according to the weight of each group and the first sample quantity, the number of network information extracted from each group as the second sample quantity corresponding to the group, wherein the sum of the second sample quantities corresponding to each group is equal to the first sample quantity;
[0027] A fourth module, configured to randomly extract, from each group, the number of network information corresponding to the second sample quantity of the group as sample network information.
[0028] Optionally, the network information processing device further includes:
[0029] A fifth module, configured to preprocess each sample network information, convert each preprocessed sample network information into a sentence vector, and perform dimensionality reduction processing on the sentence vector to obtain sample data corresponding to the sample network information;
[0030] A sixth module, configured to perform clustering processing on all sample data to obtain a plurality of clusters, and calculate the internal similarity index and the inter-cluster distance of each cluster, wherein each cluster includes one or more data points, one of the data points is a core point, and each data point corresponds to a piece of sample data.
[0031] Optionally, the network information processing device further includes:
[0032] A seventh module, configured to determine the source weight according to the source category of each sample network information, and determine the weight of the sample network information according to its group weight and source weight;
[0033] An eighth module, configured to determine the analysis type corresponding to the cluster according to the sample network information corresponding to the core point of each cluster, and determine the weight and weight ratio of the analysis type corresponding to the cluster according to the weights of the sample network information corresponding to each data point in each cluster.
[0034] Compared with the prior art, the present application provides a method, apparatus, and device for processing network information. The method includes: obtaining multi-source network information for a preset statistical period related to a preset theme, and performing preprocessing to obtain a number of network information items, and determining the source category of each network information item; grouping all the network information items, and determining the weight of each group according to the number of all the network information items and the number of network information items included in each group; based on preset parameters, determining the number of network information items to be extracted from all the network information items as the first sample quantity, and determining the number of network information items to be extracted from each group as the second sample quantity corresponding to the group according to the weight of each group and the first sample quantity, where the sum of the second sample quantities corresponding to each group is equal to the first sample quantity; randomly extracting the second sample quantity of network information items corresponding to each group from each group as sample network information items. Further, the method further includes: performing preprocessing on each sample network information item, converting each preprocessed sample network information item into a sentence vector, and performing dimensionality reduction processing on the sentence vector to obtain sample data corresponding to the sample network information item; performing clustering processing on all the sample data to obtain a number of clusters, and calculating the internal similarity index and the inter-cluster distance of each cluster, where each cluster includes one or more data points, and one of the data points is a core point, and each data point corresponds to one sample data item. Through this method, high-precision sampling of a large amount of multi-source network information can be achieved. Further, efficient aggregation of the sampled network information can be achieved. Applied to actual application scenarios, the depth of understanding of related themes and the strength of decision support can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Other features, objects, and advantages of the present invention will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings:
[0036] Figure 1 FIG. shows a schematic flowchart of a method for processing network information according to an aspect of the present application;
[0037] Figure 2 FIG. shows a schematic diagram of an apparatus for processing network information according to another aspect of the present application;
[0038] The same or similar reference numerals in the drawings represent the same or similar components. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] The present invention will be further described in detail below with reference to the drawings.
[0040] In a typical configuration of each embodiment of the present application, each device, system trusted party, and / or each module of the apparatus may include one or more processors (CPUs), input / output interfaces, network interfaces, and memories.
[0041] The memory may include non - permanent memory in the form of computer - readable media, such as random access memory (RAM) and / or non - volatile memory, such as read - only memory (ROM) or flash RAM. The memory is an example of computer - readable media.
[0042] Computer - readable media includes both permanent and non - permanent, removable and non - removable media that can store information by any method or technology. The information can be computer - readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase - change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read - only memory (ROM), electrically erasable programmable read - only memory (EEPROM), flash memory or other memory technologies, compact disc read - only memory (CD - ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non - transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer - readable media does not include transitory media, such as modulated data signals and carrier waves.
[0043] In view of the unique characteristics of the massive multi - source network information on the Internet, such as its dynamicity and real - time nature. For example, the network information of a hot event is generated in real - time, and in different public opinion stages of a hot event, the distribution, source, attributes, etc. of the related network information have different characteristics. The present application provides a method, apparatus, and device for processing network information. In view of the characteristics of the existing massive multi - source network information, it is designed to accurately sample and obtain network information related to relevant topics from the massive multi - source network information, and then perform efficient aggregation and analysis and utilization to meet the requirements of in - depth understanding of network information and improving decision - making ability based on network information in the actual application scenarios of the big data era. To achieve the efficient utilization of network information, it can serve multiple fields such as public opinion monitoring, brand management, and market research.
[0044] A method for processing network information provided by the present application obtains a large amount of multi-source network information related to a certain theme within a certain period of time from the Internet, filters out irrelevant information, and obtains several network information with clear sources; then groups all the obtained network information and determines the weight of each group; then according to the sampling strategy, determines the number of network information to be extracted from all the network information, and according to the weight of each group, determines the number of network information to be extracted from each group; finally, randomly extracts the relevant number of network information from each group as sample network information. It realizes high-precision sampling of network information related to relevant themes. Further, after efficiently aggregating the network information obtained after sampling, it is used for analysis and utilization.
[0045] To further elaborate on the technical means adopted by the present application and the achieved effects, the technical solution of the present application will be clearly and completely described below in conjunction with the accompanying drawings and preferred embodiments.
[0046] Figure 1 The figure shows a schematic flowchart of a method for processing network information according to an aspect of the present application. Among them, the method of one embodiment includes:
[0047] S101 Obtain multi-source network information of a preset statistical period related to a preset theme, and perform preprocessing to obtain several network information, and determine the source category of each network information;
[0048] S102 Group all the network information, and determine the weight of each group according to the number of all the network information and the number of network information included in each group;
[0049] S103 Based on preset parameters, determine the number of network information to be extracted from all the network information as the first sample quantity, and according to the weight of each group and the first sample quantity, determine the number of network information to be extracted from each group as the second sample quantity corresponding to the group, where the sum of the second sample quantities corresponding to each group is equal to the first sample quantity;
[0050] S104 Randomly extract the second sample quantity of network information corresponding to the group from each group as sample network information.
[0051] In this application, each method embodiment / optional embodiment can be implemented or executed by device 100, where device 100 is a computer device with corresponding software and hardware environments. Among them, the computer device includes, but is not limited to, personal computers, laptop computers, industrial computers, servers, network hosts, single network servers or network server clusters. Here, the computer device is only an example, and other existing or future devices and / or resource platforms applicable to this application should also be included in the protection scope of this application and are hereby incorporated by reference.
[0052] In this embodiment, in step S101, the target theme to be analyzed and processed in the actual application scenario is used as the preset theme. Device 100 can obtain multi-source network information of a preset statistical period related to the preset theme from the Internet, perform preprocessing to obtain several pieces of network information, and determine the source category of each piece of network information.
[0053] Among them, the target theme can be a certain hot event, topic or social phenomenon, covering different enterprises, social institutions, brands, figures, industries, public policies, laws and regulations, social phenomena, etc. According to the preset theme name, methods such as combined keyword retrieval can be used to collect multi-source network information related to the preset theme within the preset statistical period from the Internet. Among them, as a theme of network hotspots, its spread on the Internet is usually time-limited. For example, for an emergency event, from the original release of relevant information on the Internet to the basic subsidence of public opinion on the Internet, it has a life cycle. Therefore, the preset statistical period can usually be set as the continuous life cycle of the information dissemination of the preset theme on the Internet; or for example, when collecting network information and analyzing the online reputation of a certain brand within a certain period of time, this period of time can usually be used as the preset statistical period. For example, the monthly period can be used as the preset statistical period. Among them, the sources of network information can cover search engines, news websites, social media platforms, forums, blogs, etc., to ensure the comprehensiveness and diversity of the sources of network information, and can reflect the multi-angle comments of different user groups and media types on the preset theme. Among them, network information can include text data and other types. If it is other types such as audio and video and other multimedia data, it should be texturized first after acquisition. Among them, in order to improve the efficiency of subsequent data processing, before using the acquired multi-source network information for subsequent processing, the acquired multi-source network information can also be preprocessed, input into a BERT (Bidirectional Encoder Representations from Transformers) classifier trained by supervised fine-tuning, to identify and remove interference items such as irrelevant advertisements, spam information, and meaningless content, and ensure the purity of the data. Among them, multi-source network information of historical themes can be collected in advance, each network information is labeled (network information including advertisements, spam information or meaningless content, etc. is labeled as negative samples, otherwise labeled as positive samples), a data set is constructed, and then the BERT classifier is trained and verified based on the data set to obtain a BERT classifier trained by supervised learning, and the BERT classifier trained by supervised learning can be called in the form of an interface to preprocess each network information related to the preset theme obtained in real time. After preprocessing, several network information can be obtained, and the source category of each network information is determined. Among them, each network information includes the release time, release source and information content, and also includes information such as the number of forwards, comments, and likes of the network information. Among them, the source category of each network information includes two parts: release source type 1 and release source type 2. Among them, release source type 1 refers to the release source platform category, that is, the network information comes from a website, client, Weibo or public account, etc., and release source type 2 refers to the release source account category, that is, the network information comes from a personal certified account, ordinary user account, media account, enterprise account or government account, etc.
[0054] Among them, in order to ensure the unique identifiability of each piece of network information, facilitate the subsequent traceability, analysis and utilization of network information, each obtained piece of network information can also be encoded with a unique identifier. Among them, the unique identifier encoding method can adopt the method of combined hash encoding. First, convert the release time in the network information into a standard timestamp (timestamp), such as the Unix timestamp (the number of seconds starting from 00:00:00 UTC on January 1, 1970). Then, determine the identifier (sourceID) in combination with the release source category (release source type 1 + release source type 2) and the release source name. An exemplary one is that if the release source category is a media account on Weibo, a part of its domain name or a preset number can be used as the identifier; if the release source category is a personal certified account or an ordinary user account on Weibo, its account ID can be directly used as the identifier. Then, based on the information content of this piece of network information, extract several keyword combinations (keyword). Finally, splice the above standard timestamp, identifier and keyword combination into a string (such as timestamp_sourceID_keyword), and a hash function (such as MD5, SHA-1, SHA-256, etc.) can be used to generate a hash value with a fixed length as the unique identifier corresponding to this piece of network information.
[0055] Among them, each piece of network information and its source category are also structured for subsequent operations to improve the processing efficiency of subsequent data related to network information.
[0056] Continuing in this embodiment, in step S102, the device 100 can group all the network information and determine the weight of each group according to the number of all the network information and the number of network information included in each group.
[0057] Among them, the device 100 can group all the network information obtained in step S101 to obtain several groups, where each group includes several pieces of network information. Then, according to the number of all the network information and the number included in each group, determine the weight of each group. Among them, the ratio of the number included in each group to the number of all the network information can be used as the weight of the group.
[0058] Among them, due to the differences in the authority and credibility of each source of network information, the sources of network information can be divided into several categories. Exemplarily, the sources of network information can be divided into three categories: news media, opinion leaders, and ordinary users. Among them, the news media category can include websites, Weibo media accounts, WeChat media accounts, video media accounts, digital newspapers, clients, etc.; the opinion leader category can include opinion leaders in professional fields, and different lists of professional users for different themes can be customized in advance according to the characteristics of different themes. It can also include opinion leaders in non-professional fields, such as Weibo big Vs (orange V / gold V / blue V) other than media accounts, WeChat institutional accounts / personal accounts, video institutional accounts, etc. Among them, all network information can be grouped according to the categories of their sources, and then the weight of each group can be determined according to the ratio of the number of network information included in each group to the total number of all network information. Among them, the sum of the weights of each group should be 1.
[0059] For complex and changeable themes, in order to achieve refined tracking of the dynamic changes of network public opinion on a preset theme, a number of obtained network information can be grouped according to different stages of network public opinion development or key event nodes within a preset statistical period.
[0060] Optionally, among them, the grouping of all network information includes:
[0061] Determine the stage of public opinion development of the preset theme in which each network information is located;
[0062] Group all network information according to the stage of public opinion development.
[0063] Among them, device 100 can determine the stage of public opinion development of the preset theme in which each obtained network information is located, and then group all network information according to different stages of public opinion development of the preset theme. Among them, the stages of public opinion development of several network information included in each group are the same. During the life cycle of a theme, its stages of public opinion development usually include the incubation period, the rising period, the outbreak period, and the decline period. Some themes also include the rebound period.
[0064] Optionally, among them, the grouping of all network information includes:
[0065] Determine the key event nodes of the preset theme to which each network information belongs;
[0066] Group all network information according to the key event nodes.
[0067] Among them, the device 100 can also determine, for each obtained network information, the key event nodes of the preset theme to which each network information belongs. Then, according to the different key event nodes of the preset theme, all network information is grouped, where several network information included in each group belong to the same key event node. Among them, different key event nodes can be determined in combination with the characteristics of the preset theme. Exemplarily, the key event nodes of a theme can include: news release, official statement, policy change, comprehensive rectification, in-depth investigation, liability determination, legal litigation and judgment, handling and accountability, victim compensation, etc.
[0068] Continuing in this embodiment, in step S103, the device 100 can determine, based on preset parameters, the number of network information extracted from all network information as the first sample quantity, and determine, according to the weight of each group and the first sample quantity, the number of network information extracted from each group as the second sample quantity corresponding to the group, where the sum of the second sample quantities corresponding to each group is equal to the first sample quantity.
[0069] Among them, a sampling strategy for sampling from all network information can be formulated in advance to determine parameters such as sampling error and confidence level. For example, for general network hot topics, a sampling error of 5% and a confidence level of 90% can be selected; while for network hot topics with great influence, a sampling error of 3% (or even 1%) and a confidence level of 95% or higher can be selected.
[0070] Among them, the following formula (1) can be used to determine the number (n) of network information extracted from all network information as the first sample quantity.
[0071]
[0072] Where: Z α / 2 is the critical value of the standard normal distribution, corresponding to the preset confidence level. For example, the Z α / 2 ≈1.96 corresponding to the 95% confidence level; σ is the estimated value of the overall standard deviation. For the network information of the preset theme in this application, it can be estimated in advance by combining the historical network information of the same type of theme; E is the preset sampling error, which can be set according to the influence degree of the preset theme, usually between 1% - 5%.
[0073] Among them, the above three parameters can directly affect the size of the finally determined sample size. Generally speaking, the higher the confidence level, or the larger σ, or the smaller E, the larger the network information to be sampled, and the higher the accuracy of the sampled sample.
[0074] Among them, after determining the first sample quantity, in combination with the weight of each group determined in step S102, the quantity of network information extracted from each group can be determined as the second sample quantity corresponding to each group. Among them, the sum of the second sample quantities corresponding to each group should be equal to the first sample quantity. Continuing with the above example, the sources of network information can be divided into three categories: news media, opinion leaders, and ordinary users. Assuming that among all the network information obtained, the ratio of news media, opinion leaders, and ordinary users is 3:2:5, and the first sample quantity calculated according to the foregoing formula (1) is n, then the second sample quantity to be extracted from all the network information in the news media group should be 0.3n; the second sample quantity to be extracted from all the network information in the opinion leader group should be 0.2n, and the second sample quantity to be extracted from all the network information in the ordinary user group should be 0.5n.
[0075] Continuing with this embodiment, in step S104, the device 100 can randomly extract the corresponding second sample quantity of network information from each group as the sample network information according to the second sample quantity corresponding to each group determined in step S103, and a total of the first sample quantity of sample network information is obtained.
[0076] In this embodiment, for the preset theme, through steps S101 to S104, multi-source network information within a preset statistical period can be obtained from the Internet, processed and sampled, and high-precision sample network information can be obtained.
[0077] In order to obtain more information from the obtained sample network information, improve the analysis efficiency and depth based on the sample network information in the subsequent process, better mine valuable information, and provide a basis for decision-making, the obtained sample network information can also be subjected to clustering processing. Optionally, this network information processing method further includes:
[0078] S105 preprocess each sample network information, convert each preprocessed sample network information into a sentence vector, and perform dimensionality reduction processing on the sentence vector to obtain the sample data corresponding to the sample network information;
[0079] S106 perform clustering processing on all the sample data to obtain several clusters, and calculate the internal similarity index and the inter-cluster distance of each cluster. Among them, each cluster includes one or more data points, one of the data points is the core point, and each data point corresponds to a piece of sample data.
[0080] In this alternative embodiment, in step S105, for the first sample quantity of sample network information obtained after steps S101 to S104, device 100 may further preprocess each piece of sample network information, convert each piece of preprocessed sample network information into a sentence vector, and perform dimensionality reduction processing on the sentence vector to obtain the sample data corresponding to the sample network information.
[0081] Among them, device 100 also preprocesses each piece of sample network information in the first sample quantity of sample network information obtained. For example, duplicate removal, filtering of null values, and / or short texts, etc., are performed to obtain the preprocessed text. Then, an open-source sentence vector model, such as the Zhipu BGE model, etc., can be used to convert the text corresponding to each piece of sample network information into a sentence vector, that is, a numerical vector. Then, dimensionality reduction processing is performed on the sentence vector corresponding to each sample network information to obtain the sample data corresponding to the sample network information. Among them, dimensionality reduction processing is a practical and commonly used step in text analysis and natural language processing. The initial sentence vector usually has a high dimension, and the data calculation cost in the high-dimensional space is very high, while the data calculation efficiency can be improved in the lower-dimensional space. Performing dimensionality reduction processing on the initial sentence vector can retain the most important and most discriminative information in it, remove or compress those dimensions that contribute little to distinguishing the meanings of different texts; it can also improve the distribution of data in the low-dimensional space, making similar texts closer in space, which is beneficial to clustering processing of the dimensionality-reduced data. After dimensionality reduction, the initial sentence vector is still a numerical representation that quantifies the text semantics, but in the new low-dimensional space, the number of elements (i.e., feature dimensions) included in each sentence vector is reduced. In this alternative embodiment, a dimensionality reduction algorithm, such as the UMAP (Uniform Manifold Approximation and Projection) algorithm, can be used to perform dimensionality reduction processing on the initial sentence vector corresponding to each sample network information, and the dimension of the output sentence vector can be specified, mapping the initial sentence vector corresponding to each sample network information to a specified-dimensional space to obtain the sample data corresponding to the sample network information after dimensionality reduction processing.
[0082] Continuing in this alternative embodiment, in step S106, device 100 may perform clustering processing on the sample data after dimensionality reduction processing corresponding to all sample network information to obtain several clusters, and calculate the internal similarity index and the inter-cluster distance of each cluster. Among them, each cluster includes one or more data points, and one of the data points is the core point, and each data point corresponds to a piece of sample data.
[0083] Among them, clustering algorithms such as the HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise) algorithm can be used to perform clustering fitting on the sample data obtained after dimensionality reduction processing of the sentence vectors corresponding to each sample network information. All sample data are clustered into several clusters, and the obtained clustering information includes at least the cluster label, that is, the label of the cluster to which each data point corresponding to the sample data belongs; the cluster size, that is, the number of data points included in the cluster, the core point (corresponding to the core sample data, and the content of the core sample data can be traced back to explain the representative meaning of the belonging cluster), etc. In an actual application scenario, for example, the open-source sklearn library can be used to perform clustering processing on all sample data, and several clusters including the above clustering information can be obtained.
[0084] Among them, the internal similarity index of each cluster can be calculated, such as the average similarity, the within-cluster variance, etc., to evaluate the clustering effect of each cluster. Among them, the higher the internal similarity index of each cluster, the more similar the sample data corresponding to each data point within the cluster are to each other. Among them, the internal similarity index of each cluster can be calculated and determined according to the clustering result. The calculation process of the internal similarity index of each cluster can be referred to as follows: First, determine a suitable similarity or distance metric as the internal similarity index to quantify the similarity or difference between data points within the cluster. Among them, common metrics include Euclidean distance, cosine similarity, Manhattan distance, etc. Then, calculate the similarity or distance between all pairs of data points within the cluster to obtain an intra-cluster distance matrix or similarity matrix. Among them, if similarity is used as the metric, calculate the average of the similarities between all pairs of data points within the cluster as the average similarity. If distance is used as the metric, calculate the mean of the sum of the squares of the distances from all data points within the cluster to the centroid or median / mean data point of the cluster as the within-cluster variance to reflect the compactness of the cluster. The silhouette coefficient (an index that comprehensively considers the within-cluster cohesion and between-cluster separation) can also be used as the internal similarity index to measure, and its value range is from -1 to 1. The closer it is to 1, the better the clustering effect. Which internal similarity index to choose specifically should be determined in combination with the specific application scenario and the selected clustering algorithm. For example, if the HDBSCAN algorithm is used to perform clustering processing on all sample data, the shapes of the obtained several clusters may be arbitrary. At this time, the silhouette coefficient may be more able to reflect the compactness and separation of the clusters. Among them, after calculating the internal similarity index of each cluster, comparing the indexes of different clusters can know the degree of similarity between the data points within each cluster.
[0085] Among them, to evaluate the clustering effect of clustering each sample data, in addition to calculating the internal similarity index of each cluster, the distance between clusters can also be calculated to evaluate the separation degree of different clusters. For example, the shortest distance, the maximum distance, the average distance, etc. between different clusters can be used to evaluate the separation degree of different clusters. The greater the distance between clusters, the more obvious the distinction between different clusters, indicating that the clustering effect is better.
[0086] In this alternative embodiment, after steps S105 and S106, the sentence vectors of the sample network information are transformed to obtain the corresponding text, and then clustering is performed based on the text similarity. After clustering, processing a small amount of sample network information can more accurately identify and classify content with similar semantics, refine the analysis granularity, obtain more accurate analysis results, and also improve the depth and fineness of the aggregation analysis. The obtained core sample network information is more typical and explanatory.
[0087] To further reflect the representativeness and importance of a single sample network information, the weight of each sample network information can also be combined when processing the clustering results.
[0088] Optionally, this network information processing method further includes:
[0089] S107 Determine the source weight of each sample network information according to the source category of each sample network information, and determine the weight of the sample network information according to its grouping weight and source weight;
[0090] S108 Determine the analysis type corresponding to the cluster according to the sample network information corresponding to the core point of each cluster, and determine the weight and weight ratio of the analysis type corresponding to the cluster according to the weights of the sample network information corresponding to each data point in each cluster.
[0091] In this alternative embodiment, in step S107, for the first sample number of sample network information obtained after steps S101 to S104, device 100 can also determine the source weight of each sample network information according to the source category of each sample network information, and determine the weight of the sample network information according to the grouping weight of the group where the sample network information is located determined in step S102 and its source weight.
[0092] Among them, when calculating the weight of each sample network information, both the grouping weight of the group where the sample network information is located and its individual weight should be considered. Among them, the individual weight of a sample network information is mainly reflected in its publishing source, that is, the weight of its information source.
[0093] Therefore, the weight of a sample network information can be considered as the product of the group weight of the group it belongs to and its individual weight. Among them, the individual weight of a sample network information represents the influence and authority of different network information sources, and can be preset by a number of authoritative and representative experts and scholars in combination with relevant theories of information communication and practical experience on relevant topics. Exemplarily, the individual weights of sample network information from some different sources can be set with reference to Table 1 below.
[0094] Table 1
[0095]
[0096]
[0097] Continuing with this alternative embodiment, in step S108, the device 100 can determine the analysis type corresponding to each cluster according to the sample network information corresponding to the core point of each cluster obtained by the clustering process in step S107, and determine the weight and weight ratio of the analysis type corresponding to the cluster according to the weights of the sample network information corresponding to each data point in each cluster.
[0098] Among them, in step S107, all sample data are clustered to obtain a clustering result, including obtaining several clusters. Since the sample data corresponding to the data points within a cluster (i.e., the corresponding sample network information) have a high text similarity, it can be considered that the sample network information corresponding to the data points within the same cluster has common analysis characteristics. The analysis type corresponding to each cluster can be determined according to the content of the sample network information corresponding to the core point of each cluster, and the weight of the analysis type corresponding to the cluster and its proportion can be determined according to the weights of the sample network information corresponding to all data points within the cluster.
[0099] Among them, the sum of the weights of the sample network information corresponding to each data point within each cluster can be determined as the weight of the analysis type corresponding to the cluster, and the weight ratio of the analysis type corresponding to each cluster can be determined according to the weight of the analysis type corresponding to each cluster and the sum of the weights of the analysis types corresponding to all clusters. Exemplarily, after steps S101 to S107, after sampling and clustering the obtained N sample network information, the clustering result includes M clusters. Among them, the weight of the jth sample network information is w j , according to the content of the sample network information corresponding to the core point of the ith cluster C i , the corresponding analysis type is determined to be K i . The sum of the weights of the sample network information corresponding to each data point within the cluster C i can be used as the weight W i of the analysis type K i corresponding to the cluster C i , that is, Wi It can be calculated according to the following formula (2):
[0100] W i = ∑(w j * I j ), j = 1, 2, ..., N (2)
[0101] where I j is an indicator function, which takes the value of 1 when the sample j belongs to the i-th cluster C i and 0 otherwise.
[0102] In this example, the total weight W total of all analysis types corresponding to these M clusters can be calculated according to the following formula (3):
[0103] W total = ∑(W i ), i = 1, 2, ..., M (3)
[0104] The weight proportion P(K i ) of the analysis type K i corresponding to the cluster C i can be calculated according to the following formula (4):
[0105] P(K i ) = W i / W total , i = 1, 2, ..., M (4)
[0106] Suppose that after steps S101 to S107, 8 sample network information is obtained, and after clustering, the clustering result includes 3 clusters C1 to C3. Among them, cluster C1 includes 2 sample network information with weights w1 = 2 and w2 = 1 respectively, and the sample network information corresponding to its core point can determine the analysis type K1; cluster C2 includes 3 sample network information with weights w3 = 3, w4 = 1, and w5 = 2 respectively, and the sample network information corresponding to its core point can determine the analysis type K2; cluster C2 includes 3 sample network information with weights w6 = 2, w7 = 2, and w8 = 2 respectively, and the sample network information corresponding to its core point can determine the analysis type K3. Then, according to the above formulas (2) to (4), it can be obtained that: the weight W1 of the analysis type K1 corresponding to cluster C1 is 3, the weight W2 of the analysis type K2 corresponding to cluster C2 is 6, the weight W3 of the analysis type K3 corresponding to cluster C3 is 6, the total weight W total of these 3 clusters is 15, the weight proportion P(K1) of the analysis type K1 corresponding to cluster C1 is 3 / 15 = 20%, the weight proportion P(K2) of the analysis type K2 corresponding to cluster C2 is 6 / 15 = 40%, and the weight proportion P(K3) of the analysis type K3 corresponding to cluster C3 is 6 / 15 = 40%.
[0107] Among them, the average value of the weights of the sample network information corresponding to all data points within each cluster can be determined as the weight of the analysis type corresponding to this cluster, and the weight ratio of the analysis type corresponding to each cluster can be determined according to the weight of the analysis type corresponding to each cluster and the sum of the weights of the analysis types corresponding to all clusters. Exemplarily, after performing sampling and clustering processing on the obtained N pieces of sample network information through steps S101 to S107, the clustering result includes M clusters. Among them, the weight of the j-th piece of sample network information is w j , according to the content of the sample network information corresponding to the core point of the i-th cluster C i , the corresponding analysis type is determined to be K i . The average value of the weights of the sample network information corresponding to all data points within cluster C i can be used as the weight W i of the analysis type K i corresponding to cluster C i , that is, W i can be calculated according to the following formula (5):
[0108] W i = ∑(w j * I j ) / ∑I j , j = 1, 2,..., N (5)
[0109] Among them, I j is an indicator function, which takes the value of 1 when sample j belongs to the i-th cluster C i , and 0 otherwise.
[0110] In this example, the total weight W total of all analysis types corresponding to these M clusters can be calculated according to the above formula (3), and the weight ratio P(K i ) of the analysis type K i corresponding to the i-th cluster C i can be calculated according to the above formula (4).
[0111] In this alternative embodiment, the analysis type corresponding to each cluster can be determined in combination with the actual application scenario. For example, based on the sampled sample network information, aggregation is performed, and then for each cluster in the aggregation result, the content of the sample network information corresponding to its core point is analyzed and summarized to determine the analysis type, such as determining the analysis type as the public opinion sentiment dimension / intensity, public opinion view, etc., which can be used in application scenarios such as sentiment analysis and public opinion view mining. The various analysis types can also be structurally processed and visually displayed in a visual manner to provide support for formulating measures.
[0112] Figure 2Schematic diagram of a network information processing device according to another aspect of the present application. In one embodiment, the device includes:
[0113] A first module 210, configured to obtain multi-source network information for a preset statistical period related to a preset theme, perform preprocessing to obtain a number of network information, and determine the source category of each network information;
[0114] A second module 220, configured to group all the network information and determine the weight of each group according to the number of all the network information and the number of network information included in each group;
[0115] A third module 230, configured to determine, based on preset parameters, the number of network information to be extracted from all the network information as the first sample quantity, and determine, according to the weight of each group and the first sample quantity, the number of network information to be extracted from each group as the second sample quantity corresponding to the group, wherein the sum of the second sample quantities corresponding to each group is equal to the first sample quantity;
[0116] A fourth module 240, configured to randomly extract the number of network information corresponding to the second sample quantity of each group from each group as sample network information.
[0117] In this embodiment, the device is deployed or integrated in the device 100 that executes the foregoing method embodiments and / or optional embodiments.
[0118] In this embodiment, through the first module 210 of the device, the target theme to be analyzed and processed in the actual application scenario is used as the preset theme, and multi-source network information for a preset statistical period related to the preset theme can be obtained from the Internet, and preprocessing is performed to obtain a number of network information, and the source category of each network information is determined.
[0119] Continuing in this embodiment, through the second module 220 of the device, all the network information obtained through the first module 210 can be grouped to obtain a number of groups, where each group includes a number of network information. Then, according to the number of all the network information and the number included in each group, the weight of each group is determined, and the ratio of the number included in each group to the number of all the network information can be used as the weight of the group.
[0120] Continuing in this embodiment, through the third module 230 of the device, based on preset parameters, the number of network information to be extracted from all the network information is determined as the first sample quantity, and according to the weight of each group and the first sample quantity, the number of network information to be extracted from each group is determined as the second sample quantity corresponding to the group, wherein the sum of the second sample quantities corresponding to each group is equal to the first sample quantity.
[0121] Continuing in this embodiment, through the fourth module 240 of the device, according to the second sample quantity corresponding to each packet determined by the third module 230, the corresponding second sample quantity of network information can be randomly selected from each packet as sample network information, and a total of the first sample quantity of sample network information is obtained.
[0122] Optionally, the network information processing device further includes:
[0123] A fifth module 250, configured to preprocess each sample network information, convert each preprocessed sample network information into a sentence vector, and perform dimensionality reduction processing on the sentence vector to obtain sample data corresponding to the sample network information;
[0124] A sixth module 260, configured to perform clustering processing on all sample data to obtain several clusters, and calculate the internal similarity index and inter-cluster distance of each cluster, where each cluster includes one or more data points, one of the data points is a core point, and each data point corresponds to one sample data.
[0125] In this optional embodiment, through the fifth module 250 of the device, for the first sample quantity of sample network information obtained through the first module 210 to the fourth module 240, each sample network information can be preprocessed, each preprocessed sample network information can be converted into a sentence vector, and dimensionality reduction processing is performed on the sentence vector to obtain sample data corresponding to the sample network information.
[0126] Continuing in this optional embodiment, through the sixth module 260 of the device, clustering processing can be performed on the sample data after dimensionality reduction corresponding to all sample network information to obtain several clusters, and calculate the internal similarity index and inter-cluster distance of each cluster, where each cluster includes one or more data points, one of the data points is a core point, and each data point corresponds to one sample data.
[0127] Optionally, the network information processing device further includes:
[0128] A seventh module 270, configured to determine the source weight according to the source category of each sample network information, and determine the weight of the sample network information according to its grouping weight and source weight;
[0129] An eighth module 280, configured to determine the analysis type corresponding to the cluster according to the sample network information corresponding to the core point of each cluster, and determine the weight and weight ratio of the analysis type corresponding to the cluster according to the weights of the sample network information corresponding to each data point in each cluster.
[0130] In this alternative embodiment, through the seventh module 270 of the device, for the first sample quantity bar of sample network information obtained through the first module 210 to the fourth module 240, the source weight can be determined according to the source category of each sample network information, and according to the group weight of the group where the sample network information is located determined by the second module 220 and its source weight, the weight of the sample network information can be determined. Among them, when calculating the weight of each sample network information, both the group weight of the group where the sample network information is located and its individual weight should be considered. Among them, the individual weight of a sample network information is mainly reflected in its publishing source, that is, the weight of its information source. Therefore, the weight of a sample network information can be considered as the product of the group weight of the group where it is located and its individual weight. Among them, the individual weight represents the influence and authority of different network information publishing sources, and can be preset by a number of authoritative and representative experts and scholars in combination with relevant theories of information communication and practical experience of relevant topics.
[0131] Continuing in this alternative embodiment, through the eighth module 280 of the device, the analysis type corresponding to each cluster can be determined according to the sample network information corresponding to the core point of each cluster obtained through the seventh module 270, and according to the weight of the sample network information corresponding to each data point in each cluster, the weight and weight ratio of the analysis type corresponding to the cluster can be determined.
[0132] In each embodiment and / or alternative embodiment of the above device, the parts not mentioned in the method steps executed by each module are the same as the foregoing relevant method embodiments and / or alternative embodiments, and will not be repeated here.
[0133] According to another aspect of the present application, a computer-readable medium is further provided. The computer-readable medium stores computer-readable instructions, and the computer-readable instructions can be executed by a processor to implement the foregoing method embodiments.
[0134] It should be noted that in each method embodiment and / or alternative embodiment of the present application, the execution order of each step may not be strictly limited. As long as each method embodiment and / or alternative embodiment can solve the defects existing in the prior art, achieve the invention purpose of the present application, and obtain beneficial effects. Each method embodiment and / or alternative embodiment of the present application can be implemented in software and / or a combination of software and hardware. The software programs involved in the present application can be executed by a processor to implement the steps or functions of the foregoing embodiments. Similarly, the software programs (including related data structures) of the present application can be stored in a computer-readable recording medium.
[0135] In addition, part or all of the present application can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can call or provide the methods and / or technical solutions according to the present application through the operations of the computer. The program instructions for calling the methods of the present application may be stored in a fixed or removable recording medium, and / or transmitted through a data stream in a broadcast or other signal-bearing medium, and / or stored in the working memory of a computer device operating according to the program instructions.
[0136] According to another aspect of the present application, there is also provided a network information processing device, which includes: a memory storing computer program instructions and one or more processors for executing the program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to run the methods and / or technical solutions of the foregoing embodiments.
[0137] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices stated in the apparatus claims can also be implemented by one unit or device through software and / or hardware. First, second, etc. are used to denote names and do not denote any particular order.
Claims
1. A method for processing network information, characterized in that The method includes: acquiring multi-source network information of a preset statistical period related to a preset theme, preprocessing the information to obtain a number of network information, and determining the source category of each piece of network information; grouping all the network information, and determining the weight of each group according to the number of all the network information and the number of network information included in each group; based on preset parameters, determining the number of network information to be extracted from all the network information as the first sample number, and determining the number of network information to be extracted from each group as the second sample number corresponding to the group according to the weight of each group and the first sample number, wherein the sum of the second sample numbers corresponding to each group is equal to the first sample number; randomly extracting the number of network information corresponding to the second sample number of each group from each group as sample network information.
2. The method according to claim 1, characterized in that The grouping of all the network information includes: determining the public opinion development stage of the preset theme in which each network information is located; grouping all the network information according to the public opinion development stage.
3. The method according to claim 1, wherein The grouping of all the network information includes: determining the key event nodes of the preset theme to which each network information belongs; grouping all the network information according to the key event nodes.
4. The method according to claim 1, characterized in that The method further includes: preprocessing each piece of sample network information, converting each preprocessed piece of sample network information into a sentence vector, and performing dimensionality reduction processing on the sentence vector to obtain sample data corresponding to the sample network information; performing clustering processing on all the sample data to obtain a number of clusters, and calculating the internal similarity index and the inter-cluster distance of each cluster, wherein each cluster includes one or more data points, one of which is the core point, and each data point corresponds to a piece of sample data.
5. The method according to claim 4, wherein The method further includes: determining the source weight according to the source category of each piece of sample network information, and determining the weight of the sample network information according to its grouping weight and source weight; determining the analysis type corresponding to the cluster according to the sample network information corresponding to the core point of each cluster, and determining the weight and weight ratio of the analysis type corresponding to the cluster according to the weight of the sample network information corresponding to each data point in each cluster.
6. A processing device for network information, characterized in that, The device includes: a first module for acquiring multi-source network information of a preset statistical period related to a preset theme, preprocessing the information to obtain a number of network information, and determining the source category of each piece of network information; a second module for grouping all the network information, and determining the weight of each group according to the number of all the network information and the number of network information included in each group; a third module for determining, based on preset parameters, the number of network information to be extracted from all the network information as the first sample number, and determining the number of network information to be extracted from each group as the second sample number corresponding to the group according to the weight of each group and the first sample number, wherein the sum of the second sample numbers corresponding to each group is equal to the first sample number; a fourth module for randomly extracting the number of network information corresponding to the second sample number of each group from each group as sample network information.
7. The device according to claim 6, characterized in that The device further includes: A fifth module, configured to preprocess each sample network information, convert each preprocessed sample network information into a sentence vector, and perform dimensionality reduction processing on the sentence vector to obtain sample data corresponding to the sample network information; A sixth module, configured to perform clustering processing on all sample data to obtain a plurality of clusters, and calculate an internal similarity index and an inter-cluster distance for each cluster, where each cluster includes one or more data points, one of the data points is a core point, and each data point corresponds to a piece of sample data.
8. The device according to claim 7, characterized in that, The apparatus further includes: A seventh module, configured to determine a source weight according to the source category of each sample network information, and determine the weight of the sample network information according to its grouping weight and source weight; An eighth module, configured to determine an analysis type corresponding to the cluster according to the sample network information corresponding to the core point of each cluster, and determine the weight and weight ratio of the analysis type corresponding to the cluster according to the weights of the sample network information corresponding to each data point in each cluster.
9. A computer-readable medium, characterized in that computer-readable instructions are stored thereon, and the computer-readable instructions are executed by a processor to implement the method according to any one of claims 1 to 5.
10. A processing device for network information, characterized in that, The device includes: one or more processors; and a memory storing computer-readable instructions, the computer-readable instructions, when executed, causing the processor to perform the operations of the method according to any one of claims 1 to 5.