Network public opinion data acquisition method and system and electronic equipment
By using preset rules and semantic recognition models for classification and dynamic adjustment in online public opinion data collection, the inefficiency of traditional online public opinion data crawling is solved, enabling efficient acquisition of timely data and improving the practicality of public opinion monitoring and analysis.
Patent Information
- Application Number
- CN202511602636.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-04
- Publication Date
- 2026-02-27
AI Technical Summary
Traditional methods of crawling online public opinion data lack specificity and efficiency, making it difficult to quickly extract effective information from massive amounts of data. Furthermore, they cannot adjust strategies in a timely manner to obtain key data from high-timeliness platforms, thus missing the best opportunity to respond to public opinion.
Online public opinion data is obtained by pre-setting crawling rules, classified using a semantic recognition model, and the data with the lowest correlation is extracted. The crawling strategy is dynamically adjusted to prioritize the acquisition of data from high-timeliness platforms.
It enables efficient crawling of multi-source online public opinion data, quickly focuses on differential data, improves the ability to respond to hot events, and provides comprehensive and accurate support for public opinion monitoring and analysis.
Smart Images

Figure CN121579759A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of data processing, and particularly relates to a network public opinion data collection method and system and an electronic device. BACKGROUND
[0002] In the era of network information explosion, network public opinion has a significant impact on social public opinion trends, corporate reputation, and public event development. Currently, traditional network public opinion data crawling methods lack pertinence and efficiency, making it difficult to quickly obtain effective information from massive data, and the data crawling of different data sources lacks reasonable planning, resulting in resource waste.
[0003] In addition, in the face of sudden public opinion events, it is not possible to timely and efficiently adjust data crawling strategies to obtain key data on high timeliness platforms, missing the best opportunity to respond to public opinion. SUMMARY
[0004] The present application provides a network public opinion data collection method, system and electronic device, which is used to solve the technical problem of being unable to obtain key data on high timeliness platforms and missing the best opportunity to respond to public opinion.
[0005] In a first aspect, the present application provides a network public opinion data collection method, comprising:
[0006] obtaining at least one network public opinion data in a first predetermined time period according to a predetermined first crawling rule, wherein one network public opinion data contains one public opinion information and one data source information corresponding to the one public opinion information;
[0007] classifying the at least one network public opinion data using a predetermined data classification strategy according to the keywords in the network public opinion data, to obtain at least one network public opinion data set, wherein one network public opinion data set contains at least one network public opinion data of the same data type;
[0008] extracting network public opinion data subsets from each network public opinion data set according to a predetermined extraction rule, wherein one network public opinion data subset contains a first network public opinion data and a second network public opinion data with the lowest data correlation in the same network public opinion data set;
[0009] determining whether the data source information of the first network public opinion data and the data source information of the second network public opinion data in a certain network public opinion data subset are the same;
[0010] if they are the same, determining the set source information of the certain network public opinion data subset according to the data source information of the first network public opinion data and the data source information of the second network public opinion data;
[0011] The collection source information of each network public opinion data subset is classified, and a target network public opinion data subset with the largest number of categories is selected from the network public opinion data subsets, wherein the target network public opinion data subset corresponds to target collection source information;
[0012] According to the target collection source information, at least one network public opinion data in a second preset time period is obtained by using a preset second crawling rule, wherein the second preset time period is a time period adjacent to the first preset time period.
[0013] In a second aspect, the present application provides a network public opinion data collection system, comprising:
[0014] The first obtaining module is configured to obtain at least one network public opinion data in a first preset time period according to a preset first crawling rule, wherein one network public opinion data contains one public opinion information and one data source information corresponding to the public opinion information;
[0015] The classification module is configured to classify the at least one network public opinion data according to a keyword in the network public opinion data by using a preset data classification strategy, so as to obtain at least one network public opinion data set, wherein one network public opinion data set contains at least one network public opinion data of the same data type;
[0016] The extraction module is configured to extract network public opinion data subsets from each network public opinion data set according to a preset extraction rule, wherein one network public opinion data subset contains first network public opinion data and second network public opinion data with the lowest data correlation in the same network public opinion data set;
[0017] The judgment module is configured to judge whether the data source information of the first network public opinion data and the data source information of the second network public opinion data in a certain network public opinion data subset are the same;
[0018] The determination module is configured to determine the collection source information of the certain network public opinion data subset according to the data source information of the first network public opinion data and the data source information of the second network public opinion data if the data source information of the first network public opinion data and the data source information of the second network public opinion data are the same;
[0019] The selection module is configured to classify the collection source information of each network public opinion data subset, and select a target network public opinion data subset with the largest number of categories from the network public opinion data subsets, wherein the target network public opinion data subset corresponds to target collection source information;
[0020] The second acquisition module is configured to acquire at least one network public opinion data in a second preset time period according to the target set source information and by using a preset second crawling rule, wherein the second preset time period is a time period adjacent to the first preset time period.
[0021] In a third aspect, an electronic device is provided, which includes at least one processor and a memory connected with the at least one processor in communication, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the network public opinion data acquisition method of any embodiment of the present application.
[0022] In a fourth aspect, the present application further provides a computer readable storage medium having a computer program stored thereon, and the program instructions are executed by a processor to enable the processor to perform the steps of the network public opinion data acquisition method of any embodiment of the present application.
[0023] The network public opinion data acquisition method, system and electronic device of the present application efficiently crawl multi-source network public opinion data, accurately classify by a semantic recognition model, quickly focus on differential data, mine potential associated public opinions, dynamically adjust the crawling strategy based on the data source, preferentially acquire high timeliness platform data, and improve the response capability to hot events. At the same time, the present application can effectively provide comprehensive and accurate data support for public opinion monitoring and analysis, meet actual needs, has outstanding practicability, and effectively promotes the development of network public opinion research. BRIEF DESCRIPTION OF DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0025] Figure 1 A flow chart of a network public opinion data acquisition method provided by an embodiment of the present application;
[0026] Figure 2 A structural block diagram of a network public opinion data acquisition system provided by an embodiment of the present application;
[0027] Figure 3 A structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0028] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0029] Referring to Figure 1 , a flowchart of a network public opinion data collection method is shown.
[0030] As Figure 1 shown, the network public opinion data collection method specifically includes the following steps:
[0031] Step S101, acquiring at least one network public opinion data in a first preset time period according to a first preset crawling rule, wherein one network public opinion data contains one public opinion information and one data source information corresponding to the one public opinion information.
[0032] In this step, the number of crawling objects is acquired, the transmission sub-queues are set according to the number of objects, and the same number of network public opinion data is acquired from each transmission sub-queue at each first time to obtain at least one network public opinion data in the first preset time period, wherein the each first time is a time in the first preset time period.
[0033] For example, there are five crawling objects, so the transmission sub-queues are set to correspond to each crawling object one by one, that is, there are five transmission sub-queues. And the same number of network public opinion data is acquired from the five transmission sub-queues at each time.
[0034] Step S102, classifying the at least one network public opinion data according to a preset data classification strategy based on the keywords in the network public opinion data to obtain at least one network public opinion data set, wherein one network public opinion data set contains at least one network public opinion data of the same data type.
[0035] In this step, at least one keyword in the network public opinion data is obtained, and the at least one keyword is input into a preset semantic recognition model. The semantic recognition model outputs a semantic category corresponding to the at least one keyword. It is determined whether the number of each semantic category is the same. If not, a target semantic category with the largest number of semantic categories is selected from each semantic category, and the target semantic category is defined as a data semantic category of the network public opinion data. If there are at least two semantic categories with the largest number of semantic categories in each semantic category, any one of the at least two semantic categories is defined as a data semantic category of the network public opinion data. Each network public opinion data with the same data semantic category is clustered to obtain at least one network public opinion set.
[0036] It should be noted that the semantic recognition model can be obtained by training a corresponding traditional neural network.
[0037] For example, in the network public opinion data A, from the beginning of the sentence to the end of the sentence, the keywords a, b, c and d are sequentially contained. The keywords a, b, c and d are input into the semantic recognition model, and the semantic recognition model outputs the semantic categories 1, 2, 3 and 4 corresponding to the keywords a, b, c and d, respectively. The semantic categories 1, 2 and 3 are the same, and the semantic category 4 is a separate semantic category. Therefore, the semantic categories 1, 2 or 3 are defined as the data semantic category of the network public opinion data A.
[0038] If the semantic categories 1, 2, 3 and 4 are not the same, the semantic category 1 is directly defined as the data semantic category of the network public opinion data A based on the order of the keywords.
[0039] In step S103, a network public opinion data sub-set is extracted from each network public opinion data set according to a preset extraction rule, wherein the network public opinion data sub-set contains the first network public opinion data and the second network public opinion data with the lowest data correlation in the same network public opinion data set.
[0040] In this step, each network public opinion data in a certain network public opinion data set is stratified based on the number of keywords, obtaining at least one layer of network public opinion data, wherein the network public opinion data in one layer of network public opinion data contains at least one network public opinion data; each network public opinion data in a certain layer of network public opinion data with the least number of keywords is defined as the first target network public opinion data, and each network public opinion data in another layer of network public opinion data with the most number of keywords is defined as the second target network public opinion data; the keyword similarity between each first target network public opinion data and each second target network public opinion data is obtained, and the smallest keyword similarity is selected from each keyword similarity, wherein the keyword similarity is obtained by superimposing the similarity between each first keyword and each second keyword, the first keyword is any keyword in the first target network public opinion data, the second keyword is any keyword in the second target network public opinion data, the similarity between a first keyword and a second keyword is 1 when the semantic categories are the same, otherwise it is 0; and a first target network public opinion data and a second target network public opinion data corresponding to a keyword similarity are extracted into a certain network public opinion data subset.
[0041] Assume:
[0042] First target data: {keywords: ["cold", "prevention and control"]} (2 keywords)
[0043] Second target data: {keywords: ["virus", "spread", "prevention and control"]} (3 keywords)
[0044] Assume semantic categories:
[0045] "Cold": Category A
[0046] "Prevention and control": Category B
[0047] "Virus": Category C
[0048] "Spread": Category D
[0049] Calculate similarity:
[0050] "Cold" vs "virus": different categories -> 0
[0051] "Cold" vs "spread": different categories -> 0
[0052] "Cold" vs "prevention and control": different categories -> 0
[0053] "Prevention and control" vs "virus": different categories -> 0
[0054] "Prevention and control" vs "spread": different categories -> 0
[0055] "Prevention and Control" vs. "Prevention and Control": Same Category → 1
[0056] Total similarity, or data correlation: 0+0+0+0+0+1=1.
[0057] In this implementation, by stratifying online public opinion data based on the number of keywords, the data layers with the fewest and most keywords can be quickly located. This allows for focusing on specific target data sets, reducing the scope of data processing, and improving processing efficiency. Furthermore, by defining keyword similarity using semantic categories, the degree of keyword association between two public opinion data sets is quantified in a binary manner of 0 or 1. This makes the similarity calculation results intuitive and has clear semantic directionality, facilitating subsequent analysis and comparison. Selecting the data pairs corresponding to the minimum keyword similarity helps to discover public opinion data with the greatest differences at the keyword level, which may reveal potential connections between different viewpoints, positions, or angles of event development, providing a new perspective for public opinion analysis.
[0058] Step S104: Determine whether the data source information of the first online public opinion data and the data source information of the second online public opinion data in a certain subset of online public opinion data are the same.
[0059] In one specific embodiment, after determining whether the data source information of the first online public opinion data and the data source information of the second online public opinion data in a certain online public opinion data subset are the same, if they are not the same, then no set set source information is set for the certain online public opinion data subset.
[0060] Step S105: If they are the same, then determine the set source information of the certain subset of online public opinion data based on the data source information of the first online public opinion data and the data source information of the second online public opinion data.
[0061] In this step, the data source information of the first network public opinion data or the data source information of the second network public opinion data is directly defined as the set source information of a certain subset of network public opinion data.
[0062] Step S106: Classify the source information of each subset of online public opinion data, and select the target subset of online public opinion data with the largest number of categories from each subset of online public opinion data, wherein the target subset of online public opinion data corresponds to the source information of the target set.
[0063] In this step, source information of the same type is aggregated to obtain at least one category, and the target category with the largest number of categories is selected from the at least one category. The subset of online public opinion data corresponding to the target category is defined as the target online public opinion data subset.
[0064] In step S107, at least one network public opinion data in a second preset time period is obtained according to the target set source information and by using a preset second crawling rule, wherein the second preset time period is a time period adjacent to the first preset time period.
[0065] In this step, a ratio of a target number of target network public opinion data sub-sets to a total number of all network public opinion data sub-sets is obtained, a target transmission sub-queue corresponding to the target set source information is obtained, a target number of network public opinion data obtained from the target transmission sub-queue at each second time is set according to the ratio, and at least one network public opinion data in the second preset time period is obtained, wherein the target number of network public opinion data obtained at a certain second time is a product of a total number of network public opinion data to be obtained at a certain second time and the ratio, and each second time is a time in the second preset time period.
[0066] In this embodiment, by obtaining at least one network public opinion data in a second preset time period according to the target set source information and by using a preset second crawling rule, the burst traffic of high timeliness platforms (such as Twitter and Douyin) can be quickly and relatively massively obtained.
[0067] In summary, the method of the present application efficiently crawls multi-source network public opinion data, accurately classifies by a semantic recognition model, quickly focuses on differential data, mines potential associated public opinions, dynamically adjusts the crawling strategy based on data sources, preferentially obtains high timeliness platform data, and improves the response capability to hot events. At the same time, it can effectively provide comprehensive and accurate data support for public opinion monitoring and analysis, meet actual needs, has outstanding practicality, and effectively promotes the research and development of network public opinion.
[0068] Please refer to Figure 2 which shows a structural block diagram of a network public opinion data collection system of the present application.
[0069] As Figure 2 shown, the first obtaining module 210, the classification module 220, the extraction module 230, the judgment module 240, the determination module 250, the selection module 260, and the second obtaining module 270.
[0070] The first acquisition module 210 is configured to acquire at least one network public opinion data in a first preset time period according to a preset first crawling rule, wherein one network public opinion data comprises one public opinion information and one data source information corresponding to the one public opinion information; the classification module 220 is configured to classify the at least one network public opinion data according to a keyword in the network public opinion data by using a preset data classification strategy to obtain at least one network public opinion data set, wherein one network public opinion data set comprises at least one network public opinion data of the same data type; the extraction module 230 is configured to extract a network public opinion data sub-set from each network public opinion data set according to a preset extraction rule, wherein one network public opinion data sub-set comprises a first network public opinion data and a second network public opinion data with the lowest data correlation degree in the same network public opinion data set; the judgment module 240 is configured to judge whether the data source information of the first network public opinion data and the data source information of the second network public opinion data in a certain network public opinion data sub-set are the same; the determination module 250 is configured to determine the set source information of the certain network public opinion data sub-set according to the data source information of the first network public opinion data and the data source information of the second network public opinion data if the data source information of the first network public opinion data and the data source information of the second network public opinion data are the same; the selection module 260 is configured to classify the set source information of each network public opinion data sub-set, and select a target network public opinion data sub-set with the largest number of categories from the each network public opinion data sub-set, wherein the target network public opinion data sub-set corresponds to a target set source information; and the second acquisition module 270 is configured to acquire at least one network public opinion data in a second preset time period according to the target set source information by using a preset second crawling rule, wherein the second preset time period is a time period adjacent to the first preset time period.
[0071] It should be understood that Figure 2 each step in the method described with reference to Figure 1 the modules in the above description correspond. Therefore, the operations and features described above for the method and the corresponding technical effects also apply to the modules in Figure 2 , and will not be described here again.
[0072] In some other embodiments, the present application also provides a computer readable storage medium having a computer program stored thereon, wherein the program instructions are executed by a processor to cause the processor to perform the network public opinion data acquisition method in any of the above method embodiments.
[0073] As an implementation form, the computer readable storage medium of the present application stores computer executable instructions, and the computer executable instructions are configured to:
[0074] According to a preset first crawling rule, at least one network public opinion data in a first preset time period is acquired, wherein one network public opinion data contains one public opinion information and one data source information corresponding to the one public opinion information;
[0075] According to a keyword in the network public opinion data, a preset data classification strategy is used to classify the at least one network public opinion data, to obtain at least one network public opinion data set, wherein one network public opinion data set contains at least one network public opinion data of the same data type;
[0076] According to a preset extraction rule, network public opinion data subsets are extracted from each network public opinion data set, wherein one network public opinion data subset contains a first network public opinion data and a second network public opinion data with the lowest data correlation degree in the same network public opinion data set;
[0077] It is judged whether the data source information of the first network public opinion data and the data source information of the second network public opinion data in a certain network public opinion data subset are the same;
[0078] If they are the same, the set source information of the certain network public opinion data subset is determined according to the data source information of the first network public opinion data and the data source information of the second network public opinion data;
[0079] The set source information of each network public opinion data subset is classified, and a target network public opinion data subset with the largest number of categories is selected from the network public opinion data subsets, wherein the target network public opinion data subset corresponds to a target set source information;
[0080] According to the target set source information, at least one network public opinion data in a second preset time period is acquired by using a preset second crawling rule, wherein the second preset time period is a time period adjacent to the first preset time period.
[0081] The computer readable storage medium can include a program storage area and a data storage area, wherein the program storage area can store an operating system and at least one application required by a function; the data storage area can store data created according to the use of the network public opinion data acquisition system. In addition, the computer readable storage medium can include a high-speed random access memory, and can also include a memory such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state memory device. In some embodiments, the computer readable storage medium can optionally include a memory remotely arranged relative to the processor, and these remote memories can be connected to the network public opinion data acquisition system through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0082] Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiment of the present invention, such as... Figure 3 As shown, the device includes a processor 310 and a memory 320. The electronic device may also include an input device 330 and an output device 340. The processor 310, memory 320, input device 330, and output device 340 can be connected via a bus or other means. Figure 3 Taking a bus connection as an example, the memory 320 is the computer-readable storage medium described above. The processor 310 executes various server functions and data processing by running non-volatile software programs, instructions, and modules stored in the memory 320, thereby implementing the network public opinion data collection method described in the above embodiment. The input device 330 can receive input digital or character information and generate key signal inputs related to user settings and function control of the network public opinion data collection system. The output device 340 may include a display screen or other display device.
[0083] The aforementioned electronic device can execute the method provided in the embodiments of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in this embodiment can be found in the method provided in the embodiments of the present invention.
[0084] In one implementation, the above-described electronic device is applied to a network public opinion data collection system for a client, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to:
[0085] At least one piece of online public opinion data within a first preset time period is obtained according to a preset first crawling rule, wherein one piece of online public opinion data contains one piece of public opinion information and a data source information corresponding to the one piece of public opinion information;
[0086] Based on the keywords in the online public opinion data, the at least one online public opinion data is classified using a preset data classification strategy to obtain at least one set of online public opinion data, wherein a set of online public opinion data contains at least one set of online public opinion data of the same data type.
[0087] According to the preset extraction rules, a subset of online public opinion data is extracted from each set of online public opinion data. Each subset of online public opinion data contains the first set of online public opinion data and the second set of online public opinion data with the lowest data correlation in the same set.
[0088] Determine whether the data source information of the first online public opinion data and the data source information of the second online public opinion data in a certain subset of online public opinion data are the same;
[0089] If they are the same, then the collection source information of the certain subset of online public opinion data is determined based on the data source information of the first online public opinion data and the data source information of the second online public opinion data;
[0090] The source information of each subset of online public opinion data is classified, and the target subset of online public opinion data with the largest number of categories is selected from each subset of online public opinion data. The target subset of online public opinion data corresponds to the source information of the target set.
[0091] Based on the source information of the target set, at least one piece of online public opinion data within a second preset time period is obtained using a preset second crawling rule, wherein the second preset time period is a time period adjacent to the first preset time period.
[0092] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.
[0093] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for collecting online public opinion data, characterized in that, include: At least one piece of online public opinion data within a first preset time period is obtained according to a preset first crawling rule, wherein one piece of online public opinion data contains one piece of public opinion information and a data source information corresponding to the one piece of public opinion information; Based on the keywords in the online public opinion data, the at least one online public opinion data is classified using a preset data classification strategy to obtain at least one set of online public opinion data, wherein a set of online public opinion data contains at least one set of online public opinion data of the same data type. According to the preset extraction rules, a subset of online public opinion data is extracted from each set of online public opinion data. Each subset of online public opinion data contains the first set of online public opinion data and the second set of online public opinion data with the lowest data correlation in the same set. Determine whether the data source information of the first online public opinion data and the data source information of the second online public opinion data in a certain subset of online public opinion data are the same; If they are the same, then the collection source information of the certain subset of online public opinion data is determined based on the data source information of the first online public opinion data and the data source information of the second online public opinion data; The source information of each subset of online public opinion data is classified, and the target subset of online public opinion data with the largest number of categories is selected from each subset of online public opinion data. The target subset of online public opinion data corresponds to the source information of the target set. Based on the source information of the target set, at least one piece of online public opinion data within a second preset time period is obtained using a preset second crawling rule, wherein the second preset time period is a time period adjacent to the first preset time period.
2. The method for collecting online public opinion data according to claim 1, characterized in that, The step of obtaining at least one piece of online public opinion data within a first preset time period according to a preset first crawling rule includes: The number of objects to be crawled is obtained, a transmission sub-queue is set according to the number of objects, and the same number of network public opinion data is obtained from each transmission sub-queue at each first moment, so as to obtain at least one network public opinion data within a first preset time period, wherein each first moment is a moment within the first preset time period.
3. The method for collecting online public opinion data according to claim 1, characterized in that, The step of classifying the at least one set of online public opinion data according to keywords in the online public opinion data using a preset data classification strategy to obtain at least one set of online public opinion data includes: Obtain at least one keyword from a certain online public opinion data, and input the at least one keyword into a preset semantic recognition model, wherein the semantic recognition model outputs the semantic category corresponding to the at least one keyword; Determine whether the number of each semantic category is the same; If they are not the same, the target semantic category with the most semantic categories is selected from all semantic categories, and the target semantic category is defined as the data semantic category of a certain network public opinion data; If at least two semantic categories are the same across all semantic categories, then any one of the at least two semantic categories is defined as the data semantic category of a certain online public opinion data. Clustering of online public opinion data with the same semantic category yields at least one set of online public opinion data.
4. The method for collecting online public opinion data according to claim 3, characterized in that, The step of extracting subsets of online public opinion data from each set of online public opinion data according to preset extraction rules includes: In a certain set of online public opinion data, each set of online public opinion data is layered based on the number of keywords to obtain at least one layer of online public opinion data, wherein each layer of online public opinion data contains at least one piece of online public opinion data; Each piece of online public opinion data in the layer with the fewest number of keywords is defined as the first target online public opinion data, and each piece of online public opinion data in the other layer with the most number of keywords is defined as the second target online public opinion data. The keyword similarity between each first target network public opinion data and each second target network public opinion data is obtained, and the keyword similarity with the smallest value is selected from each keyword similarity. The keyword similarity is obtained by superimposing the similarity between each first keyword and each second keyword. The first keyword is any keyword in the first target network public opinion data, and the second keyword is any keyword in the second target network public opinion data. When a first keyword and a second keyword have the same semantic category, the similarity between the first keyword and the second keyword is 1, otherwise it is 0. Then, a first target online public opinion data and a second target online public opinion data corresponding to the keyword similarity are extracted into a certain online public opinion data subset.
5. The method for collecting online public opinion data according to claim 1, characterized in that, After determining whether the data source information of the first online public opinion data and the data source information of the second online public opinion data in a certain subset of online public opinion data are the same, the method further includes: If they are different, then the source information of the collection will not be set for the certain subset of online public opinion data.
6. The method for collecting online public opinion data according to claim 1, characterized in that, The step of determining the collection source information of a certain subset of online public opinion data based on the data source information of the first online public opinion data and the data source information of the second online public opinion data includes: The data source information of the first network public opinion data or the data source information of the second network public opinion data is directly defined as the set source information of the certain network public opinion data subset.
7. The method for collecting online public opinion data according to claim 1, characterized in that, The process of classifying the source information of each subset of online public opinion data and selecting the target subset of online public opinion data with the largest number of categories from each subset includes: Aggregate source information of the same type to obtain at least one category, select the target category with the largest number of categories in the at least one category, and define the subset of online public opinion data corresponding to the target category as the target online public opinion data subset.
8. A method for collecting online public opinion data according to claim 2, characterized in that, The step of obtaining at least one piece of online public opinion data within a second preset time period based on the source information of the target set and using a preset second crawling rule includes: Obtain the ratio of the target number of the target online public opinion data subset to the total number of all online public opinion data subsets; Obtain the target transmission sub-queue corresponding to the source information of the target set, and set the target number of network public opinion data to be obtained from the target transmission sub-queue at each second time according to the ratio, so as to obtain at least one network public opinion data within the second preset time period. The target number of network public opinion data to be obtained at a certain second time is the product of the total number of network public opinion data to be obtained at a certain second time and the ratio. Each second time is a time within the second preset time period.
9. A network public opinion data collection system, characterized in that, include: The first acquisition module is configured to acquire at least one piece of online public opinion data within a first preset time period according to a preset first crawling rule, wherein one piece of online public opinion data includes one piece of public opinion information and a data source information corresponding to the one piece of public opinion information; The classification module is configured to classify the at least one piece of online public opinion data according to the keywords in the online public opinion data using a preset data classification strategy, so as to obtain at least one set of online public opinion data, wherein a set of online public opinion data contains at least one piece of online public opinion data of the same data type; The extraction module is configured to extract a subset of online public opinion data from each set of online public opinion data according to a preset extraction rule. Each subset of online public opinion data contains the first set of online public opinion data and the second set of online public opinion data with the lowest data correlation in the same set. The judgment module is configured to determine whether the data source information of the first network public opinion data and the data source information of the second network public opinion data in a certain subset of network public opinion data are the same; The module is configured to determine the set source information of a certain subset of online public opinion data based on the data source information of the first online public opinion data and the data source information of the second online public opinion data if they are the same. The selection module is configured to classify the source information of each subset of online public opinion data, and select the target subset of online public opinion data with the largest number of categories from each subset of online public opinion data, wherein the target subset of online public opinion data corresponds to the source information of the target set; The second acquisition module is configured to acquire at least one piece of online public opinion data within a second preset time period based on the source information of the target set and using a preset second crawling rule, wherein the second preset time period is a time period adjacent to the first preset time period.
10. An electronic device, characterized in that, include: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.