Data acquisition method and device, equipment and storage medium
By determining the scheduling priority based on web page feature information and allocating distributed acquisition nodes, the problem of low data acquisition efficiency in the prior art is solved, and automated and efficient data acquisition is achieved.
Patent Information
- Application Number
- CN202311714245.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-13
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, the data acquisition process requires a lot of manual intervention, resulting in low efficiency.
By determining the scheduling priority of the data collection task based on the web page characteristic information, and assigning distributed acquisition nodes to collect web page data.
It improves the efficiency of data acquisition, reduces manual intervention, and realizes an automated data acquisition process.
Smart Images

Figure CN120144241A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet technologies, and in particular, to a data collection method, apparatus, device, and storage medium. Background Art
[0002] With the development of network information technologies, there is an increasing amount of web information such as websites, forums, blogs, etc. Technologies such as search engines, content analysis, and public opinion analysis all analyze and process this information. Before analyzing and processing this information, it is often necessary to collect data information in web pages.
[0003] In existing solutions, data can be collected from the Internet through web crawler technology. However, collecting data through web crawler technology generally requires a large amount of manual intervention, resulting in low data collection efficiency.
[0004] The above content is only used to assist in understanding the technical solution of the present invention, and does not represent an admission that the above content is prior art. Summary of the Invention
[0005] The main objective of the present invention is to provide a data collection method, apparatus, device, and storage medium, aiming to solve the technical problem that manual intervention is required in the data collection process in the prior art, resulting in low data collection efficiency.
[0006] To achieve the above objective, the present invention provides a data collection method, and the method includes the following steps:
[0007] Determine the web pages to be collected according to the data collection tasks, and obtain the web page feature information of each web page to be collected;
[0008] Determine the scheduling priorities of each data collection task according to the web page feature information;
[0009] Allocate distributed collection nodes to each data collection task according to the scheduling priorities, so that the distributed collection nodes collect web page data according to the received data collection tasks.
[0010] Optionally, the step of determining the scheduling priorities of each data collection task according to the web page feature information includes:
[0011] Determine the web page weight corresponding to the web page to be collected according to the web page feature information;
[0012] Obtain the web page quality and web page update time of the web page to be collected based on a preset prediction algorithm;
[0013] Determine the scheduling priorities of each data collection task based on the web page weight, the web page quality, and the web page update time.
[0014] Optionally, the step of determining the web page weight corresponding to the web page to be collected according to the web page feature information includes:
[0015] Classify the web page to be collected according to the web page feature information and a preset classification algorithm;
[0016] Determine the web page score of the web page to be collected based on the classification result, a preset scoring algorithm, and the web page feature information;
[0017] Determine the web page weight corresponding to the web page to be collected according to the web page score.
[0018] Optionally, after the step of allocating distributed collection nodes to each data collection task according to the scheduling priority, so that the distributed collection nodes perform web page data collection according to the received data collection tasks, the method further includes:
[0019] When the web page data collection is completed, obtain the element density of the target element in the web page to be collected;
[0020] Determine the element type corresponding to the target element according to the element density;
[0021] Determine the web page structure corresponding to the web page to be collected based on the element type.
[0022] Optionally, after the step of allocating distributed collection nodes to each data collection task according to the scheduling priority, so that the distributed collection nodes perform web page data collection according to the received data collection tasks, the method further includes:
[0023] When the web page data collection is completed, determine the target display dimension corresponding to the web page to be collected, where the target display dimension includes at least one of a time dimension, a space dimension, and an attribute dimension;
[0024] Visually display the web page data of the web page to be collected through the data display method corresponding to the target display dimension.
[0025] Optionally, before the step of determining the web page to be collected according to the data collection task and obtaining the web page feature information of each web page to be collected, the method further includes:
[0026] Configure a preset data collection mechanism, where the preset data collection mechanism includes: a stable collection mechanism, a reliable collection mechanism, a flexible collection mechanism, and a secure collection mechanism;
[0027] The step of allocating distributed collection nodes to each data collection task according to the scheduling priority, so that the distributed collection nodes perform web page data collection according to the received data collection tasks, includes:
[0028] Distribute distributed acquisition nodes to each data acquisition task according to the scheduling priority, so that the distributed acquisition nodes perform web data acquisition based on the preset data acquisition mechanism and the received data acquisition tasks.
[0029] Optionally, before the step of determining the web pages to be acquired according to the data acquisition tasks and obtaining the web page feature information of each web page to be acquired, the following steps are further included:
[0030] Obtain the target web page data, target web page quality, and target update time corresponding to the target web page through web crawler technology;
[0031] Construct a preset classification algorithm, a preset scoring algorithm, and a preset prediction algorithm based on the target web page data, the target web page quality, and the target update time.
[0032] In addition, to achieve the above object, the present invention further provides a data acquisition device, the device includes: a memory, a processor, and a data acquisition program stored on the memory and executable on the processor, the data acquisition program is configured to implement the steps of the data acquisition method as described above.
[0033] In addition, to achieve the above object, the present invention further provides a storage medium, on which a data acquisition program is stored, and when the data acquisition program is executed by a processor, it implements the steps of the data acquisition method as described above.
[0034] In addition, to achieve the above object, the present invention further provides a data acquisition device, the device includes: an information acquisition module, a priority determination module, and a data acquisition module;
[0035] The information acquisition module is used to determine the web pages to be acquired according to the data acquisition tasks and obtain the web page feature information of each web page to be acquired;
[0036] The priority determination module is used to determine the scheduling priority of each data acquisition task according to the web page feature information;
[0037] The data acquisition module is used to distribute distributed acquisition nodes to each data acquisition task according to the scheduling priority, so that the distributed acquisition nodes perform web data acquisition according to the received data acquisition tasks.
[0038] In the present invention, a web page to be collected is determined according to a data collection task, and web page feature information of each web page to be collected is obtained; a scheduling priority of each data collection task is determined according to the web page feature information; distributed collection nodes are allocated to each data collection task according to the scheduling priority, so that the distributed collection nodes perform web page data collection according to the received data collection tasks; compared with the prior art in which data collection by means of a web crawler technology requires manual intervention and has low efficiency, since the present invention determines the scheduling priority of each data collection task according to the web page feature information of the web page to be collected, and allocates distributed collection nodes to each data collection task according to the scheduling priority, so that the distributed collection nodes perform web page data collection, thereby solving the technical problem in the prior art that manual intervention is required in the data collection process, resulting in low data collection efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 FIG. is a schematic structural diagram of a data collection device for a hardware operating environment related to the solution of an embodiment of the present invention;
[0040] Figure 2 FIG. is a schematic flowchart of a first embodiment of a data collection method of the present invention;
[0041] Figure 3 FIG. is a schematic flowchart of a second embodiment of a data collection method of the present invention;
[0042] Figure 4 FIG. is a schematic flowchart of a third embodiment of a data collection method of the present invention;
[0043] Figure 5 FIG. is a structural block diagram of a first embodiment of a data collection device of the present invention.
[0044] The implementation, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0046] Refer to Figure 1 , Figure 1 FIG. is a schematic structural diagram of a data collection device for a hardware operating environment related to the solution of an embodiment of the present invention.
[0047] As Figure 1As shown in the figure, the data acquisition device may include: a processor 1001, such as a Central Processing Unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen and an input unit such as a keyboard. Optionally, the user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a Wireless-Fidelity (Wi-Fi) interface). The memory 1005 may be a high-speed Random Access Memory (RAM) or a stable Non-Volatile Memory (NVM), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0048] Those skilled in the art can understand that Figure 1 the structure shown in does not constitute a limitation on the data acquisition device, and it may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.
[0049] As Figure 1 shown, the memory 1005, as a storage medium, may include an operating system, a network communication module, a user interface module, and a data acquisition program.
[0050] In Figure 1 the data acquisition device shown, the network interface 1004 is mainly used for data communication with a network server; the user interface 1003 is mainly used for data interaction with a user; the processor 1001 and the memory 1005 in the data acquisition device of the present invention may be arranged in the data acquisition device. The data acquisition device calls the data acquisition program stored in the memory 1005 through the processor 1001 and executes the data acquisition method provided by the embodiments of the present invention.
[0051] The embodiments of the present invention provide a data acquisition method. Referring to Figure 2 , Figure 2 it is a schematic flowchart of the first embodiment of the data acquisition method of the present invention.
[0052] In this embodiment, the data acquisition method includes the following steps:
[0053] Step S10: Determine the web pages to be acquired according to the data acquisition task, and obtain the web page feature information of each web page to be acquired.
[0054] It should be noted that the execution entity of the method in this embodiment can be the master node that can schedule the distributed acquisition nodes for acquiring data in the intranet environment in the data acquisition architecture, or other data acquisition devices that can implement the same or similar functions and include the master node.
[0055] It should be understood that the distributed acquisition nodes can be nodes for acquiring data in the intranet environment. The data acquisition architecture in the method of this embodiment can adopt the master node and the distributed acquisition nodes to implement a distributed acquisition system (hereinafter referred to as the system). Among them, the master node can adopt a dual-active highly available architecture, that is, two master nodes are deployed simultaneously, one primary and one standby, to achieve seamless switching and ensure the high availability of the system. When one of the master nodes fails, the system will automatically switch to the standby node, thereby ensuring the continuity and stability of the system. The distributed acquisition nodes can be elastically scaled as needed. According to the needs of the acquisition task, the acquisition nodes can be dynamically increased or decreased to meet the data acquisition requirements of different scales, thereby ensuring the flexibility and scalability of the system.
[0056] It can be understood that the master node and the distributed acquisition nodes can communicate using the distributed message queue technology to achieve efficient data transmission and task scheduling. At the same time, the master node can also monitor and manage the distributed acquisition nodes to ensure the stable operation of the system. In practical applications, the master node can divide the data acquisition task into different distributed acquisition nodes for data acquisition according to the acquisition performance of a specific website. Among them, the load balancing technology can be adopted between the distributed acquisition nodes, and the master node can distribute the data acquisition requests to different acquisition nodes, thereby improving the concurrent processing ability and response speed of the system. In addition, the distributed acquisition nodes can be deployed in different containers to achieve rapid deployment and management, thereby improving the portability and scalability of the system.
[0057] It should be understood that the above data acquisition task can be a task for acquiring data of each web page.
[0058] It can be understood that the web page to be acquired can be a page that needs to be acquired for data. In practical applications, each web page to be acquired can correspond to a data acquisition task, so that the master node can determine the corresponding web page to be acquired according to the data acquisition task.
[0059] It should be noted that the above web page feature information can be information used to characterize the web page features. For example: the URL (Uniform Resource Locator) of the web page, title, keywords, etc. This embodiment does not limit this. This embodiment can extract the features of the web page to be collected through a pre-trained feature extraction algorithm, so as to obtain the web page feature information of the web page to be collected. Among them, the feature extraction algorithm in this embodiment can be algorithms such as TF-IDF, Word2Vec, Doc2Vec, etc.
[0060] In specific implementation, when it is necessary to collect web page data, a corresponding data collection task can be generated. At this time, the master node in the data collection architecture can determine the corresponding web page to be collected according to the data collection task, and after determining the web page to be collected, extract the web page feature information such as the URL, title, and keywords of the web page to be collected from the web page to be collected through the feature extraction algorithm.
[0061] Step S20: Determine the scheduling priority of each data collection task according to the web page feature information.
[0062] It should be understood that the above scheduling priority can be the priority for the master node to schedule the web page to be collected to the distributed collection node for data collection.
[0063] In specific implementation, this embodiment can identify the web page weights of each web page to be collected through an intelligent data collection scheduling algorithm, predict the web page quality and update frequency, and distinguish the scheduling priorities of each web page to be collected according to these features, and then schedule each web page to be collected according to the priority. Specifically, the feature extraction algorithm in the data collection scheduling algorithm can be used to extract the web page feature information such as the URL, title, and keywords of each web page to be collected, and determine the web page weights of each web page to be collected based on these web page feature information, so as to determine the scheduling priorities of each web page to be collected, and schedule the corresponding web page to be collected according to the scheduling priority.
[0064] Step S30: Allocate distributed collection nodes for each data collection task according to the scheduling priority, so that the distributed collection nodes collect web page data according to the received data collection tasks.
[0065] In specific implementation, the master node can allocate distributed collection nodes for each data collection task according to the scheduling priority of each data collection task for data collection, that is, according to the order of the scheduling priorities, first allocate distributed collection nodes for the data collection tasks with high priorities. When the distributed collection node receives the data collection task, it can collect the web page data of the web page to be collected corresponding to the data collection task. At the same time, scheduling according to the scheduling priority can also ensure that all web pages to be collected can be scheduled and downloaded within the specified period.
[0066] Further, in order to identify the web page structure of the collected web page, after the step S30, the method further includes: when the web page data collection is completed, obtaining the element density of the target element in the to-be-collected web page; determining the element type corresponding to the target element according to the element density; and determining the web page structure corresponding to the to-be-collected web page based on the element type.
[0067] It should be noted that the above-mentioned target element may be an element in the to-be-collected web page that has been collected, such as: text, table, etc. in the web page. Correspondingly, the element type corresponding to the above-mentioned target element may be body text, advertisement, etc., and this embodiment does not limit this.
[0068] It should be understood that the above-mentioned element density may be the degree of tightness of the element arrangement in the web page.
[0069] In specific implementation, when the web page data collection is completed, the structure of the collected web page can be automatically identified and extracted by a data processing and analysis module. Specifically, an algorithm based on web page density can be used to obtain the web page structure of the collected to-be-collected web page. This embodiment can judge the type of an element according to the density of different elements in the collected web page. For example: if the density of a certain paragraph of text in a web page is large, it can be judged that this paragraph of text is body text; if the density of a certain paragraph of text in a web page is small, it can be judged that this paragraph of text is an advertisement, so that the structure of the web page can be automatically identified and extracted in this way. In addition, in specific implementation, web page parsing technologies based on HTML (Hyper Text Markup Language) and CSS (Cascading Style Sheets) can be used to perform structured extraction on the collected web page. Among them, HTML and CSS are the basic languages of the web page, and the structure and style of the web page are controlled by HTML and CSS. This embodiment can identify elements such as titles, body texts, and lists by parsing the tags in HTML, and identify attributes such as fonts and colors by parsing the styles in CSS, so as to realize the automatic identification and extraction of the web page structure.
[0070] Further, in order to visualize the analysis result of the data in the web page, after the step S30, the method further includes: when the web page data collection is completed, determining the target display dimension corresponding to the to-be-collected web page, where the target display dimension includes at least one of a time dimension, a space dimension, and an attribute dimension; and visually displaying the web page data of the to-be-collected web page by using the data display method corresponding to the target display dimension.
[0071] It should be noted that the above target display dimension can be a dimension for displaying the collected web page data or the analysis result of the web page data.
[0072] It should be understood that the above data display method can be a method of displaying data according to each dimension. Among them, the time dimension is to visualize the data according to the time dimension (for example, drawing a time series chart, a heat map, etc.); the space dimension is to visualize the data according to the space dimension (for example, drawing a map, a scatter plot, etc.); the attribute dimension is to visualize the data according to the attribute dimension (for example, drawing a bar chart, a pie chart, etc.).
[0073] Furthermore, the step of determining the target display dimension corresponding to the web page to be collected when completing the web page data collection can specifically include: when completing the web page data collection, obtaining the user requirements corresponding to the target user; if the user requirement is to display the data change trend, then determining the target display dimension corresponding to the web page to be collected as the time dimension; correspondingly, visualizing and displaying the web page data of the web page to be collected through the time series chart representing the time dimension.
[0074] It should be noted that the above user requirements can be to display the data change trend, the regional distribution trend of the data, and the attribute distribution of the data, etc. Before visualizing and displaying the data in this embodiment, the user requirements can be obtained first, and the data display dimension corresponding to the web page can be determined according to the user requirements, and the data can be visualized and displayed through the corresponding display form in the data display dimension.
[0075] Furthermore, if the user requirement is to display the regional distribution of the data, then determine the target display dimension corresponding to the web page to be collected as the space dimension, and visualize and display the web page data of the web page to be collected through the data map representing the space dimension.
[0076] Furthermore, if the user requirement is to display the attribute distribution of the data, then determine the target display dimension corresponding to the web page to be collected as the attribute dimension, and visualize and display the web page data of the web page to be collected through the data bar chart representing the attribute dimension.
[0077] In specific implementation, when completing web page data collection, the user requirements of each user can be obtained. If the user requirement is to display the data change trend, the corresponding data display dimension of the web page can be determined as the time dimension, and the web page data or the analysis result of the web page data can be visually displayed through a time series graph representing the time dimension of the data, so that the user can understand the data change trend and pattern. If the user requirement is to display the regional distribution request of the data, the corresponding data display dimension of the web page can be determined as the space dimension, and the web page data or the analysis result of the web page data can be visually displayed through a data map representing the space dimension of the data, so that the user can understand the data distribution in different regions. If the user requirement is to display the attribute distribution of the data, the corresponding data display dimension of the web page can be determined as the attribute dimension, and the web page data or the analysis result of the web page data can be visually displayed through a data bar graph representing the attribute dimension of the data, so that the user can understand the data distribution on different data. Through the visual analysis platform, the user can intuitively understand the data distribution, trend and pattern, so as to conduct data analysis and decision-making. For example, the user can understand the change trend of a certain indicator through the time series graph, understand the data distribution in a certain region through the map, and understand the distribution of a certain attribute through the bar graph, etc.
[0078] It should be noted that this embodiment can also optimize the distributed collection system through visual configuration, so that the system can work more efficiently under the mechanism of manual assistance. In practical applications, the above optimization mechanism can provide a variety of configuration options, including collection frequency, number of threads, proxy settings, etc., and users can configure according to actual needs to achieve the best collection effect. At the same time, the optimization mechanism can also provide real-time monitoring and logging functions to facilitate users to troubleshoot and optimize performance. The specific methods of system optimization can include: by adjusting parameters such as collection frequency and number of threads, the collection efficiency can be improved, the collection time can be reduced, and the real-time nature of the data can be improved; by means of proxy settings, anti-spider strategies, etc., being blocked by the website can be avoided and the collection quality can be improved; by adjusting parameters such as number of threads and memory allocation, the system load can be reduced and the stability and reliability of the system can be improved; by adjusting parameters such as collection frequency and data format, the user experience can be improved and it is easier for users to obtain the required data.
[0079] It should be understood that before optimizing the system, it is also necessary to analyze the content to be optimized. The specific methods may include: by monitoring system metrics such as CPU, memory, and disk, the system load can be understood, and thus it can be determined whether the system configuration needs to be adjusted; by analyzing metrics such as collection time and data volume, the collection efficiency of the system can be understood, and thus it can be determined whether parameters such as collection frequency and number of threads need to be adjusted; by analyzing metrics such as the accuracy and integrity of the collected data, the collection quality of the system can be understood, and thus it can be determined whether parameters such as proxy settings and anti-spider strategies need to be adjusted; by collecting user feedback, the needs and feedback of users can be understood, and thus it can be determined whether parameters such as data format and collection frequency need to be adjusted. Through the above methods, system operators can understand the bottlenecks and problems of the system, and thus perform targeted optimization to improve the performance and stability of the system.
[0080] This embodiment discloses determining web pages to be collected according to data collection tasks, and obtaining web page feature information of each web page to be collected; determining the scheduling priorities of each data collection task according to the web page feature information; allocating distributed collection nodes to each data collection task according to the scheduling priorities, so that the distributed collection nodes perform web page data collection according to the received data collection tasks; compared with the prior art in which data collection is performed through web crawler technology and requires manual intervention with low efficiency, since this embodiment determines the scheduling priorities of each data collection task according to the web page feature information of the web pages to be collected, and allocates distributed collection nodes to each data collection task according to the scheduling priorities, so that the distributed collection nodes perform web page data collection, thereby solving the technical problem of low data collection efficiency caused by manual intervention in the data collection process of the prior art. At the same time, when the web page data collection is completed, the element type corresponding to the target element is determined according to the element density of the target element in the web page to be collected, and the web page structure corresponding to the web page to be collected is determined based on the element type, so that the automatic recognition and extraction of the web page structure can be realized. In addition, when the web page data collection is completed, the web page data is visually displayed through the data display method corresponding to the target display dimension of the web page to be collected, so that users can intuitively understand the distribution, trend, and law of the data, improving the user experience.
[0081] Reference Figure 3 , Figure 3 is a flowchart of the second embodiment of the data collection method of the present invention.
[0082] Based on the above first embodiment, in order to achieve accurate scheduling of data collection tasks, in this embodiment, the step S20 includes:
[0083] Step S201: Determine the web page weight corresponding to the web page to be collected according to the web page feature information.
[0084] It should be noted that the above web page weight can represent the importance of the web page to be crawled. In practical applications, the web page feature information such as the URL, title, and keywords in the web page to be crawled can be extracted through a feature extraction algorithm first, so as to obtain the web page weight corresponding to the web page to be crawled based on the web page feature information.
[0085] Furthermore, in order to accurately obtain the web page weights of each web page to be crawled, step S201 may include: classifying the web page to be crawled according to the web page feature information and a preset classification algorithm; determining the web page score of the web page to be crawled based on the classification result, a preset scoring algorithm, and the web page feature information; and determining the web page weight corresponding to the web page to be crawled according to the web page score.
[0086] It should be understood that the preset classification algorithm can be an algorithm for classifying the web page to be crawled, such as: Naive Bayes, Support Vector Machine, Decision Tree, etc. This embodiment does not limit this.
[0087] It can be understood that the preset scoring algorithm can be an algorithm for scoring the web page to be crawled, such as: PageRank, HITS, TF-IDF, etc. This embodiment does not limit this. In practical applications, the web page to be crawled can be scored through the web page feature information of each web page to be crawled. For example, the web page can be scored according to the update time of the web page to be crawled, whether there is a title in the web page, the length of the web page content, whether there are pictures and texts in the web page, and other web page feature information.
[0088] Step S202: Obtain the web page quality and web page update time of the web page to be crawled based on a preset prediction algorithm.
[0089] It should be noted that the above preset prediction algorithm can be an algorithm for predicting the update time and web page quality of the web page to be crawled, such as: time series analysis, regression analysis, etc. This embodiment does not limit this.
[0090] Step S203: Determine the scheduling priorities of each data collection task based on the web page weight, the web page quality, and the web page update time.
[0091] It should be understood that in this embodiment, the scheduling priorities of each data collection task can be distinguished according to the web page weight, web page quality, and web page update time or update frequency and other characteristics of each web page to be crawled, so as to schedule the web pages to be crawled corresponding to each data collection task according to the scheduling priorities.
[0092] In a specific implementation, first, web page feature information such as the URL, title, and keywords of the web page to be collected can be extracted through a feature extraction algorithm to facilitate subsequent classification and scoring of the web page. After feature extraction of the web page to be collected, a preset classification algorithm can be used to classify each web page to be collected to obtain the category to which each web page to be collected belongs, so as to score and schedule the web pages to be collected in each category subsequently. After completing the classification of the web pages to be collected, according to the classification results, a preset scoring algorithm can be used to score the web pages to be collected in each category to obtain the web page score corresponding to each web page to be collected, and the web page weight of each web page to be collected can be obtained according to the web page score, so as to better schedule the web pages to be collected. After obtaining the web page weights of each web page to be collected, a preset prediction algorithm can be used to predict the web page update time and web page quality of each web page to be collected, so that the scheduling priorities of each data collection task can be distinguished according to the web page weight, web page update time, and web page quality of each web page to be collected, and distributed collection nodes can be assigned to the data collection tasks according to the scheduling priorities, so that the distributed collection nodes can collect web page data of the corresponding web page to be collected according to the received data collection tasks.
[0093] Further, in order to pre-train the classification algorithm, scoring algorithm, and prediction algorithm to determine the scheduling priorities of each data collection task and improve data processing efficiency, before the step S10, the method further includes: obtaining target web page data, target web page quality, and target update time corresponding to the target web page through web crawler technology; constructing a preset classification algorithm, a preset scoring algorithm, and a preset prediction algorithm based on the target web page data, the target web page quality, and the target update time.
[0094] It should be noted that the above web crawler technology can be a technology for obtaining web page data from the Internet. Correspondingly, the above target web page data, target web page quality, and target update time can be the web page data, web page quality, and web page update time obtained from the Internet. In practical applications, in order to improve data processing efficiency, the above feature extraction algorithm, preset classification algorithm, preset scoring algorithm, and preset prediction algorithm can be pre-trained. Since training a model requires a large amount of data, in this embodiment, a large amount of web page data can be first crawled from the Internet by using the web crawler technology, including data information such as the URL, title, keywords, historical data, and update time of the web page, and a preset classification algorithm, preset scoring algorithm, and preset prediction algorithm can be constructed based on the target web page data, target web page quality, and target update time. At the same time, a large amount of web page data can also be obtained from publicly available datasets (such as Wikipedia, IMDB, etc.) for model training. In addition, the data feedback by users (such as the click-through rate and page view volume of the web page) can be used to evaluate the quality and update frequency of the web page. The data acquisition method for training the model in this embodiment is not limited. After obtaining the web page data, the data can be cleaned and preprocessed so that subsequent modeling and training can be performed based on the processed data to further improve data processing efficiency. Among them, preprocessing the data can include: removing duplicate data, handling missing values, and feature extraction, etc.
[0095] It should be understood that after establishing the above algorithms, algorithm optimization methods can be used to optimize the algorithms to improve the accuracy of the obtained scheduling priority. The algorithm optimization methods in this embodiment can include: feature selection, model selection, parameter adjustment, and ensemble learning, etc. Among them, feature selection means that for a large number of extracted features, a feature selection method can be used to select the features that have the greatest impact on the model prediction effect; model selection means that for different algorithm models, the most suitable model for the current dataset can be selected to improve the prediction effect; parameter adjustment means that for the parameters in the algorithm model, the optimal parameter combination can be selected through methods such as grid search to improve the prediction effect; ensemble learning means that multiple models can be integrated to improve the prediction effect. Common ensemble learning methods include Bagging, Boosting, etc. Through these algorithm optimization methods, the above algorithms can be optimized to improve the prediction effect of the algorithms.
[0096] In this embodiment, the web pages to be collected are classified according to web page feature information and a preset classification algorithm. Based on the classification result, a preset scoring algorithm, and the web page feature information, the web page score of the web pages to be collected is determined. Then, according to the web page score, the web page weight of the web pages to be collected is determined. Next, based on a preset prediction algorithm, the web page quality and web page update time of the web pages to be collected are obtained. Based on the web page weight, web page quality, and web page update time, the scheduling priorities of each data collection task are determined, thereby improving the collection efficiency of data collection for the web pages to be collected.
[0097] Reference Figure 4 , Figure 4 is a schematic flowchart of the third embodiment of the data collection method of the present invention.
[0098] Based on the above embodiments, in order to achieve efficient, stable, and reliable data collection, in this embodiment, before step S10, the method further includes:
[0099] Step S01: Configure a preset data collection mechanism, where the preset data collection mechanism includes: a stable collection mechanism, a reliable collection mechanism, a flexible collection mechanism, and a secure collection mechanism.
[0100] It should be noted that the above preset data collection mechanism can be a mechanism for data collection.
[0101] It should be understood that the stable collection mechanism can be a mechanism for ensuring the stability of data collection, such as: error handling, retry mechanism, resume from breakpoint, etc. During the data collection process, if an abnormal situation occurs, this embodiment can perform corresponding processing according to the preset error handling strategy, thereby ensuring the stability and reliability of data collection. At the same time, this embodiment also adopts a resume from breakpoint mechanism, which can automatically resume the collection progress after the collection is interrupted or fails, thereby avoiding duplicate data collection and wasting resources.
[0102] It can be understood that the reliable collection mechanism can be a mechanism for ensuring the reliability of data collection, such as: data verification, data deduplication, data cleaning, etc. During the data collection process, this embodiment can verify and deduplicate the collected data to ensure the accuracy and integrity of the collected data. At the same time, this embodiment can also clean and transform the collected data to adapt to different data formats and requirements.
[0103] It should be noted that the flexible collection mechanism can be a mechanism for improving the flexibility of data collection. This embodiment can adopt configurable parameters and strategies, and can be flexibly configured and adjusted according to different collection requirements and environments. Users can customize parameters such as the priority, collection frequency, and data format of the collection task according to actual needs to meet different collection requirements and scenarios.
[0104] It should be understood that the security collection mechanism can be a mechanism for improving the security of data collection, such as: identity authentication, data encryption, access control, etc. During the data collection process, this embodiment can perform identity authentication and encryption on the collected data to protect the security and confidentiality of the data. At the same time, this embodiment can also perform access control on the collected data to protect the privacy and security of the data.
[0105] Correspondingly, the step S30 includes:
[0106] Step S30': Allocate distributed collection nodes for each data collection task according to the scheduling priority, so that the distributed collection nodes perform web page data collection based on the preset data collection mechanism and the received data collection tasks.
[0107] In a specific implementation, this embodiment can first configure a preset data collection mechanism. When collecting web page data, the preset data collection mechanism can be used to collect web page data to ensure the stability and reliability of data collection, and improve the flexibility and security of data collection, thereby providing a better data collection experience and service for users.
[0108] This embodiment configures a preset data collection mechanism and allocates distributed collection nodes for each data collection task according to the scheduling priority, so that the distributed collection nodes perform web page data collection based on the preset data collection mechanism and the received data collection tasks, thereby ensuring the stability and reliability of data collection, and improving the flexibility and security of data collection, providing a better data collection experience and service for users.
[0109] In addition, an embodiment of the present invention also proposes a storage medium, on which a data collection program is stored. When the data collection program is executed by a processor, the steps of the data collection method described above are implemented.
[0110] Referring to Figure 5 , Figure 5 is the structural block diagram of the first embodiment of the data collection device of the present invention.
[0111] As Figure 5 shown, the data collection device proposed by the embodiment of the present invention includes: an information acquisition module 501, a priority determination module 502, and a data collection module 503;
[0112] The information acquisition module 501 is used to determine the web pages to be collected according to the data collection tasks and obtain the web page feature information of each web page to be collected;
[0113] The priority determination module 502 is used to determine the scheduling priority of each data collection task according to the web page feature information;
[0114] The data acquisition module 503 is configured to allocate distributed acquisition nodes for the respective data acquisition tasks according to the scheduling priority, so that the distributed acquisition nodes perform web page data acquisition according to the received data acquisition tasks.
[0115] Further, the data acquisition module 503 is further configured to, when the web page data acquisition is completed, obtain the element density of the target element in the to-be-acquired web page; determine the element type corresponding to the target element according to the element density; and determine the web page structure corresponding to the to-be-acquired web page based on the element type.
[0116] Further, the data acquisition module 503 is further configured to, when the web page data acquisition is completed, determine the target display dimension corresponding to the to-be-acquired web page, where the target display dimension includes at least one of a time dimension, a space dimension, and an attribute dimension; and perform visual display on the web page data of the to-be-acquired web page through the data display manner corresponding to the target display dimension.
[0117] The data acquisition device of this embodiment discloses determining a to-be-acquired web page according to a data acquisition task and obtaining the web page feature information of each to-be-acquired web page; determining the scheduling priority of each data acquisition task according to the web page feature information; allocating distributed acquisition nodes for the respective data acquisition tasks according to the scheduling priority, so that the distributed acquisition nodes perform web page data acquisition according to the received data acquisition tasks; compared with the prior art in which data acquisition is performed through a web crawler technology and requires manual intervention with low efficiency, since in this embodiment, the scheduling priority of each data acquisition task is determined according to the web page feature information of the to-be-acquired web page, and distributed acquisition nodes are allocated for the respective data acquisition tasks according to the scheduling priority, so that the distributed acquisition nodes perform web page data acquisition, thereby solving the technical problem in the prior art that manual intervention is required in the data acquisition process, resulting in low data acquisition efficiency. At the same time, when the web page data acquisition is completed, the element type corresponding to the target element is determined according to the element density of the target element in the to-be-acquired web page, and the web page structure corresponding to the to-be-acquired web page is determined based on the element type, so that automatic recognition and extraction of the web page structure can be realized. In addition, when the web page data acquisition is completed, visual display is performed on the web page data through the data display manner corresponding to the target display dimension of the to-be-acquired web page, so that the user can intuitively understand the data distribution, trend, and law, improving the user experience.
[0118] Based on the first embodiment of the data acquisition device of the present invention, a second embodiment of the data acquisition device of the present invention is proposed.
[0119] In this embodiment, the priority determination module 502 is further configured to determine the web page weight corresponding to the web page to be crawled according to the web page feature information; obtain the web page quality and web page update time of the web page to be crawled based on a preset prediction algorithm; and determine the scheduling priority of each data collection task based on the web page weight, the web page quality, and the web page update time.
[0120] Further, the priority determination module 502 is further configured to classify the web page to be crawled according to the web page feature information and a preset classification algorithm; determine the web page score of the web page to be crawled based on the classification result, a preset scoring algorithm, and the web page feature information; and determine the web page weight corresponding to the web page to be crawled according to the web page score.
[0121] Further, the information acquisition module 501 is further configured to obtain the target web page data, target web page quality, and target update time corresponding to the target web page through web crawler technology; and construct a preset classification algorithm, a preset scoring algorithm, and a preset prediction algorithm based on the target web page data, the target web page quality, and the target update time.
[0122] In this embodiment, the web page to be crawled is classified according to the web page feature information and a preset classification algorithm, the web page score of the web page to be crawled is determined based on the classification result, a preset scoring algorithm, and the web page feature information, the web page weight corresponding to the web page to be crawled is determined according to the web page score, the web page quality and web page update time of the web page to be crawled are obtained based on a preset prediction algorithm, and the scheduling priority of each data collection task is determined based on the web page weight, the web page quality, and the web page update time, thereby improving the collection efficiency of data collection for the web page to be crawled.
[0123] Based on the above device embodiments, a third embodiment of the data collection device of the present invention is proposed.
[0124] In this embodiment, the information acquisition module 501 is further configured to configure a preset data collection mechanism, and the preset data collection mechanism includes: a stable collection mechanism, a reliable collection mechanism, a flexible collection mechanism, and a secure collection mechanism;
[0125] The data collection module 503 is further configured to allocate distributed collection nodes for each data collection task according to the scheduling priority, so that the distributed collection nodes perform web page data collection based on the preset data collection mechanism and the received data collection tasks.
[0126] In this embodiment, by configuring a preset data collection mechanism and allocating distributed collection nodes to each data collection task according to the scheduling priority, the distributed collection nodes perform web data collection based on the preset data collection mechanism and the received data collection tasks, thereby ensuring the stability and reliability of data collection, improving the flexibility and security of data collection, and providing users with a better data collection experience and service.
[0127] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or system including that element.
[0128] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.
[0129] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as a read-only memory / random access memory, magnetic disk, optical disk), and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0130] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
[0131] The present invention discloses A1, a data collection method, and the data collection method includes:
[0132] Determine the web pages to be collected according to the data collection tasks, and obtain the web page feature information of each web page to be collected;
[0133] Determine the scheduling priority of each data collection task according to the web page feature information;
[0134] Allocate distributed acquisition nodes for each data acquisition task according to the scheduling priority, so that the distributed acquisition nodes perform web page data acquisition according to the received data acquisition tasks.
[0135] A2. The data acquisition method as described in A1, the step of determining the scheduling priority of each data acquisition task according to the web page feature information includes:
[0136] Determine the web page weight corresponding to the web page to be acquired according to the web page feature information;
[0137] Obtain the web page quality and web page update time of the web page to be acquired based on a preset prediction algorithm;
[0138] Determine the scheduling priority of each data acquisition task based on the web page weight, the web page quality, and the web page update time.
[0139] A3. The data acquisition method as described in A2, the step of determining the web page weight corresponding to the web page to be acquired according to the web page feature information includes:
[0140] Classify the web page to be acquired according to the web page feature information and a preset classification algorithm;
[0141] Determine the web page score of the web page to be acquired based on the classification result, a preset scoring algorithm, and the web page feature information;
[0142] Determine the web page weight corresponding to the web page to be acquired according to the web page score.
[0143] A4. The data acquisition method as described in A1, after the step of allocating distributed acquisition nodes for each data acquisition task according to the scheduling priority, so that the distributed acquisition nodes perform web page data acquisition according to the received data acquisition tasks, further includes:
[0144] When the web page data acquisition is completed, obtain the element density of the target element in the web page to be acquired;
[0145] Determine the element type corresponding to the target element according to the element density;
[0146] Determine the web page structure corresponding to the web page to be acquired based on the element type.
[0147] A5. The data acquisition method as described in A4, after the step of allocating distributed acquisition nodes for each data acquisition task according to the scheduling priority, so that the distributed acquisition nodes perform web page data acquisition according to the received data acquisition tasks, further includes:
[0148] When completing the web page data collection, determine the target display dimension corresponding to the web page to be collected, where the target display dimension includes at least one of a time dimension, a space dimension, and an attribute dimension;
[0149] Visualize the web page data of the web page to be collected through the data display method corresponding to the target display dimension.
[0150] A6. For the data collection method as described in A1 to A5, before the step of determining the web page to be collected according to the data collection task and obtaining the web page feature information of each web page to be collected, it further includes:
[0151] Configure a preset data collection mechanism, where the preset data collection mechanism includes: a stable collection mechanism, a reliable collection mechanism, a flexible collection mechanism, and a secure collection mechanism;
[0152] The step of allocating distributed collection nodes to each data collection task according to the scheduling priority so that the distributed collection nodes perform web page data collection according to the received data collection task includes:
[0153] Allocate distributed collection nodes to each data collection task according to the scheduling priority so that the distributed collection nodes perform web page data collection based on the preset data collection mechanism and the received data collection task.
[0154] A7. For the data collection method as described in A3, before the step of determining the web page to be collected according to the data collection task and obtaining the web page feature information of each web page to be collected, it further includes:
[0155] Obtain the target web page data, target web page quality, and target update time corresponding to the target web page through web crawler technology;
[0156] Build a preset classification algorithm, a preset scoring algorithm, and a preset prediction algorithm based on the target web page data, the target web page quality, and the target update time.
[0157] A8. For the data collection method as described in A5, the step of determining the target display dimension corresponding to the web page to be collected when completing the web page data collection includes:
[0158] When completing the web page data collection, obtain the user requirements corresponding to the target user;
[0159] If the user requirement is to display the data change trend, then determine the target display dimension corresponding to the web page to be collected as the time dimension;
[0160] The step of visualizing the web page data of the web page to be collected through the data display method corresponding to the target display dimension includes:
[0161] Visualize the web page data of the web page to be collected through a time series graph representing the time dimension.
[0162] A9. The data collection method as described in A8. After the step of obtaining the user requirements corresponding to the target user when the web page data collection is completed, the method further includes:
[0163] If the user requirement is to display the regional distribution of data, then determine the target display dimension corresponding to the web page to be collected as the spatial dimension;
[0164] The step of visualizing the web page data of the web page to be collected through the data display method corresponding to the target display dimension includes:
[0165] Visualize the web page data of the web page to be collected through a data map representing the spatial dimension.
[0166] A10. The data collection method as described in A8. After the step of obtaining the user requirements corresponding to the target user when the web page data collection is completed, the method further includes:
[0167] If the user requirement is to display the attribute distribution of data, then determine the target display dimension corresponding to the web page to be collected as the attribute dimension;
[0168] The step of visualizing the web page data of the web page to be collected through the data display method corresponding to the target display dimension includes:
[0169] Visualize the web page data of the web page to be collected through a data bar graph representing the attribute dimension.
[0170] The present invention also discloses B11. A data collection device, the device includes: a memory, a processor, and a data collection program stored on the memory and executable on the processor, and the data collection program is configured to implement the steps of the data collection method as described above.
[0171] The present invention also discloses C12. A storage medium, on which a data collection program is stored, and when the data collection program is executed by a processor, it implements the steps of the data collection method as described above.
[0172] The present invention also discloses D13. A data collection device, the data collection device includes: an information acquisition module, a priority determination module, and a data collection module;
[0173] The information acquisition module is used to determine the web page to be collected according to the data collection task and acquire the web page feature information of each web page to be collected;
[0174] The priority determination module is configured to determine the scheduling priorities of the data collection tasks according to the web page feature information;
[0175] The data collection module is configured to allocate distributed collection nodes to the data collection tasks according to the scheduling priorities, so that the distributed collection nodes perform web page data collection according to the received data collection tasks.
[0176] D14. The data collection device as described in D13, wherein the priority determination module is further configured to determine the web page weight corresponding to the web page to be collected according to the web page feature information;
[0177] The priority determination module is further configured to obtain the web page quality and web page update time of the web page to be collected based on a preset prediction algorithm;
[0178] The priority determination module is further configured to determine the scheduling priorities of the data collection tasks based on the web page weight, the web page quality, and the web page update time.
[0179] D15. The data collection device as described in D14, wherein the priority determination module is further configured to classify the web page to be collected according to the web page feature information and a preset classification algorithm;
[0180] The priority determination module is further configured to determine the web page score of the web page to be collected based on the classification result, a preset scoring algorithm, and the web page feature information;
[0181] The priority determination module is further configured to determine the web page weight corresponding to the web page to be collected according to the web page score.
[0182] D16. The data collection device as described in D13, wherein the data collection module is further configured to obtain the element density of the target element in the web page to be collected when the web page data collection is completed;
[0183] The data collection module is further configured to determine the element type corresponding to the target element according to the element density;
[0184] The data collection module is further configured to determine the web page structure corresponding to the web page to be collected based on the element type.
[0185] D17. The data collection device as described in D16, wherein the data collection module is further configured to determine the target display dimension corresponding to the web page to be collected when the web page data collection is completed, and the target display dimension includes at least one of a time dimension, a space dimension, and an attribute dimension;
[0186] The data acquisition module is further configured to visually display the web page data of the to-be-acquired web page by means of the data display mode corresponding to the target display dimension.
[0187] D18. For the data acquisition device as described in D13 - D17, the information acquisition module is further configured to configure a preset data acquisition mechanism, and the preset data acquisition mechanism includes: a stable acquisition mechanism, a reliable acquisition mechanism, a flexible acquisition mechanism, and a secure acquisition mechanism.
[0188] The data acquisition module is further configured to allocate distributed acquisition nodes for the respective data acquisition tasks according to the scheduling priority, so that the distributed acquisition nodes perform web page data acquisition based on the preset data acquisition mechanism and the received data acquisition tasks.
[0189] D19. For the data acquisition device as described in D15, the information acquisition module is further configured to obtain the target web page data, the target web page quality, and the target update time corresponding to the target web page through web crawler technology.
[0190] The information acquisition module is further configured to construct a preset classification algorithm, a preset scoring algorithm, and a preset prediction algorithm based on the target web page data, the target web page quality, and the target update time.
[0191] D20. For the data acquisition device as described in D17, the data acquisition module is further configured to obtain the user requirements corresponding to the target user when the web page data acquisition is completed.
[0192] The data acquisition module is further configured to, if the user requirement is to display the data change trend, determine the target display dimension corresponding to the to-be-acquired web page as the time dimension.
[0193] The data acquisition module is further configured to visually display the web page data of the to-be-acquired web page through a time series graph representing the time dimension.
Claims
1. A data collection method, characterized in that, the data collection method includes: determining the web pages to be collected according to the data collection tasks, and obtaining the web page feature information of each web page to be collected; determining the scheduling priorities of each data collection task according to the web page feature information; allocating distributed collection nodes to each data collection task according to the scheduling priorities, so that the distributed collection nodes collect web page data according to the received data collection tasks.
2. The data collection method according to claim 1, characterized in that, the step of determining the scheduling priorities of each data collection task according to the web page feature information includes: determining the web page weight corresponding to the web page to be collected according to the web page feature information; obtaining the web page quality and web page update time of the web page to be collected based on a preset prediction algorithm; determining the scheduling priorities of each data collection task based on the web page weight, the web page quality and the web page update time.
3. The data collection method according to claim 2, characterized in that, the step of determining the web page weight corresponding to the web page to be collected according to the web page feature information includes: classifying the web page to be collected according to the web page feature information and a preset classification algorithm; determining the web page score of the web page to be collected based on the classification result, a preset scoring algorithm and the web page feature information; determining the web page weight corresponding to the web page to be collected according to the web page score.
4. The data collection method according to claim 1, characterized in that, after the step of allocating distributed collection nodes to each data collection task according to the scheduling priorities, so that the distributed collection nodes collect web page data according to the received data collection tasks, it further includes: when the web page data collection is completed, obtaining the element density of the target element in the web page to be collected; determining the element type corresponding to the target element according to the element density; determining the web page structure corresponding to the web page to be collected based on the element type.
5. The data collection method according to claim 4, characterized in that, after the step of allocating distributed collection nodes to each data collection task according to the scheduling priorities, so that the distributed collection nodes collect web page data according to the received data collection tasks, it further includes: when the web page data collection is completed, determining the target display dimension corresponding to the web page to be collected, and the target display dimension includes at least one of a time dimension, a space dimension and an attribute dimension; visually displaying the web page data of the web page to be collected by the data display method corresponding to the target display dimension.
6. The data collection method according to any one of claims 1 to 5, characterized in that, before the step of determining the web pages to be collected according to the data collection tasks, and obtaining the web page feature information of each web page to be collected, it further includes: configuring a preset data collection mechanism, and the preset data collection mechanism includes: a stable collection mechanism, a reliable collection mechanism, a flexible collection mechanism and a secure collection mechanism; The step of allocating distributed acquisition nodes for each data acquisition task according to the scheduling priority, so that the distributed acquisition nodes perform web page data acquisition according to the received data acquisition tasks, includes: Allocating distributed acquisition nodes for each data acquisition task according to the scheduling priority, so that the distributed acquisition nodes perform web page data acquisition based on the preset data acquisition mechanism and the received data acquisition tasks.
7. The data acquisition method according to claim 3, wherein, Before the step of determining the web pages to be acquired according to the data acquisition tasks and obtaining the web page feature information of each web page to be acquired, it further includes: Obtaining the target web page data, target web page quality, and target update time corresponding to the target web page through web crawler technology; Constructing a preset classification algorithm, a preset scoring algorithm, and a preset prediction algorithm based on the target web page data, the target web page quality, and the target update time.
8. A data acquisition device, wherein, The device includes: a memory, a processor, and a data acquisition program stored on the memory and executable on the processor, and the data acquisition program is configured to implement the steps of the data acquisition method according to any one of claims 1 to 7.
9. A storage medium, wherein, A data acquisition program is stored on the storage medium, and when the data acquisition program is executed by a processor, it implements the steps of the data acquisition method according to any one of claims 1 to 7.
10. A data acquisition device, wherein, The data acquisition device includes: an information acquisition module, a priority determination module, and a data acquisition module; The information acquisition module is configured to determine the web pages to be acquired according to the data acquisition tasks and obtain the web page feature information of each web page to be acquired; The priority determination module is configured to determine the scheduling priority of each data acquisition task according to the web page feature information; The data acquisition module is configured to allocate distributed acquisition nodes for each data acquisition task according to the scheduling priority, so that the distributed acquisition nodes perform web page data acquisition according to the received data acquisition tasks.