Self-adaptive data acquisition method and device
Through the feature extraction and fusion processing of web page source code and crawler information, a crawler strategy generation model is generated, which solves the problem of insufficient dynamic adaptability of data acquisition methods in the existing technology, and realizes efficient and accurate data acquisition, reducing maintenance costs and improving collection quality.
Patent Information
- Application Number
- CN202510450850.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-11
AI Technical Summary
Existing data acquisition methods lack dynamic adaptability to target data sources, resulting in waste of resources, redundancy of data or omission of key information, especially in the collection of web text data, it is difficult to obtain the required information efficiently and accurately.
By obtaining the pending web page source code information, historical web page source code collection and crawler information collection, feature extraction and fusion processing are performed, crawler strategy generation model is generated, acquisition strategy is dynamically adjusted, the operation frequency, access depth and algorithm type of crawler program are optimized, and machine learning models are used for training and prediction, and finally an efficient acquisition strategy is generated.
Adaptive adjustment of the acquisition strategy is achieved according to actual conditions, improving the efficiency and quality of data acquisition, and reducing the daily maintenance costs of technicians.
Smart Images

Figure CN120386907A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text data collection, and particularly to an adaptive data collection method and device. Background Art
[0002] With the rapid development of information technology, data collection plays a crucial role in many fields such as industrial monitoring, environmental monitoring, medical diagnosis, financial analysis, intelligent transportation, and smart agriculture. However, most of the existing data collection methods adopt fixed collection frequencies and static collection strategies, lacking the dynamic adaptation ability to target data sources, resulting in problems such as resource waste, data redundancy, or omission of key information. Especially in web page text data collection, due to differences in factors such as web page structure, content update frequency, and data value density, traditional methods often struggle to efficiently and accurately obtain the required information. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide an adaptive data collection method and device that can adaptively adjust the collection strategy according to the actual situation, and can perform data collection efficiently and accurately, thereby helping to reduce the daily maintenance cost of technicians and improve the efficiency and quality of data collection.
[0004] To solve the above technical problem, in the first aspect of the embodiments of the present invention, an adaptive data collection method is disclosed, and the method includes:
[0005] S1, obtaining the source code information of the web page to be processed, the historical web page source code set, and the crawler information set; the historical web page source code set includes several historical web page source code information; the crawler information set includes several crawler program information;
[0006] S2, performing fusion processing on the historical web page source code set and the crawler information set to obtain crawler strategy generation model information;
[0007] S3, using the crawler strategy generation model information to process the source code information of the web page to be processed to obtain web page collection information.
[0008] As an optional implementation manner, in the first aspect of the embodiments of the present invention, the performing fusion processing on the historical web page source code set and the crawler information set to obtain crawler strategy generation model information includes:
[0009] S21, performing feature extraction processing on the historical web page source code set and the crawler information set to obtain a web page feature vector set and a crawler information feature vector set; the web page feature vector set includes several web page feature vectors; the crawler information feature vector set includes several crawler information feature vectors;
[0010] S22. Process the web page feature vector set and the crawler information feature vector set to obtain a web page training data set;
[0011] S23. Process the web page training data set to obtain crawler strategy generation model information.
[0012] As an optional implementation manner, in the first aspect of the embodiments of the present invention, the feature extraction and processing of the historical web page source code set and the crawler information set to obtain a web page feature vector set and a crawler information feature vector set includes:
[0013] S211. Process the historical web page source code set to obtain a web page feature vector set;
[0014] S212. Process the crawler information set to obtain a crawler information feature vector set.
[0015] As an optional implementation manner, in the first aspect of the embodiments of the present invention, the processing of the crawler information set to obtain a crawler information feature vector set includes:
[0016] S2121. Perform an extraction operation on any crawler program information in the crawler information set to obtain the running frequency feature information and access depth feature information of the crawler program information;
[0017] S2122. Perform statistical processing on the crawler program information to obtain the data volume feature information and algorithm type feature information of the crawler program information;
[0018] S2123. Perform quantization processing on the running frequency feature information, the access depth feature information, the data volume feature information, and the algorithm type feature information to obtain quantization feature information;
[0019] S2124. Use a crawler vectorization calculation model to perform vector conversion processing on the quantization feature information to obtain a crawler information feature vector corresponding to the crawler program information
[0020] Among them, the crawler vectorization calculation model is:
[0021]
[0022] In the formula, v i is the i-th component of the crawler information feature vector, QZ i is the weight value corresponding to the quantization information of the i-th feature in the quantization feature information, LH i 、LH j and LH kThey are respectively the quantization information of the i-th feature, the quantization information of the j-th feature, and the quantization information of the k-th feature in the quantization feature information. α is the quantization feature information adjustment factor, XG i,j is the correlation coefficient between the i-th feature and the j-th feature in the quantization feature information, XG i,k is the correlation coefficient between the i-th feature and the k-th feature in the quantization feature information. N is the number of features in the quantization feature information.
[0023] As an optional implementation manner, in the first aspect of the embodiments of the present invention, the processing of the web training data set to obtain the crawler policy generation model information includes:
[0024] S231. Perform missing value processing on the web training data set to obtain a first training data set;
[0025] S232. Perform outlier processing on the first training data set to obtain a second training data set;
[0026] S233. Use the second training data set to perform training processing on the initial crawler policy model information to obtain the crawler policy generation model information.
[0027] As an optional implementation manner, in the first aspect of the embodiments of the present invention, the performing outlier processing on the first training data set to obtain a second training data set includes:
[0028] S2321. Use the local sample density calculation model to perform calculation processing on the first training data set to obtain sample local density information; the sample local density information includes a number of sample local density values;
[0029] Among them, the local sample density calculation model is:
[0030]
[0031] In the formula, ρ i3 is the i3-th sample local density value in the sample local density information, YBB i3 and YBB j3 are respectively the i3-th first training sample information and the j3-th first training sample information in the first training data set. M3 is the number of the first training sample information in the first training data set, and ‖·‖ is the Euclidean norm;
[0032] S2322. Perform an averaging process on all the sample local density values in the sample local density information to obtain a sample average density value;
[0033] S2323. Process the sample average density value and a preset density threshold coefficient to obtain a sample density threshold;
[0034] S2324. Preset s = 1;
[0035] S2325. Determine whether the s-th sample local density value in the sample local density information is less than the sample density threshold to obtain a first judgment result;
[0036] When the first judgment result is negative, add the s-th first training sample information in the first training data set to the second training data set;
[0037] When the first judgment result is positive, execute S2326;
[0038] S2326. Determine whether s is greater than the number of the first training sample information in the first training data set to obtain a second judgment result;
[0039] When the second judgment result is negative, increment s by 1 and execute S2324;
[0040] When the second judgment result is positive, execute S233.
[0041] As an optional implementation manner, in the first aspect of the embodiments of the present invention, the generating model information by using the crawler strategy to process the to-be-processed web page source code information to obtain web page collection information includes:
[0042] S31. Generate model information by using the crawler strategy to process the to-be-processed web page source code information to obtain collection strategy encoding information;
[0043] S32. Perform parameter configuration processing on the initial crawler program information according to the collection strategy encoding information to obtain target crawler program information;
[0044] S33. Use the target crawler program information to process the to-be-processed web page source code information to obtain collection data information;
[0045] S34. Process the collection data information to obtain web page collection information.
[0046] The second aspect of the embodiments of the present invention discloses an adaptive data collection device, and the device includes:
[0047] An acquisition module, configured to acquire to-be-processed web page source code information, a historical web page source code set, and a crawler information set; the historical web page source code set includes a plurality of historical web page source code information; the crawler information set includes a plurality of crawler program information;
[0048] A first computing module, configured to perform fusion processing on the historical web page source code set and the crawler information set to obtain crawler policy generation model information;
[0049] A second computing module, configured to process the to-be-processed web page source code information by using the crawler policy generation model information to obtain web page collection information.
[0050] A third aspect of the embodiments of the present invention discloses another adaptive data collection device, and the device includes:
[0051] A processor;
[0052] A memory coupled to the processor and storing executable program code;
[0053] The processor calls the executable program code stored in the memory to execute some or all of the steps of the adaptive data collection method disclosed in the first aspect of the embodiments of the present invention.
[0054] A fourth aspect of the embodiments of the present invention discloses a computer-readable storage medium, and the computer-readable storage medium stores computer instructions, and when the computer instructions are called, they are used to execute some or all of the steps of the adaptive data collection method disclosed in the first aspect of the embodiments of the present invention.
[0055] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0056] In the embodiments of the present invention, to-be-processed web page source code information, a historical web page source code set, and a crawler information set are obtained; the historical web page source code set includes a plurality of historical web page source code information; the crawler information set includes a plurality of crawler program information; fusion processing is performed on the historical web page source code set and the crawler information set to obtain crawler policy generation model information; the to-be-processed web page source code information is processed by using the crawler policy generation model information to obtain web page collection information. It can be seen that this embodiment can adaptively adjust the collection policy according to the actual situation, can perform data collection efficiently and accurately, thereby helping to reduce the daily maintenance cost of technicians and improve the efficiency and quality of data collection. Description of the Drawings
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0058] Figure 1Schematic flowchart of an adaptive data acquisition method disclosed in an embodiment of the present invention;
[0059] Figure 2 Schematic structural diagram of an adaptive data acquisition device disclosed in an embodiment of the present invention;
[0060] Figure 3 Schematic structural diagram of another adaptive data acquisition device disclosed in an embodiment of the present invention. Detailed implementation manners
[0061] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0062] The terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or equipment.
[0063] Referring to "embodiment" herein means that a specific feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0064] The present invention discloses an adaptive data acquisition method and device, which can adaptively adjust the acquisition strategy according to the actual situation, and can perform data acquisition efficiently and accurately, thereby facilitating the reduction of the daily maintenance cost of technical personnel and improving the efficiency and quality of data acquisition. The following will be described in detail respectively.
[0065] Embodiment 1
[0066] Please refer to Figure 1 , Figure 1 is a schematic flowchart of an adaptive data acquisition method disclosed in an embodiment of the present invention. Among them, Figure 1The described adaptive data acquisition method is applied to an adaptive data acquisition device, such as a local server or a cloud server for optimizing and managing adaptive data acquisition, etc., which is not limited in the embodiments of the present invention. As Figure 1 shown, the adaptive data acquisition method may include the following operations:
[0067] S1. Obtain the source code information of the web page to be processed, the historical web page source code set, and the crawler information set; the historical web page source code set includes several historical web page source code information; the crawler information set includes several crawler program information;
[0068] It should be noted that the above historical web page source code set is a web page source code library: collect no less than 100 representative web page source code samples (historical web page source code information), covering four major categories: e-commerce platforms (accounting for 40%), news portals (accounting for 30%), social networks (accounting for 20%), and government agencies (accounting for 10%); the crawler information set is a crawler program library: each web page source code corresponds to at least 1 verified and effective crawler program source code (crawler program information), and a total of more than 100 web page-crawler matching pairs are formed.
[0069] S2. Perform fusion processing on the historical web page source code set and the crawler information set to obtain crawler strategy generation model information;
[0070] S3. Use the crawler strategy generation model information to process the source code information of the web page to be processed to obtain web page acquisition information.
[0071] It can be seen that implementing the adaptive data acquisition method described in the embodiments of the present invention can adaptively adjust the acquisition strategy according to the actual situation, and can perform data acquisition efficiently and accurately, thus helping to reduce the daily maintenance cost of technicians and improve the efficiency and quality of data acquisition.
[0072] In an optional embodiment, the performing fusion processing on the historical web page source code set and the crawler information set to obtain crawler strategy generation model information includes:
[0073] S21. Perform feature extraction processing on the historical web page source code set and the crawler information set to obtain a web page feature vector set and a crawler information feature vector set; the web page feature vector set includes several web page feature vectors; the crawler information feature vector set includes several crawler information feature vectors;
[0074] S22. Process the web page feature vector set and the crawler information feature vector set to obtain a web page training data set;
[0075] It should be noted that the above processing is to merge the web page feature vector set and the crawler information feature vector set. For example, if the web page feature vector set is [a, b, c] and the crawler information feature vector set is [d, e, f, g], after the above merging process, the obtained web page training data set is [a, b, c, d, e, f, g].
[0076] S23. Process the web page training data set to obtain crawler policy generation model information.
[0077] It can be seen that implementing the adaptive data acquisition method described in the embodiments of the present invention can adaptively adjust the acquisition strategy according to the actual situation, and can perform data acquisition efficiently and accurately, thereby facilitating the reduction of the daily maintenance cost of technicians and improving the efficiency and quality of data acquisition.
[0078] In another optional embodiment, the feature extraction process for the historical web page source code set and the crawler information set to obtain a web page feature vector set and a crawler information feature vector set includes:
[0079] S211. Process the historical web page source code set to obtain a web page feature vector set;
[0080] S212. Process the crawler information set to obtain a crawler information feature vector set.
[0081] It can be seen that implementing the adaptive data acquisition method described in the embodiments of the present invention can adaptively adjust the acquisition strategy according to the actual situation, and can perform data acquisition efficiently and accurately, thereby facilitating the reduction of the daily maintenance cost of technicians and improving the efficiency and quality of data acquisition.
[0082] In another optional embodiment, the process of processing the historical web page source code set to obtain a web page feature vector set includes:
[0083] S2111. Perform syntax parsing on any historical web page source code information in the historical web page source code set to obtain DOM tree structure information corresponding to the historical web page source code information;
[0084] It should be noted that for the above syntax parsing process, the BeautifulSoup library, lxml library or html5lib library of python can be used for processing. Specifically, the embodiments of the present invention do not make any limitations.
[0085] It should be noted that through DOM structure parsing, the hierarchical structure of the web page can be extracted. In subsequent solutions, the crawler program information can dynamically adjust XPath or CSS Selector according to the changes in the web page DOM structure to improve the crawling success rate.
[0086] S2112, traverse the DOM tree structure information to obtain hierarchical node list information; the hierarchical node list information includes a number of hierarchical node information; the hierarchical node information includes tag type information, node text information, attribute information, and hierarchical depth;
[0087] It should be noted that the above-mentioned tag type information, node text information, attribute information, and hierarchical depth are as follows:
[0088] Tag type information: refers to the tag name of each HTML element in the web page DOM tree, such as 、 、 、 。
[0089] Node text information: Refers to the pure text content contained in an HTML element, which is the actual readable text after removing the HTML tags.
[0090] Attribute information: Refers to the attributes of an HTML tag (Attributes), such as id, class, href, src, etc.
[0091] The hierarchical depth refers to the nested level of a certain DOM element in the HTML structure, that is, the distance from the root node to this element.
[0092] It should be noted that the above traversal can use depth-first traversal or breadth-first traversal. Specifically, the embodiments of the present invention do not make any limitations.
[0093] It should be noted that through traversal, tag type information, node text information, attribute information, and hierarchical depth can be obtained. These information can be input into the initial model information of the subsequent crawler strategy, which can improve the accuracy of model training, so that the finally trained model can effectively improve the accuracy of data collection.
[0094] S2113. Perform path weight calculation processing on the hierarchical node list information to obtain Xpath weight information; the Xpath weight information includes several Xpath weight values.
[0095] It should be noted that Xpath weight calculation can optimize the selection of crawler paths and improve the success rate of subsequent data collection.
[0096] It should be noted that Xpath weight information is a set of weight values calculated based on the Xpath path, which is used to measure the importance of different nodes in the web page DOM structure.
[0097] S2114. Perform calculation processing on the attribute information to obtain CSS strength information.
[0098] It should be noted that the processing of S2113 and S2114 above can be performed using the BeautifulSoup library, lxml library, or html5lib library of python. Specifically, the embodiments of the present invention do not make any limitations.
[0099] CSS strength information mainly refers to the style hierarchy, weight, and stability of web page elements, and can measure the importance of web page elements by calculating attribute information.
[0100] It should be noted that relying solely on the DOM structure and Xpath for web page feature extraction may lead to when the web page structure changes slightly (such as adding The content of the package), the Xpath fails, resulting in the failure of crawling. Through the CSS strength information, it can help measure the relative importance of web page elements, enabling the crawler to still identify the main content when the structure changes.
[0101] S2115, using the text content feature calculation model, calculate and process the node text information to obtain content feature information;
[0102] Among them, the text content feature calculation model is:
[0103]
[0104] In the formula, NR is the content feature information, wd j2 is the structural complexity value of the j2th sentence in the node text information, cp i2 is the frequency of the i2th word in the node text information in the node text information, cx i2 is the part-of-speech weight of the i2th word in the node text information, β is the sentence complexity coefficient, M2 is the number of sentences in the node text information, and N2 is the number of words in the node text information;
[0105] It should be noted that the sentence complexity coefficient can be set by the user or obtained from historical data. Specifically, the embodiments of the present invention do not make limitations.
[0106] It should be noted that through the text content feature calculation model, the content feature information NR of the web page node text is calculated to judge the complexity and importance of the node text information, so as to screen out more valuable text content when extracting web page features and optimizing the crawler strategy.
[0107] It should be noted that the sentence complexity value is used to reflect the syntactic complexity of the sentence and measure whether the text contains more long sentences, clauses, modified structures, etc. It can be calculated by Stanford NLP, spaCy, etc. Generally, its value range is: 0 to 1 for general sentences; 0.1 to 0.3 for simple sentences; 0.3 to 0.6 for medium complexity; 0.6 to 1.0 for high complexity;
[0108] It should be noted that the part-of-speech weight reflects the importance of different words in the text content. For example, the part-of-speech weight distribution of various words is: noun is 1.0, verb is 0.9, adjective is 0.8, adverb is 0.6, preposition is 0.3, pronoun is 0.2. Specifically, the embodiments of the present invention do not make limitations.
[0109] It should be noted that the value range of the sentence complexity coefficient is between [0.5, 2]. When the value range is between [0.5, 1], it is applicable to ordinary texts such as news and blogs; when the value range is between [1, 1.5], it is applicable to complex texts such as papers and technical documents; when the value range is between [1.5, 2], it is applicable to professional literature that requires key analysis of long sentences. The sentence complexity coefficient determines the influence degree of sentence structure complexity on content feature information. If the value of the sentence complexity coefficient is too high, it may cause the crawler to only focus on complex texts and ignore simple but important information. If it is too low, it may cause the crawler to capture meaningless short sentences.
[0110] S2116. Normalize and splice the Xpath weight information, the CSS strength information, and the content feature information to obtain a web page feature vector corresponding to the historical web page source code information.
[0111] It should be noted that for the above normalization process, Min-Max normalization can be used. For the above splicing process, different features are spliced together. For example, if the normalized Xpath weight information is represented as a1, the normalized CSS strength information is represented as a2, and the normalized content feature information is represented as a3, then the spliced web page feature vector is [a1, a2, a3].
[0112] It should be noted that through normalization and splicing, multiple different types of features can be fused together, enabling the model to utilize structural information, style information, and text content information simultaneously, thereby improving the overall grasping ability of web page features.
[0113] It can be seen that implementing the adaptive data acquisition method described in the embodiments of the present invention can adaptively adjust the acquisition strategy according to the actual situation, and can perform data acquisition efficiently and accurately, thereby helping to reduce the daily maintenance cost of technicians and improve the efficiency and quality of data acquisition.
[0114] In an optional embodiment, the processing of the crawler information set to obtain a crawler information feature vector set includes:
[0115] S2121. Extract any crawler program information in the crawler information set to obtain the running frequency feature information and access depth feature information of the crawler program information.
[0116] It should be noted that for the above extraction operation, Scrapy of Python can be used for processing. Specifically, the embodiments of the present invention do not make limitations.
[0117] It should be noted that by using the running frequency feature information and the access depth feature information as part of model training, it can help the model understand the running modes of different crawler strategies, thereby predicting better crawler strategies and improving the accuracy and efficiency of web data collection.
[0118] Among them, the running frequency feature information reflects the number of access requests of the crawler per unit time and is used to measure the crawling speed of the crawler. The access depth feature information reflects the number of link jump levels that the crawler needs to pass through from the home page to the target page during the crawling process.
[0119] S2122. Statistically process the crawler program information to obtain the data volume feature information and the algorithm type feature information of the crawler program information;
[0120] It should be noted that the above statistical processing can be performed using Scrapy in Python. Specifically, the embodiments of the present invention do not make limitations.
[0121] It should be noted that the combination of the data volume feature information and the algorithm type feature information with the previously mentioned running frequency feature information and access depth feature information can be used to optimize the crawling strategy and improve the success rate and adaptability of web data collection.
[0122] Among them, the data volume feature information measures the data acquisition volume of the crawler within a certain time or a certain crawling task. The algorithm type feature information is used to identify the strategies used by the crawler (such as BFS / DFS, multithreading, asynchronous, etc.).
[0123] S2123. Quantify the running frequency feature information, the access depth feature information, the data volume feature information, and the algorithm type feature information to obtain quantified feature information;
[0124] It should be noted that the quantization process can be performed using Scikit-learn, Pandas, or Numpy. Specifically, the embodiments of the present invention do not make limitations. For categorical features, such as crawler algorithm types (BFS, DFS, Async, Multithread, etc.), they can be converted into numerical features through One-Hot encoding. For some non-numerical features, discretization processing may be required to convert continuous features into discrete features. For example, the access depth may be discretized into different levels (such as shallow, medium, deep). Discretization helps to process discontinuous features and enables them to be combined with other features in the model. For numerical types, the features can be scaled to a specific range, usually [0,1] or [-1,1], through normalization, so that different feature values can be compared on the same scale.
[0125] Through the quantization operation, the unstructured feature information of different scales is converted into digital features that can be processed in the machine learning model. These quantized feature information can effectively serve as the input of the model, ensuring that different features can participate in model training, optimization, and inference at the same scale.
[0126] S2124, use the crawler vectorization calculation model to perform vector conversion processing on the quantized feature information to obtain the crawler information feature vector corresponding to the crawler program information.
[0127] Among them, the crawler vectorization calculation model is:
[0128]
[0129] In the formula, v i is the i-th component of the crawler information feature vector, QZ i is the weight value corresponding to the quantization information of the i-th feature in the quantized feature information, LH i , LH j and LH k are respectively the quantization information of the i-th feature, the quantization information of the j-th feature, and the quantization information of the k-th feature in the quantized feature information, α is the quantized feature information adjustment factor, XG i,j is the correlation coefficient between the i-th feature and the j-th feature in the quantized feature information, XG i,k is the correlation coefficient between the i-th feature and the k-th feature in the quantized feature information, and N is the number of features in the quantized feature information.
[0130] It should be noted that the value of the above-mentioned quantized feature information adjustment factor can be set by the user or obtained according to historical data. Specifically, the embodiments of the present invention do not make any limitations.
[0131] It should be noted that the quantized feature information adjustment factor mainly controls the relative importance of the quantized feature information in the calculation. It can affect the output of the model by adjusting the contribution of the quantized feature to the crawler information vector. By adjusting the quantized feature information adjustment factor, the influence of each feature in the final result can be balanced, thereby improving the prediction accuracy of the model, and its value range is between (0, 1).
[0132] It should be noted that the correlation coefficient is an index to measure the correlation degree between two features, and its value range is between [-1, 1]. By using the correlation coefficient, it is possible to identify which features have a strong dependence relationship, so as to appropriately adjust their contribution degrees in the calculation of the feature vector. Features with high correlation may cause redundancy in the model and reduce the calculation efficiency. Appropriate adjustment can help improve the accuracy and stability of the model.
[0133] Weight value QZ i It represents the importance of the i-th feature in the entire feature set. Its main function is to assign different relative weights to different features and determine the contribution degree of features in the model. The larger the weight value, the stronger the influence of the feature. The value range of the weight value is usually [0, 1].
[0134] It should be noted that through the above crawler vectorization calculation model, the original crawler features are converted into an efficient feature vector through quantization, weighting, correlation adjustment, etc. for the model to use for training or inference, thereby improving the accuracy of the model's data parsing, collection strategy generation, and dynamic adjustment.
[0135] It can be seen that implementing the adaptive data collection method described in the embodiments of the present invention can adaptively adjust the collection strategy according to the actual situation, and can perform data collection efficiently and accurately, thereby helping to reduce the daily maintenance cost of technicians and improve the efficiency and quality of data collection.
[0136] In an alternative embodiment, the processing of the web page training data set to obtain the crawler strategy generation model information includes:
[0137] S231, perform missing value processing on the web page training data set to obtain a first training data set;
[0138] It should be noted that the above missing value processing can be performed using python libraries such as pandas and missingno or K-nearest neighbor filling. Specifically, the embodiments of the present invention do not make limitations.
[0139] S232, perform outlier processing on the first training data set to obtain a second training data set;
[0140] S233, use the second training data set to perform training processing on the initial model information of the crawler strategy to obtain the crawler strategy generation model information.
[0141] It can be seen that implementing the adaptive data collection method described in the embodiments of the present invention can adaptively adjust the collection strategy according to the actual situation, and can perform data collection efficiently and accurately, thereby helping to reduce the daily maintenance cost of technicians and improve the efficiency and quality of data collection.
[0142] In an alternative embodiment, the performing outlier processing on the first training data set to obtain a second training data set includes:
[0143] S2321. Use the local sample density calculation model to perform calculation processing on the first training data set to obtain sample local density information; the sample local density information includes a number of sample local density values.
[0144] Among them, the local sample density calculation model is:
[0145]
[0146] In the formula, ρ i3 is the i3th sample local density value in the sample local density information, YBB i3 and YBB j3 are the i3th first training sample information and the j3th first training sample information in the first training data set respectively, M3 is the number of the first training sample information in the first training data set, and ‖·‖ is the Euclidean norm.
[0147] It should be noted that the local sample density calculation model is used to calculate the local density of each sample in the training data set. Its role is to measure the distribution of data points around a sample, which helps to discover the aggregation areas of data. By calculating, the sample local density value is obtained. If the sample local density value is small, it is a data sparse area, and may even be an outlier (abnormal point).
[0148] S2322. Perform an averaging process on all the sample local density values in the sample local density information to obtain a sample average density value.
[0149] S2323. Process the sample average density value and a preset density threshold coefficient to obtain a sample density threshold.
[0150] It should be noted that the value of the preset density threshold coefficient is 0.2. The above processing is to multiply the sample average density value by the preset density threshold coefficient to obtain the sample density threshold.
[0151] It should be noted that the sample density threshold is used to screen abnormal samples, optimize the training data, and improve the adaptive ability of the crawler strategy.
[0152] S2324. Preset s = 1.
[0153] S2325. Judge whether the s-th sample local density value in the sample local density information is less than the sample density threshold to obtain a first judgment result.
[0154] When the first judgment result is negative, add the s-th first training sample information in the first training data set to the second training data set.
[0155] When the first judgment result is yes, execute S2326;
[0156] S2326, determine whether s is greater than the number of the first training sample information in the first training dataset, and obtain a second judgment result;
[0157] When the second judgment result is no, increment s by 1 and execute S2324;
[0158] When the second judgment result is yes, execute S233.
[0159] It can be seen that implementing the adaptive data acquisition method described in the embodiments of the present invention can adaptively adjust the acquisition strategy according to the actual situation, and can perform data acquisition efficiently and accurately, thereby facilitating the reduction of the daily maintenance cost of technicians and improving the efficiency and quality of data acquisition.
[0160] In an optional embodiment, the training the initial crawler strategy model information by using the second training dataset to obtain the crawler strategy generation model information includes:
[0161] S2331, train the initial crawler strategy model information by using the second training dataset to obtain the crawler strategy training result information and the crawler strategy training model information;
[0162] It should be noted that the above initial crawler strategy model information is an LSTM model.
[0163] S2332, perform a calculation process on the crawler strategy training result information to obtain the crawler strategy loss function value;
[0164] S2333, determine whether the crawler strategy loss function value is less than a preset crawler strategy loss function threshold to obtain a third judgment result;
[0165] It should be noted that the value range of the preset crawler strategy loss function threshold is between [0.01, 0.05]. Specifically, the embodiments of the present invention do not make a limitation.
[0166] When the third judgment result is no, determine that the crawler strategy training model information is the initial crawler strategy model information, update the model parameter information of the initial crawler strategy model information, and execute S2331;
[0167] It should be noted that the above update process for the model parameter information of the crawler strategy initial model information is performed by combining the backpropagation algorithm with the gradient descent method. The model parameter information includes the weights of the forget gate, input gate, candidate cell state, output gate, learning rate, and the number of samples used for training, etc. Specifically, the embodiments of the present invention do not make any limitations in this regard.
[0168] When the third judgment result is yes, it is determined that the crawler strategy training model information is the crawler strategy generation model information.
[0169] It can be seen that implementing the adaptive data acquisition method described in the embodiments of the present invention can adaptively adjust the acquisition strategy according to the actual situation, and can perform data acquisition efficiently and accurately, which is beneficial to reducing the daily maintenance cost of technicians and improving the efficiency and quality of data acquisition.
[0170] In an optional embodiment, the calculation process for the crawler strategy training result information to obtain the crawler strategy loss function value includes:
[0171] Using the crawler strategy loss function calculation model, the crawler strategy training result information is calculated to obtain the crawler strategy loss function value;
[0172] Among them, the crawler strategy loss function calculation model is:
[0173]
[0174] In the formula, SH is the crawler strategy loss function value, JG is the true label information corresponding to the crawler strategy training result information, ZZ is the crawler strategy training result information, M1 and N1 are respectively the number of features in the crawler strategy training result information and the number of crawler strategy training result values corresponding to each feature, JG j1 i1 and ZZ j1 i1 are respectively the i1-th true label value of the j1-th feature in the true label information corresponding to the crawler strategy training result information and the i1-th crawler strategy training result value of the j1-th feature in the crawler strategy training result information, and δ1 is the first weight parameter.
[0175] It should be noted that the first weight parameter can be set by the user or obtained according to historical data. Specifically, the embodiments of the present invention do not make any limitations in this regard.
[0176] It should be noted that the value range of the first weight parameter is between [0, 1], which is used to balance the prediction error and stability. When the first weight parameter is relatively large, the model pays more attention to the relative error and is suitable for the situation where the data distribution is relatively unbalanced. When the first weight parameter is relatively small, the model relies more on the cross-entropy loss and is suitable for the situation where the data is relatively balanced.
[0177] It should be noted that by calculating the model through the crawler strategy loss function, the generalization ability of the model can be improved, enabling it to still adapt when the web page changes, and optimizing the accuracy of the model's prediction of the crawler strategy.
[0178] It can be seen that implementing the adaptive data acquisition method described in the embodiments of the present invention can adaptively adjust the acquisition strategy according to the actual situation, perform data acquisition efficiently and accurately, thereby helping to reduce the daily maintenance cost of technicians and improve the efficiency and quality of data acquisition.
[0179] In an optional embodiment, the generating model information using the crawler strategy to process the source code information of the web page to be processed to obtain web page acquisition information includes:
[0180] S31. Using the model information generated by the crawler strategy to process the source code information of the web page to be processed to obtain acquisition strategy coding information;
[0181] It should be noted that the acquisition strategy coding information refers to the structured crawling instructions generated by the crawler strategy generation model based on the source code information of the web page to be processed, which is used to guide the crawler program to efficiently access, parse, and extract web page data.
[0182] The above processing takes the source code information of the web page to be processed as the input of the crawler strategy generation model information, and performs calculation processing through the crawler strategy generation model information to output the acquisition strategy coding information.
[0183] S32. According to the acquisition strategy coding information, perform parameter configuration processing on the initial crawler program information to obtain the target crawler program information;
[0184] It should be noted that the initial crawler program information is obtained. The initial crawler program information refers to the basic crawler program that has not been optimized or customized, usually a general crawler framework or template, which contains basic crawling logic, such as:
[0185] Basic request module (such as making HTTP requests using requests or Scrapy);
[0186] Data parsing module (such as parsing using BeautifulSoup, lxml, or XPath);
[0187] Data storage module (such as JSON, CSV, or database);
[0188] Exception handling mechanisms (such as timeout retry, IP rotation).
[0189] It should be noted that the above parameter configuration processing can be performed through tools such as Jinja2 and ConfigParser. Specifically, the embodiments of the present invention do not make limitations in this regard.
[0190] It should be noted that specifically for the above parameter configuration processing, the initial crawler program information is a basic version of the crawler, which has general crawling capabilities but is not optimized for specific web pages. The collection strategy coding information provides crawling rules (such as access frequency, URL structure, parsing method, anti-crawling measures, etc.). Using the collection strategy coding information for parameter configuration processing means adjusting the information of the initial crawler program so that it can accurately collect web page data according to the strategy, and finally obtaining the target crawler program information.
[0191] S33, using the target crawler program information to process the to-be-processed web page source code information to obtain collection data information;
[0192] It should be noted that the above collection data information is obtained by the target crawler program information collecting data from the to-be-processed web page source code information.
[0193] S34, processing the collection data information to obtain web page collection information.
[0194] It can be seen that implementing the adaptive data collection method described in the embodiments of the present invention can adaptively adjust the collection strategy according to the actual situation, and can perform data collection efficiently and accurately, thereby helping to reduce the daily maintenance cost of technicians and improve the efficiency and quality of data collection.
[0195] In an alternative embodiment, the processing the collection data information to obtain web page collection information includes:
[0196] S341, using a collection success rate calculation model to perform calculation processing on the collection data information to obtain a collection success rate;
[0197] Wherein, the collection success rate calculation model is:
[0198]
[0199] Wherein, cg is the collection success rate, YX is the number of target fields successfully captured in the collected data information, YXZ is the total number of target fields in the collected data information, BL is the number of requests rejected by the target website in the collected data information, QQZ is the total number of requests in the collected data information, JZ is the number of successfully loaded web page dynamic elements in the collected data information, and DTZ is the minimum number of web page dynamic elements required for the normal display of the page; δ2, δ3, and δ4 are the second weight parameter, the third weight parameter, and the fourth weight parameter respectively;
[0200] It should be noted that through the collection success rate calculation model, the collection quality of the current crawler is evaluated, the effectiveness of the crawler strategy is judged, and adaptive adjustment is performed when necessary.
[0201] It should be noted that the value range of the second weight parameter is between [0.4, 0.6], the value range of the third weight parameter is between [0.2, 0.4], and the value range of the fourth weight parameter is between [0.2, 0.4]. Among them, when the value of the second weight parameter is large, more attention is paid to the integrity of data capture (suitable for scenarios with high requirements for data accuracy). When the value of the third weight parameter is large, more attention is paid to the anti-blocking ability of the crawler (suitable for websites with strong anti-crawling). When the value of the fourth weight parameter is large, more attention is paid to the page integrity (suitable for scenarios where AJAX / JS rendered data needs to be captured).
[0202] It should be noted that the number of successfully captured target fields refers to the number of data entries that meet the definition of the target field in the web page data successfully collected by the crawler; the total number of target fields refers to the total number of target fields that should be collected theoretically, that is, the number of all data entries that meet the rules in the web page; the number of requests rejected by the target website refers to the number of requests rejected by the anti-crawling mechanism (such as IP blocking, User-Agent blocking) when the crawler accesses the web page; the total number of requests refers to the total number of all HTTP requests sent by the crawler, including successful and failed requests; the number of successfully loaded web page dynamic elements refers to the number of JavaScript dynamic elements (such as data loaded by AJAX) successfully loaded during page rendering; the minimum number of web page dynamic elements required for the normal display of the page refers to the minimum number of dynamic elements required for the page to fully display data.
[0203] S342. Determine whether the collection success rate is less than a preset collection success threshold to obtain a fourth determination result;
[0204] It should be noted that the value range of the preset collection success threshold is [0.9, 0.95]. Specifically, the embodiments of the present invention do not make limitations.
[0205] When the fourth determination result is yes, execute S31;
[0206] When the fourth judgment result is yes, it indicates that the current crawler strategy fails to effectively collect web page data and the data integrity is not high enough. At this time, it indicates that there are problems such as web page structure changes, crawling rule errors, or anti-crawling mechanism interferences in the current web page. It is necessary to execute S31 to optimize or correct the crawler strategy, so as to achieve the purpose of adaptive data collection, ensure that the finally collected data is complete enough, and reduce the subsequent data cleaning and supplementary collection costs.
[0207] When the fourth judgment result is no, determine that the collected data information is web page collection information.
[0208] It can be seen that implementing the adaptive data collection method described in the embodiments of the present invention can adaptively adjust the collection strategy according to the actual situation, and can perform data collection efficiently and accurately, which is beneficial to reducing the daily maintenance costs of technicians and improving the efficiency and quality of data collection.
[0209] Embodiment 2
[0210] Please refer to Figure 2 , Figure 2 is a schematic structural diagram of an adaptive data acquisition device disclosed in an embodiment of the present invention. Among them, Figure 2 the described adaptive data acquisition device is applied to an adaptive data acquisition optimization system, such as a local server or a cloud server for adaptive data acquisition, etc., which is not limited in the embodiments of the present invention. As Figure 2 shown, the adaptive data acquisition device includes:
[0211] An acquisition module 201, configured to acquire to-be-processed web page source code information, a historical web page source code set, and a crawler information set; the historical web page source code set includes several historical web page source code information; the crawler information set includes several crawler program information;
[0212] A first calculation module 202, configured to perform fusion processing on the historical web page source code set and the crawler information set to obtain crawler strategy generation model information;
[0213] A second calculation module 203, configured to use the crawler strategy generation model information to process the to-be-processed web page source code information to obtain web page acquisition information.
[0214] It can be seen that implementing the adaptive data acquisition device described in the embodiments of the present invention can adaptively adjust the acquisition strategy according to the actual situation, and can perform data acquisition efficiently and accurately, thereby facilitating the reduction of the daily maintenance cost of technicians and improving the efficiency and quality of data acquisition.
[0215] Embodiment Three
[0216] Please refer to Figure 3 , Figure 3 is a schematic structural diagram of another adaptive data acquisition device disclosed in an embodiment of the present invention. Among them, Figure 3 the described adaptive data acquisition device is applied to an adaptive data acquisition optimization system, such as a local server or a cloud server for adaptive data acquisition, etc., which is not limited in the embodiments of the present invention. As Figure 3 shown, the adaptive data acquisition device includes:
[0217] A processor 301;
[0218] A memory 302 coupled to the processor 301 and storing executable program code;
[0219] The processor 301 calls the executable program code stored in the memory 302 to execute some or all of the steps of the adaptive data acquisition method in Embodiment One.
[0220] It can be seen that implementing the adaptive data acquisition device described in the embodiments of the present invention can adaptively adjust the acquisition strategy according to the actual situation, and can perform data acquisition efficiently and accurately, thereby helping to reduce the daily maintenance cost of technicians and improve the efficiency and quality of data acquisition.
[0221] Embodiment 4
[0222] The embodiments of the present invention disclose a computer-readable storage medium. The computer-readable storage medium stores computer instructions, and when the computer instructions are called, they are used to execute some or all of the steps of the adaptive data acquisition method in Embodiment 1.
[0223] Embodiment 5
[0224] The embodiments of the present invention disclose a computer program product. The computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute some or all of the steps of the adaptive data acquisition method described in Embodiment 1.
[0225] The system embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0226] Through the specific descriptions of the above embodiments, those skilled in the art can clearly understand that each implementation can be achieved by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, and the storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium capable of carrying or storing data.
[0227] Finally, it should be noted that: the adaptive data acquisition method and device disclosed in the embodiments of the present invention only disclose the preferred embodiments of the present invention, and are only used to illustrate the technical solutions of the present invention, rather than limiting them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An adaptive data acquisition method, characterized in that, The method includes: S1. Obtain the source code information of the web page to be processed, the historical web page source code set, and the crawler information set; the historical web page source code set includes several historical web page source code information; the crawler information set includes several crawler program information; S2. Perform fusion processing on the historical web page source code set and the crawler information set to obtain crawler strategy generation model information; S3. Use the crawler strategy generation model information to process the source code information of the web page to be processed to obtain web page collection information.
2. The adaptive data acquisition method according to claim 1, wherein The performing fusion processing on the historical web page source code set and the crawler information set to obtain crawler strategy generation model information includes: S21. Perform feature extraction processing on the historical web page source code set and the crawler information set to obtain a web page feature vector set and a crawler information feature vector set; the web page feature vector set includes several web page feature vectors; the crawler information feature vector set includes several crawler information feature vectors; S22. Process the web page feature vector set and the crawler information feature vector set to obtain a web page training data set; S23. Process the web page training data set to obtain crawler strategy generation model information.
3. The adaptive data acquisition method according to claim 2, wherein The performing feature extraction processing on the historical web page source code set and the crawler information set to obtain a web page feature vector set and a crawler information feature vector set includes: S211. Process the historical web page source code set to obtain a web page feature vector set; S212. Process the crawler information set to obtain a crawler information feature vector set.
4. The adaptive data acquisition method according to claim 3, wherein The processing the crawler information set to obtain a crawler information feature vector set includes: S2121. Perform extraction operations on any crawler program information in the crawler information set to obtain the running frequency feature information and access depth feature information of the crawler program information; S2122. Perform statistical processing on the crawler program information to obtain the data volume feature information and algorithm type feature information of the crawler program information; S2123. Perform quantization processing on the running frequency feature information, the access depth feature information, the data volume feature information, and the algorithm type feature information to obtain quantization feature information; S2124. Using a crawler vectorization calculation model, perform vector transformation processing on the quantization feature information to obtain a crawler information feature vector corresponding to the crawler program information Among them, the crawler vectorization calculation model is as follows: where v i is the i-th component of the crawler information feature vector, QZ i is the weight value corresponding to the quantization information of the i-th feature in the quantization feature information, LH i , LH j and LH k are the quantization information of the i-th feature, the quantization information of the j-th feature, and the quantization information of the k-th feature in the quantization feature information respectively, α is the quantization feature information adjustment factor, XG i,j is the correlation coefficient between the i-th feature and the j-th feature in the quantization feature information, XG i,k is the correlation coefficient between the i-th feature and the k-th feature in the quantization feature information, and N is the number of features in the quantization feature information.
5. The adaptive data acquisition method according to claim 2, wherein The processing the web page training data set to obtain crawler strategy generation model information includes: S231. Perform missing value processing on the web page training data set to obtain a first training data set; S232. Perform outlier processing on the first training data set to obtain a second training data set; S233. Use the second training data set to perform training processing on the initial crawler strategy model information to obtain crawler strategy generation model information.
6. The adaptive data acquisition method according to claim 5, wherein, The performing outlier processing on the first training data set to obtain a second training data set includes: S2321. Use the local sample density calculation model to perform calculation processing on the first training data set to obtain sample local density information; the sample local density information includes several sample local density values; Among them, the local sample density calculation model is: where ρ i3 is the i3-th sample local density value in the sample local density information, YBB i3 and YBB j3 are the i3-th first training sample information and the j3-th first training sample information in the first training dataset respectively, M3 is the number of the first training sample information in the first training dataset, and ‖·‖ is the Euclidean norm; S2322. Calculate the average value of all the sample local density values in the sample local density information to obtain the sample average density value; S2323. Process the sample average density value and a preset density threshold coefficient to obtain the sample density threshold; S2324. Preset s = 1; S2325. Determine whether the s-th sample local density value in the sample local density information is less than the sample density threshold to obtain a first judgment result; When the first judgment result is no, add the s-th first training sample information in the first training data set to the second training data set; When the first judgment result is yes, execute S2326; S2326. Determine whether s is greater than the number of the first training sample information in the first training data set to obtain a second judgment result; When the second judgment result is no, increment s by 1 and execute S2324; When the second judgment result is yes, execute S233.
7. The adaptive data acquisition method according to claim 1, wherein The processing of the to-be-processed web page source code information by using the crawler strategy generation model information to obtain the web page collection information includes: S31. Process the to-be-processed web page source code information by using the crawler strategy generation model information to obtain the collection strategy coding information; S32. Perform parameter configuration processing on the initial crawler program information according to the collection strategy coding information to obtain the target crawler program information; S33. Process the to-be-processed web page source code information by using the target crawler program information to obtain the collection data information; S34. Process the collection data information to obtain the web page collection information.
8. An adaptive data acquisition device, characterized in that, The device includes: An acquisition module, configured to acquire to-be-processed web page source code information, a historical web page source code set, and a crawler information set; the historical web page source code set includes several historical web page source code information; the crawler information set includes several crawler program information; A first calculation module, configured to perform fusion processing on the historical web page source code set and the crawler information set to obtain the crawler strategy generation model information; A second calculation module, configured to process the to-be-processed web page source code information by using the crawler strategy generation model information to obtain the web page collection information.
9. An adaptive data acquisition device, characterized in that, The device includes: A processor; A memory coupled to the processor and storing executable program code; The processor calls the executable program code stored in the memory and executes the adaptive data acquisition method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, which are used to execute the adaptive data acquisition method according to any one of claims 1-7 when called.
Citation Information
Patent Citations
Focusing-oriented Web page acquisition and information extraction method
CN106970938A
Universal network crawler model implementation method and system
CN107391775A
Local representation coefficient-based nearest neighbor classification method, storage medium and terminal
CN110288012A
Crawler behavior detection method and device, equipment and storage medium
CN115730112A
Data acquisition method and device, medium and equipment
CN118606535A
Cited By
Data acquisition processing method, equipment and medium
CN120849688A
A data acquisition processing method, device and medium
CN120849688B
Acquisition method and device based on dynamic threshold value and medium
CN120849689A
A dynamic threshold-based acquisition method, device, and medium
CN120849689B