Multi-task collaborative incremental webpage data acquisition method and system
By conducting initial and in-depth evaluations before scheduling web page collection tasks, generating structural consistency and content update indexes, and combining pre-trained models to select collection paths, the problems of resource waste and failure in multi-source web page collection are solved, and the efficiency and stability of multi-task collaborative scheduling are improved.
Patent Information
- Application Number
- CN202510888339.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-06-30
AI Technical Summary
Existing incremental web page collection methods lack the ability to coordinate judgment and dynamic scheduling between tasks, and are unable to cope with the frequent changes in multi-source web page structures and irregular content updates, resulting in resource waste, repeated crawling or collection failure.
By performing an initial assessment before scheduling web page collection tasks, calculating the structural uncertainty score and entering the in-depth evaluation process, generating the page structure consistency index and content update index, and using the pre-trained collection scheduling prediction model to output the collection scheduling coefficient, the appropriate collection path is selected, and resource allocation and scheduling time window adjustments are performed according to the nature of the task.
It achieves pre-awareness of the dynamic changes in the two-dimensional state of web page structure and content, reduces collection failures and information loss, and improves the system's multi-task processing efficiency and stability in complex web page collection.
Smart Images

Figure CN120780428A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of web page data collection, more particularly, the present application relates to a multi-task cooperative incremental web page data collection method and system. BACKGROUND
[0002] With the increasing demand for information acquisition, web page data collection, as a basic link of data mining and knowledge construction, is widely used in various intelligent systems. Especially in the scenes of public opinion monitoring, commodity price comparison, policy tracking, content aggregation, etc., the rapid, accurate and stable collection of web page content is put forward with higher requirements.
[0003] Traditional web page collection methods are mostly based on static rules or timed full-volume crawling methods, which are difficult to deal with frequent changes in web page structure and irregular updates in content. Especially in the environment of processing multiple sources of web pages at the same time, the structural variation frequency and content refresh cycle of different web pages have obvious heterogeneity. If a unified strategy is used for scheduling, it is easy to cause resource waste, repeated crawling or collection failure. In recent years, incremental web page collection methods have gradually attracted attention. The core idea is to collect only the changed web page content since the last collection. However, the existing incremental collection methods are mostly based on single web pages, lack of cooperative judgment and dynamic scheduling ability between tasks, and still have the following defects in actual deployment:
[0004] The existing scheme often relies on time interval or content summary to determine whether to collect, and does not establish a systematic evaluation process based on web page structure level and visual change behavior, resulting in insufficient response to high-frequency structural disturbance pages. For changed web pages, a simple "whether changed" is often used as a basis to trigger collection, and there is a lack of comprehensive modeling of structural evolution trend and content evolution rule, making it difficult to develop accurate collection paths. Therefore, a multi-task cooperative incremental web page data collection method and system are proposed to solve the above problems. SUMMARY
[0005] To achieve the above purpose, the present application provides the following technical scheme:
[0006] The multi-task cooperative incremental web page data collection method comprises the following steps:
[0007] Before scheduling the web page collection task, a preliminary evaluation operation is performed on the target web page. The evaluation operation includes obtaining the structural level change range of the web page in a preset historical period, the reconstruction frequency of the visual area and the rendering position fluctuation, calculating the structural uncertainty score, and determining whether to enter the depth evaluation process according to the score;
[0008] When the structural uncertainty score exceeds a preset disorder threshold, a deep evaluation operation is performed, which includes analyzing the historical variation trend of the webpage structure nodes, the element rearrangement distribution, and the area stability pattern, to generate a page structure consistency index; meanwhile, multiple historical content snapshots are compared to calculate the field update density, semantic variation amplitude, and time distribution law, to generate a content update index;
[0009] The page structure consistency index and the content update index are input into a pre-trained collection scheduling prediction model as input parameters, and a collection scheduling coefficient is output, which is used to evaluate the execution priority and strategy tendency of the current task in the current scheduling period;
[0010] According to the numerical interval of the collection scheduling coefficient, the collection path of the task is selected, which includes a marker tracking and positioning process based on structural reconstruction tolerance enhancement, or an incremental comparison and snapshot update process based on content difference identification;
[0011] After determining the collection path, the combination relationship of the page structure consistency index and the content update index is combined to classify the execution strategy of the collection task, and the classification result is structure control type, content control type, or scheduling response type. According to the classification result and the preset cooperative strategy, the resource allocation order and the scheduling time window of the task are adjusted to realize the cooperative scheduling and strategy adaptation among multiple collection tasks.
[0012] In a preferred embodiment, the calculation of the structural uncertainty score includes the following steps:
[0013] The structural hierarchy information of the target webpage obtained in the last three consecutive collection periods is compared, and the structural node set in each period is extracted, each node in the structural node set including three indicators: tag name, level depth, and number of sibling nodes;
[0014] The structural node sets of any two periods are mapped one by one, and the number of label changes, the number of level migrations, and the number of added or missing nodes are compared. The inter-period structural variation quantity formed by any two periods is obtained by weighting and summing according to a preset proportion; the average value of all inter-period structural variation quantities is taken as the structural variation reference value, and the reference value is divided by the total number of nodes of the current webpage structure. The quotient multiplied by the preset floating weight coefficient corresponding to the current webpage type is the structural uncertainty score.
[0015] In a preferred embodiment, the generation process of the page structure consistency index includes the following steps:
[0016] The structural hierarchy information of the target webpage at five consecutive historical collection time points is extracted, and the webpage structure at each time point is converted into a node path set, each path being represented by a combination of tag sequence and level depth;
[0017] Taking the path set corresponding to the fifth time point as the reference benchmark, the alignment operation is sequentially performed on each path in the first four time points, and the path alignment is based on the longest common substring of the label sequence as the matching basis. If the length of the common substring exceeds 60% of the total length of the path, it is determined as successful alignment;
[0018] After successful alignment, the hierarchical distance difference between the path pairs is calculated, that is, the hierarchical difference value of the label end node. The average value of the hierarchical difference values of all path pairs is calculated to form the hierarchical difference value of the time point relative to the benchmark structure;
[0019] The concentration of the structure path of each time point is extracted, and the concentration is defined as the percentage of the number of paths in the same layer to the total number of paths. The hierarchical difference value of each time point is multiplied by its concentration, and the product value of the five time points is averaged to finally obtain the page structure consistency index.
[0020] In a preferred embodiment, the generation process of the content update index includes the following steps:
[0021] Extract the content snapshots of the target webpage at the consecutive five historical collection time points consistent with the structure hierarchical sampling. Divide each snapshot content into multiple content blocks according to the paragraph boundary, and perform difference analysis on the content blocks in the same position of adjacent snapshots in each time point;
[0022] For each pair of adjacent content blocks, the text editing span is calculated, which is defined as the ratio of the minimum editing distance between the two versions of the content block to the total number of words in the block in the current snapshot;
[0023] The keyword group change amplitude is calculated, which is defined as the ratio of the symmetric difference set size between the keyword sets extracted from the two versions of the content to the size of the union set of the two sets;
[0024] The publication time arrangement offset is calculated, which is defined as the difference between the standard deviation of the publication time of all paragraphs in the content block and the standard deviation of the average publication time of the paragraphs in the whole time point;
[0025] The above three values of each group of paragraphs, i.e. the text editing span, the keyword group change amplitude and the publication time arrangement offset, are normalized to the interval of zero to one, and the normalized result is multiplied by the display density of the paragraph in the webpage. The display density is the percentage of the paragraph in the total content length of the webpage. The product values of all paragraphs in each pair of snapshots are summed, and the average value of the sum between the five snapshots is finally taken to generate the content update index.
[0026] In a preferred embodiment, the collection scheduling prediction model for generating the collection scheduling coefficient is pre-trained by supervised learning, and the training samples include the page structure consistency index, the content update index, the actual collection success rate, the scheduling delay time, and the resource utilization rate in the historical collection task; during the training process, the model aims to minimize the comprehensive loss function of the scheduling error rate, the scheduling delay time, and the resource waste rate; finally, the model is constructed by using an integrated regression algorithm, and the output collection scheduling coefficient is a continuous value between zero and one, which is used to represent the comprehensive urgency and resource adaptation difficulty of the current collection task in the scheduling period.
[0027] In a preferred embodiment, the marker tracking and positioning process includes the following steps:
[0028] Obtain the stored structure label path set from the last successful collection version, each path set containing a sequence of label names and its corresponding hierarchical index; perform path scanning on the current web page structure tree, calculate the matching similarity of each stored path and the current structure path, and mark the path as an offset path if the similarity is lower than the preset ratio threshold; reconstruct the hierarchical pointer for each offset path and update it by replacing the old path with the new path.
[0029] In a preferred embodiment, the incremental comparison and snapshot update process includes the following steps:
[0030] Obtain the content area of the current web page, and perform paragraph division on the content area corresponding to the structure position in the last version web page, each paragraph being a content block, and the content block being indexed by double coordinates of paragraph start position and length;
[0031] Perform bidirectional mapping on the content block indexes in the two versions, and the matching success condition is that the start position of the current version paragraph is offset by no more than fifty characters and the length change is no more than twenty percent, if the condition is not met, it is determined as a failed block;
[0032] Extract key word groups from all failed content blocks, each key word group being selected from the title field and the first fifty words in the paragraph by a preset text analysis algorithm, and the key word group difference being defined as the ratio of the symmetric difference set number between the two version key word sets to the union set number; if the difference value is greater than zero point six, the content block is determined as changed content; all paragraphs determined as changed content are used as update paragraphs to replace the original paragraph content in the corresponding position of the snapshot, and the timestamp and key word group of this change are recorded.
[0033] In a preferred embodiment, after determining the collection path, the collection task is classified into execution strategies through a fuzzy logic classifier in combination with the combined relationship between the page structure consistency index and the content update index. The fuzzy logic classifier uses the page structure consistency index and the content update index as input variables, and generates classification labels based on preset membership functions and decision rules. The classification results are structure-dominated, content-dominated, or scheduling-responsive.
[0034] In a preferred embodiment, the multi-task collaborative incremental web page data collection system specifically includes:
[0035] The initial structural evaluation module is used to perform an initial evaluation of the target web page before the web page collection task is scheduled. The evaluation operation includes obtaining the structural level change range of the web page within a preset historical period, the reconstruction frequency of the visible area, and the rendering position fluctuation, calculating the structural uncertainty score, and judging whether to enter the in-depth evaluation process based on the score;
[0036] The deep perception module is used to perform a deep assessment when the structural uncertainty score exceeds a preset disorder threshold. The deep assessment includes analyzing the historical change trends of web page structure nodes, element rearrangement distribution, and regional stability patterns to generate a page structure consistency index. It also compares multiple historical content snapshots to calculate field update density, semantic variation amplitude, and time distribution patterns to generate a content update index.
[0037] The scheduling prediction module is used to input the page structure consistency index and content update index as input parameters into the pre-trained collection scheduling prediction model, and output the collection scheduling coefficient, which is used to evaluate the execution priority and policy tendency of the current task in the current scheduling cycle;
[0038] The path execution module is used to select the task acquisition path according to the numerical range of the acquisition scheduling coefficient. The acquisition path includes a marker tracking and positioning process based on structural reconstruction tolerance enhancement, or an incremental comparison and snapshot update process based on content difference recognition;
[0039] The strategy dispatch module is used to classify the execution strategy of the collection task after determining the collection path, combining the combined relationship between the page structure consistency index and the content update index. The classification results are structure-dominated, content-dominated or scheduling-responsive. According to the classification results and the preset collaboration strategy, the resource allocation order and scheduling time window of the task are adjusted.
[0040] The technical effects and advantages of the present invention are as follows:
[0041] The present invention performs an initial evaluation operation before scheduling a web page acquisition task, and calculates a structural uncertainty score based on the range of structural level changes, the frequency of visual area reconstruction, and the fluctuation of rendering position, so as to determine whether the web page structure is stable before acquisition. If the web page exhibits obvious structural fluctuation characteristics, the system will automatically enter a deep evaluation process to further extract historical structural change trends and element rearrangement distributions, and generate a page structure consistency index. At the same time, it combines historical snapshots to analyze the field density, semantic variation, and time patterns of the content to form a content update index. This method realizes the pre-awareness of the dynamic change status of the two dimensions of structure and content, so that the acquisition system can complete adaptation preparations before the web page changes significantly.
[0042] After obtaining the page structure consistency index and content update index, the present invention inputs the two into a pre-trained acquisition scheduling prediction model, outputs an acquisition scheduling coefficient, and dynamically selects an appropriate acquisition path based on the numerical range of the coefficient. The acquisition path includes two processing flows: one is a marker tracking and positioning flow for enhanced structural reconstruction tolerance, and the other is an incremental comparison and snapshot update flow based on content difference identification. This path selection mechanism can match the optimal execution mode based on the actual changing characteristics of the web page, effectively reducing acquisition failures or information loss caused by path incompatibility, thereby enhancing the system's adaptability to scenarios with heterogeneous web page structures and frequent content changes.
[0043] After determining the acquisition path, the present invention further classifies the execution strategy of each acquisition task based on the combined relationship between the page structure consistency index and the content update index. The results are classified as structure-dominated, content-dominated, or scheduling-responsive. Based on the classification results and the preset collaborative strategy, the system adjusts the resource allocation order and scheduling time window of the task, thereby achieving policy-differentiated scheduling when multiple acquisition tasks run in parallel. This mechanism not only avoids task resource competition and scheduling congestion, but also provides an orderly and reasonable execution rhythm based on task characteristics, improving the efficiency and stability of the overall scheduling system in handling multiple tasks in complex web page acquisition. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings;
[0045] Figure 1 This is a schematic diagram of the principle of the multi-task collaborative incremental web page data collection method in the present invention.
[0046] Figure 2 This is a schematic diagram of the multi-task collaborative incremental web page data acquisition system in the present invention. DETAILED DESCRIPTION
[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0048] Reference Figure 1 - Figure 2 The following examples were obtained:
[0049] Example 1: A multi-task collaborative incremental web page data collection method includes the following steps:
[0050] Before scheduling a web page collection task, an initial assessment is performed on the target web page. This assessment involves obtaining the extent of the web page's structural level changes within a preset historical period, the frequency of visual area reconstruction, and fluctuations in rendering position. A structural uncertainty score is then calculated, and based on this score, a decision is made whether to proceed to a deeper assessment process. This step is crucial because, before system resources are allocated to the collection task, preliminary information from the structural and visual levels can be collected to determine whether the web page exhibits structural fluctuations or unstable layout behavior. By analyzing structural level statistics, visual area change frequency, and rendering offset, the system can obtain a comprehensive picture of the web page's stability. The calculated structural uncertainty score serves as a quantitative indicator, determining whether further analysis and processing are necessary, thereby enabling risk assessment and screening before the collection task begins.
[0051] When the structural uncertainty score exceeds the preset disorder threshold, a deep assessment is performed. This includes analyzing the historical change trends of webpage structural nodes, element rearrangement distribution, and regional stability patterns to generate a page structural consistency index. Simultaneously, multiple historical content snapshots are compared to calculate field update density, semantic variation amplitude, and temporal distribution patterns to generate a content update index. This step is crucial for triggering a more precise analysis process for webpages whose initial assessment results indicate structural anomalies or frequent adjustments. Historical tracking and rearrangement analysis of structural nodes help capture deep-level change patterns and establish a structural consistency assessment benchmark. Comparisons between content snapshots reveal the trajectory of content evolution. These two indices, reflecting the dynamic behavior of the structural and content dimensions, respectively, serve as core input variables for subsequent decision-making and reasoning.
[0052] The page structure consistency index and content update index are fed as input parameters into a pre-trained collection scheduling prediction model, which outputs a collection scheduling coefficient. This coefficient is used to assess the execution priority and strategic orientation of the current task within the current scheduling cycle. This step is significant because, by introducing a learning-based scheduling mechanism, the structure and content change indicators are mapped into a one-dimensional scheduling coefficient, thereby uniformly assessing the urgency and complexity of the current collection task. The scheduling prediction model is trained based on historical collection performance and can predict whether a task deserves immediate execution, whether it requires high-priority resources, and which collection path should be prioritized based on different index combinations, demonstrating the system's intelligent scheduling decision-making capabilities.
[0053] Based on the numerical range of the acquisition scheduling coefficient, the task's acquisition path is selected. This includes either a marker tracking and positioning process based on enhanced structural reconstruction tolerance or an incremental comparison and snapshot update process based on content difference identification. This step is crucial because the system selects specific execution path strategies based on the scheduling coefficients, achieving adaptive matching of structure and content. If the webpage structure fluctuates significantly, marker tracking is preferred to enhance position recovery capabilities. If the content changes frequently, a content comparison process is used to quickly identify incremental data. Path decisions directly impact the success rate and accuracy of acquisition operations and are a critical transition from scheduling to execution.
[0054] After determining the acquisition path, the execution strategy of the acquisition task is classified based on the combined relationship between the page structure consistency index and the content update index. The classification results are structure-dominated, content-dominated, or scheduling-responsive. According to the classification results and the preset collaborative strategy, the resource allocation order and scheduling time window of the task are adjusted to achieve collaborative scheduling and strategy adaptation between multiple tasks. The significance of this step is that after the acquisition path is confirmed, the system intelligently classifies the nature of the task according to the relative relationship between the two indices, so as to perform hierarchical management in the concurrent scheduling of multiple tasks. The three types of strategies are mapped to different resource allocation priorities and time window settings, so that the system can adaptively handle high-frequency content update tasks, structurally variable web page tasks, and ordinary tasks, ensuring the parallel efficiency and strategy consistency of the overall acquisition system.
[0055] To achieve collaborative scheduling and strategy adaptation among multiple tasks, the system implements the following two-level scheduling control process in each scheduling cycle based on the execution strategy classification results of all acquisition tasks and the corresponding page structure consistency index and content update index: The preset collaborative strategy refers to a set of scheduling rules that the system pre-configures to adjust resource allocation priorities and timing peak-shifting strategies based on the characteristic differences between different types of acquisition tasks. Collaborative strategies are divided into three categories according to the task classification results:
[0056] Structure-master task strategy: The default scheduling priority is high. Due to frequent structural changes and high mark reconstruction costs, the system prioritizes allocating execution threads with strong structural tolerance. If there are too many structure-master tasks in the same batch, the system limits their concentrated execution by staggering the peak hours to avoid system concurrency bottlenecks caused by structural mark reconstruction operations.
[0057] Content-driven task strategy: Priority is automatically adjusted based on the content update index. If the index is higher than the preset high-change threshold, the task will be scheduled earlier. Similar paragraph snapshots between content-driven tasks can be merged for processing, and the system will aggregate similar tasks into parallel batch processing groups to improve data processing efficiency.
[0058] Scheduling responsive task strategy: This type of task does not trigger emergency collection due to significant structural or content features; the system regards it as a filler task in the scheduling idle period; its execution order is based on the dynamic priority sorting of the collection scheduling coefficient in the low range, and it dynamically selects the opportunity to enter the execution window.
[0059] Resource allocation order adjustment mechanism: Before entering the scheduling engine, each collection task is inserted and sorted in the resource allocation table according to its classification label (structure master, content master or scheduling response type) and the value of the collection scheduling coefficient. The specific logic is as follows: first, group according to the classification label and queue in the order of structure master > content master > scheduling response; within each group, sort from high to low according to the collection scheduling coefficient. The higher the collection scheduling coefficient, the higher the priority of the task; for resource contention conflicts, the system sets a thread limit threshold. For example, the structure master class can process no more than 40% of the total number of threads at the same time, and the rest is occupied by the content master type and scheduling response type.
[0060] Scheduling time window adjustment mechanism: The system divides each scheduling cycle into multiple time windows (such as minute-level and second-level granularity). The scheduling time window of the task is jointly determined by the following parameters: Classification label: Structure-controlled types are given priority in early windows; Index combination: Tasks with low structural consistency index and high content update index are assigned earlier; Coefficient density control: The system evaluates the distribution density of scheduling coefficients in the same window to avoid scheduling congestion; Cluster merging: For tasks with high similarity in structural paths or content snapshots, the scheduling windows are merged and processed in parallel in batches.
[0061] The implementation process for multi-task collaboration: After initial and in-depth evaluations, all tasks receive a structural consistency index and a content update index. These two indices are input into the scheduling prediction model to generate the acquisition scheduling coefficient. A fuzzy logic classifier strategically categorizes tasks, labeling them as one of three categories. The task classification result + index value + coefficient are then input into the preset collaborative strategy engine. The collaborative strategy engine outputs the task's resource allocation priority and scheduling time window. The system initiates tasks based on this schedule, dynamically monitoring and adjusting strategy adaptation. For example, suppose there are 10 acquisition tasks in the current scheduling cycle. The system assessment finds that: Task A is structure-dominated, with an acquisition scheduling coefficient of 0.89; Tasks B and C are content-dominated, with coefficients of 0.75 and 0.68, respectively; Task DJ is scheduling-responsive, with coefficients ranging from 0.32 to 0.56. The system schedules A for execution in the first window, assigning it a high-performance structure tagging thread. Tasks B and C are aggregated into the second window, performing snapshot difference comparisons. DJ is sorted and inserted into subsequent idle windows in batches. If resources are scarce, Task J is postponed to the next cycle.
[0062] In the present invention, the initial evaluation is the "pre-judgment mechanism" for data collection decisions. The specific logic is as follows: the purpose of the initial evaluation is to determine whether the current web page has sufficiently significant structural fluctuation behavior; if the structural uncertainty score is lower than the "preset disorder threshold", it means that the recent structural changes of the web page are not significant, and the system determines that no deep evaluation is needed; because the incremental collection strategy of the present invention completely relies on the "page structure consistency index" and "content update index" produced by the deep evaluation as input, therefore: if the deep evaluation is not entered → the two indexes will not be generated; the prediction model cannot evaluate the collection scheduling coefficient; the system will not execute the incremental collection task of the web page in the current scheduling cycle to save resources and avoid redundant collection. The present invention is clearly applicable to the scheduling and strategy collaboration scenarios of parallel collection tasks of multiple web pages. Multiple web page collection tasks are scheduled simultaneously in the current scheduling cycle; each task undergoes individual evaluation → model prediction → classification decision; finally, the system constructs a task index combination matrix to perform operations such as resource allocation, parallel window staggering, and strategy clustering in the preset collaboration strategy.
[0063] The calculation of the structural uncertainty score includes the following steps: comparing the structural hierarchy information obtained from the target web page in the last three consecutive acquisition cycles, extracting the set of structural nodes in each cycle, and each node in the set of structural nodes contains three indicators: tag name, hierarchy depth and number of sibling nodes; in the present invention, the three behavioral characteristics in the initial evaluation operation are "structural hierarchy change range", "visual area reconstruction frequency" and "rendering position fluctuation", and their judgment and calculation are based on the analysis of the web page structural node set, and the structural node set is constructed from the structural hierarchy information extracted from the target web page in the last three consecutive acquisition cycles. Each structural node contains three core indicators: tag name, hierarchy depth and number of sibling nodes. Among them, "structural hierarchy change range" is mainly obtained by vertically comparing the hierarchy depth of the node. The hierarchy depth reflects the nesting level of the node in the entire DOM structure tree. When the hierarchy depth of the node at the same position in consecutive cycles changes significantly, or the deep structure is reconstructed as a whole, it means that the structural hierarchy has fluctuated, thereby reflecting the instability of the structural hierarchy range.
[0064] The frequency of visual area reconstruction can be determined by the increase or decrease in the number of sibling nodes between consecutive cycles. Frequent additions and deletions of sibling nodes under the same parent node indicate frequent local structural updates in that area, particularly those occurring within the visible area. This may lead to content reflow or re-rendering. This type of structural fluctuation indirectly reflects frequent changes in the composition of the visible area of the webpage.
[0065] "Rendering position fluctuation" is related to the change of tag name. Changes in tag name often lead to changes in the corresponding element's appearance, CSS style binding or behavior attributes on the web page, such as<di v> becomes<secti on> 、 becomes etc., which can cause the display position, layering order and visual logic of elements to be adjusted. By the number of changes in the tag name and the distribution of positions in the continuous period, it can be inferred whether the rendering position is stable. Although the three indicators in the structural node set are not directly equivalent to the three behavior characteristics described above, they provide underlying data support for quantifying and analyzing the behavior of changes in the structure of the webpage. By comparing and analyzing these indicators between different collection periods and combining the preset weighted calculation rules, the structural level fluctuation trend, the reconstruction frequency change and the stability of the rendering structure can be effectively derived, thereby generating a structure uncertainty score.
[0066] The structural node set refers to the collection of all identifiable markup elements that make up the Document Object Model of the webpage. Each structural node corresponds to an HTML tag, such as , <section>, etc. Tag name refers to the specific type of the node in the HTML structure, which is used to identify the function of the structure; hierarchical depth refers to the number of nesting layers of the node from the root node, and the larger the value, the deeper the nesting; the number of sibling nodes refers to the number of nodes that exist side by side with the node in the same level and under the same parent node, reflecting the horizontal density of the structure. In this step, the system respectively captures the complete structure information of the target webpage in three adjacent collection cycles, and constructs three sets of structure node sets, each set maintaining the complete node attributes inside. For example, for a news list page, the first cycle contains 20 Nodes, two new in the second cycle due to content updates node and remove an ad The node set changes in label category, hierarchy, and quantity, reflecting structural change characteristics.
[0067] The structural node sets of any two periods are mapped one by one, and the number of label changes, the number of hierarchy migrations, and the number of newly added or missing nodes are compared. After weighted summation according to a preset ratio, the structural change amount between any two periods is obtained. The one-by-one mapping of structural nodes means that for nodes with the same label name or similar structural position in the two periods before and after, a corresponding relationship is established, and the attribute changes thereof are compared. The number of label changes refers to the number of changes in the label name of the node at the same position. The number of hierarchy migrations represents the number of times the depth of the same label node in the hierarchical structure changes. The number of newly added or missing nodes represents the total number of new nodes appearing in the next period or the total number of old nodes disappearing. The system sets a weighting coefficient for each of the above three changes, for example, the label change weight is 0.5, the hierarchy migration is 0.3, and the newly added or missing is 0.2. The weighted values of the three are summed to obtain the structural change amount between the two periods. The two-by-two combination of the three periods will produce three inter-period change values. Continuing with the above example, if there are 5 label changes, 3 hierarchy migrations, and 2 nodes added from the first to the second period, then the inter-period structural change amount = 5x0.5+3x0.3+2x0.2=2.5+0.9+0.4=3.8. The average value of all inter-period structural change amounts is taken as the structural variation reference value, and the reference value is divided by the total number of current webpage structure nodes. The quotient multiplied by the preset floating weight coefficient corresponding to the current webpage type is the structural uncertainty score.
[0068] The structural variation reference value is the arithmetic mean of the structural change amounts of the three periods, reflecting the overall trend of structural changes. The total number of current webpage structure nodes refers to the total number of structural nodes extracted in the latest collection period, which is used to standardize the variation value to the scale of the structure size. The preset floating weight coefficient is a weighting factor set in advance for different webpage types (such as news, e-commerce, and portal), which is used to improve the adaptability of the score in different application scenarios. For example: the floating weight of a news page is 1.0, the floating weight of an e-commerce page is 0.8, and the floating weight of a portal information page with relatively stable content is 0.5. Finally, the ratio of the structural variation reference value to the total number of structures is taken as the change proportion, multiplied by the corresponding floating weight, and the structural uncertainty score is obtained. The score is a floating-point number between 0 and 1, and the higher the value, the stronger the uncertainty of the webpage structure in the short term. If the structural change amounts of the above three times are 3.8, 2.6, and 4.1 respectively, the average value is 3.5, the total number of current webpage nodes is 120, and the webpage type is news, the floating weight is 1.0, then the structural uncertainty score = (3.5 / 120) x 1.0 = 0.029.
[0069] The generation process of the page structure consistency index includes the following steps: extracting the structure level information of the target webpage at five consecutive historical collection time points, converting the webpage structure at each time point into a node path set, and each path is represented by a combination of label sequence and its level depth; the structure level information refers to the node nesting situation of the webpage in the DOM (Document Object Model) structure, and each path records from the root node to the last level structure unit. The node path set represents the sum of all paths in the webpage, and each path is a chain composed of a sequence of label names, such as / / , the label sequence is the label arrangement order in the path, and the hierarchical depth is the nesting number of the terminal node of the path in the DOM tree (the root node is the first layer). Each time point corresponds to a complete web page collection, and the system extracts its structure and converts it into the above path set form to form a unified comparable data unit.
[0070] Taking the path set corresponding to the fifth time point as the reference benchmark, the alignment operation is sequentially performed on each path in the previous four time points. The path alignment is based on the longest common substring of the label sequence, and if the length of the common substring exceeds 60% of the total length of the path, it is considered to be successfully aligned. The reference benchmark refers to selecting the latest web page structure among the five historical collection time points as the structure alignment target. Path alignment is to compare any two paths at different time points to determine whether they represent the same or similar structure unit. The longest common substring of the label sequence refers to the longest same label segment in the two label sequences in sequence. The similarity is obtained by dividing the length of the longest common substring by the total path length. If the proportion exceeds 0.6 (i.e. 60%), it is considered that the two paths remain stable in the structure sense, and it is considered to be successfully aligned. For example, a path A is / <section> / / , path B is / / , the longest common substring between the two is / / / / , its length is 5, the total length of path A is 6, and the similarity is 5 / 6≈0.83, which meets the matching criteria.
[0071] After successful alignment, the hierarchical distance difference between the path pairs is calculated, that is, the hierarchical difference of the label end nodes, and the hierarchical difference values of all path pairs are averaged to form the hierarchical difference value relative to the reference structure at that time point; a path pair refers to a combination of path pairs that are successfully matched in the alignment operation. The hierarchical distance difference refers to the absolute difference between the levels of the terminal nodes (i.e., the path end nodes) in the two aligned paths. For example, if the terminal node of one path is at the 6th level and the other is at the 5th level, the hierarchical distance difference is |6-5|=1. The system calculates the difference for all successfully aligned path pairs at each time point and the fifth time point, and takes the average to generate the hierarchical difference value at the current time point. The smaller the value, the higher the temporal stability of the path hierarchy. The concentration of the structural path at each time point is extracted. The concentration is defined as the percentage of the number of paths at the same level divided by the total number of paths. The hierarchical difference value at each time point is multiplied by its concentration, and the product value of the five time points is averaged to finally obtain the page structure consistency index.
[0072] The concentration of a structural path refers to the proportion of the number of paths that fall on the same level (such as the third level, fourth level, etc.) to the total number of all paths on a web page at a certain point in time. The concentration reflects whether the structure of a web page exhibits hierarchical aggregation characteristics, for example, the content is densely concentrated in the third to fifth levels, while the other levels are relatively sparse. The system uses the concentration at each time point as a weighting factor for the stability of the structural hierarchy at that time point, and multiplies it with the corresponding hierarchical difference value to obtain the weighted hierarchical deviation. The weighted hierarchical deviations of the five time points are then averaged to obtain the page structure consistency index. For example, assuming that the hierarchical difference values corresponding to the five time points are 1.0, 1.2, 0.9, 1.3, and 0.8, respectively, and the corresponding concentrations are 0.7, 0.6, 0.8, 0.5, and 0.9, the weighted values are 0.7, 0.72, 0.72, 0.65, and 0.72, respectively, with an average value of 0.702. The smaller the value, the more stable the structural path in the time series and the more consistent the page structure.
[0073] The process of generating the content update index includes the following steps: extracting content snapshots of the target web page at five consecutive historical collection time points that are consistent with the structural level sampling, dividing the content of each snapshot into multiple content blocks according to paragraph boundaries, and performing difference analysis on the content blocks at the same position of adjacent snapshots at each time point; a content snapshot refers to a static copy of the content of the web page body area extracted at a specific collection time point, maintaining the original text order, paragraph structure and positioning identifiers. Each snapshot is divided into content blocks according to paragraph boundaries, and paragraph boundaries are the logical paragraph separation tags (such as 、 、 <section>) or a line break node (such as ). A content block is each paragraph unit in the segmentation result, and has text continuity and page position stability. The system matches blocks in the same content area in snapshots at two adjacent time points, using the position offset of no more than 20 characters and consistent paragraph numbers as the judgment criteria. After a successful match, a difference analysis operation is performed on the block pair.
[0074] For each pair of adjacent content blocks, the text edit span is calculated. This is defined as the ratio of the minimum edit distance between the two versions of the content block to the total number of words in the block in the current snapshot. The text edit span measures the extent of textual changes to the content at the same location in two snapshots. The minimum edit distance refers to the minimum number of character operations required to transform one version of content into another, including insertions, deletions, and substitutions. The total number of words in the block in the current snapshot is the total number of characters in the corresponding paragraph in the snapshot at the current time point. The ratio of the two constitutes the edit span metric, which represents the average degree of change per character. After normalization, it facilitates uniform comparison between paragraphs of different lengths. For example, a paragraph whose previous version reads "The system supports web scraping" and whose current version reads "The system can perform incremental web scraping" has a minimum edit distance of 5 and a total word count of 15, resulting in an edit span of 5 / 15 = 0.33.
[0075] The magnitude of keyword group change is calculated as the ratio of the symmetric difference between the keyword sets extracted from the two versions of the content to the union of the two sets. A keyword set refers to important nouns, verb phrases, or entity phrases extracted from a paragraph using algorithms such as word frequency statistics, part-of-speech tagging, and TF-IDF. It is typically used to describe the semantic core of the paragraph. After extracting keyword sets from the two versions of the content, their symmetric difference is calculated, which is the sum of the unique words in each set. The union size represents the total number of unique words in the two sets. The ratio between the two reflects the proportion of word replacement or disappearance in the content semantics and is a measure of paragraph semantic change. For example, if set A = {collection, webpage, increment} and set B = {collection, page, update}, then the symmetric difference = {webpage, increment, page, update}, with a size of 4; the union = {collection, webpage, increment, page, update}, with a size of 5, then the magnitude of keyword group change is 4 / 5 = 0.8. Calculate the release time arrangement offset, defined as the difference between the standard deviation of the release times of all paragraphs in a content block and the standard deviation of the average release time of all paragraphs at that time point. Release time refers to the explicit timestamp carried by a webpage paragraph, such as the time of news, comment release time, update mark time, etc. A set of time data associated with each paragraph can be extracted. The standard deviation of this group measures whether the time distribution within the paragraph is concentrated. Compared with the average standard deviation of the release times of all paragraphs at that time point, if the difference is large, it means that the time distribution of the paragraph is significantly abnormal, which may be a hot update area or a backend batch replacement area. This item quantifies the degree of abnormality in the time characteristics of the content block and helps to identify implicit change trends. For example, if the average standard deviation at a certain time point is 2 hours and the standard deviation within a paragraph is 5 hours, the offset is 5-2=3 hours, which is normalized and used in the update index calculation.
[0076] The three values for each paragraph group—text edit span, keyword change amplitude, and publication time offset—are normalized to a range between zero and one. The normalized result is then multiplied by the paragraph's display density on the webpage, which is the percentage of the paragraph's total length. The product of these values for all paragraphs in each snapshot pair is summed, and the sums across the five snapshots are averaged to generate the content update index. Normalization uses the min-max normalization method to uniformly map the raw values to the range [0,1]. Display density represents the relative weight of the paragraph on the webpage and is calculated by dividing the number of words in the paragraph by the total number of words on the webpage. For example, if a paragraph has 300 words and the total page has 3,000 words, the display density is 0.1. The normalized result of the three metrics is multiplied by the density to create a weighted index, emphasizing that changes in the core area of the webpage have a greater impact than changes in the peripheral areas. For each snapshot pair, the weighted change values of all paragraphs are summed to form the change score for that snapshot pair. The four change scores across the five snapshots are averaged to form the content update index. The index ranges from [0,1]. The closer it is to 1, the more frequently the webpage content is updated and the greater the change.
[0077] The collection scheduling prediction model for generating the collection scheduling coefficient is pre-trained by a supervised learning method. The supervised learning method refers to a machine learning method for optimizing model parameters based on known input and output sample pairs, and requires providing explicit labels as training targets. In the technical solution, the training samples are derived from historical collection tasks, and the collection tasks are used as basic sample units, each sample containing a group of input features and an output result. The training samples include the following five input feature variables: page structure consistency index: used to quantify the stability of the page structure in the time series, derived from the structure level path trend analysis; content update index: used to measure the update frequency and variation degree of the text area in the historical snapshot; actual collection success rate: refers to the ratio of the number of times of successfully extracting target data to the total number of execution times of the page of this type in the historical task within the execution period, reflecting the task accessibility; scheduling delay time: refers to the average time interval between the planning and execution of the collection task, in seconds, used to describe the degree of task delay by the system; resource usage rate: refers to the proportion of the consumed computing resources (such as CPU core number, memory occupation, etc.) of the task to the system allowed upper limit, reflecting the resource pressure.
[0078] During training, the model objective is to minimize a combined loss function consisting of scheduling error rate, scheduling delay, and resource waste rate. The following are definitions and calculations for these three loss dimensions: The scheduling error rate refers to the rate of scheduling failures caused by the discrepancy between the model's output of the acquisition scheduling coefficient in historical tasks and the actual execution results. It is calculated as: Scheduling error rate = number of tasks that exceeded the scheduling coefficient threshold but failed to complete successfully / total number of acquisition tasks. The system sets certain thresholds based on the output acquisition scheduling coefficient (for example, a value above 0.8 recommends immediate scheduling). If a task's acquisition scheduling coefficient exceeds the threshold but fails to complete successfully (for example, due to rapid web page changes or resource conflicts), it is counted as a scheduling failure. By counting these "scheduling prediction deviations," the scheduling error rate is calculated. Scheduling delay: The difference between the scheduled task completion time and the scheduling trigger time is averaged across all training samples. A larger value indicates a system bottleneck or insufficient prediction accuracy, leading to inefficient resource allocation. Resource waste rate: This indicates the rate of resource waste resulting from the model's prediction of resource usage for a task but the task fails to fully utilize the allocated resources. The calculation method is: Resource waste rate = (total allocated resources - actual consumed resources) / total allocated resources. The comprehensive loss function is a weighted combination of the above three indicators, constructed as an objective function for model optimization. Assuming the three loss terms are L1 (error rate), L2 (delay time), and L3 (waste rate), the comprehensive loss function can be expressed as: Loss = α*L1+β*L2+γ*L3, where α, β, and γ are non-zero weight coefficients set by the user or system based on scheduling sensitivity. For example, for a highly real-time acquisition system, β>α>γ may be set.
[0079] The final model is constructed using an integrated regression algorithm, and the output collection scheduling coefficient is a continuous value between zero and one. The integrated regression algorithm refers to a prediction model that obtains higher accuracy and stability by integrating multiple basic prediction models (such as decision tree regressors, linear regressors, gradient boosting regressors, etc.). Common methods include random forest regression and XGBoost. The output of this model is the collection scheduling coefficient, which is a continuous value in the range of [0,1] after normalization. This coefficient is used to represent the comprehensive urgency and resource adaptation difficulty of the current collection task within the scheduling cycle. The closer the value is to 1, the more urgent the task, the more drastic the change in structure or content, and the more necessary it is for the system to prioritize resource allocation for processing. For example, the characteristics of a historical task sample are as follows: Page Structure Consistency Index: 0.85, Content Update Index: 0.76, Actual Collection Success Rate: 0.92, Scheduling Delay: 120 seconds, and Resource Utilization: 0.55. This five-dimensional vector is input into the model training process as an input feature. The model uses its actual performance to update the comprehensive loss function, allowing the final model to predict a continuous value output (e.g., 0.89) for new task inputs. This collection scheduling coefficient itself serves as the urgency score used by the collection scheduling system. A higher value indicates that the webpage structure and content are highly volatile, the historical success rate is high, and current delays and resource utilization are within reasonable ranges, thus ensuring the necessity of priority scheduling.
[0080] After the acquisition scheduling prediction model outputs the acquisition scheduling coefficient, the system selects the acquisition path for the task based on the numerical range of the coefficient. Based on the preset path switching threshold, the acquisition paths are divided into the following two types: (1) Tag tracking and positioning process (when the scheduling coefficient is ≥0.6). When the acquisition scheduling coefficient is greater than or equal to the structural path activation threshold (for example, 0.6), the system selects the tag tracking and positioning process based on structural reconstruction tolerance enhancement as the acquisition path for the task. This process calculates the similarity of the tag sequence in the structural path, the path offset and the level tolerance, rebuilds the tag index and completes the data positioning. It is suitable for web pages with high structural change frequency and fast structural evolution speed. (2) Incremental comparison and snapshot update process (when the scheduling coefficient is <0.6). When the acquisition scheduling coefficient is less than the structural path activation threshold, the system selects the incremental comparison and snapshot update process based on content difference recognition as the acquisition path. This process focuses on the rapid analysis of the text editing span, keyword group differences and release time arrangement offset between content blocks, accurately extracts the changed parts and performs snapshot local replacement. It is suitable for web pages with relatively stable structure and frequently updated content.
[0081] The tag tracking and location process is used to re-identify and adjust the structural path of the target data on a page when the webpage structure changes frequently, thereby achieving accurate tracking of data tags. This process includes the following steps: First, a set of stored structural tag paths is obtained from the last successfully collected version. A structural tag path set refers to the set of paths in the webpage structure tree that point to the target data location. Each path consists of a sequence of tag names and their corresponding hierarchical index in the structure tree. The tag name is the element name in the HTML structure, such as div or span, and the hierarchical index is its depth position in the DOM tree, such as the number of levels from the root node. Subsequently, a path scanning operation is performed on the current webpage structure tree. The structure tree is the hierarchical representation of the Document Object Model (DOM) of the current webpage. Starting from the root node, the system traverses the nodes at each level, extracting all identifiable paths in the current page and matching them one by one with the stored paths. For each stored path, the system calculates the matching similarity between it and the current structural path. This similarity is calculated based on factors such as the consistency of the tag sequence, the degree of hierarchical alignment, and the change in path length. If the similarity value of a path is lower than the preset ratio threshold (for example, 0.7), the path is identified as an offset path. An offset path indicates that the page structure has changed, causing the element position pointed to by the original path to become invalid or drift. For each offset path, the system reconstructs the hierarchical pointer based on the structural node of the current web page. The pointer is a new path from the root node to the current matching position, which is used to replace the old structural path. After completing the path reconstruction, the system replaces the old path with the new path for update, thereby ensuring the accuracy and stability of data labeling in subsequent collection tasks.
[0082] The incremental comparison and snapshot update process is used to accurately update and partially replace snapshots by detecting content differences when webpage content changes frequently but the structure is relatively stable. This process involves the following steps: First, the system extracts the content area from the current webpage. The content area refers to the main text area associated with the target data, typically the central body, excluding structural elements such as navigation bars and footers. The system also extracts the corresponding content area from the previous version of the webpage at the same structural location. Each content area is divided into multiple content blocks based on paragraph boundaries. Paragraph boundaries can be determined by line breaks in the HTML structure or by text density. Each content block is treated as the minimum comparison unit and indexed by its paragraph start position and length on the page to facilitate subsequent mapping and comparison. Subsequently, a bidirectional mapping operation is performed on the content block indexes in the two versions. A successful match is determined when the start position of the paragraph in the current version deviates by no more than 50 characters from the start position of the corresponding paragraph in the previous version, and the length change does not exceed 20%. Content blocks that meet these conditions are considered "matched blocks"; those that do not meet these conditions are considered "unmatched blocks." For each content block that fails to match, the system extracts the keyword group corresponding to the block.
[0083] Keyword groups refer to a set of representative core vocabulary extracted from web page content blocks, which are used to measure the degree of semantic deviation in content changes. In the "Incremental Comparison and Snapshot Update Process", for content blocks that fail to match, the system extracts keyword groups through a preset text analysis algorithm. This algorithm is a keyword extraction algorithm based on term frequency-inverse document frequency (TF-IDF), and is optimized in combination with a part-of-speech screening mechanism. The specific steps are as follows: The keyword extraction of each content block comes from two semantic positions: The title field to which the paragraph belongs: refers to the title tag content to which the content block belongs in the web page structure. The title tag includes the HTML <h1>、< / h1> <h2>、< / h2> <h3>The system uses the nearest title field above the content block as its semantic label. The first fifty words of a paragraph refer to the first fifty consecutive words at the beginning of the main body of the current content block. This section is considered the most representative and the semantic core with high information density. The text content from these two sources is combined to form the original candidate corpus for keyword extraction. A TF-IDF score is calculated for each term in the candidate corpus. The TF-IDF score is then calculated for each term and sorted in descending order of score. After sorting, the system selects the final keyword group from the top 20% of terms. This screening process incorporates a part-of-speech (POS) judgment mechanism to retain only terms belonging to the following part-of-speech categories: nouns (e.g., subject terms, object names, organization names); verbs (e.g., key actions, operating methods); and proper nouns (e.g., brand, product, and technical terms). The system removes stop words, punctuation, numbers, and non-semantically loaded words (e.g., "的," "是," and "了"), and merges duplicate or inflected words. The final set of keywords is used as the "keyword group" of the content block, which is used for semantic variation comparison between subsequent versions. The difference in keyword groups is defined as the ratio of the number of symmetric difference sets of the keyword sets of two versions to the number of unions of the two sets. The symmetric difference set is the number of non-overlapping parts in the two sets, reflecting the degree of semantic change. If the difference value is greater than 0.6, the system marks the content block as changed content. All paragraphs that are judged to be changed content will be used as updated paragraphs to replace the original paragraph content at the corresponding position in the snapshot. At the same time, the system records the timestamp of this change (i.e. the time when the current version was collected) and the corresponding keyword group for subsequent snapshot version management and change trend analysis. For example, a web page content block is a product introduction paragraph, and its title is "H1: Product Function Description". The first fifty words in the text include "intelligent recognition", "image enhancement", "edge detection", "automatic adjustment" and other words. After TF-IDF calculation and part-of-speech filtering, the system selects terms with high TF-IDF values and valid parts of speech as keyword groups: → The keyword group is {"intelligent recognition", "image enhancement", "edge detection", "automatic adjustment"}. If the keyword group between the two versions changes significantly, that is, the symmetric difference ratio is greater than 0.6, the system determines that the block has a semantic change.
[0084] After determining the collection path, the system uses a fuzzy logic classifier to classify and judge the execution strategy of the current collection task. In this solution, the fuzzy logic classifier uses the page structure consistency index and content update index as input variables, generates classification labels through preset membership functions and decision rules, and realizes the classification of the execution strategy of the collection task. Page structure consistency index: a value generated after in-depth evaluation, which indicates the stability of the target web page structure in the historical period. The higher the value, the smoother the structural evolution; content update index: a value generated based on historical content snapshots, which reflects the update frequency and change amplitude of the target web page content in the time series dimension. The higher the value, the faster the content changes. Both are continuous real-valued inputs between [0,1].
[0085] The system defines multiple membership functions for each of the two input variables, representing the fuzzy sets to which they semantically belong. For example, for the "Page Structure Consistency Index," three membership functions are defined: low consistency (Low), medium consistency (Medium), and high consistency (High); for the "Content Update Index," three membership functions are defined: low update frequency, moderate update frequency, and high update frequency. Each input value can belong to multiple membership functions simultaneously, depending on its position in the numerical range. The system uses triangular or trapezoidal membership functions for modeling, meeting the principles of continuity and overlap. The system establishes a fuzzy rule base based on task characteristics. For example, if the structure consistency index is "High" and the content update index is "Low," the task is classified as "Structure-Dominated"; if the structure consistency index is "Low" and the content update index is "High," the task is classified as "Content-Dominated"; if both the structure consistency index and the content update index are "Medium," or if one is high and the other is low, the task is classified as "Scheduling Response." Each rule in the rule base is composed of "if (IF) - then (THEN)" form, which is used to describe the correspondence between the input fuzzy set and the output classification label.
[0086] The fuzzy logic classifier performs fuzzy reasoning calculation according to the membership value of all activated rules, and adopts the maximum membership principle as the output mechanism, that is, the classification label with the highest membership is selected as the final strategy category of the task. The classification result includes: structure master type, indicating that the current task acquisition is significantly affected by structural changes, and a structural reconstruction tolerant marker tracking strategy should be used; content master type, indicating that the task acquisition mainly relies on content variation detection, and an incremental comparison and snapshot update process is suitable; scheduling response type, indicating that the task is not affected by obvious structural or content dominant factors, and the execution priority and time window should be adjusted according to the acquisition scheduling coefficient. Assuming that a web page has a page structure consistency index of 0.82 and a content update index of 0.18 after deep evaluation. According to the membership function, the structure consistency value has a "high consistency" membership of 0.9 and a "medium consistency" of 0.1; the content update value has a membership of 0.85 under "low update frequency" and a membership of 0.15 under "medium update frequency". The system matches the fuzzy rule "if the structure consistency is high and the content update is low, then the task is structure master type", which is activated to the highest degree, and finally outputs the classification as structure master type. The system maps this task to the resource scheduling and acquisition path execution plan matching the structure dominant strategy.
[0087] Embodiment 2: A multi-task cooperative incremental web page data acquisition system, specifically comprising:
[0088] A structure preliminary evaluation module is configured to perform a preliminary evaluation operation on the target web page before scheduling the web page acquisition task, and the evaluation operation includes obtaining the structural level change range, the reconstruction frequency of the visible area, and the rendering position fluctuation of the web page in a preset historical period, calculating a structure uncertainty score, and determining whether to enter a deep evaluation process according to the score;
[0089] A deep perception module is configured to perform a deep evaluation operation when the structure uncertainty score exceeds a preset disorder threshold, and the deep evaluation includes analyzing the historical variation trend of the web page structure node, the element rearrangement distribution, and the area stability mode, generating a page structure consistency index; and comparing multiple historical content snapshots to calculate the field update density, the semantic variation amplitude, and the time distribution rule, and generate a content update index;
[0090] A scheduling prediction module is configured to input the page structure consistency index and the content update index into a pre-trained acquisition scheduling prediction model as input parameters, and output an acquisition scheduling coefficient, which is used to evaluate the execution priority and strategy tendency of the current task in the current scheduling period;
[0091] A path execution module is configured to select the acquisition path of the task according to the numerical interval of the acquisition scheduling coefficient, and the acquisition path includes a marker tracking and positioning process based on structural reconstruction tolerance enhancement, or an incremental comparison and snapshot update process based on content difference identification.
[0092] The strategy dispatch module is used to classify the execution strategy of the collection task after determining the collection path, combining the combined relationship between the page structure consistency index and the content update index. The classification results are structure-dominated, content-dominated or scheduling-responsive. According to the classification results and the preset collaboration strategy, the resource allocation order and scheduling time window of the task are adjusted.
[0093] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.
[0094] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0095] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0096] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0097] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.< / h3> < / section> < / section>
Claims
1. A multi-task collaborative incremental web page data collection method, characterized in that: The following steps are involved: Before scheduling a webpage scraping task, perform an initial assessment of the target webpage. This includes obtaining the structural level change range, visual area reconstruction frequency, and rendering position fluctuations of the webpage within a preset historical period, calculating a structural uncertainty score, and determining whether to enter the in-depth assessment process based on the score. When the structural uncertainty score exceeds the preset disorder threshold, a deep assessment is performed. This deep assessment includes analyzing the historical change trends of web page structure nodes, element rearrangement distribution, and regional stability patterns to generate a page structure consistency index. Simultaneously, multiple historical content snapshots are compared to calculate field update density, semantic variation amplitude, and time distribution patterns to generate a content update index. The page structure consistency index and content update index are fed into the pre-trained acquisition scheduling prediction model as input parameters, and the acquisition scheduling coefficient is output. This coefficient is used to evaluate the execution priority and policy tendency of the current task in the current scheduling cycle. Based on the numerical range of the acquisition scheduling coefficient, the acquisition path of the task is selected. The acquisition path includes a marker tracking and positioning process based on structural reconstruction tolerance enhancement, or an incremental comparison and snapshot update process based on content difference recognition. After determining the collection path, the execution strategy of the collection task is classified based on the combined relationship between the page structure consistency index and the content update index. The classification results are structure-dominated, content-dominated, or scheduling-responsive. According to the classification results and the preset collaborative strategy, the resource allocation order and scheduling time window of the task are adjusted to complete the collaborative scheduling and strategy adaptation between multiple tasks.
2. The multi-task collaborative incremental web page data collection method according to claim 1, characterized in that: The calculation of the structural uncertainty score involves the following steps: Compare the structural hierarchical information of the target web page obtained in the last three consecutive collection cycles, and extract the set of structural nodes in each cycle. Each node in the set of structural nodes contains three indicators: label name, hierarchical depth and number of sibling nodes; The sets of structural nodes between any two cycles are mapped one by one, and the number of label changes, the number of hierarchical migrations, and the number of new or missing nodes are compared. The inter-cyclic structural variation of any two cycles is obtained by weighted summing according to the preset ratio; the average value of the inter-cyclic structural variation is used as the structural variation benchmark value, and the benchmark value is divided by the total number of structural nodes of the current web page. The quotient is multiplied by the preset floating weight coefficient corresponding to the current web page type to obtain the structural uncertainty score.
3. The multi-task collaborative incremental web page data collection method according to claim 2, characterized in that: The process of generating the page structure consistency index includes the following steps: Extract the structural hierarchical information of the target web page at five consecutive historical collection time points, and convert the web page structure at each time point into a set of node paths. Each path is represented by a combination of a label sequence and its hierarchical depth. Using the path set corresponding to the fifth time point as a reference, perform alignment on each path from the first four time points. Path alignment is based on matching the longest common substring of the tag sequence. If the length of the common substring exceeds 60% of the total path length, the alignment is considered successful. After successful alignment, the hierarchical distance difference between the path pairs is calculated, that is, the hierarchical difference value of the label end node, and the hierarchical difference values of all path pairs are averaged to form the hierarchical difference value relative to the reference structure at that time point; The concentration of the structural path at each time point is extracted. The concentration is defined as the percentage of the number of paths on the same level divided by the total number of paths. The hierarchical difference value at each time point is multiplied by its concentration, and the product values at the five time points are averaged to finally obtain the page structure consistency index.
4. The multi-task collaborative incremental web page data collection method according to claim 3, characterized in that: The process of generating the content update index includes the following steps: Extract content snapshots of the target webpage at five consecutive historical sampling time points consistent with the structural level sampling, divide the content of each snapshot into multiple content blocks based on paragraph boundaries, and perform difference analysis on the content blocks at the same position in adjacent snapshots at each time point; For each pair of adjacent content blocks, the text edit span is calculated, which is defined as the ratio of the minimum edit distance between the two versions of the content block to the total number of words in the block in the current snapshot; Calculate the change range of keyword groups, which is defined as the ratio of the size of the symmetric difference between the keyword sets extracted from the two versions of the content to the size of the union of the two sets; Calculate the publishing time alignment offset, defined as the difference between the standard deviation of the publishing time of all paragraphs in the content block and the standard deviation of the average publishing time of all paragraphs at that time point; The above three values of each group of paragraphs, namely the text editing span, the change amplitude of keyword groups and the release time arrangement offset, are normalized to the range of zero to one, and the normalized result is multiplied by the display density of the paragraph in the web page, where the display density is the percentage of the paragraph in the total content length of the web page; the product values of all paragraphs in each pair of snapshots are summed up, and finally the sum between the five snapshots is averaged to generate the content update index.
5. The multi-task collaborative incremental web page data collection method according to claim 4, characterized in that: The collection scheduling prediction model used to generate the collection scheduling coefficient is pre-trained through supervised learning. The training samples include the page structure consistency index, content update index, actual collection success rate, scheduling delay time and resource utilization rate in historical collection tasks. During the training process, the model goal is to minimize the comprehensive loss function of scheduling error rate, scheduling delay time and resource waste rate. The final model is constructed using an integrated regression algorithm, and the output collection scheduling coefficient is a continuous value between zero and one, which is used to represent the comprehensive urgency and resource adaptation difficulty of the current collection task within the scheduling cycle.
6. The multi-task collaborative incremental web page data collection method according to claim 5, characterized in that: The marker tracking and positioning process includes the following steps: Obtain the stored structure tag path set from the last successful acquisition version. Each path set contains a tag name sequence and its corresponding hierarchical index. Perform a path scan on the current web page structure tree and calculate the matching similarity between each stored path and the current structure path. Paths with a similarity lower than the preset ratio threshold will be marked as offset paths. Reconstruct the hierarchical pointer for each offset path and replace the old path with the new path for update.
7. The multi-task collaborative incremental web page data collection method according to claim 6, characterized in that: The incremental comparison and snapshot update process includes the following steps: Get the content area of the current webpage and divide it into paragraphs according to the structural position of the content area in the previous version of the webpage. Each paragraph is a content block, and the content block is indexed by the double coordinates of the paragraph starting position and length; A bidirectional mapping is performed on the content block indexes in the two versions. The conditions for successful matching are that the starting position of the paragraph in the current version is offset by no more than 50 characters and the length change does not exceed 20%. If this condition is not met, the block is considered to have failed to match. Keyword groups are extracted from all content blocks that fail to match. Each group of keywords is screened from the title field and the first fifty words of the paragraph using a preset text analysis algorithm. The difference in keyword groups is defined as the ratio of the number of symmetric difference sets between the keyword sets of two versions to the number of their unions. If the difference value is greater than 0.6, the content block is determined to be changed content. All paragraphs determined to be changed content are treated as updated paragraphs to replace the original paragraph content at the corresponding position in the snapshot, and the timestamp and keyword group of this change are recorded.
8. The multi-task collaborative incremental web page data collection method according to claim 7, characterized in that: After determining the collection path, the execution strategy of the collection task is classified through the fuzzy logic classifier based on the combined relationship between the page structure consistency index and the content update index. The fuzzy logic classifier uses the page structure consistency index and the content update index as input variables, and generates classification labels based on the preset membership function and decision rules. The classification results are structure-dominated, content-dominated, or scheduling-responsive.
9. A multi-task collaborative incremental web page data collection system based on the multi-task collaborative incremental web page data collection method according to any one of claims 1 to 8, characterized in that: Specifically include: The initial structural evaluation module is used to perform an initial evaluation of the target web page before the web page collection task is scheduled. The evaluation operation includes obtaining the structural level change range of the web page within a preset historical period, the reconstruction frequency of the visible area, and the rendering position fluctuation, calculating the structural uncertainty score, and judging whether to enter the in-depth evaluation process based on the score; The deep perception module is used to perform a deep assessment when the structural uncertainty score exceeds a preset disorder threshold. The deep assessment includes analyzing the historical change trends of web page structure nodes, element rearrangement distribution, and regional stability patterns to generate a page structure consistency index. It also compares multiple historical content snapshots to calculate field update density, semantic variation amplitude, and time distribution patterns to generate a content update index. The scheduling prediction module is used to input the page structure consistency index and content update index as input parameters into the pre-trained collection scheduling prediction model, and output the collection scheduling coefficient, which is used to evaluate the execution priority and policy tendency of the current task in the current scheduling cycle; The path execution module is used to select the task acquisition path according to the numerical range of the acquisition scheduling coefficient. The acquisition path includes a marker tracking and positioning process based on structural reconstruction tolerance enhancement, or an incremental comparison and snapshot update process based on content difference recognition; The strategy dispatch module is used to classify the execution strategy of the collection task after determining the collection path, combining the combined relationship between the page structure consistency index and the content update index. The classification results are structure-dominated, content-dominated or scheduling-responsive. According to the classification results and the preset collaboration strategy, the resource allocation order and scheduling time window of the task are adjusted.
Citation Information
Patent Citations
Deep Web-oriented self-adaptive incremental data acquisition method
CN109977285A
Data acquisition monitoring method and device based on InfluxDB database, and medium
CN113485885A
Internet data intelligent acquisition method and system
CN119322899A
Multi-dimensional resource dynamic game optimization control method and system based on multi-agent system
CN120197716A
Extraction and comparison method for text of webpage
WO2017080090A1
Cited By
Dynamic page generation method and device based on metadata management and medium
CN121070509A
Non-intrusive webpage adaptive processing method and system based on instantaneous environment situation
CN122087214A
Non-intrusive web page adaptive processing method and system based on instantaneous environmental situation
CN122087214B