System browser page comparison method and system based on DOM (Document Object Model) superposition algorithm
By constructing interactive behavior sequences and aggregated DOM structures, combined with graph neural network analysis, filtering invalid nodes, the problems of redundancy and low accuracy of DOM structures in the existing technology are solved, and efficient and accurate browser page comparison is achieved.
Patent Information
- Application Number
- CN202510868812.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-26
AI Technical Summary
In the existing system browser page comparison method based on DOM overlay algorithm, there are many redundant nodes in the DOM structure and lack an effective filtering mechanism, which leads to high noise and low accuracy in comparison results, making it difficult to accurately capture structural changes at the visual perception and semantic levels.
By constructing interactive behavior sequences, obtaining DOM structure snapshots, extracting common node features, building aggregated DOM structures, and filtering invalid nodes, combining graph neural networks for topological analysis and multimodal fusion scores to identify substantive structural change areas.
It significantly improves the detection accuracy and comparison efficiency of complex page structure changes, reduces DOM snapshot redundancy, improves the identification accuracy of key structural changes and the coverage of the comparison process.
Smart Images

Figure CN120372226A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of dynamic analysis of web page structures. More specifically, the present invention relates to a method and system for comparing system browser pages based on a DOM overlay algorithm. Background Art
[0002] With the rapid development of Internet technology, the complexity and interactivity of web applications have increased day by day, and the web page structures and presentation forms have become more diverse and dynamic. As an important technical means in the fields of web page testing, automated verification, front-end regression detection, etc., system browser page comparison is widely used in scenarios such as version update difference analysis, visual regression detection, and function verification. Accurately and efficiently identifying the structural differences between different versions of web pages and ensuring user experience and system stability are one of the core objectives of current web page comparison technologies.
[0003] For example, a method for extracting information of web page structured data disclosed in the invention patent with the publication number: CN107423391B first preprocesses the web page code to remove noise information, constructs its DOM tree based on web page layout tags as nodes through the nested and hierarchical relationships of layout tags, and stores it in a List. Pruning the DOM tree by judging whether the branches are the same to form a DOM reconstruction tree; then marking the nodes through node paths, comparing the DOM reconstruction trees corresponding to two web pages, determining the feature paths where the target objects are located, and generating corresponding wrappers to achieve automatic extraction. The present invention can automatically and quickly process a large amount of WEB content and extract correct information.
[0004] For example, a method for automatically extracting web page information disclosed in the invention patent with the publication number: CN111428444B, characterized by including the following steps: preprocessing the web page information; constructing a block DOM tree; positioning the body area; and extracting the web page body; wherein, constructing the block DOM tree includes the following steps: performing fault tolerance compensation and DOM parsing on the web page source code; constructing a block DOM structure based on the DOM in combination with HTML block layout elements; counting the number of basic theme elements of the DOM block in combination with display features; and performing weighted calculation on the basic theme elements of the DOM block; wherein, when positioning the body area, positioning the body area according to the theme weight value obtained by weighted calculation. The beneficial effect of the present invention is that it takes into account both the efficiency and accuracy of web page information extraction. On the basis of not significantly reducing the traditional web page extraction method, it considers the layout features of the web page and some visual features of HTML, and effectively improves the accuracy of web page information extraction.
[0005] In the above-disclosed technical solutions, there are at least the following technical problems: In the existing system browser page comparison method based on the DOM overlay algorithm, there are many redundant nodes in the DOM structure and a lack of an effective filtering mechanism, resulting in large noise and low accuracy in the comparison results. During the comparison process of the DOM overlay algorithm, traditional comparison often relies on simple node difference determination, making it difficult to accurately capture structural changes at the visual perception and semantic levels. To address the above problems, the present invention proposes a solution. Summary of the Invention
[0006] To overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a system browser page comparison method and system based on the DOM overlay algorithm. Through the aggregation of the DOM generation mechanism and the purified overlay comparison method, it can effectively solve problems such as insensitivity to dynamic structural changes, interference from invalid nodes, and inaccurate positioning of version differences in existing browser page comparison, and significantly improve the detection accuracy and comparison efficiency of complex page structure changes.
[0007] To achieve the above object, the present invention provides the following technical solutions: A system browser page comparison method based on the DOM overlay algorithm includes: performing structural analysis on the interaction graph of the initial DOM structure, constructing an interaction behavior sequence, and obtaining a DOM structure snapshot corresponding to the current page state to construct a DOM snapshot sequence that is spatio-temporally aligned with the interaction behavior sequence; performing structural preprocessing on the DOM snapshot sequence, extracting common node features of all snapshots, and constructing an aggregated DOM structure; evaluating the nodes in the aggregated DOM structure, and filtering out invalid nodes according to the evaluation results to generate a purified DOM structure; performing DOM overlay comparison on the purified DOM structure, and identifying substantial structural change regions through hierarchical topology matching.
[0008] In a preferred embodiment, the construction of the interaction behavior sequence is as follows: obtaining the initial DOM structure of the target page and identifying the set of interactive elements therein; constructing an interaction graph based on the set of interactive elements and extracting the attribute feature vectors of each vertex; inputting the interaction graph into a neural network for topological analysis to output the structural mutation sensitivity of each node; constructing a joint scoring function based on the structural mutation sensitivity and the attributes of the vertices to generate an action priority strategy; sequentially selecting and executing interaction actions according to the action priority strategy, obtaining a snapshot of the DOM structure of the page after interaction and performing structural difference analysis with the pre-interaction structure; updating the vertex attributes and edge weights of the interaction graph according to the difference features, and looping until the termination condition is met, and outputting the page interaction behavior sequence.
[0009] In a preferred embodiment, updating the vertex attributes and edge weights of the interaction graph according to the difference features, and looping until the termination condition is met, and outputting the page interaction behavior sequence, specifically as follows: updating the interaction graph according to the structural difference features after the interaction operation, including adding interactive nodes, vertex attributes generated by the structural differences, and adjusting the associated edge weights; recording the information of the current interaction node and the structural difference feedback as a behavior item; determining the termination condition after each round of interaction, and outputting the final interaction behavior sequence.
[0010] In a preferred embodiment, inputting the interaction graph into a neural network for topological analysis, and outputting the structural variation sensitivity of each node, specifically as follows: extracting the attribute features of each node in the interaction graph, generating a node feature matrix, and extracting the connection relationships in the graph structure to construct an adjacency matrix; inputting the feature matrix and the adjacency matrix into a graph neural network, performing feature aggregation and embedding calculations, and fusing with neighbor information through multi-layer graph neural network propagation to obtain the embedding vector of each node; inputting the embedding vector into the structural variation sensitivity prediction layer, and outputting the structural variation sensitivity score of each node.
[0011] In a preferred embodiment, obtaining a DOM structure snapshot corresponding to the current page state, and constructing a DOM snapshot sequence that is spatio-temporally aligned with the interaction behavior sequence, specifically as follows: configuring an independent page rendering view for each interaction behavior, and capturing the corresponding rendered frame image; comparing the current frame with the historical frame through a visual similarity algorithm, and activating the DOM change listener when the structural change exceeds the threshold; recording the event context when the listener is triggered, and judging whether to trigger snapshot collection based on a preset scoring strategy; for changes that reach the collection threshold, triggering the incremental snapshot mechanism and dynamically aggregating them into a DOM snapshot sequence.
[0012] In a preferred embodiment, the scoring strategy is specifically as follows: responding to the page structure change event triggered by user interaction, and identifying the DOM area range of the structural change; respectively rendering the states of the area before and after the interaction as image frames, generating a pre-state image block and a post-state image block; inputting the image blocks into a pre-trained lightweight multi-layer visual model, extracting perceptual features and outputting a perceptual feature vector; extracting the semantic attributes of the DOM nodes in the changed area, constructing a semantic vector, and aligning and fusing it with the perceptual feature vector to generate a changed area representation vector; inputting the representation vector into the importance scoring model, and outputting the importance score of the changed area.
[0013] In a preferred embodiment, the structural preprocessing of the DOM snapshot sequence is performed to extract the common node features of all snapshots and construct an aggregated DOM structure as follows: Generate a unique structural positioning path for each node in the interaction line sequence to construct a path space; Extract the node content under the same path and classify the nodes through a hash function. The classification includes redundant nodes, newly added nodes, and state-variant nodes; Insert the state-variant nodes and newly added nodes into the aggregated structure according to the path index; Perform topological optimization on the inserted aggregated structure, mark the node change types, and generate the final aggregated DOM structure.
[0014] In a preferred embodiment, the nodes in the aggregated DOM structure are evaluated as follows: Traverse all nodes in the aggregated DOM structure and extract the multi-dimensional data feature set of each node; Normalize the feature set to generate a standardized node feature vector; Input the node feature vector into the node validity scoring model to output the node validity score.
[0015] In a preferred embodiment, the nodes in the aggregated DOM structure are evaluated, and invalid nodes are filtered according to the evaluation results to generate a purified DOM structure as follows: Set a scoring threshold according to the task requirements, traverse all nodes in the aggregated DOM structure, screen out the nodes that meet the scoring threshold and their subtrees to generate a filtered aggregated DOM structure; Based on the structural nesting position relationship, perform path mapping operations on different versions of the filtered aggregated DOM structures to establish a node mapping set; Perform superposition comparison according to the mapping relationship and output a structure difference report.
[0016] A system for a system browser page comparison method based on the DOM superposition algorithm includes a snapshot sequence module, a structure construction module, a filtering module, and a superposition comparison module, and there are connections between the modules; The snapshot sequence module is used to perform structural analysis on the interaction graph of the initial DOM structure, construct an interaction behavior sequence, and obtain the DOM structure snapshot corresponding to the current page state to construct a DOM snapshot sequence that is spatio-temporally aligned with the interaction behavior sequence; The structure construction module is used to perform structural preprocessing on the DOM snapshot sequence, extract the common node features of all snapshots, and construct an aggregated DOM structure; The filtering module is used to evaluate the nodes in the aggregated DOM structure and filter out invalid nodes according to the evaluation results to generate a purified DOM structure; The superposition comparison module is used to perform DOM superposition comparison on the purified DOM structure and identify the substantial structure change area through hierarchical topology matching.
[0017] The technical effects and advantages of a system browser page comparison method and system based on the DOM superposition algorithm of the present invention: 1. By constructing an interaction behavior sequence, the present invention can simulate operations such as clicks, scrolls, and expansions of real users on a page, and drive the evolution of the page structure in a controllable manner. This sequence not only captures the interaction actions themselves, but also forms a two-way association between the interaction link and the structural change through a differential feedback mechanism, significantly enhancing the modeling ability of dynamic page state transitions, and solving the problem that traditional static DOM snapshots cannot reflect the real interaction path of the page. And a graph neural network is introduced to perform topological analysis and structural sensitivity modeling on the interaction graph, effectively characterizing the degree of influence of each interaction node in the structural evolution. By constructing a joint scoring function based on structural mutation sensitivity and node attribute features, dynamic prioritization of interaction actions can be achieved, improving the goal orientation of interaction path construction and the structural coverage rate of page exploration, and reducing the number of redundant interactions during the comparison process.
[0018] 2. Through the importance scoring strategy of multimodal fusion, the present invention determines whether the page change is sufficient to capture a snapshot after each interaction, no longer relying on the full-scale acquisition mechanism, significantly reducing the redundancy of DOM snapshots. At the same time, by combining three types of features: structural semantics, visual perception, and style consistency, a unified representation vector of the changed area is constructed, and then input into the scoring model to output the importance score, improving the recognition accuracy and interpretability of key structural changes. The DOM structure aggregation method based on path space and node content hash comparison can unify and fuse DOM snapshots in multiple versions and multiple states. By identifying redundant nodes, newly added nodes, and state mutation nodes, and performing structure insertion and annotation operations with paths as indexes, a sustainable evolving aggregated DOM structure is constructed, providing unified and standardized structural semantic support for subsequent comparison. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a schematic flowchart of a method for comparing system browser pages based on the DOM overlay algorithm of the present invention.
[0020] Figure 2 It is a schematic structural diagram of a system browser page comparison system based on the DOM overlay algorithm of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0022] Embodiment 1 Figure 1 A method for comparing system browser pages based on the DOM overlay algorithm of the present invention is given, including: S1. Perform a structural analysis on the interaction diagram of the initial DOM structure, construct an interaction behavior sequence, and obtain a DOM structure snapshot corresponding to the current page state to construct a DOM snapshot sequence.
[0023] In this embodiment, the construction of the interaction behavior sequence is as follows: Obtain the initial DOM structure of the target page, and identify the interactive elements in the initial DOM structure through a rule engine. The interactive elements include tag nodes with capabilities such as click, scroll, and hover, such as <button> etc.; Based on the interactive elements, construct an interaction graph, where the nodes of the graph represent each interactive DOM element, and the edges represent the page jump, partial update or structural change path that may be caused by the interaction operation, forming a graph structure model for analyzing the user behavior path; For each node in the interaction graph, extract the node attribute features, and the node attribute features include but are not limited to: element operability labels (such as clickable, hoverable), CSS style significance scores (such as z-index, occupied area), event binding types, and depth position information, etc., for subsequent graph learning and modeling; Perform topological analysis on the interaction graph through a graph neural network (such as GraphSAGE or GAT), and output the structural variation sensitivity of each node, where the structural variation sensitivity represents the possibility of significant page structure update after the node is triggered; Construct a combined scoring function based on the structural variation sensitivity and node attribute features; Sort the candidate interaction nodes in descending order according to the scoring function, generate an action priority list, and construct an action priority strategy based on the action priority list. Specifically, if the combined score is lower than the threshold, the node is excluded from the action pool; if the node is highly similar to the recently executed node (such as approximate XPath structure, same parent node), it is regarded as a "duplicate node", and the score is reduced or the selection is postponed; if the node has been executed, set a cooling time to avoid repeated clicks immediately and improve the information coverage diversity; Select the currently highest-scoring interaction node from the action priority list as the target node to be executed, and simulate the user behavior in combination with the operation type of the node (such as click, scroll, hover), and execute the interaction action; After each interaction operation is completed, obtain a snapshot of the DOM structure of the page after the interaction, compare it with the pre-interaction structure, analyze the structural differences, and the structural differences include changes in the number of nodes, changes in the visible area, changes in the focus element, etc., and generate structural difference features; Based on the structural difference features, update the interaction graph, and form a behavior record by combining the current interaction action and its structural feedback result, and append it to the interaction behavior sequence according to the time sequence until the action pool is exhausted or the preset termination condition is met, so as to construct a page interaction behavior sequence with high coverage and sensitive to structural changes; The preset termination conditions include: the action pool is empty, that is, all interactive nodes have been processed or excluded; no new changes have occurred in the page interaction graph, and there is no significant structural feedback after N consecutive interactions; the interaction behavior sequence has reached the set maximum step length; the time window is exhausted or the user-defined termination logic is triggered.
[0024] The specific form of the combined scoring function is as follows:
[0025]
[0026] In the formula: represents the combined score, represents the node 's structural variation sensitivity score, represents the node 's weighted total score of attribute features, and are adjustable weight coefficients used to balance the importance of the two types of indicators, represents whether the node is bound to an interactive event. If bound, it is recorded as 1; if not bound, it is 0, represents the proportion of the visible area of the node in the page viewport, and the value range is [0, 1], represents the depth of the node in the DOM tree. After taking the reciprocal, it is normalized to prefer the upper-level structure, represents the historical access frequency of the node in the user click heat map (such as the normalized number of clicks), represents the complexity score of the event bound to the node, such as whether it contains asynchronous operations, nested functions, etc., , , , and are the weighted coefficients of each feature item.
[0027] Based on the structural difference features, update the interaction graph, and combine the current interaction action with its structural feedback result to form a behavior record, which is appended to the interaction behavior sequence according to the time sequence until the action pool is exhausted or the preset termination condition is met, so as to construct a page interaction behavior sequence with high coverage and sensitive to structural changes, specifically as follows: According to the structural difference features identified after the current interaction operation is completed, judge whether there are new interactive nodes in the page, or whether the state of the original node has changed (such as changing from hidden to visible, binding a new event, etc.). If a new or changed node is identified, add it as a new node to the interaction graph and establish an edge connection with the current interaction node to reflect the possible structural evolution path caused by this interaction, thereby dynamically updating the interaction graph; According to the updated interaction graph, combine the information of the current interaction node (including node ID, interaction type, execution time, trigger path) with its corresponding structural difference feedback (such as DOM tree difference summary, structural variation label) to form a behavior record item, which is used as an item in the interaction behavior sequence; After each round of interaction, it is determined whether the action pool is exhausted or a preset termination condition is met. If either condition is satisfied, the process of constructing the behavior sequence is stopped, and the final interaction behavior sequence is output. The behavior sequence consists of multiple interaction records arranged in chronological order, which can effectively depict the evolution path of the page structure and the actual feedback of interaction behaviors. This sequence can be applied to application scenarios such as user path simulation, page robustness testing, automated regression testing, and usability analysis.
[0028] In this embodiment, topological analysis is performed on the interaction graph through a graph neural network (such as GraphSAGE or GAT), and the structural variation sensitivity of each node is output as follows: Based on the historical interaction behavior logs, some nodes in the interaction graph are marked with structural changes, that is, each node is assigned a structural variation label to indicate whether it has triggered significant page structure updates (such as DOM node rearrangement, addition, visual area mutation, etc.) during historical interactions, thereby constructing a training data set for supervised learning; According to the constructed interaction graph, the attribute features of each node in the graph are extracted to generate a node feature matrix , and the connection relationships in the graph structure are extracted to construct an adjacency matrix , where N represents the number of nodes, F represents the feature dimension of each node, and each row represents the feature vector of a node; The feature matrix and the adjacency matrix are input into the graph neural network to perform feature aggregation and embedding calculation, and through multi-layer graph neural network propagation and neighbor information fusion, the embedding vector of each node is gradually obtained , where represents the node 's structural semantic representation in the low-dimensional space; The embedding vector is input into the structural variation sensitivity prediction layer, which can be a simple feed-forward neural network or a linear classifier with sigmoid activation, used to output the structural variation probability score of each node, that is, the structural variation sensitivity. The higher the score, the stronger the structural variation sensitivity of the node, and its interaction value should be given priority in the follow-up; The structural variation sensitivity results of all nodes are output as a vector, which is used in the construction of the joint scoring function and the sorting strategy of interaction action priorities.
[0029] In this embodiment, a DOM structure snapshot corresponding to the current page state is obtained, and a DOM snapshot sequence is constructed as follows: Start the target browser, construct a virtual container environment that supports multi-instance rendering, and configure the page rendering perspectives at two different time points before and after each interaction behavior (such as click, scroll), and record the corresponding rendered frame images; Based on the rendered frame image, the rendered images are compared according to a visual similarity algorithm (such as the Structural Similarity Index SSIM or the perceptual difference metric model based on VGG features) to determine whether the interactive operation has caused a significant change in the visual structure; If it is recognized as a visual change, immediately inject a DOM change listener into the page. The listener runs based on the MutationObserver mechanism and monitors in real time events such as the addition, deletion, and modification of the DOM structure, attribute changes, and subtree adjustments; When the listener triggers an event, record its event context information, including but not limited to the triggering node identifier, timestamp, DOM range of the changed area, change type, and combine it with a preset importance scoring strategy to determine whether the event reaches the snapshot collection threshold; For change events that meet the collection conditions, trigger the "incremental snapshot mechanism", that is, instead of re-collecting the entire DOM tree, locate the structure change area, extract its corresponding DOM subtree structure (Sub-DOM), and based on the tree structure difference algorithm (such as Tree-Diff), fuse the subtree differences into the previous version of the DOM snapshot to complete the dynamic aggregation update, and finally form a DOM snapshot sequence reflecting the page evolution process.
[0030] The specific importance scoring strategy is as follows: Whenever the user performs an interactive operation (such as clicking, scrolling, etc.) on the page, triggering a page structure change event, the system uses a preset DOM listening mechanism (such as MutationObserver) or DOM snapshot difference mechanism to identify the area where the structure has changed; Locate the changed DOM sub-region in the page, extract the corresponding structural range of the changed area, and render the states before and after the interaction of this area as image frames respectively, denoted as state image blocks, where the state image blocks include the current state image block and the previous state image block; Extract the structural style features of the DOM nodes involved in the changed area. The structural style features include but are not limited to font size, color, position, border style, transparency, etc., and map the structural style features to a unified standard style space; Input the image frames of the current state image block and the previous state image block into a pre-trained lightweight multi-layer visual model respectively to extract corresponding perceptual features such as texture gradients, edge contours, and significant regions, forming perceptual feature vectors. The perceptual features include the perceptual feature vectors of the current area before and after the interaction; Extract the semantic attributes of the DOM nodes in the changed area, including tag types (such as button, nav), ARIA role, class names, etc., and construct semantic vectors; Align the semantic vectors and fuse the perceptual feature vectors to construct a representation vector for the changed region , where is the perceptual feature vector of the current region after interaction, is the perceptual feature vector of the current region before interaction, is the structural semantic vector of the current region after interaction, is the structural semantic vector of the current region before interaction, is the style change feature difference vector; Input the representation vector of the changed region into a preset importance scoring model to output an importance score, which represents the degree of influence of the changed region on the user's visual or interaction experience in the current environment; Judge whether the event reaches the snapshot collection threshold according to the importance score.
[0031] The specific form of the importance scoring model is as follows:
[0032] In the formula: is the importance score, is the Sigmoid function, which is used to compress the output value into the interval [0,1] as the probability score of the change importance, is the weight parameter matrix, which represents the weight effect of each feature dimension in the scoring, is the representation vector of the changed region, is the bias term, which is used to adjust the overall shift of the score to prevent underfitting.
[0033] S2. Perform structural preprocessing on the DOM snapshot sequence, extract the common node features of all snapshots, and construct an aggregated DOM structure.
[0034] In this embodiment, performing structural preprocessing on the DOM snapshot sequence, extracting the common node features of all snapshots, and constructing an aggregated DOM structure are specifically as follows: Obtain the DOM structure snapshots corresponding to each key interaction node in the interaction behavior sequence. Based on XPath, CSSSelector, or a custom path encoding mechanism, generate a unique structural positioning path for each node in each DOM snapshot to construct a unified path space, so that even if the same node in different snapshots has a slight change in the tree structure position, it can be logically aligned and mapped; Extract the node content of the node paths that appear multiple times in the path space in each snapshot. The node content includes: tag name, attribute set (such as class, id, style), embedded text content (innerText, innerHTML), event binding, and interactive attributes; Use a hash function (such as SHA-1, MD5) or a structure signature algorithm on the node content to generate a content digest, and determine whether the content of the nodes under this path in each snapshot is consistent; Identify redundant nodes, newly added nodes, and state-variant nodes based on the judgment results, as follows: If the node content in multiple snapshots under the same path is consistent, it is identified as a redundant node, and only its first version is retained; If there are differences, mark it as a "state-variant node" and enter the subsequent structure merging process; If the path only exists in some snapshots, it is identified as a "newly added node"; Based on the path indexing results, allocate insertion positions for state-variant nodes and newly added nodes in the aggregated DOM structure, as follows: If the path does not originally exist in the aggregated DOM, insert it directly; If the path already exists but the content is different, coexist the structures by introducing state markers (such as data-state="afterClick") or version tags; If there are conflicts in the parent node structure, determine the final insertion position according to context consistency (such as parent node type, child node order); After completing the fusion and insertion operations of all state-variant nodes and newly added nodes, further aggregate and annotate the merged DOM structure to construct the final aggregated DOM structure.
[0035] It should be noted that this aggregated structure not only completely covers all valid DOM nodes that have appeared during the interaction process, but also attaches the source snapshot number and the corresponding triggered interaction action label to each node to retain its generation context. At the same time, add a visual state marker to each node to distinguish the visibility and activity of the node in different interaction steps, so as to support subsequent page state restoration, coverage rate statistics, or automated interaction playback and other applications. As the final output result, this structure has high information fidelity and state reducibility, significantly improving the ability to model and analyze complex page interaction structures.
[0036] The aggregation and annotation of the merged DOM structure to construct the final aggregated DOM structure are as follows: Perform aggregation processing on the initial merged DOM structure that has integrated DOM snapshots at each stage to eliminate redundant nodes and unify node identifiers, specifically including: Based on the XPath or CSS selector path method, map DOM nodes with the same path in different snapshots to a unified path space; For nodes with the same path, calculate their structure hash (structure tree layout + key attribute set) and content hash (text content + data attribute value); If both the structural hash and the content hash are the same, the nodes are regarded as duplicate nodes, and only one copy is retained; For nodes with the same structure but different styles (such as CSS properties, visibility, class names), on the basis of retaining the copy of the main nodes, a data-style-delta attribute is added to record the difference vector between the styles of each snapshot (such as display / block and none, color contrast, size change, etc.) for subsequent style restoration or playback; If a node exists in all snapshots and its structure, content, and style remain the same, it is determined as a stable node, and a data-stable="true" attribute is added to this node to indicate that it does not change during the interaction and can be optimized for rendering or cached during playback; Add the source context information to each node in the aggregated merged structure; According to the DOM rendering results and view screenshot analysis results corresponding to each interaction stage, mark the visible state of each node in the aggregated DOM; Based on the above aggregation and annotation results, construct and output the complete aggregated DOM structure.
[0037] The addition of the source context information to each node in the aggregated merged structure specifically includes: data-snapshot-id: Identifies the occurrence position of the node in multiple snapshots. For example, data-snapshot-id="S1,S3" indicates that the node exists in snapshots S1 and S3; data-trigger-action: Records the interaction trigger action that the node depends on, such as click, scroll, etc. For example, data-trigger-action="click#btn_submit"; data-interaction-depth: Records the level of the interaction path where the node is located, used to analyze the interaction stage in which it appears.
[0038] The marking of the visible state of each node in the aggregated DOM includes, but is not limited to: data-visible: Used to identify in which snapshots the node is in a visible state, such as data-visible="S1,S2"; data-occluded: Used to identify whether the node is in an invisible state due to occlusion, scrolling, page switching, etc. in some stages; data-theme-variant: Used to distinguish the visual representation of nodes in different theme modes (such as light / dark), for example, data-theme-variant="dark".
[0039] S3. Evaluate the nodes in the aggregated DOM structure, and filter out the invalid nodes according to the evaluation results to generate a purified DOM structure.
[0040] In this embodiment, the evaluation of the nodes in the aggregated DOM structure is as follows: Traverse all the nodes in the aggregated DOM structure, and extract the node data features, where the node data features include the visual contribution degree, structural position, interaction response situation, and historical co-occurrence frequency of the nodes; Normalize the node data features to form a node feature vector; Input the node feature vector into the node validity scoring model to output the node validity score.
[0041] The node validity scoring model is specifically as follows:
[0042]
[0043] In the formula: is the updated representation of node v at the l+1 layer, that is, the new feature vector obtained after node v aggregates its neighbors, is the feature vector representation of node u at the l layer, is the set of adjacent nodes of node v, is the representation vector of node v in the l layer, is the feature aggregation function (which can take the mean of neighbor features, weighted sum, attention weighting, etc.), is the validity score of the node, and the value range is between 0 and 1, is the weight vector of the output layer, is the activation function (usually Sigmoid, which compresses the value to 0-1), The final representation of node v after L-layer propagation, is the bias term.
[0044] It should be noted that for the visual contribution degree index: Through the browser viewport rendering analysis tool, obtain the actual visible area of each node in the user window, and use this ratio to represent its visual occupancy weight; The structural position score: According to the information such as the depth of the node in the DOM tree, the parent-child structure weight, and the position symmetry, calculate the structural weight through the structural heuristic rule or the position mapping network; Interaction response situation: Detect whether the node is bound with JavaScript event handlers (such as onclick, onhover), and combined with the historical interaction behavior log, extract the trigger frequency and interaction popularity of the node; Historical co-occurrence frequency: By analyzing the occurrence frequency of the node in multiple past versions of the page, judge its structural stability and generality.
[0045] S4. Filter the invalid nodes in the aggregated DOM structure according to the evaluation results, and perform DOM overlay comparison on the filtered aggregated DOM structure.
[0046] In this embodiment, filtering the invalid nodes in the aggregated DOM structure according to the evaluation results, and performing DOM overlay comparison on the filtered aggregated DOM structure are specifically as follows: Set a scoring threshold according to the task requirements, traverse all aggregated DOM nodes, screen out the nodes and their subtrees that meet the scoring threshold, and form a filtered aggregated DOM structure; If some parent nodes have low scores but are necessary parents of valid nodes, they shall be retained but marked as "weak nodes"; Based on the comparison of node paths (such as XPath) or structural nesting positions, perform path mapping on the filtered aggregated DOM structures between different versions to ensure that the corresponding relationships between nodes are one-to-one matched, and handle the addition, deletion, and movement of nodes; Perform overlay comparison on the filtered aggregated DOM structure according to the mapping relationship, and identify the differences in node attributes. The differences include node content, attributes, style changes, etc.; Generate a structured difference report according to the difference results, clearly identify the DOM structure differences and change points between page versions, and facilitate subsequent page change analysis and automated testing.
[0047] Embodiment 2 Figure 2 A system browser page comparison system based on the DOM overlay algorithm of the present invention is provided, which is characterized by including a snapshot sequence module, a structure construction module, and a filtering module, and there are connections between the modules; The snapshot sequence module is used to perform structural analysis on the interaction graph of the initial DOM structure, construct an interaction behavior sequence, and obtain a DOM structure snapshot corresponding to the current page state to construct a DOM snapshot sequence; The structure construction module is used to construct an aggregated DOM structure according to the preprocessed DOM snapshot sequence; The filtering module is used to evaluate the nodes in the aggregated DOM structure, filter the invalid nodes in the aggregated DOM structure according to the evaluation results, and perform DOM overlay comparison on the filtered aggregated DOM structure.
[0048] The above formulas are all dimensionless and take their numerical calculations. The formula is a formula obtained by collecting a large amount of data for software simulation to approximate the real situation. The preset parameters in the formula are set by those skilled in the art according to the actual situation.
[0049] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product.
[0050] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0051] In addition, in each embodiment of the present application, the various functional modules can be integrated into a processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.
[0052] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0053] Finally, the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.< / button>
Claims
1. A method for comparing system browser pages based on the DOM overlay algorithm, characterized in that, including: Conduct a structural analysis on the interaction graph of the initial DOM structure, construct an interaction behavior sequence, and obtain a DOM structure snapshot corresponding to the current page state, and construct a DOM snapshot sequence that is spatio-temporally aligned with the interaction behavior sequence; Perform structural preprocessing on the DOM snapshot sequence, extract the common node features of all snapshots, and construct an aggregated DOM structure; Evaluate the nodes in the aggregated DOM structure, and filter out invalid nodes according to the evaluation results to generate a purified DOM structure; Perform DOM overlay comparison on the purified DOM structure, and identify the substantial structure change area through hierarchical topology matching.
2. The method for comparing system browser pages based on the DOM overlay algorithm according to claim 1, characterized in that The construction of the interaction behavior sequence is specifically as follows: Obtain the initial DOM structure of the target page and identify the set of interactive elements therein; Construct an interaction graph based on the set of interactive elements and extract the attribute feature vectors of each vertex; Input the interaction graph into a neural network for topological analysis and output the structural mutation sensitivity of each node; Construct a joint scoring function based on the structural mutation sensitivity and the attribute features of the vertices to generate an action priority strategy; Select and execute interactive actions in sequence according to the action priority strategy, obtain a snapshot of the DOM structure of the page after interaction, and perform a structural difference analysis with the pre-interaction structure; Update the vertex attributes and edge weights of the interaction graph according to the difference features, and loop until the termination condition is met, and output the page interaction behavior sequence.
3. The method for comparing system browser pages based on the DOM overlay algorithm according to claim 2, characterized in that, The updating of the vertex attributes and edge weights of the interaction graph according to the difference features, and looping until the termination condition is met, and outputting the page interaction behavior sequence is specifically as follows: Update the interaction graph according to the structural difference features after the interaction operation, including adding interactive nodes, vertex attributes generated by the structural differences, and adjusting the associated edge weights; Record the information of the current interaction node and the structural difference feedback as a behavior item; Execute the determination of the termination condition after each round of interaction and output the final interaction behavior sequence.
4. The method for comparing system browser pages based on the DOM overlay algorithm according to claim 3, wherein, The inputting of the interaction graph into a neural network for topological analysis and outputting the structural mutation sensitivity of each node is specifically as follows: Extract the attribute features of each node in the interaction graph to generate a node feature matrix, and extract the connection relationships in the graph structure to construct an adjacency matrix; Input the feature matrix and the adjacency matrix into a graph neural network, perform feature aggregation and embedding calculations, and fuse with neighbor information through multi-layer graph neural network propagation to obtain the embedding vector of each node; Input the embedding vector into the structural mutation sensitivity prediction layer and output the structural mutation sensitivity score of each node.
5. The method for comparing system browser pages based on the DOM overlay algorithm according to claim 4, characterized in that, The obtaining of the DOM structure snapshot corresponding to the current page state and the construction of the DOM snapshot sequence that is spatio-temporally aligned with the interaction behavior sequence are specifically as follows: Configure an independent page rendering view for each interaction behavior and capture the corresponding rendered frame image; Compare the current frame with the historical frame through a visual similarity algorithm, and activate the DOM change listener when the structure changes exceed the threshold; Record the event context when the listener is triggered, and judge whether to trigger snapshot collection based on a preset scoring strategy; For changes that reach the collection threshold, trigger the incremental snapshot mechanism and dynamically aggregate it into a DOM snapshot sequence.
6. The method for comparing system browser pages based on the DOM overlay algorithm according to claim 5, characterized in that, The scoring strategy is specifically as follows: Respond to the page structure change event triggered by user interaction and identify the range of the DOM area with structural changes; Render the states of the said regions before and after interaction as image frames respectively, and generate a pre-state image block and a post-state image block; Input the said image blocks into a pre-trained lightweight multi-layer vision model to extract perceptual features and output perceptual feature vectors; Extract the semantic attributes of the DOM nodes in the changed region, construct semantic vectors, and align and fuse them with the perceptual feature vectors to generate a representation vector for the changed region; Input the said representation vector into an importance scoring model to output the importance score of this changed region.
7. The method for comparing system browser pages based on the DOM overlay algorithm according to claim 6, characterized in that, The structural preprocessing of the DOM snapshot sequence, extracting the common node features of all snapshots, and constructing an aggregated DOM structure are as follows: Generate unique structural positioning paths for the nodes of each DOM snapshot in the interaction row sequence and construct a path space; Extract the node contents under the same path and classify the nodes through a hash function. The said classification includes redundant nodes, newly added nodes, and state-varied nodes; Insert the state-varied nodes and newly added nodes into the aggregated structure according to the path index; Perform topological optimization on the inserted aggregated structure, label the node change types, and generate the final aggregated DOM structure.
8. The method for comparing system browser pages based on the DOM overlay algorithm according to claim 7, characterized in that, The evaluation of the nodes in the aggregated DOM structure is as follows: Traverse all the nodes in the aggregated DOM structure and extract the multi-dimensional data feature sets of each node; Normalize the said feature sets to generate normalized node feature vectors; Input the node feature vectors into a node validity scoring model to output the node validity scores.
9. The method for comparing system browser pages based on the DOM overlay algorithm according to claim 8, characterized in that The evaluation of the nodes in the aggregated DOM structure, filtering out invalid nodes according to the evaluation results to generate a purified DOM structure, is as follows: Set a scoring threshold according to the task requirements, traverse all the nodes in the aggregated DOM structure, and screen out the nodes that meet the scoring threshold and their subtrees to generate a filtered aggregated DOM structure; Based on the structural nesting position relationship, perform path mapping operations on the filtered aggregated DOM structures of different versions to establish a node mapping set; Perform superposition comparison according to the mapping relationship and output a structure difference report.
10. A system using the system browser page comparison method of the DOM overlay algorithm according to any one of claims 1-9, characterized in that, It includes a snapshot sequence module, a structure construction module, a filtering module, and a superposition comparison module, and there are connections between the modules; The snapshot sequence module is used to perform structural analysis on the interaction graph of the initial DOM structure, construct an interaction behavior sequence, and obtain the DOM structure snapshot corresponding to the current page state, and construct a DOM snapshot sequence that is spatio-temporally aligned with the interaction behavior sequence; The structure construction module is used to perform structural preprocessing on the DOM snapshot sequence, extract the common node features of all snapshots, and construct an aggregated DOM structure; The filtering module is used to evaluate the nodes in the aggregated DOM structure and filter out invalid nodes according to the evaluation results to generate a purified DOM structure; The superposition comparison module is used to perform DOM superposition comparison on the purified DOM structure and identify the substantial structure change regions through hierarchical topology matching.
Citation Information
Patent Citations
Information extraction methods from structured web page data
CN107423391B
Automatic Webpage Information Extraction Method
CN111428444B
Brower page switching method and browser
CN108052539A
A method of automatically extracting list pages
CN109144513A
Efficient evaluation for diff of XML documents
US20070240035A1
Cited By
Page comparison detection method and computer program product
CN121301975A