Method, system and device for intelligent classification of browser front-end components based on element perception

By capturing DOM dynamic behavior in real time and constructing loading trajectory data streams, and utilizing cross-modal graph neural network models and contextual semantic graphs, we solve the problems of component recognition accuracy and classification consistency in Web systems and achieve high-precision component semantic classification.

CN120540743BActive Publication Date: 2025-09-23HEFEI D2S INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511033306.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-09-23
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

In existing technologies for SPA, lazy loading, and dynamic rendering Web systems, it is difficult to accurately identify components that have not yet been rendered or lazy loaded, resulting in poor component identification accuracy and high classification inconsistency.

Method used

By capturing DOM dynamic behavior in real time, building a loading trajectory data stream, generating an enhanced DOM structure graph, using a cross-modal graph neural network model to extract structural, behavioral and visual features, combining the contextual semantic graph to perform component semantic classification, and using a graph propagation algorithm to correct the results.

Benefits of technology

It significantly improves the ability to fully recognize the structure of asynchronous rendering and dynamic components, enhances the accuracy of structural restoration and classification consistency of component areas, and achieves high-precision and robust semantic classification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120540743B_ABST
    Figure CN120540743B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for intelligent classification of browser front-end components based on element perception, which relates to the technical field of intelligent classification of front-end components, and includes the following steps: performing dynamic structure completion on incompletely rendered DOM areas through a preset structure prediction model based on a trajectory data stream, and generating an enhanced DOM structure graph; extracting first data of each candidate component area in the enhanced DOM structure graph, wherein the first data includes structural features, behavioral features, and visual features, and generating a component semantic embedding vector; constructing a contextual semantic graph based on the component semantic embedding vector, outputting an initial classification result, and performing contextual correction on the initial classification result; mapping the corrected initial classification result to the front-end interface, and outputting a real-time classification result. The present invention significantly enhances the structural restoration accuracy and modeling robustness of complex component areas in the front-end page by introducing a loading trajectory data stream modeling and structure dynamic completion mechanism based on the browser operating environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent classification of front-end components, and more specifically, to a method, system and device for intelligent classification of browser front-end components based on element perception. Background Art

[0002] With the development of modern front-end technologies, mechanisms such as single-page applications (SPAs), lazy loading, and dynamic rendering have been widely adopted in web systems to improve page responsiveness and user experience. However, while these mechanisms offer performance improvements, they also pose significant challenges to the structural identification and semantic classification of components within front-end pages. Specifically, because route switching in SPA architectures typically does not refresh the entire page structure, but rather dynamically inserts DOM nodes through partial rendering, the page structure exhibits significant instability at different points in time.

[0003] At the same time, the lazy loading mechanism relies on user scrolling, viewport changes, or interactive events to trigger content loading. This results in a large number of components not appearing in the DOM tree during initial page rendering, making their true structure and functionality difficult to capture in a timely manner. Furthermore, the inherent delay between asynchronous data loading and DOM rendering prevents traditional component classification methods based on static DOM snapshots or predefined rules from accurately restoring the true semantic boundaries of components, leading to problems such as reduced recognition accuracy and inconsistent classification.

[0004] Modern web pages widely use SPA, lazy loading, and dynamic rendering technologies, which means that elements are not in the DOM when initially loaded. Due to the unpredictable dynamic content triggering mechanism, such as SPA route switching and lazy loading related to user behavior, and the timing difference between asynchronous data acquisition and DOM rendering, it is difficult for the model to accurately identify components that have not yet been rendered or lazy loaded, resulting in poor component recognition accuracy and high classification inconsistency.

[0005] In view of the above problems, the present invention proposes a solution. Summary of the Invention

[0006] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides an intelligent classification method and system for browser front-end components based on element perception. By capturing DOM dynamic behavior in real time, constructing a loading trajectory data stream and fusing a graph neural network model of structure-behavior-visual multimodal information, the semantic types of components are accurately classified, and dynamic correction of the classification results is achieved based on the contextual semantic graph, so as to solve the problems in the prior art of poor component recognition accuracy and high classification inconsistency caused by component lazy loading, asynchronous rendering and dynamic changes in page structure.

[0007] In the first aspect, the present invention provides an intelligent classification method for browser front-end components based on element perception, including: based on the loading trajectory data stream, performing dynamic structure completion on the incompletely rendered DOM area through a preset structure prediction model, and generating an enhanced DOM structure diagram; extracting first data of each candidate component area in the enhanced DOM structure diagram, the first data including structural features, behavioral features and visual features; constructing a cross-modal graph neural network model based on the first data, and generating a component semantic embedding vector; constructing a contextual semantic graph based on the component semantic embedding vector, outputting the initial classification result through a graph propagation algorithm, and performing contextual correction on the initial classification result based on the component topological relationship; mapping the corrected initial classification result to the front-end interface, and outputting real-time classification results.

[0008] In a preferred embodiment, the loading trajectory data stream of the component is generated as follows: by injecting a monitoring script, the MutationObserver interface is called to monitor the dynamic changes of the DOM structure, the identification information and DOM snapshot of the changed node are recorded, and a structure change event is generated; the visual state of the candidate node of the component is monitored through the IntersectionObserver interface, and a visual state event is generated when the node enters the window for the first time or the visual area ratio changes by ≥10%; the structure change event and the visual state event are merged in ascending order of timestamps to construct a unified time sequence trajectory chain to generate a structured loading trajectory data structure.

[0009] In a preferred embodiment, the method performs dynamic structure completion on the incompletely rendered DOM area based on the loading trajectory data stream through a preset structure prediction model and generates an enhanced DOM structure graph, specifically: based on the loading trajectory data stream, constructs a loading behavior graph with time synchronization features; identifies component areas with abnormal insertion sequences based on the loading behavior graph; through a multi-path behavior-induced structure generation network, integrates the structural context of components with consistent behavior in other pages, and completes the component structure of the abnormal area; uses a structural adversarial discrimination module to evaluate the confidence of the completed structure, and when the structural inconsistency score exceeds a preset threshold, automatically adjusts the label type, hierarchical relationship and style candidate set of the completed node through an enhanced feedback mechanism; writes the adjusted completed structure into the component perception graph to generate an enhanced DOM structure graph, and updates the input feature view of the semantic classification module.

[0010] In a preferred embodiment, the component area with abnormal insertion sequence is identified based on the loading behavior map, specifically: the loading behavior of each node in the loading behavior map is analyzed to extract the behavior feature vector; the behavior feature vector is input into a preset anomaly detection model for cluster analysis to identify the abnormal node set with significant deviation in loading timing or structural semantic break; the abnormal node set is structurally traced back and aggregated with the lower-level child nodes to generate the abnormal insertion area to be completed.

[0011] In a preferred embodiment, the first data of each candidate component area in the enhanced DOM structure diagram is extracted and input into the cross-modal graph neural network model to generate a component semantic embedding vector, specifically: the first data of the candidate component area in the enhanced DOM structure diagram is extracted, the first data including structural features, behavioral features and visual features; data processing is performed on the first data, including constructing graph node topology information of the component area based on structural features, calculating node loading behavior vectors based on behavioral features, and extracting style embedding and screenshot image vectors based on visual features; constructing a cross-modal graph neural network model based on graph node topology information, loading behavior vectors, style embedding and screenshot image vectors, and outputting the component semantic embedding vector.

[0012] In a preferred embodiment, the cross-modal graph neural network model is constructed based on graph node topology information, loading behavior vectors, style embeddings and screenshot image vectors, and the component semantic embedding vector is output. Specifically, a heterogeneous graph containing three types of nodes: structure, behavior, and vision is constructed based on the enhanced DOM structure graph; the modal feature vector of each node in the heterogeneous graph structure is initialized; the information contribution of the modal feature vector is dynamically calculated through the cross-modal attention fusion module, and weighted fusion is performed to generate a unified node representation; node-level feature propagation and aggregation are performed in the graph neural network, and a node semantic representation vector is output; aggregation operations are performed on the semantic representation vectors of all nodes in the component area to generate a component semantic embedding vector.

[0013] In a preferred embodiment, the context semantic graph is constructed based on the component semantic embedding vector, the initial classification result is output through the graph propagation algorithm, and the initial classification result is contextually corrected based on the component topological relationship, specifically: based on the semantic embedding vector of the candidate component, initial semantic discrimination is performed through a preset classification model, and the initial classification result and classification confidence of each component are output; the candidate component area is used as a node, and the semantic embedding vector of the component area is used as the node feature, and a context semantic graph is constructed based on the relationship within the page; high-confidence component nodes are marked as context anchors according to the classification confidence; based on the graph connection relationship between the anchor point and the low-confidence node, multiple rounds of semantic propagation are performed: category information is transmitted along the edge direction; and the classification results of low-confidence nodes are dynamically adjusted according to the edge weight.

[0014] In a preferred embodiment, the corrected initial classification results are mapped to the front-end interface and real-time classification results are output, specifically: based on the final component category label and the position information of the component in the enhanced DOM structure diagram, the corresponding front-end DOM node is located; a semantic label mapping table is established to store the binding relationship of label ID-node reference-category label; the semantic label data is injected into the rendering process through the browser data communication interface to drive the interface layer display logic; the component semantic label is visually displayed on the front-end interface according to the mapping table; DOM change events are monitored, and when the binding node update is detected, the corresponding semantic label is synchronously updated.

[0015] In the second aspect, the present invention provides an element-aware browser front-end component intelligent classification system, comprising the following modules: a structure prediction and completion module: for performing dynamic structure completion on incompletely rendered DOM areas based on a loading trajectory data stream through a preset structure prediction model, and generating an enhanced DOM structure graph; a feature extraction module: for extracting first data of each candidate component area in the enhanced DOM structure graph, wherein the first data includes structural features, behavioral features, and visual features; a cross-modal semantic embedding module: for constructing a cross-modal graph neural network model based on the first data, and generating a component semantic embedding vector; a contextual semantic correction module: for constructing a contextual semantic graph based on the component semantic embedding vector, outputting the initial classification result through a graph propagation algorithm, and performing contextual correction on the initial classification result based on the component topological relationship; a semantic label mapping module: for mapping the corrected initial classification result to the front-end interface, and outputting real-time classification results.

[0016] In a third aspect, the present invention provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to execute the element-aware browser front-end component intelligent classification method according to the first aspect or any corresponding embodiment thereof.

[0017] As can be seen from the above scheme, the present invention, by introducing a loading trajectory data flow modeling and structure dynamic completion mechanism based on the browser operating environment, can perceive the changes in the DOM tree structure and the lazy loading process in real time, effectively improve the structural integrity recognition ability of asynchronous rendering and dynamic components, and significantly enhance the structural restoration accuracy and modeling robustness of complex component areas in the front-end page. The present invention uses a cross-modal graph neural network and contextual semantic graph collaborative fusion mechanism, combined with the dynamic weighted fusion of structure, behavior and visual information and anchor-driven semantic propagation, to effectively improve the consistency and contextual adaptability of component semantic classification, and can still achieve high-precision and strong robustness semantic classification result output under the conditions of diversified page layouts and similar component appearances. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 Schematic diagram of the process of the method for intelligent classification of browser front-end components based on element perception of the present invention;

[0019] Figure 2 This is a structural diagram of the browser front-end component intelligent classification system based on element perception of the present invention. DETAILED DESCRIPTION

[0020] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0021] Example 1, Figure 1 The present invention provides an intelligent classification method for browser front-end components based on element perception, which includes the following steps:

[0022] S1, based on the loading trajectory data stream, performs dynamic structure completion on the incompletely rendered DOM area through the preset structure prediction model and generates an enhanced DOM structure diagram. The specific steps are as follows:

[0023] S11: Based on the loading trajectory data stream, a loading behavior graph with time synchronization characteristics is constructed; the loading behavior graph includes DOM node insertion time, first visible time, parent loading time, and the average loading value and structural relationship of the same-level components;

[0024] S12, identifying component regions with abnormal insertion sequences based on the loading behavior profile;

[0025] S13, through the structure generation network induced by multi-path behavior, integrates the structural context of components with consistent behavior in other pages to complete the component structure of abnormal areas;

[0026] S14, uses the structural adversarial discrimination module to evaluate the confidence of the completed structure. When the structural inconsistency score exceeds the preset threshold, the label type, hierarchical relationship and style candidate set of the completed node are automatically adjusted through the enhanced feedback mechanism;

[0027] S15, write the optimized completion structure into the component perception graph to obtain an enhanced DOM structure graph, and synchronously update the input feature view of the semantic classification module.

[0028] S16, the structure prediction model is obtained by integrating the structural context of components with consistent behaviors in other pages through a structure generation network induced by multi-path behaviors.

[0029] The loading trajectory data stream is obtained by capturing the dynamic change events of the front-end DOM tree and the lazy loading status of elements in the visible area in real time;

[0030] The DOM tree described above converts an HTML document into a hierarchical object model, representing the document structure of a webpage. It organizes all elements, attributes, and content within a page in a tree-like structure, with each node being an object containing the corresponding element's properties, methods, and relationships. Lazy loading refers to a strategy that optimizes resource loading, with the core goal of reducing resource consumption during initial loading and improving page performance.

[0031] The specific method for obtaining the loaded trajectory data stream is as follows:

[0032] By injecting a monitoring script into the browser runtime environment, the system monitors dynamic changes in the DOM structure in real time based on the MutationObserver interface, including child node insertion, deletion, and node attribute changes. It also records the identification information of the changed node and the DOM snapshot as a structure change event.

[0033] Node attribute changes include changes in the value of the node's identification attributes (such as id), style attributes (such as class, style), function attributes (such as src, href, value, type), state attributes (such as disabled, checked), and custom data attributes (such as data-*).

[0034] Register a visual detector based on the IntersectionObserver interface, monitor the window visibility of component candidate nodes in the front-end page, and generate a visual state event when the component first enters the visible area or the visible area ratio changes by ≥10%;

[0035] In this step, the visual detector uses the window as the root node and sets multi-level visibility thresholds to capture the transition process of components from invisible to fully visible.

[0036] Component candidate nodes include HTML elements with preset semantic tags, where the tags include div, section, button, input, img, and canvas.

[0037] Merge the structure change events and the visual status events in ascending timestamp order to build a unified time series trace chain, generating a standardized loading trace data structure containing the component unique identifier, loading behavior sequence, lazy loading status flag, and first visible time field;

[0038] After the loading trajectory data is constructed, this embodiment further sets termination monitoring conditions, including the duration of the page stable state, the stable visual time of the component, or the user page jump event.

[0039] The duration of the page stable state is: the page enters the stable state and no DOM changes occur for a period of time (e.g., 1 second).

[0040] Component stable visibility time: The monitored component enters the view and remains stable and visible for more than the set time (such as 5 seconds);

[0041] User page jump event: The user jumps to a new page or uninstalls a window.

[0042] In this embodiment, component regions with abnormal insertion sequences are identified based on the loading behavior graph, specifically:

[0043] Based on the loading behavior of each node in the loading behavior graph, behavioral feature vectors including insertion delay, visual delay, structural isolation and layout deviation are extracted;

[0044] Clustering node behavior feature vectors based on a preset anomaly detection model to identify abnormal node sets with significant loading sequence deviations or structural semantic breaks.

[0045] The abnormal node set is structurally traced upward and aggregated with its subordinate child nodes to generate the abnormal insertion area to be completed. Structural tracing upward means starting from the detected abnormal node, searching upward along the DOM structure hierarchy for its parent and ancestor nodes, looking for potential logical affiliations or semantic upper bounds to determine which larger component or container the abnormal node may belong to. Subordinate child node aggregation involves starting from the abnormal node and aggregating its child nodes downward along the DOM structure, including the complete subtree belonging to the abnormal area in the completion analysis scope.

[0046] Furthermore, node behavior feature vectors are clustered based on a preset anomaly detection model to identify significant deviations in loading timing or structural semantic breaks. Specifically:

[0047] The outlier degree of each node's behavior feature vector is scored using a preset anomaly detection model based on local outlier factors.

[0048] Nodes with outlier scores higher than the preset outlier threshold are initially marked as abnormal candidate nodes;

[0049] Perform loading timing analysis on abnormal candidate nodes. If the loading time of an abnormal candidate node deviates from the loading time of its structural superior or adjacent node by more than a preset first deviation threshold, it is determined to be a node with significantly deviated loading timing.

[0050] Perform structural semantic consistency comparison on abnormal candidate nodes. If the structural hierarchy depth or attribute type similarity between its parent node, sibling node and it is lower than the preset structural consistency threshold, it is determined to be a structural semantic fracture node.

[0051] Nodes that meet any abnormal condition are included in the abnormal node set.

[0052] The behavioral feature vector ,in Insertion delay (the delay when a node is first inserted into the DOM (relative to the start of page load)), is the visual delay (the time when the node first enters the visual area), is the layout deviation (the degree to which the node position deviates from the main direction of its parent node), is the structural isolation (the connection density between a node and its peers), is the structural depth (the depth of the node in the DOM tree).

[0053] In this embodiment, the structural adversarial discrimination module is used to evaluate the confidence of the completed structure. When the structural inconsistency score exceeds a preset threshold, the label type, hierarchical relationship, and style candidate set of the completed node are automatically adjusted through the enhanced feedback mechanism. Specifically,

[0054] Based on the structural adversarial discrimination module, the completed component structure is compared with the structural consistency samples in the real page, and the structural inconsistency score is output;

[0055] Compare the structural inconsistency score with the preset confidence threshold. When the score exceeds the threshold, the structural completion and correction process is triggered.

[0056] Based on the enhanced feedback mechanism, we collect mismatch features of completion nodes in the context structure, including label distribution deviation, parent-child hierarchy breaks, and style boundary conflicts.

[0057] Using mismatched features as input, the system updates the priority, hierarchical nesting strategy, and style inheritance rules of the candidate tag set during the completion process, automatically generating a new structural completion solution.

[0058] The updated completion result is re-sent into the structural adversarial discrimination module for evaluation until the structural inconsistency is lower than the confidence threshold or the preset maximum number of iterations is reached.

[0059] The structural inconsistency score is scored using the following formula:

[0060]

[0061]

[0062] in, Score structural consistency. 、 、 are the preset weighting coefficients for label, level and style similarity, To complete the set of structure nodes, is the set of real structure sample nodes, is the node label distribution similarity (using Jaccard similarity to measure the degree of overlap between the HTML tag sets in the completed structure and the real structure), is the hierarchical structure similarity (using tree edit distance or average path length difference to measure structural nested consistency), is the style similarity (obtained based on the cosine similarity between the completion node and the context style attribute), Score the structural inconsistency.

[0063] Furthermore, the anomaly detection model is an unsupervised clustering model or an anomaly score evaluation function based on a preset behavior feature vector, which is used to identify anomalies in the loading behavior of component nodes.

[0064] S2, extracting first data of each candidate component area in the enhanced DOM structure diagram, wherein the first data includes structural features, behavioral features, and visual features;

[0065] S3: Build a cross-modal graph neural network model based on the first data to generate component semantic embedding vectors. The specific steps are as follows:

[0066] S31: extracting first data of each candidate component area in the enhanced DOM structure graph, wherein the first data includes structural features, behavioral features, and visual features;

[0067] S32: Processing the first data includes constructing graph node topology information of the component area based on the structural features, calculating the node loading behavior vector based on the behavior features, and extracting style embedding and screenshot image vector based on the visual features;

[0068] S33: Build a cross-modal graph neural network model based on graph node topology information, loading behavior vectors, style embeddings, and screenshot image vectors to generate component semantic embedding vectors.

[0069] Among them, graph node topology information is used to define the node connection relationship in the graph neural network and maintain the structural hierarchical context of the component; loading behavior vectors are used to reflect the dynamic presentation and user interaction characteristics of the component and enhance the node distinguishability under the behavior dimension; style embedding and screenshot image vectors are used as visual modality input channels to assist the graph neural network in identifying component areas with similar appearance but different functions.

[0070] In step S33, the cross-modal graph neural network model is constructed based on the graph node topology information, loading behavior vector, style embedding and screenshot image vector, specifically:

[0071] Based on the enhanced DOM structure graph, nodes are represented as DOM elements in the candidate component area, and edges are represented as heterogeneous graph structures of DOM structure hierarchical relationships or behavioral similarity relationships;

[0072] Initialize the structural embedding vector of each node in the heterogeneous graph structure separately , behavior embedding vector and visual embedding vectors ;

[0073] Add a cross-modal attention fusion module to the preset graph neural network model to calculate the information contribution of structure, behavior and visual embedding vectors, and dynamically weight and fuse them to generate a unified node representation ;

[0074] The node representation is input into the graph neural network, and the node connection relationship in the graph structure is combined to perform feature propagation and aggregation to obtain the node-level semantic representation vector;

[0075] An aggregation operation is performed on the semantic representations of the nodes in each candidate component region to obtain a unified semantic embedding representation vector of the component region.

[0076] The embedding representation vector is specifically:

[0077]

[0078] in, is the output of the l+1th layer, i.e., the updated node feature, which represents the semantic embedding vector of the component. is an adjacency matrix with self-loops added, which represents the connection relationship between components. for The degree matrix of is a diagonal matrix formed by summing the rows of the adjacency matrix. is the input feature matrix of the l-th layer graph neural network, each row is the feature vector of a node, is the preset weight matrix for the lth layer, is the activation function.

[0079] In this embodiment, the node representation is input into the graph neural network, and feature propagation and aggregation are performed based on the node connection relationship in the graph structure. Specifically,

[0080] In the lth layer of the graph neural network, for each node, feature information is collected from its neighboring node set and the following propagation process is performed:

[0081] Combine the neighbor information with the current node representation to generate a new node representation:

[0082] in, is the neighbor feature aggregation result of the i-th node in the l-th layer graph neural network, is mean pooling, weighted summation or attention weighted summation, is the set of neighbor nodes, is the representation vector of neighbor node j in the l-1th layer, is the new feature representation vector of node i in layer l, is the preset nonlinear activation function, is the vector concatenation operation, is the weight matrix of the lth layer, is the self-embedding representation of node i in the l-1th layer, is the aggregate information collected from neighboring nodes, is the bias vector of the lth layer.

[0083] It should be noted that the cross-modal graph neural network model construction method significantly enhances the structural integrity perception, behavioral dynamic adaptability and visual consistency expression capabilities in the component identification process through the embedded collaborative expression and dynamic fusion of the three modalities of structure, behavior and vision, providing support for high-accuracy component classification and identification.

[0084] The information contribution calculation is specifically as follows:

[0085]

[0086] The node representation is specifically:

[0087]

[0088] in, , is the information contribution of node i under mode m, is the transpose of the preset shared attention weight vector, is the weight matrix of mode m, is the embedding vector of node i under mode m, is the hyperbolic tangent activation function, To traverse all modal index variables, s represents the structural embedding vector, b represents the behavioral embedding vector, and v represents the visual embedding vector. is the node representation of node i, is the embedding vector of node i in modality m.

[0089] S4: Build a contextual semantic graph based on the component semantic embedding vectors, output the initial classification results through the graph propagation algorithm, and perform contextual correction on the initial classification results based on the component topological relationships. The specific steps are as follows:

[0090] Based on the semantic embedding vectors of the candidate component regions, the preset semantic classification model is used to perform initial semantic category discrimination on each component, and the initial classification result and classification confidence of each component are obtained;

[0091] Each candidate component area is used as a node, and the semantic embedding vector of the component area is used as the node feature. A contextual semantic graph is constructed based on the relationships within the page; the relationships include spatial adjacency, loading timing, and interaction frequency.

[0092] High-confidence component nodes are marked as context anchors based on the classification confidence, and semantic propagation operations are performed based on the connection relationship and edge weights between anchor nodes and low-confidence nodes in the graph structure;

[0093] Dynamically adjust the classification results of low-confidence nodes during several rounds of propagation to correct the consistency of the context of the initial classification results;

[0094] The corrected classification result is used as the final component category label.

[0095] Furthermore, the initial classification results include the semantic category labels corresponding to each candidate component area (navigation bar, sidebar, text block, advertising area, bottom footer, function pop-up window).

[0096] The classification confidence value is the maximum category probability value of the component area in the semantic classification model output, specifically:

[0097]

[0098]

[0099] in, is the classification confidence, The probability value of the i-th component predicted as category c, is a set of semantic categories, is the category with the largest predicted probability value for component i, which is the initial semantic classification result.

[0100] The contextual semantic graph uses candidate component areas as graph nodes, and node features as corresponding semantic embedding vectors; edge connections between nodes are established through spatial adjacency relationships (such as sibling or parent-child structures), loading timing relationships (such as loading order), and interaction frequency relationships (such as click linkage) to form a structured graph relationship; among them, spatial edge weights can be weighted according to the physical distance between components, temporal edges can be weighted according to the loading time interval, and interaction edges can be weighted based on click or visit frequency.

[0101] The semantic propagation operation is specifically as follows: dividing node roles based on confidence scores, marking component nodes with confidence scores higher than a preset threshold in the initial classification results as context anchor nodes, and marking the remaining nodes as nodes to be adjusted; constructing a context semantic graph, using the semantic embedding vectors of each candidate component area as node feature representation, and the edges in the graph represent semantic connection relationships with spatial adjacency, similar loading timing, or frequent user interaction behaviors; performing semantic propagation operations based on the graph structure, using graph convolution or graph attention mechanisms, transferring the classification category information of the context anchor node to its adjacent nodes to be adjusted according to the connection weight, and generating a node representation after propagation; fusing the propagated information with the original prediction, and for each node to be adjusted, performing a weighted fusion of its initial classification result and the category confidence vector obtained by propagation to obtain a revised category distribution; updating the classification label, using the category with the highest probability in the fusion result as the revised classification result of the component, and recursively propagating to other relevant nodes in the context semantic graph until all nodes to be adjusted are corrected or converged.

[0102] The consistency correction is specifically as follows:

[0103] Initial classification results:

[0104] Each round of propagation updates the semantic vector:

[0105] Reclassify after each round of propagation:

[0106] If exists ,and , then the node is considered to be successfully corrected to a more reasonable classification result.

[0107] in, is the initial classification result, is the semantic representation vector of component node i in propagation round t+1, is the set of contextual adjacent nodes of component node i, is the edge weight between node i and its neighbor node j, which is calculated by multimodal similarity.

[0108] S5, maps the corrected initial classification results to the front-end interface and outputs the real-time classification results.

[0109] The modified initial classification results are mapped to the front-end interface to output the real-time classification results. The specific steps are as follows:

[0110] Based on the final component category label and the position information of the component area in the enhanced DOM structure diagram, determine the corresponding front-end DOM node or front-end component instance;

[0111] Establish a mapping table between semantic tags and DOM nodes, and bind each tag to the corresponding front-end node;

[0112] Through the data communication interface of the browser front end, the semantic tag information is synchronously injected into the front-end rendering process to drive the interface layer display logic;

[0113] Based on the mapping relationship, the semantic classification results of each component are displayed on the front-end interface;

[0114] Listen to the front-end DOM structure change events. When it detects that the component node is updated or redrawn, the label is updated synchronously according to the original mapping relationship table to ensure that the classification result display is consistent with the latest structure.

[0115] A browser is a software application used to access and display Internet information. It can convert the URL entered by the user into a request to the network server, obtain web page data (such as HTML, CSS, JavaScript, etc.), and parse and render it into a visual page for users to browse text, pictures, videos and other content. It also supports interactive operations (such as clicking links, filling out forms, etc.).

[0116] The front-end interface (FUI) is the part of a website, application, or software that users directly see and interact with. It typically refers to the user interface layer, consisting of visual elements and interactive components. It serves as the "bridge" between the user and the system, presenting back-end data in an intuitive manner and responding to user actions like clicks, input, and swipes.

[0117] Semantic classification results are the output of natural language processing (NLP) technology that analyzes text, sentences, or words and categorizes them into pre-set categories. Its core is to use algorithms to understand the semantic connotations of text (rather than relying solely on surface vocabulary) and match them to corresponding classification labels. This is widely used in scenarios such as sentiment analysis, intent recognition, and content moderation.

[0118] Example 2, Figure 2 The present invention provides an intelligent classification system for browser front-end components based on element perception, which includes the following modules:

[0119] Structure prediction and completion module: It is used to perform dynamic structure completion on the incompletely rendered DOM area based on the loading trajectory data stream through the preset structure prediction model and generate an enhanced DOM structure diagram;

[0120] Feature extraction module: used to extract first data of each candidate component area in the enhanced DOM structure diagram, wherein the first data includes structural features, behavioral features and visual features;

[0121] A cross-modal semantic embedding module: used to build a cross-modal graph neural network model based on the first data and generate component semantic embedding vectors;

[0122] Contextual semantic correction module: This module is used to construct a contextual semantic graph based on component semantic embedding vectors, output the initial classification results through the graph propagation algorithm, and perform contextual correction on the initial classification results based on the component topological relationships.

[0123] Semantic label mapping module: used to map the corrected initial classification results to the front-end interface and output real-time classification results.

[0124] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.

[0125] The above embodiments may be implemented in whole or in part through software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product.

[0126] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0127] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0128] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0129] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. An intelligent classification method for browser front-end components based on element perception, characterized in that: The following steps are involved: Based on the loading trajectory data stream, dynamic structure completion is performed on the incompletely rendered DOM area through the preset structure prediction model, and an enhanced DOM structure diagram is generated; Extracting first data of each candidate component area in the enhanced DOM structure graph, wherein the first data includes structural features, behavioral features, and visual features; Building a cross-modal graph neural network model based on the first data to generate component semantic embedding vectors; Build a contextual semantic graph based on component semantic embedding vectors, output the initial classification results through the graph propagation algorithm, and perform contextual correction on the initial classification results based on the component topological relationships; Map the corrected initial classification results to the front-end interface and output real-time classification results.

2. The method for intelligent classification of browser front-end components based on element perception according to claim 1 is characterized in that: The specific method for obtaining the loaded trajectory data stream is as follows: By injecting a monitoring script and calling the MutationObserver interface to monitor dynamic changes in the DOM structure, the identification information of the changed nodes and the DOM snapshot are recorded, and a structure change event is generated. Monitor the visual state of candidate nodes of a component through the IntersectionObserver interface, and generate a visual state event when a node first enters the viewport or when the visible area ratio changes by ≥10%. The structural change events and visual status events are merged in ascending order of timestamps to build a unified temporal trajectory chain and generate a structured loading trajectory data structure.

3. The method for intelligent classification of browser front-end components based on element perception according to claim 2 is characterized in that: Based on the loading trajectory data stream, the preset structure prediction model is used to perform dynamic structure completion on the incompletely rendered DOM area and generate an enhanced DOM structure diagram, specifically: Based on the loading trajectory data stream, a loading behavior map with time synchronization characteristics is constructed; Identify component regions with aberrant insertion sequences based on loading behavior profiles; Through the structure generation network induced by multi-path behavior, the structural context of components with consistent behavior in other pages is integrated to complete the component structure of abnormal areas; The confidence of the completed structure is evaluated using a structural adversarial discrimination module. When the structural inconsistency score exceeds a preset threshold, the label type, hierarchical relationship, and style candidate set of the completed node are automatically adjusted through an enhanced feedback mechanism. The adjusted completion structure is written into the component perception graph to generate an enhanced DOM structure graph, and the input feature view of the semantic classification module is updated.

4. The method for intelligent classification of browser front-end components based on element perception according to claim 3 is characterized in that: The component regions where abnormal insertion sequences are present are identified based on the loading behavior map, specifically: Analyze the loading behavior of each node in the loading behavior graph and extract the behavior feature vector; Input the behavior feature vector into the preset anomaly detection model for cluster analysis to identify abnormal node sets with significant loading sequence deviations or structural semantic breaks. The abnormal node set is structurally traced upward and aggregated with the lower-level child nodes to generate the abnormal insertion area to be completed.

5. The method for intelligent classification of browser front-end components based on element perception according to claim 4 is characterized in that: The cross-modal graph neural network model is constructed based on the first data to generate component semantic embedding vectors, specifically: Extracting first data of a candidate component area in the enhanced DOM structure graph, wherein the first data includes structural features, behavioral features, and visual features; Processing the first data includes constructing graph node topology information of the component area based on the structural features, calculating the node loading behavior vector based on the behavior features, and extracting style embedding and screenshot image vector based on the visual features; A cross-modal graph neural network model is constructed based on graph node topology information, loading behavior vectors, style embeddings, and screenshot image vectors, and the component semantic embedding vectors are output.

6. The method for intelligent classification of browser front-end components based on element perception according to claim 5 is characterized in that: The cross-modal graph neural network model is constructed based on graph node topology information, loading behavior vectors, style embeddings, and screenshot image vectors, and the component semantic embedding vectors are output, specifically: Based on the enhanced DOM structure graph, a heterogeneous graph containing three types of nodes: structure, behavior, and visual; Initialize the modal feature vector of each node in the heterogeneous graph structure; The cross-modal attention fusion module dynamically calculates the information contribution of modal feature vectors, performs weighted fusion, and generates a unified node representation. Perform node-level feature propagation and aggregation in graph neural networks, and output node semantic representation vectors; Aggregate the semantic representation vectors of all nodes in the component area to generate the component semantic embedding vector.

7. The method for intelligent classification of browser front-end components based on element perception according to claim 6 is characterized in that: The context semantic graph is constructed based on the component semantic embedding vector, the initial classification result is output through the graph propagation algorithm, and the initial classification result is contextually corrected based on the component topological relationship, specifically: Based on the semantic embedding vectors of candidate components, the preset classification model is used to perform initial semantic discrimination and output the initial classification results and classification confidence of each component; The candidate component areas are used as nodes, the semantic embedding vectors of the component areas are used as node features, and a contextual semantic graph is constructed based on the relationships within the page. High-confidence component nodes are marked as context anchors based on classification confidence; Based on the graph connection relationship between anchor points and low-confidence nodes, multiple rounds of semantic propagation are performed: category information is transmitted along the edge direction; and the classification results of low-confidence nodes are dynamically adjusted according to the edge weights.

8. The method for intelligent classification of browser front-end components based on element perception according to claim 7 is characterized in that: The modified initial classification results are mapped to the front-end interface to output real-time classification results, specifically: Based on the final component category label and the component's position information in the enhanced DOM structure diagram, locate the corresponding front-end DOM node; Establish a semantic label mapping table to store the binding relationship between label ID, node reference and category label; Inject semantic tag data into the rendering process through the browser data communication interface to drive the interface layer display logic; Visually display component semantic labels on the front-end interface based on the mapping table; Listen to DOM change events, and when the binding node is detected to be updated, the corresponding semantic tag is updated synchronously.

9. A system using the method for intelligent classification of browser front-end components based on element perception according to any one of claims 1 to 8, characterized in that: Includes the following modules: Structure prediction and completion module: It is used to perform dynamic structure completion on the incompletely rendered DOM area based on the loading trajectory data stream through the preset structure prediction model and generate an enhanced DOM structure diagram; Feature extraction module: used to extract first data of each candidate component area in the enhanced DOM structure diagram, wherein the first data includes structural features, behavioral features and visual features; A cross-modal semantic embedding module: used to build a cross-modal graph neural network model based on the first data and generate component semantic embedding vectors; Contextual semantic correction module: This module is used to construct a contextual semantic graph based on component semantic embedding vectors, output the initial classification results through the graph propagation algorithm, and perform contextual correction on the initial classification results based on the component topological relationships. Semantic label mapping module: used to map the corrected initial classification results to the front-end interface and output real-time classification results.

10. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the element-aware browser front-end component intelligent classification method according to any one of claims 1 to 8 by executing the computer instructions.

Citation Information

Patent Citations

  • Semantic component-based academic knowledge question and answer method, system and equipment and storage medium

    CN115344714A

  • Component attribute form generation method and electronic equipment

    CN120066476A