Page processing method and device, electronic equipment and computer readable medium

By combining recall rules with node prediction models, ad elements that affect the user browsing experience are identified and blocked, solving the problem of ad elements' impact on user experience on mobile devices and improving page loading speed and ecosystem security.

CN111353112BActive Publication Date: 2026-07-21BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
Filing Date
2020-02-27
Publication Date
2026-07-21

Smart Images

  • Figure CN111353112B_ABST
    Figure CN111353112B_ABST
Patent Text Reader

Abstract

The method comprises: determining a plurality of layout object nodes of a page according to an obtained HyperText Markup Language (HTML) file; after layout of any layout object node of the page, screening the any layout object node by using a preset recall rule to obtain a layout object node in the plurality of layout object nodes that meets the recall rule; predicting, based on a preset node prediction model, whether the layout object node that meets the recall rule is a specified target node; performing shielding processing on the specified target node, and generating a page after the shielding processing by using the layout object nodes remaining after the shielding processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a page processing method, apparatus, electronic device, and computer-readable medium. Background Technology

[0002] With the widespread adoption of mobile internet, more and more websites are conducting advertising and application promotion in mobile scenarios. On the one hand, limited by the screen size of mobile devices, the impact of advertising and other elements on the user's browsing experience is becoming increasingly apparent; on the other hand, some websites, in order to maximize short-term profits, display a large number of false, pornographic, and deceptive advertising elements, seriously affecting the user's browsing experience and undermining mobile ecosystem security.

[0003] Therefore, it is necessary to filter the page content displayed on websites to ensure the security of the mobile search ecosystem and thus improve the user browsing experience. Summary of the Invention

[0004] This disclosure provides a page processing method, apparatus, electronic device, and computer-readable medium.

[0005] In a first aspect, embodiments of this disclosure provide a page processing method, comprising: determining multiple layout object nodes of a page based on an acquired Hypertext Markup Language (HTML) file; after laying out any layout object node of the page, filtering the any layout object node using a preset recall rule to obtain layout object nodes that conform to the recall rule from among the multiple layout object nodes; predicting whether the layout object node that conforms to the recall rule is a specified target node based on a preset node prediction model; performing masking processing on the specified target node, and generating a masked page using the remaining layout object nodes after masking processing.

[0006] Secondly, embodiments of this disclosure provide a page processing apparatus, comprising: a node determination module, configured to determine multiple layout object nodes of a page based on an acquired Hypertext Markup Language (HTML) file; a node filtering module, configured to, after laying out any layout object node of the page, filter the any layout object node using a preset recall rule to obtain layout object nodes that conform to the recall rule among the multiple layout object nodes; a model prediction module, configured to, based on a preset node prediction model, predict whether the layout object node that conforms to the recall rule is a specified target node; and a masking processing module, configured to mask the specified target node and generate a masked page using the remaining layout object nodes after the masking processing.

[0007] Thirdly, embodiments of this disclosure provide an electronic device comprising: one or more processors; a memory storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors perform any of the page processing methods described above; and one or more I / O interfaces connected between the processors and the memory, configured to enable information interaction between the processors and the memory.

[0008] Fourthly, embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements any of the page processing methods described above.

[0009] The page processing method, apparatus, electronic device, and computer-readable medium provided in this disclosure combine recall rules with node prediction models to process pages. For layout object nodes filtered by recall rules, the node prediction model is used to determine whether they affect the browsing experience. In this way, layout object nodes that are predicted to affect the browsing experience are blocked, thereby optimizing the overall page browsing experience and providing protection for the security of the mobile search ecosystem. Attached Figure Description

[0010] The accompanying drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed description of exemplary embodiments with reference to the accompanying drawings, in which:

[0011] Figure 1 A schematic diagram of a page processing architecture provided in this disclosure embodiment;

[0012] Figure 2 This is a flowchart of a page processing method according to an embodiment of the present disclosure;

[0013] Figure 3 This is a schematic diagram of the recall rules in an exemplary embodiment of this disclosure;

[0014] Figure 4 This is a flowchart of a page processing method according to another embodiment of the present disclosure;

[0015] Figure 5 This is a schematic diagram illustrating the effect of the page processing method disclosed in this publication;

[0016] Figure 6 A block diagram of a page processing apparatus provided in an embodiment of this disclosure;

[0017] Figure 7 A block diagram of an electronic device provided in an embodiment of this disclosure;

[0018] Figure 8 This is a block diagram illustrating the composition of a computer-readable medium provided in an embodiment of the present disclosure. Detailed Implementation

[0019] To enable those skilled in the art to better understand the technical solutions of this disclosure, the page processing methods, apparatus, electronic devices and computer-readable media provided in this disclosure will be described in detail below with reference to the accompanying drawings.

[0020] Exemplary embodiments will be described more fully below with reference to the accompanying drawings; however, these exemplary embodiments may be embodied in different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will enable those skilled in the art to fully understand the scope of this disclosure. Unless otherwise specified, the various embodiments of this disclosure and the features thereof may be combined with each other.

[0021] Figure 1 This is a schematic diagram of the page processing architecture according to an embodiment of this disclosure. Figure 1 As shown, the architecture may include a mobile device 20 and a website 30, wherein the mobile device 20 may include a browser kernel 21, memory 22 and a display screen 23; the website 30 may include multiple pages 31.

[0022] The mobile device 20 may include, but is not limited to, personal computers, smartphones, tablets, personal digital assistants, servers, etc. They can all have various applications (Apps) installed, such as email apps.

[0023] Page 31 in this embodiment includes, but is not limited to, a landing page. A landing page can represent an independent webpage and can be used for marketing or advertising campaigns, such as a page accessed by a user or visitor through clicking on an advertisement that appears in a search result or through a paid search channel.

[0024] In one embodiment, when user 10 accesses website 30 via mobile device 20 and clicks on the Uniform Resource Locator (URL) of a page 31 on website 30, the browsing kernel 21 initiates the download of a Hypertext Markup Language (HTML) file based on the URL, parses the downloaded HTML file to obtain a DOM (Document Object Model) tree, and simultaneously initiates the download of CSS and JS files when parsing resource links such as Cascading Style Sheets (CSS) and scripting language (JavaScript, JS) files on the HTML file. The downloaded CSS and JS files are stored in memory 22.

[0025] Because website behavior changes very rapidly, it's impossible to exhaustively list all types and pages using configured rule sets. Furthermore, not all advertisements negatively impact the user experience; when ad elements are positioned without interfering with the main page content and do not engage in misleading or deceptive practices, they represent normal business behavior. Widespread mis-filtering of such ads would disrupt the normal internet ecosystem. However, many current solutions cannot distinguish between ads representing normal business behavior and those impacting the user experience. If page elements on a website are filtered based on rule sets, page loading speed can be significantly affected when the rule set becomes too large.

[0026] This disclosure provides a page processing method that, before displaying page 31 on the screen 23 of the mobile device 20, intelligently identifies the types of page elements in page 31 during the rendering stage of the browser kernel 21 and automatically blocks page elements that affect the user's browsing experience. After page 31 is rendered, the user 10 sees an optimized page, which greatly improves the user's browsing experience and provides security for the mobile search ecosystem.

[0027] All of the following embodiments can be applied to the system architecture of this embodiment. For the sake of brevity, the following embodiments can be referenced and cited in turn.

[0028] Figure 2 This is a flowchart illustrating a page processing method according to an embodiment of this disclosure. Figure 2 As shown, the page processing method may include the following steps.

[0029] S110, Based on the obtained Hypertext Markup Language (HTML) file, determine multiple layout object nodes of the page.

[0030] S120: After laying out any layout object node on the page, use the preset recall rules to filter the any layout object node and obtain the layout object nodes that meet the recall rules from multiple layout object nodes.

[0031] S130, based on a preset node prediction model, predicts whether a layout object node that meets the recall rules is the specified target node.

[0032] S140: Mask the specified target node, and use the remaining layout object nodes after masking to generate the masked page.

[0033] According to the page processing method of this disclosure, a combination of recall rules and node prediction models is used to process the page. After the layout object nodes are filtered by the recall rules, the node prediction model is used to determine whether they affect the browsing experience. The layout object nodes that are predicted to affect the browsing experience are then blocked, and the page after the blocking process is generated. This optimizes the overall page browsing experience and provides protection for the security of the mobile search ecosystem.

[0034] In this embodiment, since the rendering kernel's process of processing web pages is very complex, choosing an appropriate time to hide target nodes is extremely important from the perspective of processing performance and user experience. The layout of layout object nodes represents the process of arranging and calculating the geometric information such as the width, height, and position of the layout object nodes. If we simply perform ad recognition and re-layout the entire page each time the overall page layout is completed, although this can achieve the recognition, the webpage needs to undergo dozens or even hundreds of layouts during display, and the entire page needs to be traversed for target node recognition. Traversal and re-layout both consume time, significantly impacting the overall page loading time and directly resulting in a perceived slowdown in webpage loading.

[0035] Therefore, in order to achieve the best performance and user experience, the page processing method of this disclosure embodiment can actively trigger partial layout without traversing the entire page or re-laying the entire page. Specifically, in step S120 above, after laying out any layout object node of the page, the layout object nodes can be filtered using preset recall rules.

[0036] In other words, in this embodiment of the disclosure, each node on the page calls its own layout method during layout, thereby avoiding traversing the DOM tree. After the node is laid out, if the node is identified as a target node that affects the browsing experience, the target node is blocked. For example, the target node's state is set to hidden, and the kernel layout state is reset. The kernel re-layout is actively initiated, so that the node can be laid out locally, avoiding re-layout at the entire page level.

[0037] In one embodiment, step S110 may specifically include: S21, parsing the HTML file to obtain the Document Object Model (DOM) and Cascading Style Sheets (CSS); S22, parsing the CSS to obtain the style data of the HTML element nodes in the DOM; and S23, determining multiple layout object nodes of the page based on the HTML element nodes that need to be rendered in the DOM and the style data.

[0038] Each layout object node corresponds to an HTML element node that needs to be rendered, and the style data of each layout object node is the style data of the corresponding HTML element node.

[0039] In this embodiment, the Document Object Model (DOM) can be a tree-structured DOM, i.e., a DOM tree; multiple layout object nodes can be nodes in a Layout Object tree; after the Layout Object tree is established and layout is performed, the nodes of the Layout Object tree can have a series of attribute information such as coordinates, width, and height.

[0040] In other words, in this embodiment, each node in the Layout Object tree corresponds to an HTML element node in the DOM that needs to be rendered. The CSS property object used to describe the HTML element node in the DOM tree is set to the layout object node in the newly created Layout Object tree so that the layout object node in the Layout Object tree can be drawn according to the style data in the CSS.

[0041] In one embodiment, if parsing the HTML file yields a script file link, then before step S23, the method may further include: S31, downloading and executing the script file corresponding to the script file link to obtain the HTML element node corresponding to the script file; S32, using the HTML element node corresponding to the script file as a layout object node that conforms to the recall rule.

[0042] In other words, in some embodiments, after determining multiple layout object nodes of the page, the page processing method may further include: if the multiple layout object nodes include layout object nodes loaded through a script file, then the layout object nodes loaded through the script file are taken as layout object nodes that meet the recall rules.

[0043] In this embodiment, since many target nodes that affect the browsing experience are dynamically loaded by JS, the HTML element nodes corresponding to the script files can be used as layout object nodes that meet the recall rules based on whether the nodes are loaded by JS. This is used to initially filter the nodes to be identified, thereby triggering node re-layout through asynchronously loaded JS resources. This effectively reduces the time required for subsequent identification of nodes that affect the browsing experience using the node prediction model.

[0044] Figure 3 This diagram illustrates the recall rules in an exemplary embodiment of the present disclosure. In this embodiment, filtering layout object nodes using preset recall rules is referred to as rule-based coarse recall.

[0045] likeFigure 3 As shown, in rule-based coarse recall, node recall conditions can be set from aspects such as node width-to-height ratio, node embedding form, node position characteristics, node content, node generation mechanism, and node structure.

[0046] In other words, recall rules may include: rules that are pre-set based on at least one of the following: node width-to-height ratio, node embedding form, node location characteristics, node content, node generation mechanism, and node structure.

[0047] In one embodiment, step S120 may specifically include: S41, laying out any layout object node of the page to obtain the attribute information of the laid-out layout object node; S42, determining whether the attribute information meets the node recall conditions defined in the recall rules; S43, taking the layout object node that meets the node recall conditions as the layout object node that meets the recall rules.

[0048] As an example, based on the rules set for the node's height-to-width ratio, nodes whose height ratio is less than a height ratio threshold and / or whose width ratio is less than a width ratio threshold are considered as nodes that meet the recall rules. In this example, nodes that affect the browsing experience rarely occupy the entire screen; they mostly exist in an interspersed or floating form on the page. Nodes whose height occupies, for example, 75% of the screen are very likely not target nodes, and child nodes of other nodes whose width ratio is less than the width ratio threshold can be filtered out.

[0049] As an example, rules based on node embedding include identifying nodes with a specified embedding format as nodes that meet the recall criteria. For instance, data analysis reveals that iframe nodes are a common habitat for target nodes, containing embedded data from many advertisers; therefore, nodes with iframes are also included in the set of suspected target nodes.

[0050] As an example, the rules set based on node location characteristics include including floating nodes as nodes that meet the recall rules. In this example, target nodes can be fixed, embedded, or floating relative to the page. Among these, floating target nodes have the worst impact on the browsing experience, as they obscure useful information and force users to close the page. Therefore, floating nodes are also included in the set of suspected target nodes.

[0051] As an example, the rules set based on node content characteristics include: selecting nodes with specified types of content as nodes that meet the recall rules. In this example, if a node contains a lot of text, images, interactive content, etc., it is highly likely to be a non-target node.

[0052] As an example, the rules set based on the node generation mechanism include: selecting nodes with a specified generation mechanism as nodes that meet the recall rules. In this example, if the nodes on the page include HTML source code and dynamically generated JS nodes, where JS-generated nodes are flexible and varied (most of the main content of the page is in HTML, while other dynamically changing content such as advertisements and related recommendations are generated by JS), then nodes generated by JS are highly likely to be the target nodes.

[0053] As an example, rules set based on node structure characteristics include: selecting nodes with a specified structure as nodes that meet the recall rules. In this example, the structural characteristics of nodes in the DOM tree can also be used as a filtering criterion. For example, in the DOM tree structure, nodes containing only plain text are mostly not target nodes (nodes that do not meet the recall rules); and block-level nodes with div / a / img formats are likely nodes promoted through images.

[0054] According to the page processing method of this disclosure, in the rule-based coarse recall, as long as the node recall condition defined by any one of the recall rules is met, it can be indicated that the node has the characteristics of a suspected target node, and the subsequent target node judgment logic can be performed; if none of the rules are met, it is regarded as a non-target node, so that a large number of normal nodes that do not affect the browsing experience can be filtered out through a recall rule filtering strategy.

[0055] In one embodiment, the following steps may be included before step S130 described above.

[0056] S51, take the layout object nodes that meet the recall rules as the layout object nodes obtained in the initial screening, and determine the node status of the layout object nodes obtained in the initial screening.

[0057] S52: After all layout object nodes on the page have completed their layout, retrieve the layout object nodes whose node states have changed.

[0058] S53, again using the preset recall rules, filters the layout object nodes whose node states have changed.

[0059] S54. The layout object nodes selected in the first screening and the layout object nodes obtained in the second screening are taken as layout object nodes that meet the recall rules.

[0060] In this embodiment, during the node layout process, since some nodes have interdependent relationships, the accurate visual information of the nodes has not yet been calculated during the initial layout, making it difficult to use a coarse recall strategy. Therefore, after the overall layout is completed, it is necessary to check the node status, such as the nodes whose visual information has changed, and to re-apply the coarse recall strategy to the node status. In this way, a batch of nodes that meet the recall rules after their status changed during the layout process can be retrieved through the recheck mechanism, thereby recalling more nodes that meet the recall rules and preventing target nodes from being missed.

[0061] In one embodiment, step S130 may specifically include the following steps.

[0062] S61, Calculate the node features of the layout object nodes that meet the recall rules based on the attribute information of the layout object nodes that meet the recall rules.

[0063] S62, using a preset node prediction model to process node features, obtains the probability value of a layout object node that conforms to the recall rule as the specified target node.

[0064] S63, based on the probability value, determine whether the layout object node that meets the recall rule is the specified target node.

[0065] In this embodiment, a machine learning model can be used to determine whether a node that meets the recall rules is a designated target node that affects the browsing experience.

[0066] In one embodiment, the layout object node that meets the recall rules is a node in the layout object tree of the page. Specifically, S61 may include the following steps.

[0067] S71, Obtain the attribute information of the layout object nodes that meet the recall rules. The attribute information is obtained during the layout process. S72, Using a depth-first traversal, the attribute information is used to perform top-down feature calculation on the layout object nodes that meet the recall rules in the layout object tree to obtain the node features of the layout object nodes that meet the recall rules.

[0068] In one embodiment, the node feature can be a specified dimension feature extracted and calculated from aspects such as node visual information, node content, and node structure. The specified dimension can be set according to actual calculation needs, for example, the specified dimension is greater than or equal to 10, and this embodiment does not make a specific limitation.

[0069] In this embodiment, bottom-up feature calculation can calculate node features and pass them to the parent node when building the layout object tree. However, in this mode, almost all page nodes have to participate in feature calculation. Since normal nodes are filtered through recall rules when the node layout is performed, top-down feature calculation can selectively calculate node features for layout object nodes that meet the recall rules (i.e., suspected target nodes) using a depth-first traversal method, thereby reducing the number of nodes for feature calculation and improving the speed of node feature calculation.

[0070] In one embodiment, the node prediction model is a model pre-trained using labeled, offline-rendered static page data, and the node prediction model is a gradient-enhanced decision tree model with a specified depth and a specified number of decision trees.

[0071] For example, since the node features processed by the browser kernel change dynamically, static data rendered offline can be used for annotation when selecting training data, and a highly accurate automated annotation tool can be set up to assist manual annotation, ultimately forming the training data.

[0072] For example, the node prediction model obtained by machine learning includes a gradient boosted decision tree (GBDT) model. The GBDT model is pre-trained using labeled data to obtain a specified depth and a specified number of decision trees, such as a model file with 100 trees at a depth of 4. Subsequently, the model file is used to predict whether the layout object node that meets the recall rule is the specified target node.

[0073] It should be understood that the depth of the node prediction model and the number of decision trees obtained by the above training are exemplary values. In actual application scenarios, model training can be completed according to the actual needs of users. This disclosure does not impose any specific limitations.

[0074] In one embodiment, step S140, which involves masking the specified target node, may specifically include the following steps.

[0075] S81, calculate the corresponding node characteristic information based on the attribute information of the specified target node.

[0076] The node characteristic information includes at least one of the following: position on the page, width, height, whether it is in the main content, and area percentage on the page.

[0077] S82, if the node characteristic information reaches the corresponding preset blocking threshold, the layout object node that affects the browsing experience is blocked by setting the state of the specified target node to hidden.

[0078] This disclosure provides a page processing method for blocking target nodes, which can adopt targeted processing mechanisms based on the characteristics of specified target nodes. After identifying the target node, the characteristics and area proportion of the target node on the overall page can be calculated. Then, based on a configurable blocking threshold, such as the node's position, width, height, or whether it is within the main content, the node can be blocked. This achieves flexible blocking of specified target nodes, maintains and ensures the ecological security of mobile search, and optimizes the overall page browsing experience.

[0079] In this embodiment, elements that affect the user's browsing experience are blocked. After the page is rendered and drawn, the user sees an optimized page, which greatly improves the user's browsing experience and provides security for the mobile search ecosystem.

[0080] The page processing method in this embodiment of the disclosure performs masking processing on specified target nodes, such as setting the node state to hidden and resetting the kernel layout state, and actively initiating kernel re-layout. The entire page processing process occurs before the node is drawn, thereby ensuring that the user does not have any jitter perception of hidden page nodes when browsing the page, thus optimizing the overall page browsing experience.

[0081] To better understand the page processing methods in this disclosure, the following will be used as an example. Figure 4 A page processing flow according to another embodiment of this disclosure is described. Figure 4 A flowchart illustrating another embodiment of the page processing method of this disclosure is shown. Figure 4 As shown, page processing methods may include the following steps.

[0082] S201, Download the Hypertext Markup Language (HTML) file based on the page URL.

[0083] S202: The parser parses the HTML file to obtain the DOM tree. When the CSS and JS file resource links on the HTML file are obtained, the CSS is downloaded and parsed, and the JS file is downloaded and executed.

[0084] In this step, the CSS is downloaded and parsed to obtain the style data of the nodes in the DOM tree; after downloading and executing the JS file, the nodes dynamically loaded by JS can be obtained, and the dynamically loaded nodes can be inserted / added to the DOM tree.

[0085] S203, construct the Layout Object tree based on the HTML element nodes that need to be rendered in the DOM tree and the style data of the nodes in the DOM tree.

[0086] S204. After the Layout Object tree is built, create the Layout Layer tree.

[0087] In this step, layer positioning and layout can be achieved based on the Layout Layer tree.

[0088] S205 filters the nodes dynamically loaded by JS in the layout object node tree and executes S209 to trigger the relayout of the dynamically loaded nodes.

[0089] exist Figure 4 In JavaScript, since dynamic loading of JS is an asynchronous resource loading process, the process of re-layouting nodes generated by dynamic loading can also be called node re-layouting triggered by asynchronous resource loading.

[0090] S206, During the process of laying out the nodes in the Layout Object tree, collect the attribute information of the layout object nodes.

[0091] S207, based on a preset node prediction model, scores whether the layout object node is a specified target node that affects the browsing experience, and predicts whether the layout object node will affect the browsing experience based on the scoring results.

[0092] In this step, the score of the layout object node is the probability value of whether the layout object node is a specified target node that affects the browsing experience.

[0093] In some embodiments, after laying out any node in the Layout Object tree, a preset recall rule can be used to filter the any layout object node to obtain layout object nodes in the Layout Object tree that conform to the recall rule. Thus, in step S207 above, based on a preset node prediction model, a score is given to whether the layout object nodes that conform to the recall rule are specified target nodes that affect the browsing experience.

[0094] S208 If the predicted target node is one that affects the browsing experience, the browser kernel sets the layout state and executes S209 to actively trigger the re-layout of the layout object node.

[0095] In step S208, the specified target node can be masked (e.g., the node state can be set to hidden) by rearranging the specified target node.

[0096] S209, perform a re-layout of the layout object node to obtain the layout object node after the re-layout masking process.

[0097] S210, Draw the page based on the layout object nodes after the masking process, so as to display the drawn page on the specified display screen.

[0098] According to the page layout method of this disclosure, a combination of recall rule strategy preprocessing and machine learning model is used to filter the nodes to be rendered, thereby blocking elements in the page that affect the browsing experience.

[0099] Figure 5 This diagram illustrates the effect of page processing in an embodiment of this disclosure. For example... Figure 5 As shown, page 1 includes multiple layout object nodes corresponding to multiple HTML object elements, such as node 1, node 2, node 3, or node 4.

[0100] exist Figure 5 In this example, each layout object node in page 1 calls its own layout method during the layout process, thus avoiding traversal of the DOM tree. The following steps can be performed for each layout object node.

[0101] like Figure 5 As shown in S301 "Rule-based coarse recall", after any layout object node on the page is laid out, the layout object node is filtered using preset recall rules to obtain the layout object nodes on the page that meet the recall rules.

[0102] Step S301 has the same processing procedure as step S120 in the above embodiment, and will not be described again in this embodiment.

[0103] like Figure 5 As shown in S302 "Recheck mechanism", after all layout object nodes on the page have completed their layout, the preset recall rules are used again to filter the layout object nodes whose node states have changed.

[0104] Step S302 has the same processing procedure as S53 in the above embodiment, and will not be described again in this disclosure.

[0105] like Figure 5 As shown in S303 "Model Recall", based on a preset node prediction model, it predicts whether the layout object node that meets the recall rules is the specified target node.

[0106] Step S303 has the same processing procedure as step S130 in the above embodiment, and will not be described again in this embodiment.

[0107] like Figure 5 As shown in S304 "Shielding Processing", a shielding process is performed on a specified target node, and the shielded layout object node is used to generate a shielded page.

[0108] Step S304 has the same processing procedure as step S140 in the above embodiment, and will not be described again in this embodiment.

[0109] like Figure 5 As shown, after page 1 is rendered, the user sees the optimized page 2, which greatly improves the user browsing experience and provides security for the mobile search ecosystem.

[0110] Figure 6 This diagram illustrates the composition of a page processing apparatus provided in an embodiment of the present disclosure. Figure 6 As shown, the page processing device includes the following modules.

[0111] The node determination module 610 is used to determine multiple layout object nodes of the page based on the obtained Hypertext Markup Language (HTML) file.

[0112] The node filtering module 620 is used to filter any layout object node on the page after layout, using a preset recall rule, to obtain layout object nodes that meet the recall rule among multiple layout object nodes.

[0113] The model prediction module 630 is used to predict whether a layout object node that meets the recall rules is a specified target node based on a preset node prediction model.

[0114] The masking module 640 is used to mask specified target nodes and generate a masked page using the remaining layout object nodes after masking.

[0115] The page processing apparatus according to embodiments of this disclosure can filter the page content displayed on a website, providing security for the mobile search ecosystem and thereby improving the user browsing experience.

[0116] In one embodiment, the node determination module 610 may include the following units.

[0117] The first parsing unit is used to parse the HTML file to obtain the Document Object Model (DOM) and Cascading Style Sheets (CSS); the second parsing unit is used to parse the CSS to obtain the style data of the HTML element nodes in the DOM; the node determination module 610 is specifically used to determine multiple layout object nodes of the page based on the HTML element nodes and style data that need to be rendered in the DOM.

[0118] Each layout object node corresponds to an HTML element node that needs to be rendered, and the style data of each layout object node is the style data of the corresponding HTML element node.

[0119] In one embodiment, if parsing the HTML file yields a script file link, the node determination module 610 may further include: a download execution unit, used to download and execute the script file corresponding to the script file link to obtain the HTML element node corresponding to the script file; the node determination module 610 is specifically used to use the HTML element node corresponding to the script file as a layout object node that conforms to the recall rules.

[0120] In one embodiment, the node filtering module 620 may further include: after determining the multiple layout object nodes of the page, if the multiple layout object nodes include layout object nodes loaded through a script file, then the layout object nodes loaded through the script file are regarded as layout object nodes that meet the recall rules.

[0121] In one embodiment, the node filtering module 620 may specifically include: an attribute information acquisition unit, used to lay out any layout object node of the page to obtain attribute information of the layout object node after the layout; a condition judgment unit, used to judge whether the attribute information meets the node recall conditions defined in the recall rules; and a recall node determination unit, used to take the layout object node that meets the node recall conditions as the layout object node that meets the recall rules.

[0122] In one embodiment, the recall rule may include: a rule pre-set based on at least one of the following: node width-to-height ratio, node embedding form, node location characteristics, node content, node generation mechanism, and node structure.

[0123] In one embodiment, the page processing apparatus may further include: a node state determination module, configured to determine the node state of the layout object nodes obtained from the initial screening, using layout object nodes that conform to the recall rules as layout object nodes obtained from the initial screening; a state change node acquisition module, configured to acquire layout object nodes whose node states have changed after all layout object nodes on the page have completed their layout; a node re-screening module, configured to re-screen layout object nodes whose node states have changed using the preset recall rules; and a filtered node determination module, configured to use the layout object nodes obtained from the initial screening and the layout object nodes obtained from the re-screening as layout object nodes that conform to the recall rules.

[0124] In one embodiment, the model prediction module 330 may include: a feature calculation unit, configured to calculate node features of layout object nodes that conform to the recall rule based on the attribute information of the layout object nodes that conform to the recall rule; a probability calculation unit, configured to process the node features using a preset node prediction model to obtain a probability value that the layout object node that conforms to the recall rule is a specified target node; and a target node determination unit, configured to determine whether the layout object node that conforms to the recall rule is a specified target node based on the probability value.

[0125] In one embodiment, the layout object node that meets the recall rules is a node in the layout object tree of the page.

[0126] In this embodiment, the feature calculation unit may include: an attribute information collection subunit, used to obtain attribute information of layout object nodes that conform to the recall rules, wherein the attribute information is information obtained during the layout process; and a feature calculation unit, specifically used to perform top-down feature calculation on layout object nodes that conform to the recall rules in the layout object tree using a depth-first traversal method and the attribute information, to obtain the node features of layout object nodes that conform to the recall rules.

[0127] In one embodiment, the node prediction model is a model pre-trained using labeled, offline-rendered static page data, and the node prediction model is a gradient-enhanced decision tree model with a specified depth and a specified number of decision trees.

[0128] In one embodiment, the blocking processing module 340 may specifically include: a feature calculation unit, used to calculate the corresponding node feature information based on the attribute information of the specified target node, the node feature information including at least one of position, width, height, whether it is in the theme content, and area ratio in the page; and a node blocking unit, used to block the specified target node by setting the state of the specified target node to hidden if the node feature information reaches the corresponding preset blocking threshold.

[0129] In one embodiment, the masking processing module 340 may further include: a drawing unit, used to re-layout the remaining layout object nodes after masking processing, and draw the re-laid-out layout object nodes to obtain the drawn masked page.

[0130] According to the page processing apparatus of this disclosure, a scheme combining rule recall and model prediction is used to mask specified target nodes. The entire page processing process occurs before the nodes are drawn, thereby ensuring that users do not experience any jitter due to hidden page nodes when browsing the page, thus optimizing the overall page browsing experience.

[0131] Figure 7This diagram illustrates a block diagram of an electronic device provided in an embodiment of the present disclosure; as shown... Figure 7 As shown, this disclosure provides an electronic device 700, including: one or more processors 701;

[0132] The memory 702 stores one or more programs, which, when executed by one or more processors, enable the one or more processors to implement any of the page handling methods described above; one or more I / O interfaces 703 are connected between the processor and the memory and configured to enable information exchange between the processor and the memory.

[0133] The processor 701 is a device with data processing capabilities, including but not limited to a central processing unit (CPU); the memory 702 is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), and flash memory (FLASH); the I / O interface (read / write interface) 703 is connected between the processor 701 and the memory 702, enabling information exchange between the processor 701 and the memory 702, including but not limited to a data bus (Bus).

[0134] In some embodiments, the processor 701, memory 702, and I / O interface 703 are interconnected via bus 704, and thus connected to other components of the electronic device 700.

[0135] Figure 8 This diagram illustrates the composition of a computer-readable medium provided in an embodiment of the present disclosure. Figure 8 As shown, this disclosure provides a computer-readable medium storing a computer program thereon, which, when executed by a processor, implements any of the page processing methods described above.

[0136] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0137] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.

Claims

1. A page processing method, comprising: Based on the obtained Hypertext Markup Language (HTML) file, determine multiple layout object nodes of the page; Actively trigger partial layout, after laying out any layout object node of the page, use preset recall rules to filter the any layout object node to obtain the layout object node that meets the recall rules among the multiple layout object nodes. Based on a preset node prediction model, predict whether the layout object node that conforms to the recall rule is the specified target node. The specified target node is masked, and the remaining layout object nodes after the masking process are used to generate the masked page. The method based on a preset node prediction model, which predicts whether a layout object node that matches the recall rules is a specified target node, includes: Obtain the attribute information of the layout object nodes that conform to the recall rules, wherein the attribute information is obtained during the layout process; Using a depth-first traversal approach, the attribute information is used to perform top-down feature calculations on the layout object nodes in the layout object tree of the page that conform to the recall rules, thereby obtaining the node features of the layout object nodes that conform to the recall rules. The node features are processed using the preset node prediction model to obtain the probability value that the layout object node that conforms to the recall rule is the specified target node; Based on the probability value, determine whether the layout object node that conforms to the recall rule is the specified target node.

2. The method according to claim 1, wherein, Following the determination of the multiple layout object nodes of the page, the following is also included: If the plurality of layout object nodes includes a layout object node loaded via a script file, then the layout object node loaded via the script file is considered as a layout object node that conforms to the recall rule.

3. The method according to claim 1, wherein, After laying out any layout object node of the page, the layout object node is filtered using a preset recall rule to obtain layout object nodes that meet the recall rule from among the plurality of layout object nodes, including: Layout any layout object node on the page to obtain the attribute information of the layout object node after layout; Determine whether the attribute information meets the node recall conditions defined in the recall rules; Layout object nodes that meet the node recall conditions will be considered as layout object nodes that conform to the recall rules.

4. The method according to claim 3, wherein, The recall rules include: rules that are pre-set based on at least one of the following: node width-to-height ratio, node embedding form, node location characteristics, node content, node generation mechanism, and node structure.

5. The method according to claim 1, wherein, Before predicting whether a layout object node that matches the recall rule is a specified target node based on a preset node prediction model, the method further includes: The layout object nodes that meet the recall rules are used as the layout object nodes obtained in the initial screening, and the node status of the layout object nodes obtained in the initial screening is determined. After all layout object nodes on the page have completed their layout, obtain the layout object nodes whose node states have changed. The preset recall rules are used again to filter layout object nodes whose node states have changed. The layout object nodes selected in the initial screening and the layout object nodes obtained in the second screening are taken as layout object nodes that meet the recall rules.

6. The method according to claim 1, wherein, The node prediction model is a model pre-trained using labeled, offline rendered static page data, and the node prediction model is a gradient-enhanced decision tree model with a specified depth and a specified number of decision trees.

7. The method according to claim 1, wherein, The process of masking the specified target node includes: Based on the attribute information of the specified target node, the corresponding node characteristic information is calculated. The node characteristic information includes at least one of the following: position, width, height, whether it is in the theme content, and area percentage on the page. If the node characteristic information reaches the corresponding preset blocking threshold, the specified target node is blocked by setting its state to hidden.

8. A page processing apparatus, comprising: The node determination module is used to determine multiple layout object nodes of the page based on the obtained Hypertext Markup Language (HTML) file. The node filtering module is used to actively trigger local layout. After laying out any layout object node on the page, it uses a preset recall rule to filter the any layout object node and obtain the layout object nodes that meet the recall rule among the multiple layout object nodes. The model prediction module is used to predict whether a layout object node that conforms to the recall rule is a specified target node based on a preset node prediction model. The masking module is used to mask the specified target node and generate the masked page using the remaining layout object nodes after the masking process. The model prediction module includes: The feature calculation unit is used to obtain the attribute information of the layout object nodes that conform to the recall rule. The attribute information is obtained during the layout process. Using a depth-first traversal method, the feature calculation is performed from top to bottom on the layout object nodes that conform to the recall rule in the layout object tree of the page using the attribute information to obtain the node features of the layout object nodes that conform to the recall rule. The probability calculation unit is used to process the node features using a preset node prediction model to obtain the probability value that the layout object node that conforms to the recall rule is the specified target node. The target node determination unit is used to determine, based on the probability value, whether the layout object node that conforms to the recall rule is the specified target node.

9. An electronic device, comprising: One or more processors; A storage device having stored one or more programs thereon, which, when executed by one or more processors, cause the one or more processors to implement the page processing method according to any one of claims 1-7; One or more I / O interfaces are connected between the processor and the memory and configured to enable information interaction between the processor and the memory.

10. A computer-readable medium having a computer program stored thereon, the program being executed by a processor to implement the page processing method according to any one of claims 1-7.