A method for classifying page elements, a parser, a medium, and a device

Through logistic regression model and single-hot encoding combined with XPath analysis, page element classification is automated, which solves the problems of low efficiency and high labor costs in the existing technology, and achieves more accurate page element classification.

CN113065600BActive Publication Date: 2025-07-08YLZ INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110378864.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-08
Publication Date
2025-07-08
Estimated Expiration
2041-07-08

AI Technical Summary

Technical Problem

In the prior art, the page element classification method is inefficient and has high labor costs, and cannot accurately reflect the element operation logic, and the manual labeling efficiency is low.

Method used

Logistic regression model is used to combine single-hot encoding and XPath analysis, and feature extraction and classification training of known page elements is carried out to establish decision boundaries to achieve automated page element classification.

Benefits of technology

It improves the accuracy of page element classification, reduces labor costs, and achieves reasonable classification results without reading the source code.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113065600B_ABST
    Figure CN113065600B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of page element classification, and particularly relates to a page element classification method, a parser, a medium and a device. Among them, a page element classification method includes extracting known page elements and classifying them, and extracting element features from the classified known page elements; establishing a logistic regression model and performing classification training on the logistic regression model; classifying page elements on a page file according to the trained logistic regression model; first classifying known page elements and extracting element features to perform classification training on the logistic regression model, and then using the trained logistic regression model to classify page elements on the page file, which improves the accuracy of page element classification, makes the classification closer to the essential features of page elements, so that developers can obtain more reasonable classification results without reading the source code during the upgrade and transformation, reducing the labor cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of page element classification, and particularly relates to a page element classification method, parser, medium and device. Background Art

[0002] With the popularization of computer technology, people's lives today have gradually entered the intelligent era. Not only computers, mobile phones, and PADs, but also all aspects of people's clothing, food, housing, and transportation have begun to apply the intelligent technology that has emerged not long ago, such as smart TVs, smart navigation, smart homes, etc. Intelligent technology will provide convenient and fast services in all aspects of people's lives. For example, intelligent voice interaction is a new generation of interaction mode based on voice input, and feedback results can be obtained by speaking.

[0003] Internet products are complex and diverse, and the designs on the pages are even more different. Among them, text and pictures are the two most basic elements that make up a web page. Text is the content of the web page, and pictures are for the beauty of the web page. In addition, the elements of the web page also include animations, music, programs, etc. By extracting and classifying page elements, users' interaction behaviors can be analyzed, which can help products and operations for subsequent optimization. For example, a Chinese patent application (publication number CN111310044A) discloses a method, device, equipment, and storage medium for extracting page element information, but this patent application does not mention how to classify page elements. The commonly used page element classification in the prior art is mainly based on html tag classification, classification by tag name, or page rendering by a page parser.

[0004] Among them, html tag classification includes element categories such as block-level elements and inline elements. However, this classification scheme is usually only applicable to the construction of layout and structure and cannot be used for accurate operation scheme design; and only classifying by tag name cannot truly and comprehensively reflect the operation logic contained in the element and is prone to ambiguity; the page parser is also only applicable to the fixed scenarios of realization and display and cannot make effective classifications. If you want to operate page elements in a unified and standardized manner, it is necessary to manually label each element on the page file one by one, which undoubtedly requires a large amount of labor costs and is not conducive to the implementation of application upgrade and transformation. Summary of the Invention

[0005] To solve the deficiency of low efficiency in manually labeling each element on the page file in the above-mentioned prior art, a page element classification method provided by the present invention can improve the accuracy rate of page element classification and reduce labor costs.

[0006] A page element classification method provided by the present invention includes the following steps:

[0007] S100: Extract known page elements and classify them, and extract element features from the classified known page elements;

[0008] S200: Establish a logistic regression model and perform classification training on the logistic regression model;

[0009] S300: Classify the page elements on the page file according to the trained logistic regression model.

[0010] Further, classify according to the functional characteristics of the known page elements, and the known page elements include but are not limited to display elements, operable elements, list elements or external elements;

[0011] By judging the influencing factors included in the element features, comparing the content, order or proportion of the influencing factors to determine the element type, and then extracting the element features according to the element type, and the element features include but are not limited to tags, structures, naming habits or attribute events.

[0012] Further, establish the logistic regression model based on the Logistic distribution function, and the Logistic distribution function is

[0013]

[0014] where μ is the location parameter and γ > 0 is the shape parameter.

[0015] Further, use one-hot encoding for feature encoding to extract the element features. When the influencing factor corresponding to the element feature is a known influencing factor, the element feature forms a classification sample; perform classification training on the logistic regression model through the classification sample, and then fit the decision boundary to establish the connection between the decision boundary and the classification training probability, so that the logistic regression model obtains the classification probability of the page elements.

[0016] Further, when the influencing factor corresponding to the element feature is an unknown influencing factor, use random logistic regression in the stability selection method for feature screening, and then add the screened supplementary element features to the logistic regression model, and backpropagate and correct the element features and the corresponding influencing factors in the logistic regression model.

[0017] Further, classifying the file page elements on the page file includes the following steps:

[0018] S301: Extract the page elements on the page file;

[0019] S302: Input the extracted page elements into the trained logistic regression model;

[0020] S303: The logical regression model outputs the page element group for which classification is completed.

[0021] Further, in step 301, fuzzy search using XPath is performed on the page file, and Dom nodes are parsed layer by layer based on the document to extract page elements.

[0022] The present invention also provides a page element classification parser, including:

[0023] An element extraction module, configured to extract and classify known page elements, and extract element features of the classified known page elements;

[0024] A modeling and training module, configured to establish a logical regression model and perform classification training on the logical regression model;

[0025] An element classification module, configured to classify page elements on a page file according to the trained logical regression model.

[0026] The present invention also provides a computer-readable storage medium storing computer instructions, and when the computer is executed by a processor, it implements a page element classification method as described in any one of the above.

[0027] The present invention also provides a computer device, including at least one processor and a memory communicatively connected to the processor, wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the processor executes a page element classification method as described in any one of the above.

[0028] Compared with the prior art, a page element classification method provided by the present invention classifies known page elements and extracts element features to form classification samples for classifying and training a logical regression model, and then uses the trained logical regression model to classify page elements on a page file; improves the accuracy of page element classification, makes the classification closer to the essential features of page elements, so that developers can obtain more reasonable classification results without reading source code during upgrade and transformation, reducing labor costs. Description of the Drawings

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for describing the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0030] Figure 1Flow chart of the page element classification method provided by the present invention;

[0031] Figure 2 Function schematic diagram of the logistic regression model provided by the present invention;

[0032] Figure 3 Flow chart of classifying file page elements provided by the present invention. Detailed implementation manners

[0033] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. In addition, the technical features designed in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0034] As Figure 1 shown, a page element classification method provided by the present invention includes the following steps:

[0035] Step 1: Extract known page elements and classify them, and extract element features of the classified known page elements; Step 2: Establish a logistic regression model and perform classification training on the logistic regression model; Step 3: Classify the page elements on the page file according to the trained logistic regression model.

[0036] Step 1: Extract known page elements and classify them, and extract element features of the classified known page elements.

[0037] Specifically, as Figure 1 shown, first extract known page elements and classify them. In this embodiment, fuzzy search of XPath is used to parse Dom nodes layer by layer based on the document to extract known page elements. After extracting the known page elements, they can be classified according to the functional characteristics of the known page elements. The classified known page elements include but are not limited to display elements, operable elements, list elements or external elements;

[0038] Among them, the display elements can be some inoperable elements on the page, such as elements only for display purposes like pictures, document texts, icons, etc.; the operable elements can be elements that can trigger corresponding logical methods through operations such as clicking, checking, sliding, etc. by the user; the list elements can be elements with indexes and embedded sub-elements, with a sequence order in the list, such as ordered lists, unordered lists, internal elements in dropdown selection boxes, etc. all belong to list elements; the external elements can be elements that are not actually displayed on the page body and are usually used to control some special operations of the entire application, such as return, close, etc.

[0039] Next, extract the element features of the classified known page elements. When extracting the element features, determine the element type by judging the influencing factors included in the element features, comparing the content, order, or proportion of the influencing factors, etc., and then extract the element features according to the element type. The element type can be a display element, an operable element, a list element, or an external element. One-hot encoding can be used for feature encoding to extract the element features. The element features include but are not limited to tags, structures, naming conventions, or attribute events, etc.

[0040] Specifically, the tags in the element features contain influencing factors such as , etc. It can be considered that such elements can be regarded as list elements, and the influencing factors included are such as , , <h1>, Text tags such as are usually regarded as display elements; when an onclick attribute is added inside the tag to become a link element, it is a type of operable element.

[0041] Therefore, it is necessary to first determine the influencing factors included in the element features, compare the content, order, or proportion of the influencing factors to determine the element type, and then extract the element features according to the element type; the tags in the element features can be used as a sufficient but unnecessary condition for element classification, which will not be listed one by one here, and important weighting is achieved through model training.

[0042] Similarly, for the structure in the element features, in addition to the tags in the element features, the document context structure where the page element is located and the embedded structure of the page element can usually distinguish whether the element type is a list element, and an important feature of a list element is that it has a series of peer sub-elements with indexes inside.

[0043] Regarding the naming convention in the element features, by building a set of common semantic analyses for application page operations, some common Chinese texts inside various page elements can be found. For example, buttons usually appear in combination with words such as "confirm", "save", "OK", etc., so the classification of page elements can be defined by searching for key Chinese characters in the document.

[0044] Regarding the attribute events in the element features, the attribute event features attached inside the element tag can be considered as empowering the element. The same element can be considered to belong to different classifications if it carries different attribute event features; for example, a general ordinary display element can be considered an operable element if it carries contextmenu, and elements with events such as onclick and onchange can also be considered operable elements.

[0045] Step 2: Establish a logistic regression model and conduct classification training on the logistic regression model.

[0046] In specific implementation, as Figure 1 、 Figure 2 shown, the logistic regression model is established based on the Logistic distribution function. The Logistic distribution is a continuous distribution defined by its location and scale parameters. The shape of the Logistic distribution is similar to that of the normal distribution, but the tail of the Logistic distribution is longer. Therefore, in this embodiment, the Logistic distribution is used for modeling, which has a data distribution with a longer tail and a higher peak than the normal distribution. The Logistic distribution function is

[0047]

[0048] where μ is the location parameter and γ > 0 is the shape parameter;

[0049] Under normal circumstances, there are hundreds or thousands of element features in the known page element dataset. In this embodiment, one-hot encoding is used for feature encoding to extract element features, and the element features are categorical values. Therefore, it is necessary to select the element features that have an obvious impact on the classification result for further modeling training;

[0050] The influence factors corresponding to the element features that have an obvious impact on the classification result are known influence factors. The element features form classification samples. The logistic regression model is classified and trained through separating the samples, and then the decision boundary is fitted to establish the connection between the decision boundary and the classification training probability, so that the logistic regression model obtains the classification probability of the page elements.

[0051] In this embodiment, obvious and predictable element features such as tags, structures, naming conventions, or attribute events are selected as the classification basis to classify the known page elements; the element features are extracted using one-hot encoding as follows,

[0052] Tags: ["div", "span", "a", "button",....N],

[0053] Structures: ["no sub-structure", "has sub-structure", "is a sub-structure", "is not a sub-structure"],

[0054] Naming conventions: ["confirm", "cancel", "xx list",...N],

[0055] Attributes, events: ["onClick", "onShow", "onBlur", "contextmenu",....N],

[0056] Among them, Y0 is the display element, Y1 is the operable element, Y2 is the list element, and Y3 is the external element;

[0057] So there are the following samples,

[0058] Y1( <i class="reduce"onclick="countminus(index)">Reduce ): ["i", "no sub-structure", "decrease", "onclick"] is represented as [10000..N 100010000...N, 100000N];

[0059] Y1( <span class="unfold"@click="openscrolldialog(index)">Expand > ): ["span", "no sub-structure", "expand", "onclick"] is represented as [00010..N1000 0000.1..N, 100000N];

[0060] Y0(<imgv-if="key=='sign.BS'"src=".. / .. / assets / images / xuetang.png" / >): ["img", "no sub - structure", "", ""] is represented as [01000..N 1000 01000...N, 0100000N];

[0061] Y2( <div v-for="(item,index)in eatdata":key="index"class="cards">): ["div", "has sub - structure", "food list", "v - for"] is represented as [00010..N 0100 00000.1..N, 000010N];

[0062] Preferably, each element feature contains various different influencing factors. When the influencing factor corresponding to the element feature is an unknown influencing factor, the F - value and P - value of each original feature are given through the F - test (abbreviated as F - test for f_regression test), so that variables can be screened, that is, select the element features with large F - value and small return P - value;

[0063] Then, the screened supplementary element features are added to the logistic regression model. When the logistic regression model is trained for classification, backpropagation is used to correct the element features and the corresponding influencing factors in the logistic regression model, thereby improving the classification accuracy of the logistic regression model.

[0064] Step 3: Classify the page elements on the page file according to the trained logistic regression model.

[0065] In specific implementation, as Figure 1 、 Figure 3 shown, when classifying the page elements on the page file after the logistic regression model training is completed, first extract the file page elements on the page file. Specifically, use the fuzzy search of XPath on the page file, parse the Dom nodes layer by layer based on the document, and extract the file page elements;

[0066] Then, input the extracted file page elements into the trained logistic regression model. Finally, the logistic regression model outputs the classified page element group. In this embodiment, the page element group includes but is not limited to a display element group, an operable element group, a list element group, or an external element group, etc.

[0067] Compared with the prior art, a page element classification method provided by the present invention classifies known page elements and extracts element features to form classification samples for training the logistic regression model, and then uses the trained logistic regression model to classify the page elements on the page file; improves the classification accuracy of the page elements, makes the classification closer to the essential features of the page elements, so that developers can obtain more reasonable classification results without reading the source code during the upgrade and transformation, reducing the labor cost.

[0068] The present invention also provides a page element classification parser, including an element extraction module, a modeling training module, and an element classification module. In this embodiment, the above - mentioned element extraction module, modeling training module, and element classification module can implement the above - mentioned page element classification method, and their implementation principles and technical effects are similar, so they will not be elaborated here.

[0069] In specific implementation, the element extraction module is used to extract and classify known page elements, and extract element features from the classified known page elements; the modeling and training module is used to establish a logistic regression model and perform classification training on the logistic regression model; the element classification module is used to classify the page elements on the page file according to the trained logistic regression model.

[0070] The present invention also provides a computer-readable storage medium storing computer instructions, and when the computer is executed by a processor, it implements a page element classification method as described in any one of the above.

[0071] In specific implementation, in this embodiment, the computer-readable storage medium is a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), etc.; the computer-readable storage medium may also include a combination of the above types of memories.

[0072] The present invention also provides a computer device including at least one processor and a memory communicatively connected to the processor, wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the processor executes a page element classification method as described in any one of the above.

[0073] In specific implementation, in this embodiment, the number of processors may be one or more, and the processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or a combination of the above types of chips. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0074] The memory and the processor can be communicatively connected through a bus or other means. The memory stores instructions executable by at least one processor. The instructions are executed by at least one processor so that the processor executes a page element classification method according to any one of the foregoing embodiments.

[0075] Compared with the prior art, a page element classification method, a parser, a medium and a device provided by the present invention classify known page elements and extract element features, so as to form classification samples to perform classification training on a logistic regression model, and then use the trained logistic regression model to classify the page elements on a page file; the accuracy of page element classification is improved, and the classification is closer to the essential features of the page elements, so that developers can obtain more reasonable classification results without reading the source code during the upgrade and transformation, reducing the labor cost.

[0076] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention. < / h1>

Claims

1. A method for classifying page elements, characterized in that, It includes the following steps: S100: Extract known page elements and classify them. For the classified known page elements, first determine the influencing factors included in the element features, compare the content, order, or proportion of the influencing factors to determine the element type, and then extract the element features according to the element type; S200: Establish a logistic regression model and perform classification training on the logistic regression model: When the influencing factor corresponding to the element feature is a known influencing factor, the element feature forms a classification sample, and the logistic regression model is classified and trained through the classification sample. Then, the decision boundary is fitted to establish the connection between the decision boundary and the classification training probability, so that the logistic regression model obtains the classification probability of the page element; When the influencing factor corresponding to the element feature is an unknown influencing factor, use random logistic regression in the stability selection method for feature screening, and then add the screened supplementary element features to the logistic regression model, and backpropagate and correct the element features and the corresponding influencing factors in the logistic regression model; S300: Classify the page elements on the page file according to the trained logistic regression model.

2. The page element classification method according to claim 1, wherein: Classify according to the functional characteristics of the known page elements. The known page elements include but are not limited to display elements, operable elements, list elements, or external elements; The element features include but are not limited to tags, structures, naming habits, or attribute events.

3. The page element classification method according to claim 2, wherein: The logistic regression model is established based on the Logistic distribution function, and the Logistic distribution function is where μ is the location parameter and γ>0 is the shape parameter.

4. A method for classifying page elements according to claim 3, characterized in that: Use one-hot encoding for feature encoding to extract the element features.

5. A method for classifying page elements according to claim 1, characterized in that Classifying the file page elements on the page file includes the following steps: S301: Extract the page elements on the page file; S302: Input the extracted page elements into the trained logistic regression model; S303: The logistic regression model outputs the classified page element group.

6. A method for classifying page elements according to claim 5, characterized in that: In step S301, use fuzzy search of XPath on the page file, and parse the Dom nodes layer by layer based on the document to extract page elements.

7. A page element classification parser for implementing a page element classification method according to any one of claims 1-6, characterized in that It includes: An element extraction module for extracting and classifying known page elements and extracting element features for the classified known page elements; A modeling and training module for establishing a logistic regression model and performing classification training on the logistic regression model; An element classification module for classifying the page elements on the page file according to the trained logistic regression model.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer is executed by a processor, it implements a page element classification method as described in any one of claims 1-6.

9. A computer device, characterized in that: It includes at least one processor and a memory communicatively connected to the processor, where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to cause the processor to execute a page element classification method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Page element information extraction method and device, equipment and storage medium

    CN111310044A

  • Webpage element classification method

    CN112507186A