Data extraction method and device for webpage elements, electronic equipment and storage medium
By analyzing web page code information to obtain style and animation data, and filtering based on the web page elements selected by the user, the problem of inflexible web page element extraction in the existing technology is solved, and efficient and user-friendly data extraction effect is achieved.
Patent Information
- Application Number
- CN202311520619.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-15
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to flexibly extract web page elements and their related information based on user needs, affecting the user experience.
By detecting the extraction operation of the target web page element, the code information of the target web page is obtained, and parsed into a lexical tree to determine the first style data and the first animation data. Then, the second style data is determined from the first style data based on the target web page element, and the second animation data is determined from the first animation data based on the second style data, and finally output the second style data and the second animation data in the target format.
It realizes web element data extraction that does not rely on browser developer tools or extensions, simplifies the data extraction process, and improves data extraction efficiency and user experience.
Smart Images

Figure CN120011669A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of computer vision technology and web page automated testing technology, and in particular to a method, device, electronic device and storage medium for extracting data from web page elements. Background Art
[0002] With the continuous development of computer technology, users can check various materials through browsers, search and view various commodity information through shopping applications, etc. Application software can interact with users through interactive pages, and interactive pages include web pages. Web pages include various elements, such as text, pictures, audio, animation, video, etc.
[0003] In the process of implementing the inventive concept of the present disclosure, the inventors found that the related technology has at least the following technical problems: the extraction of web page elements depends on the browser's own development tools or extensions, and it is difficult to flexibly extract web page elements and their related information based on user needs, which affects the user experience. Summary of the invention
[0004] In view of the above problems, the present disclosure provides a method, device, electronic device and storage medium for extracting data of web page elements.
[0005] According to a first aspect of the present disclosure, a method for extracting data from a web page element is provided, comprising:
[0006] In response to detecting an extraction operation on a target webpage element in a target webpage, acquiring code information of the target webpage;
[0007] Determine first style data and first dynamic effect data of the target webpage according to the code information, wherein the first dynamic effect data is used to realize the dynamic effect of the webpage element;
[0008] Determining second style data from the first style data according to the target web page element;
[0009] Determining second motion effect data from the first motion effect data based on the second style data; and
[0010] The second style data and the second animation data are output in a target format.
[0011] According to an embodiment of the present disclosure, determining first style data and first motion effect data of a target webpage according to code information includes:
[0012] Parsing the code information into a lexical tree, wherein the lexical tree includes M first nodes and N second nodes, the first nodes include grammar rule nodes, the second nodes include declaration grammar rule nodes, N is a positive integer, and M is a positive integer;
[0013] Based on the M first nodes, determining first pattern data; and
[0014] Based on the N second nodes, first motion effect data is determined.
[0015] According to an embodiment of the present disclosure, determining the first style data based on the M first nodes includes:
[0016] Determine a selector and style definition information corresponding to each first node, wherein the selector is used to control the style of the web page element; and
[0017] The selector and the style definition information are processed into first style data of a target data structure.
[0018] According to an embodiment of the present disclosure, determining the selector and style definition information corresponding to each first node includes:
[0019] Determining a selector corresponding to each first node according to the selector attribute of the first node; and
[0020] The first sub-function of the first node is called to obtain the style definition information of each first node.
[0021] According to an embodiment of the present disclosure, determining the first motion effect data based on the N second nodes includes:
[0022] Based on the identification name of the second node, L second nodes including the dynamic effect keyword are screened out from the N second nodes, where L is less than or equal to N and is a positive integer;
[0023] According to the identification parameter value of the second node, obtaining the first animation name of L second nodes, wherein the first animation name is the animation name of the webpage element in the target webpage;
[0024] Calling the second sub-function of the second node to convert the first motion effect rule definitions of the L second nodes into second motion effect rule definitions in a string form; and
[0025] The first animation name and the second animation rule definition are processed into first animation data of the target data structure.
[0026] According to an embodiment of the present disclosure, the first style data includes a selector and style definition information related to a target web page element and its sub-elements, the sub-elements represent elements located inside the target web page element in the target web page, and the selectors are used to control the styles of the web page element and the sub-elements;
[0027] Determining second style data from first style data according to a target web page element includes:
[0028] Obtaining a sub-element list of the target webpage element, wherein the sub-element list includes E sub-elements, where E is an integer; and
[0029] According to the target webpage element and the E sub-elements, the style data of the valid selector is screened out from the first style data to obtain the second style data, wherein the valid selector represents the selector related to the target webpage element or sub-element in the first style data.
[0030] According to an embodiment of the present disclosure, the style definition information in the second style data includes style attributes and style attribute values, the style attributes include dynamic effect style attributes, and the first dynamic effect data includes a second dynamic effect rule definition;
[0031] Determining second motion effect data from the first motion effect data based on the second style data includes:
[0032] Determine a style attribute value of the animation style attribute from the second style data;
[0033] According to the style attribute value, a second animation name is obtained by parsing, wherein the second animation name is an animation name of a target web page element and its sub-elements;
[0034] Based on the second animation name, a third animation rule definition is obtained from the second animation rule definition.
[0035] According to an embodiment of the present disclosure, outputting the second style data and the second motion effect data in a target format includes:
[0036] Convert the second style data and the second dynamic effect data into third style data and third dynamic effect data in the form of cascading style sheets respectively; and
[0037] The formatting processing function is called to format the third style data and the third motion effect data, and the formatted third style data and the third motion effect data are output.
[0038] According to an embodiment of the present disclosure, it also includes:
[0039] Based on the webpage attributes of the target webpage element, acquiring code information in the form of hypertext markup language, wherein the code information in the form of hypertext markup language includes codes related to the target webpage element; and
[0040] The formatting processing function is called to format the code information in the form of hypertext markup language, and the formatted code information in the form of hypertext markup language is output.
[0041] According to an embodiment of the present disclosure, obtaining code information of a target webpage includes:
[0042] Obtain code information based on the style tags of the target web page; and / or
[0043] At least one style file in the form of a cascading style sheet is obtained according to a request address related to a resource request operation of a target webpage; and code information is obtained from the at least one style file.
[0044] A second aspect of the present disclosure provides a data extraction device for web page elements, comprising:
[0045] an acquisition module, configured to acquire code information of a target webpage in response to detecting an extraction operation on a target webpage element in the target webpage;
[0046] A determination module, used to determine first style data and first dynamic effect data of a target web page according to the code information, wherein the first dynamic effect data is used to realize a dynamic effect of a web page element;
[0047] A style determination module, used to determine second style data from the first style data according to the target web page element;
[0048] A motion effect determination module, configured to determine second motion effect data from the first motion effect data based on the second style data; and
[0049] The output module is used to output the second style data and the second motion effect data in a target format.
[0050] A third aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the above-mentioned data extraction method for web page elements.
[0051] A fourth aspect of the present disclosure further provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the above-mentioned method for extracting data from web page elements.
[0052] The fifth aspect of the present disclosure further provides a computer program product, including a computer program, which implements the above-mentioned data extraction method for web page elements when executed by a processor.
[0053] In the embodiments of the present disclosure, since the first style data and the first dynamic effect data of the target web page are determined from the code information of the target web page, and the code information is obtained from the original resource file, the acquisition operation of the first style data and the second dynamic effect data does not rely on the browser's developer tools or extensions, and can flexibly obtain style data and dynamic effect data from the original resource file, simplifying the data extraction process, and improving data extraction efficiency and user experience. In addition, in the process of extracting style data and dynamic effect data, after extracting the second style data from the first style data based on the target web page elements, the first dynamic effect data is filtered based on the filtered second style data, and there is no need to re-traverse the code information during the dynamic effect data screening process, thereby reducing the amount of data processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The above contents and other purposes, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0055] Figure 1 The application scenario of the data extraction method of the web page element according to the embodiment of the present disclosure is schematically shown;
[0056] Figure 2 A flowchart of a method for extracting data from a web page element according to an embodiment of the present disclosure is schematically shown;
[0057] Figure 3 A data flow diagram of a method for extracting data from a web page element according to an embodiment of the present disclosure is schematically shown;
[0058] Figure 4 A flowchart for determining first style data and first motion effect data according to code information according to a specific embodiment of the present disclosure is schematically shown;
[0059] Figure 5 A schematic diagram schematically shows a method of determining second style data from first style data according to an embodiment of the present disclosure;
[0060] Figure 6 A schematic diagram of determining second motion effect data from first motion effect data according to an embodiment of the present disclosure is schematically shown;
[0061] Figure 7 A structural block diagram schematically shows a device for extracting data from web page elements according to an embodiment of the present disclosure; and
[0062] Figure 8 A block diagram of an electronic device suitable for a method for extracting data from web page elements according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0063] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present disclosure. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0064] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise", "include", etc. used herein indicate the existence of the features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.
[0065] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.
[0066] When using expressions such as "at least one of A, B, and C, etc.", they should generally be interpreted according to the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0067] In the technical solution of the present invention, the user information (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of relevant data comply with the relevant laws, regulations and standards of relevant countries and regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0068] In the related art, an element in a web page is generally extracted based on the browser's own developer tools or extensions. For example, by opening the "developer tools" in the browser, the web page content and the rendered code can be displayed in different display frames of the same web page; the user can select an element in the web page and obtain the style information of the element from the rendered code. Alternatively, the user can also export the style information of the selected element through browser extensions, such as CSSViewer, CSS Peeper, etc.
[0069] However, the "developer tools" and extensions can only extract the style information of a certain element, and cannot extract the style information of multiple elements in batches, which requires users to repeatedly perform the selection-export operation, resulting in complex data extraction operations for web page elements, low efficiency, and poor user experience. In addition, the "developer tools" and extensions can only extract the style information of the selected element, and cannot extract other information related to the element, such as animation effect information (dynamic effect information), and cannot extract the style information and / or animation effect information of multiple elements in batches, which results in complex data extraction operations for web page elements, low efficiency, and poor user experience.
[0070] In order to at least partially solve the above technical problems, an embodiment of the present disclosure provides a data extraction method for a web page element, comprising: in response to detecting an extraction operation for a target web page element in a target web page, obtaining code information of the target web page; determining first style data and first animation data of the target web page according to the code information, wherein the first animation data is used to achieve a dynamic effect of the web page element; determining second style data from the first style data according to the target web page element; determining second animation data from the first animation data based on the second style data; and outputting the second style data and the second animation data in a target format.
[0071] Figure 1 The application scenario of the method for extracting data from web page elements according to an embodiment of the present disclosure is schematically illustrated.
[0072] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or optical fiber cables, etc.
[0073] The user may use at least one of the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0074] For example, a user uses one of the first terminal device 101, the second terminal device 102, and the third terminal device 103 to browse web pages, and performs operations such as shopping, searching, instant messaging, sending and receiving emails, browsing social platforms, etc. on the browser. The server 105 sends or receives information to one of the first terminal device 101, the second terminal device 102, and the third terminal device 103 through the network 104 to support operations such as shopping, searching, instant messaging, sending and receiving emails, browsing social platforms, etc.
[0075] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0076] The server 105 may be a server that provides various services, such as a background management server (only as an example) that provides support for websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process the received data such as user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0077] It should be noted that the data extraction method of the web page element provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the data extraction device of the web page element provided in the embodiment of the present disclosure can generally be set in the server 105. The data extraction method of the web page element provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the data extraction device of the web page element provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.
[0078] It should be understood that Figure 1The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to the implementation requirements.
[0079] The following will be based on Figure 1 The scene described by Figure 2 to Figure 6 The data extraction method of the web page element of the disclosed embodiment is described in detail.
[0080] Figure 2 The flowchart of the method for extracting data from web page elements according to an embodiment of the present disclosure is schematically shown.
[0081] like Figure 2 As shown, the method 200 includes operations S210 to S250.
[0082] In operation S210, in response to detecting an extraction operation on a target web page element in a target web page, code information of the target web page is acquired.
[0083] In operation S220, first style data and first dynamic effect data of the target webpage are determined according to the code information, wherein the first dynamic effect data is used to realize a dynamic effect of the webpage element.
[0084] In operation S230, second style data is determined from the first style data according to the target web page element.
[0085] In operation S240 , second animation data is determined from the first animation data based on the second style data.
[0086] In operation S250, the second style data and the second motion effect data are output in a target format.
[0087] According to an embodiment of the present disclosure, the target webpage includes the currently displayed webpage. The target webpage includes various types of webpages displayed in the browser application, such as a login webpage of a shopping application, a product details webpage, a login webpage of a social platform, a browsing webpage, etc.
[0088] According to an embodiment of the present disclosure, the target web page includes various types of web page elements, such as icon elements, text elements, animation elements, web page link elements, background elements, text box elements, area boundary elements, etc. The target web page element may be a web page element selected by a user in the target web page.
[0089] According to an embodiment of the present disclosure, the target webpage element may be an element representing a certain area in the target webpage.
[0090] For example, the search box element includes the search box itself, and also includes the text element "search" inside the search box, the search icon element, the blank input box element, etc., wherein the text element "search" inside the search box, the search icon element, and the blank input box element can be understood as sub-elements of the search box element. Alternatively, the target web page element is a recommendation theme element, and the recommendation body element also includes multiple recommended picture elements and price elements of products.
[0091] According to an embodiment of the present disclosure, the target webpage element in the target webpage can be determined by a user clicking on the target webpage. The extraction operation for the target webpage element in the target webpage includes: after determining the target element, the user clicks on the "Export" button; or, the extraction operation for the target webpage element can also be realized by operations such as determining in a prompt box, dragging the target element to a target area, etc.
[0092] According to an embodiment of the present disclosure, the code information is obtained from the original resource file of the target webpage. For example, the code information may be code obtained from an original resource file in Cascading Style Sheets (CSS) format, such as an original resource file with a suffix of .css.
[0093] According to an embodiment of the present disclosure, the code information includes style data, dynamic effect data and other data of multiple web page elements in the target web page. Among them, the style data is used to define the size, position, color and other information of the web page elements, and the dynamic effect data is used to define the dynamic effect of the implemented web page elements, such as "floating appearance", "background music" and other information. Other data includes data related to the specific content and execution logic of the web page elements.
[0094] According to an embodiment of the present disclosure, first style data and first animation data related to the style and animation can be obtained from the code information by distinguishing keywords, distinguishing the control mode of web page elements, and the like.
[0095] For example, the first style data and the first motion effect data may be obtained from the code information by analyzing the implementation process of the code information, such as a syntax tree, a lexical tree, etc.
[0096] According to an embodiment of the present disclosure, after obtaining the first style data and the first animation data from the code information of the target web page, the second style data related to the target web page element and its sub-elements can be selected from the first style data based on the target web page element selected by the user; then, based on the second style data, the second animation data can be determined from the first animation data.
[0097] According to an embodiment of the present disclosure, since a screening operation is performed on a target web page element during the screening process of the first style data, the first motion effect data can be screened based on the filtered second style data during the screening process of the first motion effect data, without having to re-traverse the code data or the target web page element during the screening process of the motion effect data, thereby reducing the amount of data processing.
[0098] According to an embodiment of the present disclosure, based on the second style data, determining the second animation data from the first animation data includes: parsing the second style data to obtain relevant information of the target web page element and its sub-elements; and based on the relevant information, determining the second animation data from the first animation data. The relevant information includes the animation name of the target web page element and its sub-elements.
[0099] According to an embodiment of the present disclosure, after determining the second style data and the second dynamic effect data, the second style data and the second dynamic effect data can be output in a target format. Specifically, the target format can be the same data format as the code information, such as css format.
[0100] According to an embodiment of the present disclosure, the target format may also be a data format that is the same as the code information after the second style data and the second animation data are formatted.
[0101] In the embodiments of the present disclosure, since the first style data and the first dynamic effect data of the target web page are determined from the code information of the target web page, and the code information is obtained from the original resource file, the acquisition operation of the first style data and the second dynamic effect data does not rely on the browser's developer tools or extensions, and can flexibly obtain style data and dynamic effect data from the original resource file, simplifying the data extraction process, and improving data extraction efficiency and user experience. In addition, in the process of extracting style data and dynamic effect data, after extracting the second style data from the first style data based on the target web page elements, the first dynamic effect data is filtered based on the filtered second style data, and there is no need to re-traverse the code information during the dynamic effect data screening process, thereby reducing the amount of data processing.
[0102] Figure 3 The data flow diagram of the method for extracting data from web page elements according to an embodiment of the present disclosure is schematically shown.
[0103] like Figure 3 As shown, for the data extraction method 300 of the web page element, in the dimension of the target web page, the code information 301 of the target web page is extracted; and the first style data 303 and the first dynamic effect data 304 are determined from the code information 301 of the target web page.
[0104] Regarding the style, second style data 305 related to the target web page element and its sub-elements may be selected from the first style data 303 based on the target web page element 302 selected by the user.
[0105] Regarding the dynamic effect, the second dynamic effect data 306 may be determined from the first dynamic effect data 304 based on the filtered second style data 305 .
[0106] According to an embodiment of the present disclosure, first style data and first animation data of a target web page are determined based on code information, including: parsing the code information into a lexical tree, wherein the lexical tree includes M first nodes and N second nodes, the first node includes a grammar rule node, the second node includes a declaration grammar rule node, N is a positive integer, and M is a positive integer; based on the M first nodes, the first style data is determined; and based on the N second nodes, the first animation data is determined.
[0107] According to an embodiment of the present disclosure, a lexical analysis tool may be used to parse code information and parse the code information into a lexical tree, wherein the lexical analysis tool includes analysis tools developed using a variety of programming languages, such as Ada, C, Java, and other programming languages.
[0108] According to an embodiment of the present disclosure, a pre-built code database may be used to parse code information into a lexical tree. For example, a postcss library may be used to parse all style codes into a lexical tree.
[0109] According to an embodiment of the present disclosure, the first node includes a grammar rule node, rule node, and the second node includes a declaration grammar rule node, atrule node. The rule node includes all CSS rules in the selector, that is, the style part of the selector; the atrule node implements multiple definitions through the identifier name in "@identifier name + identifier parameter value", for example, charset is used to prompt the string encoding method used by the CSS file, import is used to introduce the CSS file, and keyframes generates a kind of data for defining animation keyframes.
[0110] According to an embodiment of the present disclosure, by traversing M rule nodes in a lexical tree and parsing all rule nodes, the first style data including all web page elements in a target web page can be obtained. By traversing N atrule nodes in a lexical tree and parsing all atrule nodes, the first dynamic effect data including all web page elements in a target web page can be obtained.
[0111] According to an embodiment of the present disclosure, two threads may be created simultaneously to traverse the rule node and the atrule node respectively, so as to simultaneously obtain the first style data and the first motion effect data.
[0112] In an embodiment of the present disclosure, by parsing the code information into a lexical tree and performing targeted analysis on different types of nodes in the lexical tree, the first style data and the first animation data can be obtained at the same time, thereby improving the data extraction efficiency and the flexibility of extracting style data and animation data from the code information.
[0113] According to an embodiment of the present disclosure, first style data is determined based on M first nodes, including: determining a selector and style definition information corresponding to each first node, wherein the selector is used to control the style of a web page element; and processing the selector and the style definition information into first style data of a target data structure.
[0114] According to an embodiment of the present disclosure, each rule node may correspond to at least one selector, and each selector controls the style of at least one web page element. The selector is stored in the code information in the form of an array. In the process of traversing M rule nodes, the selector corresponding to each node is determined by obtaining all selector arrays corresponding to each rule node.
[0115] According to an embodiment of the present disclosure, the rule node also includes a decl subtype, and at least one style definition information corresponding to each rule node is obtained by traversing all decl subtypes under each node.
[0116] According to an embodiment of the present disclosure, the style definition information includes a style attribute prop and a style attribute value value.
[0117] According to an embodiment of the present disclosure, after obtaining the selectors and style definition information of M rule nodes, the style definition information of all rule nodes is sorted out in units of selectors to obtain the first style data. The first style data is a full object, and the full selector object is subsequently processed without processing the local selector object of each rule node.
[0118] For example, when only selector A is included in the first style data, the first style data may be: {selector A: {style definition A prop: style definition A value; style definition B prop: style definition B value; style definition C prop: style definition C value}}. Selector A is used to control three styles of a web page element, such as style A, style B, and style C.
[0119] According to an embodiment of the present disclosure, determining the selector and style definition information corresponding to each first node includes: determining the selector corresponding to each first node according to the selector attribute of the first node; and calling the first sub-function of the first node to obtain the style definition information of each first node.
[0120] According to an embodiment of the present disclosure, a rule node includes a selector attribute, and the selector attribute is used to determine at least one selector corresponding to the rule node. According to an embodiment of the present disclosure, the selector attribute may include a name or an identifier of the selector. The selector attribute may be a selectors attribute.
[0121] According to an embodiment of the present disclosure, the first sub-function is used to traverse the decl sub-type under the rule node and obtain at least one style definition information corresponding to each rule node.
[0122] For example, call the walkRules function of the lexical tree to traverse all rule nodes, and then obtain all selector arrays corresponding to the node through the selectors attribute of the rule node; at the same time, through the walkDecls function of the rule node, that is, the first subfunction, traverse all style definition information. Among them, walkDecls can obtain the prop and value of each style declaration.
[0123] In an embodiment of the present disclosure, according to the selector attribute of the first node, the selector corresponding to the first node is determined by traversing the first node, and the decl subtype under the first node is traversed by calling the first subfunction to obtain the style definition information. The first style data can be obtained based on the attribute parameters of the rule node itself without the need for complicated data acquisition operations, thereby improving the efficiency of the first style data extraction operation.
[0124] According to an embodiment of the present disclosure, first animation data is determined based on N second nodes, including: based on the identification name of the second node, L second nodes including animation keywords are screened out from the N second nodes, wherein L is less than or equal to N and L is a positive integer; according to the identification parameter value of the second node, the first animation name of the L second nodes is obtained, wherein the first animation name is the animation name of the web page element in the target web page; the second sub-function of the second node is called to convert the first animation rule definition of the L second nodes into the second animation rule definition in the form of a string; and the first animation name and the second animation rule definition are processed into the first animation data of the target data structure.
[0125] According to an embodiment of the present disclosure, since the second node can define different forms of rules for web page elements based on the identification name, L nodes whose identification names include the animation keyword "keyframes" can be filtered out from N atrule nodes, thereby filtering out other second nodes that are not related to the animation.
[0126] According to an embodiment of the present disclosure, after filtering out L atrule nodes whose identification names include the animation keyword "keyframes", the animation names related to the L atrule nodes, that is, the first animation name, can be obtained according to the identification parameter values params of the L atrule nodes.
[0127] According to an embodiment of the present disclosure, the second sub-function may be a walkRules function, which is used to convert all first motion rule definitions of the atrule node into a motion definition string, that is, a second motion rule definition in string form.
[0128] According to an embodiment of the present disclosure, after determining the first animation name and the second animation rule definition, the first animation name and the second animation rule definition can be assembled into first animation data of a target data structure using the first animation name as a unit.
[0129] For example, the first animation data including only one animation name A may be: {Animation name A: [Animation definition A, animation definition B, ...]}. Animation definition A and animation definition B are two animation rule definitions for web page elements.
[0130] In the embodiment of the present disclosure, firstly, L second nodes related to the animation are selected from N second nodes by the animation keyword, and then the animation name and the animation rule definition are parsed based on the identification parameter values of the L second nodes, and the first animation data including the animation name and the specific animation definition is obtained. The first animation data is obtained based on the identification parameters of the atrule node itself, without the need for complex data acquisition operations, thereby improving the efficiency of the first animation data extraction operation.
[0131] For ease of understanding, a specific embodiment is used below to describe the process of determining the styles and animations of all web page elements in a target web page.
[0132] Figure 4 A flowchart for determining first style data and first motion effect data based on code information according to a specific embodiment of the present disclosure is schematically shown.
[0133] like Figure 4 As shown, a method 400 for determining first style data and first motion effect data according to code information is performed by parsing code information 401 of a target web page to obtain a lexical tree 402 .
[0134] The lexical tree 402 includes a plurality of first nodes 403 and a plurality of second nodes 404. The first nodes 403 and the second nodes 404 have different code forms and different functions. The first nodes 403 are related to the style, and the second nodes 404 are related to the animation.
[0135] By traversing all first nodes 403 in the lexical tree 402, a selector 405 for controlling multiple web page elements in the target web page is obtained based on the selector attribute of the first node 403, and style definition information 406 for implementing a specific style. Afterwards, the selector 405 and the style definition information 406 are combined to obtain the first style data 409.
[0136] By traversing all second nodes 404 in the lexical tree 402, the second nodes 404 are first filtered based on their identification names, and then the identification parameter values of the filtered second nodes are parsed to obtain the animation names 407 of all web page elements in the target web page; at the same time, the animation rule definition 408 is obtained by calling the second sub-function for parsing. Afterwards, the first animation data 410 can be obtained by combining the animation name 407 and the animation rule definition 408. Figure 4 The middle animation rule definition 408 is a second animation rule definition in the form of a character string.
[0137] According to an embodiment of the present disclosure, since a web page element may include sub-elements, the first style data of the target web page also includes selectors and style definition information related to the target web page element and its sub-elements.
[0138] According to an embodiment of the present disclosure, second style data is determined from first style data based on a target web page element, including: obtaining a sub-element list of the target web page element, wherein the sub-element list includes E sub-elements, and E is an integer; and based on the target web page element and the E sub-elements, filtering out style data of a valid selector from the first style data to obtain second style data, wherein the valid selector represents a selector in the first style data that is related to the target web page element or sub-element.
[0139] According to an embodiment of the present disclosure, the sub-element list is used to store sub-element information of the target web page element. The target web page element may include at least one sub-element at the same time, and the at least one sub-element may be a sub-element of multiple types, such as an animation element, a sound effect element, etc. The target web page element may also not include a sub-element, in which case the sub-element list is empty.
[0140] According to an embodiment of the present disclosure, a list of sub-elements in a target web page element may be obtained by calling a querySelectorAll('*') method of the target web page element.
[0141] According to an embodiment of the present disclosure, the first style data includes style data of all web page elements in the target web page. Therefore, after determining the target web page elements and their sub-element list, the selectors in the first style data can be filtered based on the target web page elements and their E sub-elements, and finally the style data of the valid selector, i.e., the second style data, is obtained.
[0142] For example, the first style data may include two selectors at the same time, namely, selector A and selector B. The first style data is: {selector A: {style definition A prop: style definition Avalue; style definition B prop: style definition B value; style definition C prop: style definition C value}, selector B: {…}}.
[0143] Selector B is not used to control the style of the target web page element and its E child elements, but is used to control the style of other web page elements in the target web page. Therefore, the second style data obtained after filtering the first style data only retains selector A and its style definition information. That is, the second style data only includes the style data of the valid selector A, and does not include the style data of the invalid selector B.
[0144] According to an embodiment of the present disclosure, the second style data is in a different form from the first style data. The second style data may be in the form of a key-value pair. Still taking the style data of the filtered valid selector A as an example, the second style data is: [{key: selector A, data:{style definition A prop: style definition A value; style definition B prop: style definition B value; style definition C prop: style definition C value}}].
[0145] According to a specific embodiment of the present disclosure, based on the target web page element and E sub-elements, the style data of the valid selector is filtered out from the first style data, and the second style data can be obtained by: traversing the target web page element and E sub-elements, and recording the currently traversed element as element X; while traversing the first style data, using document.querySelectorAll (selector) to obtain all elements matched by the current selector. Then, it is determined whether element X is in the matched element list. If element X is in the matched list, it means that the selector is a valid selector.
[0146] In the embodiment of the present disclosure, the first style data including the selector is screened according to the target web page element and its sub-elements, which can not only screen out all style data related to multiple elements such as the target web page element and its sub-elements, but also screen out other irrelevant element styles in the target web page, thereby reducing the amount of data exported for the second style data, and realizing batch export of style data of multiple elements, thereby improving data extraction efficiency. In addition, since the data extraction of the target web page element and its associated sub-elements can be completed through the extraction operation of the target web page element, the data usage scenarios are expanded, user needs are better met, and user experience is improved.
[0147] According to an embodiment of the present disclosure, the style definition information in the second style data includes style attributes and style attribute values, the style attributes include dynamic effect style attributes, and the first dynamic effect data includes a second dynamic effect rule definition.
[0148] For example, the animation style attributes included in the style attribute prop may be animation and animation-name.
[0149] According to an embodiment of the present disclosure, second animation data is determined from first animation data based on second style data, including: determining style attribute values of animation style attributes from the second style data; parsing to obtain a second animation name based on the style attribute values, wherein the second animation name is an animation name of a target web page element and its sub-elements; and obtaining a third animation rule definition from the second animation rule definition based on the second animation name.
[0150] According to an embodiment of the present disclosure, determining the style attribute value of the dynamic style attribute from the second style data includes: traversing all style definition information in the second style data, and filtering out a plurality of style definition information whose style attribute is "dynamic style attribute". Thereafter, obtaining the corresponding style attribute value from the plurality of style definition information whose style attribute is "dynamic style attribute".
[0151] For example, the second style data is traversed to determine the style definition information whose style attribute prop is animation and animation-name; then, the value of the style definition information whose style attribute prop is animation and animation-name is obtained.
[0152] According to an embodiment of the present disclosure, parsing the second animation name according to the style attribute value includes: parsing the second animation name from the style attribute value by character string splitting.
[0153] According to an embodiment of the present disclosure, since the second animation name filtered out from the second style data is the animation name of the target web page element and its sub-elements, and the first animation data includes the animation name, there is no need to re-filter the first animation data based on the target web page element and its sub-elements. The first animation data can be filtered based on the second animation name to obtain the second animation data related to the target web page element and its sub-elements.
[0154] In the embodiments of the present disclosure, since the process of filtering the second animation data from the first animation data only requires the second animation name in the second style data, there is no need to re-traverse the code information or re-process the target web page elements and their sub-elements during the animation data filtering process, which simplifies the animation filtering process and improves data extraction efficiency.
[0155] According to an embodiment of the present disclosure, outputting second style data and second animation data in a target format includes: converting the second style data and the second animation data into third style data and third animation data in the form of cascading style sheets, respectively; and calling a formatting processing function to format the third style data and the third animation data, and outputting the formatted third style data and the third animation data.
[0156] According to an embodiment of the present disclosure, the second style data and the second animation data may be in a data storage format of key-value pairs.
[0157] According to the embodiment of the present disclosure, since the code information is in the form of CSS, in the data extraction process, the second style data and the second dynamic effect data are processed into the form of key-data for the convenience of screening. Therefore, before outputting the second style data and the second dynamic effect data, it is also necessary to convert the second style data and the second dynamic effect data into the third style data and the third dynamic effect data in the form of cascading style sheets, format the third style data and the third dynamic effect data, and output the formatted third style data and the third dynamic effect data.
[0158] According to an embodiment of the present disclosure, the third style data in the form of CSS may be: selector {style definition prop: style definition value; style definition prop: style definition value; ...}, and the third animation data in the form of CSS may be: @keyframes animation name {animation definition; animation definition; ...}.
[0159] According to an embodiment of the present disclosure, the formatting processing function may be a js-beautify function.
[0160] According to an embodiment of the present disclosure, the third style data and the third motion effect data in the formatted CSS form may be output and displayed to the user in a pop-up window, so that the user can copy and export the third style data and the third motion effect data from the pop-up window.
[0161] According to an embodiment of the present disclosure, the third style data and the third motion effect data may also be directly exported in the form of a resource file.
[0162] The third style data and the third dynamic effect data output in the embodiment of the present disclosure are CSS codes, which are data obtained from the original resource file and the style code file, and are not the rendered code directly obtained based on the developer tool, nor are they the element style data obtained through getComputedStyle. Therefore, the embodiment of the present disclosure can flexibly export the styles of multiple elements and the dynamic effect data associated with the elements from the original resource file.
[0163] Figure 5The following schematic diagram schematically illustrates a method of determining second style data from first style data according to an embodiment of the present disclosure.
[0164] like Figure 5 As shown, in the style processing scenario 500, the first style data 501 may be: {selector A: {style definition A prop: style definition A value; style definition B prop: style definition B value; style definition C prop: style definition C value; ...}, selector B: {...}, ...}.
[0165] The second style data 502 may be: [{key: selector A, data: {style definition A prop: style definition A value; style definition B prop: style definition B value; style definition C prop: style definition C value; ...}}, {key: selector B, data: {...}}, ...].
[0166] Figure 6 The following schematic diagram shows a method for determining second motion effect data from first motion effect data according to an embodiment of the present disclosure.
[0167] like Figure 6 As shown, in the motion effect processing scene 600, the first motion effect data 601 may be: {motion effect name A: [motion effect definition A, motion effect definition B, ...], motion effect name B: [...], ...}.
[0168] The second animation data 602 may be: [{key: animation name A, data: {animation definition A, animation definition b, ...}}, {key: animation name B, data: {...}}, ...].
[0169] According to an embodiment of the present disclosure, the data extraction method of the web page element also includes: obtaining code information in the form of hypertext markup language based on the web page attributes of the target web page element, wherein the code information in the form of hypertext markup language includes code related to the target web page element; and calling a formatting processing function to format the code information in the form of hypertext markup language, and outputting the formatted code information in the form of hypertext markup language.
[0170] According to an embodiment of the present disclosure, while processing the second style data and the second motion effect data, the target web page code in the form of Hypertext Markup Language (HTML) can also be displayed to the user.
[0171] According to an embodiment of the present disclosure, all HTML codes including the target web page element itself can be obtained through the outerHTML attribute of the target web page element; then, the js-beautify function is called to format and output the obtained HTML code.
[0172] According to an embodiment of the present disclosure, the third style data, the third animation data and the HTML code may be output simultaneously.
[0173] According to an embodiment of the present disclosure, obtaining code information of a target web page includes: obtaining code information based on a style tag of the target web page; and / or obtaining at least one style file in the form of a cascading style sheet based on a request address related to a resource request operation of the target web page; and obtaining code information from at least one style file.
[0174] According to an embodiment of the present disclosure, all style tags of the target web page can be traversed through the injected Javascript script, and the style code in the style tag can be obtained. Alternatively, all resource file request URLs recorded can be traversed through the injected Javascript script, and all style files with the .css suffix can be filtered out; then, the style code in the file can be downloaded through the URL request. Thus, by obtaining the style code in the style tag and / or downloading the style code through the URL request, the code information can be obtained from the original resource file.
[0175] In the embodiments of the present disclosure, code information can be obtained in various forms through style tags and / or through URL requests, so that the embodiments of the present disclosure are applicable to various forms of web pages, thereby improving the wide applicability of the solution.
[0176] According to an embodiment of the present disclosure, operations S210 to S250 can be implemented by a developed element style dynamic effect extraction tool. The following will use a specific embodiment to describe in detail the process of obtaining and outputting style data and dynamic effect data by using the element style dynamic effect extraction tool.
[0177] First, before using the Element Style Animation Extraction Tool, access the browser and install the "Element Style Animation Extraction Tool" by "loading the unpacked extension". After successful installation, enable the tool and access any web page. When the browser loads a web page, the tool will inject the Javascript script integrated in the tool into the web page; at the same time, it will record all resource file requests initiated by the current web page and record the URL of the corresponding resource file.
[0178] In the process of extracting the animation data and style data for the target web page elements in the target web page, the user can browse any web page and select the elements in the web page whose style and animation they want to export, and click the export button provided by the "Element Style and Animation Extraction Tool". After the user clicks, the "Element Style and Animation Extraction Tool" will record the currently selected element as element A.
[0179] In the stage of obtaining code information, the "Element Style Animation Extraction Tool" can traverse all style tags of the current web page through the injected Javascript script to obtain the style code in the tag, and then traverse all recorded resource file request URLs, filter out all style files with the .css suffix, and download and obtain the style code in the file through URL request.
[0180] The postcss library has been integrated into the "Element Style Animation Extraction Tool". Through the integrated postcss library, all style codes are parsed into lexical trees, all parsed lexical trees are traversed, and walkRules of the corresponding lexical tree is called to traverse all rule nodes. Then, the selectors attribute of the rule node is used to obtain the array of all selectors corresponding to the node. At the same time, all style declaration definitions are traversed through the walkDecls method of the rule node. The prop and value of each style declaration can be obtained in walkDecls. Finally, the node selector, style declaration prop and style declaration value are combined to assemble and store them as a full selector object B.
[0181] By calling the walkAtRules method of the corresponding lexical tree, we can traverse the atRule nodes and filter out all atRule nodes whose names do not contain the keyframes keyword. Then we get the animation name through the params of the node and call the walkRules method of atRule to convert all the animation rule definitions into animation definition strings. Finally, we combine the animation name and animation definition to assemble the following data result and store it as the animation object C.
[0182] The "Element Style Animation Extraction Tool" can also send the parsed full selector object B and animation object C to the script object injected into the web page, and then call the querySelectorAll('*') method of element A to obtain a list of all sub-elements in element A. It will traverse element A and all sub-element lists obtained in turn, and each time it traverses, it will record the currently traversed element as element X. At the same time, it will traverse the full selector object B once, use document.querySelectorAll(selector) to obtain all elements matched by the current selector, and then determine whether element X is in the list of matched elements. If element X is in the matching list, it means that the selector is a valid selector and will be written into the selector list D.
[0183] Traverse all style definitions whose props are animation and animation-name in the selector list D, get the corresponding style definition value, and then parse the animation name by string splitting. Then get the corresponding animation definition data from the animation object C based on the animation name and store it as the animation list E.
[0184] Before displaying, use the outerHTML attribute of element A to obtain all HTML codes of the selected element A, including itself. Traverse the selector list D and restore the selector list D to CSS code; traverse the animation list E and restore the animation list E to CSS code. Finally, use js-beautify to format the HTML and CSS finally obtained and generated, and then use a pop-up box to display it for easy selection, copying and exporting, completing the data extraction operation of web page elements.
[0185] The embodiment of the present disclosure obtains the original style definition code and parses the style code into a lexical tree, traverses the lexical tree to find the style of the target element and obtains the corresponding dynamic effect definition based on the style parsing, thereby exporting the style and dynamic effect definition of the web page element and all its child elements. It is possible to flexibly obtain style data and dynamic effect data from the original resource file, simplify the data extraction process, and improve data extraction efficiency and user experience.
[0186] Figure 7 The structural block diagram of the data extraction device for web page elements according to an embodiment of the present disclosure is schematically shown.
[0187] like Figure 7 As shown, the data extraction device 700 of the web page element of this embodiment includes an acquisition module 710, a determination module 720, a style determination module 730, a motion effect determination module 740 and an output module 750.
[0188] The acquisition module 710 is used to acquire code information of the target webpage in response to detecting an extraction operation on a target webpage element in the target webpage.
[0189] The determination module 720 is used to determine the first style data and the first dynamic effect data of the target web page according to the code information, wherein the first dynamic effect data is used to realize the dynamic effect of the web page element.
[0190] The style determination module 730 is used to determine the second style data from the first style data according to the target web page element.
[0191] The motion effect determination module 740 is used to determine the second motion effect data from the first motion effect data based on the second style data.
[0192] The output module 750 is used to output the second style data and the second motion effect data in a target format.
[0193] According to an embodiment of the present disclosure, the determination module 720 includes a parsing submodule, a first determination submodule and a second determination submodule.
[0194] The parsing submodule is used to parse the code information into a lexical tree, wherein the lexical tree includes M first nodes and N second nodes, the first nodes include grammar rule nodes, the second nodes include declaration grammar rule nodes, N is a positive integer, and M is a positive integer.
[0195] The first determination submodule is used to determine first style data based on M first nodes.
[0196] The second determination submodule is used to determine the first motion effect data based on the N second nodes.
[0197] According to an embodiment of the present disclosure, the first determining submodule includes a first determining unit and a second determining unit.
[0198] The first determining unit is used to determine the selector and style definition information corresponding to each first node, wherein the selector is used to control the style of the web page element.
[0199] The second determining unit is used to process the selector and the style definition information into first style data of a target data structure.
[0200] According to an embodiment of the present disclosure, the first determining unit includes a determining subunit and an acquiring subunit.
[0201] The determination subunit is used to determine the selector corresponding to each first node according to the selector attribute of the first node.
[0202] The acquisition subunit is used to call the first subfunction of the first node to acquire the style definition information of each first node.
[0203] According to an embodiment of the present disclosure, the second determining submodule includes a third determining unit, a fourth determining unit, a fifth determining unit, and a sixth determining unit.
[0204] The third determining unit is used to filter out L second nodes including the animation keyword from N second nodes based on the identification name of the second node, where L is less than or equal to N and is a positive integer.
[0205] The fourth determining unit is used to obtain the first animation name of L second nodes according to the identification parameter value of the second node, wherein the first animation name is the animation name of the web page element in the target web page.
[0206] The fifth determining unit is used to call the second sub-function of the second node to convert the first motion effect rule definitions of the L second nodes into second motion effect rule definitions in the form of a character string.
[0207] The sixth determining unit is used to process the first animation name and the second animation rule definition into first animation data of the target data structure.
[0208] According to an embodiment of the present disclosure, the first style data includes selectors and style definition information related to a target web page element and its sub-elements. The sub-elements represent elements located inside the target web page element in the target web page. The selectors are used to control the styles of the web page elements and sub-elements.
[0209] According to an embodiment of the present disclosure, the style determination module 730 includes a sub-element acquisition sub-module and a screening sub-module.
[0210] The sub-element acquisition sub-module is used to acquire a sub-element list of the target web page element, wherein the sub-element list includes E sub-elements, where E is an integer.
[0211] The screening submodule is used to screen out style data of valid selectors from the first style data according to the target web page element and E sub-elements to obtain second style data, wherein the valid selector represents the selector in the first style data related to the target web page element or sub-element.
[0212] According to an embodiment of the present disclosure, the style definition information in the second style data includes style attributes and style attribute values, the style attributes include dynamic effect style attributes, and the first dynamic effect data includes a second dynamic effect rule definition.
[0213] According to an embodiment of the present disclosure, the motion effect determination module 740 includes an attribute value acquisition submodule, a motion effect name acquisition submodule and a motion effect acquisition submodule.
[0214] The attribute value acquisition submodule is used to determine the style attribute value of the dynamic style attribute from the second style data.
[0215] The animation name acquisition submodule is used to parse and obtain a second animation name according to the style attribute value, wherein the second animation name is the animation name of the target web page element and its sub-elements.
[0216] The motion effect acquisition submodule is used to acquire the third motion effect rule definition from the second motion effect rule definition based on the second motion effect name.
[0217] According to an embodiment of the present disclosure, the output module 750 includes a conversion submodule and a first formatting submodule.
[0218] The conversion submodule is used to convert the second style data and the second motion effect data into third style data and third motion effect data in the form of cascading style sheets respectively.
[0219] The first formatting submodule is used to call the formatting processing function to format the third style data and the third motion effect data, and output the formatted third style data and the third motion effect data.
[0220] According to an embodiment of the present disclosure, the output module 750 further includes a hypertext code acquisition submodule and a second formatting submodule.
[0221] The hypertext code acquisition submodule is used to acquire code information in the form of hypertext markup language based on the webpage attributes of the target webpage element, wherein the code information in the form of hypertext markup language includes codes related to the target webpage element.
[0222] The second formatting submodule is used to call the formatting processing function to format the code information in the form of hypertext markup language, and output the formatted code information in the form of hypertext markup language.
[0223] According to an embodiment of the present disclosure, the acquisition module 710 includes a first acquisition submodule and / or a second acquisition submodule.
[0224] The first acquisition submodule is used to acquire code information based on the style tag of the target web page.
[0225] The second acquisition submodule is used to acquire at least one style file in the form of a cascading style sheet according to a request address related to a resource request operation of a target webpage; and acquire code information from the at least one style file.
[0226] According to an embodiment of the present disclosure, any multiple modules among the acquisition module 710, the determination module 720, the style determination module 730, the motion effect determination module 740 and the output module 750 can be combined into one module for implementation, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module.
[0227] According to an embodiment of the present disclosure, at least one of the acquisition module 710, the determination module 720, the style determination module 730, the motion effect determination module 740, and the output module 750 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or by any one of the three implementation methods of software, hardware, and firmware, or by a suitable combination of any of them. Alternatively, at least one of the acquisition module 710, the determination module 720, the style determination module 730, the motion effect determination module 740, and the output module 750 may be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding function may be executed.
[0228] Figure 8 A block diagram of an electronic device suitable for a method for extracting data from web page elements according to an embodiment of the present disclosure is schematically shown.
[0229] like Figure 8 As shown, the electronic device 800 according to an embodiment of the present disclosure includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage part 808 to a random access memory (RAM) 803. The processor 801 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 801 may also include an onboard memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0230] In RAM 803, various programs and data required for the operation of electronic device 800 are stored. Processor 801, ROM 802 and RAM 803 are connected to each other via bus 804. Processor 801 performs various operations of the method flow according to the embodiment of the present disclosure by executing the program in ROM 802 and / or RAM 803. It should be noted that the program can also be stored in one or more memories other than ROM 802 and RAM 803. Processor 801 can also perform various operations of the method flow according to the embodiment of the present disclosure by executing the program stored in the one or more memories.
[0231] According to an embodiment of the present disclosure, the electronic device 800 may further include an input / output (I / O) interface 805, which is also connected to the bus 804. The electronic device 800 may further include one or more of the following components connected to the input / output I / O interface 805: an input portion 806 including a keyboard, a mouse, etc.; an output portion 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 808 including a hard disk, etc.; and a communication portion 809 including a network interface card such as a LAN card, a modem, etc. The communication portion 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed, so that a computer program read therefrom is installed into the storage portion 808 as needed.
[0232] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present disclosure is implemented.
[0233] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, an apparatus or a device. For example, according to an embodiment of the present disclosure, a computer-readable storage medium may include the ROM 802 and / or RAM 803 described above and / or one or more memories other than ROM 802 and RAM 803.
[0234] The embodiment of the present disclosure also includes a computer program product, which includes a computer program, and the computer program contains program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the above method provided by the embodiment of the present disclosure.
[0235] The above functions defined in the system / device of the embodiment of the present disclosure are executed when the computer program is executed by the processor 801. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0236] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of signals on a network medium, and downloaded and installed through the communication part 809, and / or installed from a removable medium 811. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0237] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 809, and / or installed from the removable medium 811. When the computer program is executed by the processor 801, the above functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the system, device, means, module, unit, etc. described above can be implemented by a computer program module.
[0238] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level process and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, Java, C++, python, "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on the remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect through the Internet).
[0239] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0240] It will be appreciated by those skilled in the art that the features described in the various embodiments and / or claims of the present disclosure may be combined and / or combined in a variety of ways, even if such combinations and / or combinations are not explicitly described in the present disclosure. In particular, the features described in the various embodiments and / or claims of the present disclosure may be combined and / or combined in a variety of ways without departing from the spirit and teachings of the present disclosure. All of these combinations and / or combinations fall within the scope of the present disclosure.
[0241] The specific embodiments described above further illustrate the purpose, technical solutions and beneficial effects of the present disclosure. It should be understood that the above description is only a specific embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present disclosure should be included in the protection scope of the present disclosure.
Claims
1. A method for extracting data from a web page element, comprising: In response to detecting an extraction operation on a target webpage element in a target webpage, acquiring code information of the target webpage; Determine first style data and first dynamic effect data of the target webpage according to the code information, wherein the first dynamic effect data is used to realize a dynamic effect of a webpage element; Determining second style data from the first style data according to the target web page element; Based on the second style data, determining second motion effect data from the first motion effect data; and The second style data and the second animation data are output in a target format.
2. The method according to claim 1, wherein: Determining the first style data and the first animation data of the target webpage according to the code information includes: Parsing the code information into a lexical tree, wherein the lexical tree includes M first nodes and N second nodes, the first nodes include grammar rule nodes, the second nodes include declaration grammar rule nodes, N is a positive integer, and M is a positive integer; Determine the first style data based on the M first nodes; and The first motion effect data is determined based on the N second nodes.
3. The method according to claim 2, wherein: The determining the first style data based on the M first nodes includes: Determining a selector and style definition information corresponding to each of the first nodes, wherein the selector is used to control the style of the web page element; and The selector and the style definition information are processed into first style data of a target data structure.
4. The method according to claim 3, wherein: The determining the selector and style definition information corresponding to each of the first nodes includes: Determining a selector corresponding to each of the first nodes according to the selector attributes of the first nodes; and The first sub-function of the first node is called to obtain the style definition information of each of the first nodes.
5. The method according to claim 2, wherein: The determining the first motion effect data based on the N second nodes includes: Based on the identification name of the second node, L second nodes including the animation keyword are screened out from the N second nodes, where L is less than or equal to N and is a positive integer; According to the identification parameter value of the second node, obtaining the first animation names of L second nodes, wherein the first animation name is the animation name of the webpage element in the target webpage; Calling the second sub-function of the second node to convert the first motion effect rule definitions of the L second nodes into second motion effect rule definitions in a string form; and The first animation name and the second animation rule definition are processed into first animation data of a target data structure.
6. The method according to any one of claims 1 to 5, wherein: The first style data includes selectors and style definition information related to the target web page element and its sub-elements, wherein the sub-elements represent elements located inside the target web page element in the target web page, and the selectors are used to control the styles of the web page element and the sub-elements; The determining the second style data from the first style data according to the target webpage element includes: Obtaining a sub-element list of the target webpage element, wherein the sub-element list includes E sub-elements, where E is an integer; and According to the target web page element and the E sub-elements, style data of valid selectors are screened out from the first style data to obtain second style data, wherein the valid selector represents a selector in the first style data related to the target web page element or the sub-element.
7. The method according to any one of claims 1 to 5, wherein: The style definition information in the second style data includes style attributes and style attribute values, the style attributes include dynamic style attributes, and the first dynamic data includes a second dynamic rule definition; The determining second motion effect data from the first motion effect data based on the second style data includes: Determine a style attribute value of a dynamic style attribute from the second style data; According to the style attribute value, a second animation name is obtained by parsing, wherein the second animation name is the animation name of the target web page element and its sub-elements; Based on the second animation effect name, a third animation effect rule definition is obtained from the second animation effect rule definition.
8. The method according to any one of claims 1 to 5, wherein: The outputting the second style data and the second motion effect data in a target format includes: Convert the second style data and the second animation data into third style data and third animation data in the form of cascading style sheets, respectively; and The formatting processing function is called to format the third style data and the third motion effect data, and the formatted third style data and the third motion effect data are output.
9. The method according to claim 8, further comprising: Based on the webpage attribute of the target webpage element, acquiring code information in the form of hypertext markup language, wherein the code information in the form of hypertext markup language includes a code related to the target webpage element; and The formatting processing function is called to format the code information in the form of hypertext markup language, and the formatted code information in the form of hypertext markup language is output.
10. The method according to any one of claims 1 to 5, wherein: The step of obtaining the code information of the target webpage includes: Based on the style tag of the target webpage, obtaining the code information; and / or At least one style file in the form of a cascading style sheet is obtained according to a request address related to a resource request operation of the target webpage; and the code information is obtained from the at least one style file.
11. A data extraction device for web page elements, comprising: an acquisition module, configured to acquire code information of a target webpage in response to detecting an extraction operation on a target webpage element in the target webpage; A determination module, configured to determine first style data and first dynamic effect data of the target webpage according to the code information, wherein the first dynamic effect data is used to realize a dynamic effect of a webpage element; A style determination module, configured to determine second style data from the first style data according to the target web page element; a motion effect determination module, configured to determine second motion effect data from the first motion effect data based on the second style data; and An output module is used to output the second style data and the second motion effect data in a target format.
12. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are enabled to execute the method according to any one of claims 1 to 10.
13. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the method according to any one of claims 1 to 10.
14. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.