Page data extraction method, page automated testing method
By obtaining and converting the xpath relative path of the tree data structure of the web page, the problem of low accuracy of manual extraction is solved, and the accuracy and coverage of automated testing is improved, and different page designs are adapted to different page designs.
Patent Information
- Application Number
- CN202211085150.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-08-29
- Filing Date
- 2022-09-06
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-09-06
AI Technical Summary
In the prior art, the accuracy of manually extracting tree-like content from web pages to tree-shaped data structures is low, and it is impossible to read tree-shaped data structures directly from the html level, resulting in insufficient accuracy and coverage of automated tests.
By obtaining the xpath relative path of the root node of the tree data structure, using the recursive call method to obtain the xpath relative path list, and convert it into a dictionary tree data structure, automatic page data extraction and testing are realized.
It improves the accuracy of page data extraction and the coverage of automated testing, can adapt to changes in different page design styles, and achieves the objectivity and completeness of automated testing.
Smart Images

Figure CN115795193B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer network technologies, and particularly to a method for extracting page data and a method for page testing. Background Art
[0002] In web page design, data in a tree-like data structure often is also displayed in a similar tree-like form. For example, LADP (Lightweight Directory Access Protocol) user organizational structures, corporate organization plans, forum directories, etc. In the field of automated testing of computer networks, it is necessary to read the tree-like content displayed on a web page, convert the read content into a corresponding tree-like data structure, and then compare it with the source data to test whether the content displayed on the web page is correct.
[0003] However, the tree-like content displayed in a web page, from the perspective of the html layer, is not equivalent to the tree-like data structure in the sense of a data structure. For example, in the html layer, there will be no corresponding pointers from a parent node to a child node, nor will there be a method to jump from a certain child node to its sibling node. Therefore, the corresponding tree-like data structure cannot be directly read from the html layer.
[0004] Currently, the manual method is adopted to extract the tree-like content displayed in a web page in the form of a tree-like data structure and then perform subsequent testing. However, the manual extraction method has low accuracy. Summary of the Invention
[0005] In order to solve the problem of low accuracy in subsequent testing when manually extracting the tree-like content displayed in a web page in the form of a tree-like data structure, this application provides a method for extracting page data and a method for automated page testing through the following aspects.
[0006] The first aspect of this application provides a method for extracting page data. The method for extracting page data is used to extract the tree-like data structure data corresponding to the tree-like content displayed on a web page; the method for extracting page data includes:
[0007] Obtain the xpath relative path of the root node in the tree-like data structure in the page to be tested;
[0008] Use the xpath relative path of the root node as the current xpath relative path, and use a preset recursive call method to obtain a list of xpath relative paths, where the list of xpath relative paths includes the xpath relative paths of each node arranged in depth-first order in the tree-like data structure;
[0009] Convert the list of xpath relative paths into a dictionary to obtain tree-structured data in dictionary form.
[0010] In some embodiments, a preset recursive call method includes:
[0011] Determine whether the current node exists in the corresponding record dictionary, where the current node is the node corresponding to the current xpath relative path, and the record dictionary is used to record the traversed nodes and the corresponding first quantity and second quantity, where the first quantity is the number of child nodes included in the target node, and the second quantity is the number of already recorded child nodes corresponding to the target node, and the target node is any one of the traversed nodes;
[0012] If the current node exists in the record dictionary, determine whether the first quantity of the current node is equal to the second quantity;
[0013] If the first quantity of the current node is not equal to the second quantity, obtain the xpath relative paths of the child nodes of the current node according to the current xpath relative path;
[0014] Record the xpath relative paths of the child nodes of the current node in the list of xpath relative paths, and increment the second quantity corresponding to the current node by 1;
[0015] Use the xpath relative paths of the child nodes of the current node as the current xpath relative path and re-execute the preset recursive call method;
[0016] If the first quantity of the current node is equal to the second quantity, determine whether the current node has a younger brother node;
[0017] If the current node has a younger brother node, obtain the xpath relative path of the younger brother node of the current node according to the current xpath relative path;
[0018] Record the xpath relative path of the younger brother node of the current node in the list of xpath relative paths, and increment the second quantity corresponding to the parent node of the current node by 1;
[0019] Use the xpath relative path of the younger brother node of the current node as the current xpath relative path and re-execute the preset recursive call method;
[0020] If the current node does not have a younger brother node, obtain the xpath relative path of the parent node of the current node according to the current xpath relative path;
[0021] Use the xpath relative path of the parent node of the current node as the current xpath relative path and re-execute the preset recursive call method.
[0022] In some embodiments, if the current node does not exist in the record dictionary, the first quantity of the current node is obtained according to the current xpath relative path;
[0023] Initialize the second quantity of the current node to 0;
[0024] Record the current node and the corresponding first quantity and second quantity in the record dictionary, and continue to execute the step of judging whether the first quantity is equal to the second quantity.
[0025] In some embodiments, if the current node exists in the record dictionary, before judging whether the first quantity is equal to the second quantity, the page data extraction method further includes:
[0026] Judge whether the current xpath relative path is the xpath relative path of the root node;
[0027] If the current xpath relative path is the xpath relative path of the root node, terminate the execution of the recursive call method;
[0028] If the current xpath relative path is not the xpath relative path of the root node, continue to execute the preset recursive call method.
[0029] In some embodiments, obtaining the xpath relative path of the child node of the current node according to the current xpath relative path includes:
[0030] Add the xpath offset value of the child node to the current xpath relative path to obtain the xpath relative path of the child node of the current node.
[0031] In some embodiments, obtaining the xpath relative path of the sibling node of the current node according to the current xpath relative path includes:
[0032] Delete the xpath offset value of the current node on the current xpath relative path to obtain the xpath relative path of the parent node of the current node;
[0033] Add the xpath offset value of the sibling node to the xpath relative path of the parent node of the current node to obtain the xpath relative path of the sibling node of the current node.
[0034] In some embodiments, obtaining the xpath relative path of the parent node of the current node according to the current xpath relative path includes:
[0035] Delete the xpath offset value of the current node on the current xpath relative path to obtain the xpath relative path of the parent node of the current node.
[0036] In some embodiments, converting a list of xpath relative paths into a dictionary to obtain tree-structured data in dictionary form includes:
[0037] Extracting the absolute path string of each node from the list of xpath relative paths;
[0038] Obtaining tree-structured data in dictionary form based on the absolute path string of each node.
[0039] A second aspect of the present application provides a page automation testing method, including: obtaining tree-structured data corresponding to tree-like content displayed on a page to be tested according to a page data extraction method provided in the first aspect of the present application, to obtain a test dictionary;
[0040] Performing page testing according to the test dictionary.
[0041] A second aspect of the present application provides a terminal device, including: at least one processor and a memory;
[0042] The memory is used to store program instructions;
[0043] The processor is used to call and execute the program instructions stored in the memory, so that the terminal device executes the page data extraction method as described in the first aspect of the present application.
[0044] The present application provides a page data extraction method and a page automation testing method. The page data extraction method is used to extract tree-structured data corresponding to tree-like content displayed on a web page, including: obtaining the xpath relative path of the root node in the tree-structured data of the page to be tested; using the root node's xpath relative path as the current xpath relative path, and using a preset recursive call method to obtain a list of xpath relative paths, where the list of xpath relative paths includes the xpath relative paths of each node arranged in depth-first order in the tree-structured data; converting the list of xpath relative paths into a dictionary. The page extraction method can automatically extract the tree-structured data corresponding to the tree-like content and return it in dictionary form, which is convenient for subsequent page automation testing. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a schematic diagram showing an example of the display of the user organization structure on the web page of a certain firewall device;
[0046] Figure 2 It is a schematic diagram showing an example of the display of the user organization structure on the web page of a certain firewall device;
[0047] Figure 3It is a schematic structural diagram of a tree - shaped data structure corresponding to tree - like content on a web page;
[0048] Figure 4 It is a schematic diagram of the working process of a page data reading method provided by an embodiment of the present application;
[0049] Figure 5 It is an html code example corresponding to tree - like content displayed on a web page;
[0050] Figure 6 It is a schematic diagram of the working process of a preset call recursive method in a page data reading method provided by an embodiment of the present application;
[0051] Figure 7 It is a schematic diagram of the function call relationship corresponding to the page data extraction method provided by an embodiment of the present application;
[0052] Figure 8 It is a schematic diagram of the execution effect of test script 1;
[0053] Figure 9 It is a schematic diagram of the execution effect of test script 2. Detailed implementation manners
[0054] Next, the technical solutions in the embodiments of the present application will be described in conjunction with the accompanying drawings in the embodiments of the present application. The terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to limit the present application.
[0055] To facilitate the description of the technical solutions of the application, some concepts related to the present application will be described first below.
[0056] See Figure 1 and Figure 2 , which is a display example of the user organization structure on the web page of a certain firewall device. Figure 1 Each "-" icon in Figure 2 can be clicked. When clicked, the content it contains will gather. After that, the "-" becomes "+", as shown in Figure 3 . Correspondingly, when each "+" icon is clicked, the part it contains will be expanded. See Figure 1-2 , which is a schematic diagram of the tree - shaped data structure read from the web page shown in
[0057] A dictionary is an unordered, mutable, and indexed collection. In Python, dictionaries are written with curly braces and have keys and values. Each key-value pair (key: value) in a dictionary is separated by a colon :, and two key-value pairs are separated by a comma ,, and the entire dictionary is enclosed in curly braces {}, in the following format: d = {key1: value1, key2: value2}. Exemplarily, as Figure 3 shown in the tree data structure, it can be represented using the following dictionary:
[0058]
[0059]
[0060] Among them, each node in the tree data structure serves as a key in the dictionary. If a node has no subsequent content, then the value corresponding to this key is an empty dictionary {}. Through this dictionary, the tree structure on the page can be presented completely, thus meeting the requirements of automated testing.
[0061] The embodiment of this application provides a page data extraction method for extracting the tree data structure data corresponding to the tree-like content displayed on a web page. See Figure 4 , the page data extraction method includes steps 101 - step 103.
[0062] Step 101, obtain the relative xpath of the root node in the tree data structure of the page to be tested. In this application, the tree data structure in the page to be tested can be a complete tree in the page to be tested, or a subtree in a complete tree in the page to be tested. Correspondingly, if the tree data structure to be extracted corresponds to a complete tree, then the root node is the root node of this complete tree; if the tree data structure to be extracted corresponds to a subtree, then the root node is the root node of this subtree. Exemplarily, the relative xpath of the root node is an input parameter, which can be obtained by the user from the html code and passed into the code program corresponding to the method; or it can be obtained by the program according to a preset rule and passed into the code program corresponding to the method.
[0063] Step 102, use the relative xpath of the root node as the current relative xpath, and use a preset recursive call method to obtain a list of relative xpaths, where the list of relative xpaths includes the relative xpaths of each node in the tree data structure arranged in depth-first order.
[0064] Step 103, convert the list of relative xpaths into a dictionary to obtain the tree data structure data in dictionary form.
[0065] See Figure 5 , which is Figure 1 an HTML code example corresponding to the class tree content displayed on the web page shown. It can be analyzed from Figure 5 the HTML code example shown that, at the HTML level, the above class tree content is roughly organized in the form of li-ul-li. Among them, in Figure 5 the HTML code example shown, the following rules can be summarized:
[0066] First, each node in the tree data structure can correspond to an li tag in HTML.
[0067] Second, if a node contains child nodes, all child nodes of this node are mounted in the ul tag under this node.
[0068] Third, the name of the node, such as " / ", "default group", "g1", "g11", etc., is not saved in a certain attribute of the li tag, but in the title attribute of the a hung under the li tag; therefore, in the way of / a[@title = ' / '] / .., first find the a tag carrying the node name downward, and then return the li tag upward to get the relative xpath of the node name.
[0069] Fourth, at the HTML code level, the names of nodes can be the same at different levels. If you want to uniquely determine a certain node, you cannot locate it only by its own name, but need a complete path to determine. Through this rule, any node can be found through the relative xpath. Exemplarily, the relative xpath corresponding to the node / g1 / g11 / g112 is:
[0070] / / ul[@id = 'areaTree'] / li / a[@title = ' / '] / .. / ul / li / a[@title = 'g1'] / .. / ul / li / a[@title = 'g11'] / ... / ul / li / a[@title = 'g112'] / ..
[0071] Correspondingly, for any node, its relative xpath for positioning can be obtained by finding the names of its ancestor nodes and then splicing them into a relative xpath. Exemplarily, abstracting the previous example, if a node path is root-nodeA-nodeB-nodeC, then its relative xpath in Figure 5 the HTML code style shown is:
[0072] / / ul[@id='areaTree'] / li / a[@title=root] / .. / ul / li / a[@title=A] / .. / ul / li / a[@title=B] / ... / ul / li / a[@title=C] / ..
[0073] Fifth, for each node except the root node, relative to its parent node, its relative path is obtained by adding the xpath offset value of the node to the xpath relative path of the parent node.
[0074] In the present application, the xpath offset value of a certain node represents the difference between the xpath relative paths between this node and its parent node. Depending on different design styles, the xpath offset values between a node and its child nodes are also different. Exemplarily, in Figure 5 the html code example shown, “ / ul / li / a[@title='g1'] / ..” is called the xpath offset value of node g1.
[0075] Exemplarily, in another web page design style, the xpath offset value of a certain node is " / / ul / li / a / span[text()='{node_name}'] / .. / ..". According to different web page design styles, only the xpath offset value of the node needs to be updated to extract the tree-like content in different web pages.
[0076] It can be seen from the above rules that a node can be located through the xpath relative path of the node, and the relationship between the xpath relative paths of the node and its parent node can be found. When the xpath relative path of a certain node is known, the xpath relative paths of its parent node, sibling nodes, and child nodes can be found.
[0077] In this way, through the page data extraction method shown in steps 101 - 103, using the xpath relative path of the node as the recursive parameter, and using the preset recursive call method, each node is traversed in depth-first order to obtain a list of xpath relative paths, and then the list of xpath relative paths is converted into a dictionary to obtain the tree-like data structure data in dictionary form.
[0078] In some embodiments, referring to Figure 6 , the preset recursive call method includes:
[0079] Step 201, determine whether the current node exists in the corresponding record dictionary, where the current node is the node corresponding to the current xpath relative path, and the record dictionary is used to record the traversed nodes and the corresponding first quantity and second quantity. The first quantity is the number of child nodes contained in the target node, and the second quantity is the number of child nodes corresponding to the target node that have been recorded. The target node is any one of the traversed nodes.
[0080] Step 202, if the current node exists in the record dictionary, determine whether the first quantity of the current node is equal to the second quantity.
[0081] Step 203, if the first quantity of the current node is not equal to the second quantity, obtain the xpath relative path of the child nodes of the current node according to the current xpath relative path.
[0082] In a possible implementation manner, add the xpath offset value of the child node to the current xpath relative path to obtain the xpath relative path of the child node of the current node. Since the depth-first order is adopted to traverse the nodes in the preset recursive call method in this embodiment, in Step 203, the child node of the current node specifically refers to the first child node of the current node.
[0083] For the convenience of description, the current node is represented by node A, and the current xpath relative path is represented by xpath_A.
[0084] Exemplarily, to calculate the xpath relative path of the first child node of node A, the idea is: find the name of the first child node, and then add the xpath offset value of the child node to the current node's xpath relative path; it can be implemented using the following python code:
[0085] title=browser.driver.find_element_by_xpath(f"{xpath} / ul / li[{n}] / a").get_attribute("title")
[0086] result=xpath+f" / ul / li / a[@title='{title}'] / .."
[0087] Step 204, record the xpath relative path of the child node of the current node into the xpath relative path list, and add 1 to the second quantity corresponding to the current node.
[0088] Step 205: Use the xpath relative path of the child node of the current node as the current xpath relative path, and re - execute the preset recursive call method.
[0089] Step 206: If the first quantity of the current node is equal to the second quantity, determine whether there is a younger brother node of the current node.
[0090] Exemplarily, according to xpath_A, try to find another adjacent li tag. If this li tag exists, then node A has a younger brother node. The following Python code can be used to implement it:
[0091]
[0092] Step 207: If the current node has a younger brother node, obtain the xpath relative path of the younger brother node of the current node according to the current xpath relative path.
[0093] In one implementation, the obtaining the xpath relative path of the younger brother node of the current node according to the current xpath relative path includes:
[0094] Delete the xpath offset value of the current node on the current xpath relative path to obtain the xpath relative path of the parent node of the current node; add the xpath offset value of the younger brother node to the xpath relative path of the parent node of the current node to obtain the xpath relative path of the younger brother node of the current node. Exemplarily, the following Python code can be used to implement it:
[0095]
[0096] Step 208: Record the xpath relative path of the younger brother node of the current node into the xpath relative path list, and increment the second quantity corresponding to the parent node of the current node by 1.
[0097] Step 209: Use the xpath relative path of the younger brother node of the current node as the current xpath relative path, and re - execute the preset recursive call method.
[0098] Step 210: If the current node does not have a younger brother node, obtain the xpath relative path of the parent node of the current node according to the current xpath relative path.
[0099] In one implementation, obtaining the relative XPath of the parent node of the current node based on the current relative XPath includes: deleting the XPath offset value of the current node on the current relative XPath to obtain the relative XPath of the parent node of the current node. Exemplarily, the following Python code can be used for implementation:
[0100]
[0101]
[0102] Step 211, use the relative XPath of the parent node of the current node as the current relative XPath, and re-execute the preset recursive call method.
[0103] In one implementation, if the current node does not exist in the record dictionary, the preset recursive call method further includes Step 212: obtaining the first quantity of the current node according to the current relative XPath; initializing the second quantity of the current node to 0; recording the current node and the corresponding first and second quantities into the record dictionary, and continue to execute Step 202.
[0104] In Figure 5 In the shown HTML code style, to obtain the number of child nodes of the current node A according to XPath_A, it can be obtained by the number of tags ul / li under the relative XPath of the current node A. Exemplarily, the following Python code can be used for implementation:
[0105] result = len(browser.driver.find_elements_by_xpath(f"{xpath_A} / ul / li"))
[0106] To prevent re-accessing node A after moving to the parent node of node A and thus entering an infinite loop. In the preset recursive call method provided in this embodiment, a record dictionary is introduced to record the traversed nodes and the corresponding first and second quantities. For node A, when node A is first accessed, calculate the first quantity of node A, initialize the second quantity of node A to 0, and record node A and the corresponding first and second quantities into the dictionary. When moving from node A to its child nodes or from one child node of A to another child node of A, the second quantity of node A will be incremented by 1 to update the second quantity. When the first quantity of node A is equal to the second quantity, it means that there are no more child nodes to be accessed under node A, and thus it will not enter the child nodes of node A again, thereby preventing the occurrence of an infinite loop.
[0107] The preset recursive call method provided in Step 201 - Step 211 processes the traversal process of a tree - like structure in a depth - first order. For each node in the tree - like structure, the processing process of Step 201 - Step 211 is executed.
[0108] Exemplarily, as Figure 3 shown in the tree - like data structure, using the preset recursive call method, its traversal order is as follows: " / --default group--g1--g11--g111--g1111--g1112--g111--g11--g112--g11--g12--g13--g1--g2--g11--g111--g11--g12--g21--g22--g222--g22--g2-- / ". It can be seen from the above traversal order that nodes in non - terminal positions are visited more than once. However, it is specified in the above - mentioned preset recursive call method that only during the process of moving from parent to child or from sibling to sibling, the relative xpath of the node will be recorded, so there will be no situation of repeated traversal. If it is the process of moving from child to parent (such as Step 210 - Step 211), the corresponding relative xpath will not be recorded.
[0109] In terms of code implementation, the preset recursive call method provided in the above - mentioned embodiment is actually the recursive function calling itself. And the recursive function always has a condition to jump out and will not recurse infinitely. Thus, in some embodiments, if the current node exists in the record dictionary, before determining whether the first quantity is equal to the second quantity, the page data extraction method further includes:
[0110] Step 213, determining whether the current relative xpath is the relative xpath of the root node. From Figure 3 the traversal order corresponding to the shown tree - like data structure, it can be seen that the traversal starts from the root node and ends at the root node. Because when starting the traversal, the record dictionary is empty and Step 212 will not be executed. Only when all traversals are completed and return to the root node again, the preset recursive call method jumps out, terminates the call, and returns the list of relative xpaths.
[0111] If the current relative xpath is the relative xpath of the root node, then terminate the execution of the recursive call method.
[0112] If the current relative xpath is not the relative xpath of the root node, then continue to execute the preset recursive call method.
[0113] Exemplarily, part of the content of the relative xpath list is as follows:
[0115] " / / ul[@id='areaTree'] / li[1] / ul / li / a[@title='g1'] / .. / ul / li / a[@title='g11'] / .."
[0116] " / / ul[@id='areaTree'] / li[1] / ul / li / a[@title='g1'] / .. / ul / li / a[@title='g11'] / .. / ul / li / a[@title='g111'] / ..",
[0117] " / / ul[@id='areaTree'] / li[1] / ul / li / a[@title='g1'] / .. / ul / li / a[@title='g11'] / .. / ul / li / a[@title='g111'] / .. / u l / li / a[@title='g1111'] / ..",
[0118] " / / ul[@id='areaTree'] / li[1] / ul / li / a[@title='g1'] / .. / ul / li / a[@title='g12'] / .."
[0120] Convert the relative XPath path into a dictionary to obtain tree-structured data in dictionary form, which may include step 301 and step 302.
[0121] Step 301: Extract the absolute path string of each node from the relative XPath path list.
[0122] Step 302: Obtain tree-structured data in dictionary form according to the absolute path string of each node.
[0123] In one implementation, use regular expressions to extract the name of the node from the XPath string of each path entry in the relative XPath path list to obtain the absolute path string of each node. Then convert it into a list. Exemplarily, the list corresponding to part of the above relative XPath path list is as follows:
[0124]
[0125] The above embodiments provide a method for extracting page data, which is used to extract the tree - structured data corresponding to the tree - like content displayed on a web page. The page data extraction method includes: obtaining the relative xpath of the root node in the tree - structured data in the page to be tested; using the relative xpath of the root node as the current relative xpath and applying a preset recursive call method to obtain a list of relative xpaths, where the list of relative xpaths includes the relative xpaths of each node in the tree - structured data arranged in depth - first order; converting the list of relative xpaths into a dictionary to obtain the tree - structured data in dictionary form. The page extraction method can automatically extract the tree - structured data corresponding to the tree - like content and return it in the form of a dictionary, which is convenient for subsequent page automated testing. Further, when the structure of the page to be tested changes, only the xpath offset value of the node needs to be adjusted to quickly adapt to the updated web page.
[0126] Regardless of how different the tree - like content presented on the page is at the html level, as long as the relationship between nodes follows a certain pattern, the corresponding tree - structured data can be parsed from the web page by the page data extraction method provided in the above embodiments.
[0127] The embodiments of the present application provide a page automated testing method. According to the page data extraction method provided in the foregoing embodiments, obtain the tree - structured data corresponding to the tree - like content displayed on the page to be tested to obtain a test dictionary; perform page testing according to the test dictionary.
[0128] Using the page automated testing method provided in this embodiment is more objective, accurate, and has a more complete test coverage compared to manual testing.
[0129] The embodiments of the present application also provide a terminal device. The terminal device includes: at least one processor and a memory. The memory is used to store program instructions; the processor is used to call and execute the program instructions stored in the memory so that the terminal executes the corresponding content in the embodiments of the page data extraction method.
[0130] To better illustrate the working process of the page data extraction method and the page automated testing method provided in the embodiments of the present application, the following combines specific python code to describe in detail the implementation process of the software device for extracting tree - structured data based on Figure 3 the web page design style shown. Even for other different web page design styles, as long as the setting of the xpath offset value is adjusted, the same solution can be reused. It should be noted that other programming languages such as Java or C can also be used to implement the corresponding software device.
[0131] See Figure 7 , which is a schematic diagram of the function call relationship corresponding to the page data extraction method. The code examples corresponding to steps 101 - 103 are as follows:
[0132]
[0133] Among them, the main function parse_user_structure_tree_locally provides four input parameters: root, root_xpath, display_trace, and output_in_dict. The parameter root allows the user to set different root nodes. The parameter root_xpath is the relative xpath of the root node. The parameter display_trace determines whether to print out the traversal process. When the parameter output_in_dict is True, the result is returned as a dictionary. When it is False, it is returned as a list of absolute paths. In the main function, a variable xpath_history is initialized. This variable will be passed as an external variable to the recursive function parse_user_structure_tree_locally to store the relative xpath of the nodes gradually traversed during the recursive process. When the recursive process is completed, the variable xpath_history will save the relative xpath of all nodes.
[0134] The value of the parameter output_in_dict being True means that it is desired to return the final tree structure in the form of a dictionary, while the value of output_in_dict being False means that it is desired to return the final tree structure in the form of a list. According to whether the value of the parameter output_in_dict is True or False, the function convert_list_to_dict or convert_xpath_to_path will be called in the main function.
[0135] The convert_xpath_to_path function is used to extract the characters identifying the node names from each xpath string in the list of relative xpaths to form a list of absolute paths. The implementation code of the convert_xpath_to_path function is as follows:
[0136]
[0137] The code corresponding to the preset call recursive method for steps 201 - 212 is as follows:
[0138]
[0139]
[0140] Among them, the recursive function path_finder provides multiple parameters. The xpath represents the relative xpath of the current node and is the most important parameter in the recursive function path_finder. The entire traversal process is completed by passing in the relative xpath of each level of node. The xpath_history passes the external variable xpath_history. In the main function parse_user_structure_tree_locally, the external variable xpath_history is initialized to an empty list []. Whenever moving to a new child node or a sibling node, the corresponding relative xpath is recorded in the external variable xpath_history. The xpath_child_number is a dictionary that stores the number of child nodes of each node and the number of accessed child nodes. In the recursive function path_finder, if the current node has not yet established a record related to its child nodes, a new key will be added to this dictionary and the relevant data will be recorded. Whenever moving to an unvisited child node and sibling node, the value of the accessed child nodes will be updated. The root_xpath is a passed-in parameter and is the relative xpath of the root node of the tree-shaped data structure to be extracted currently, which is used to determine whether to jump out of the current recursive logic.
[0141] The cal_child_numbers function is used to obtain the first quantity of the current node. The cal_child_numbers function takes the relative xpath of the current node as an input and returns the number of child nodes. If the current node has no child nodes, it returns 0. The implementation code of the cal_child_numbers function is as follows:
[0142]
[0143] The jump_to_child function is used to obtain the relative xpath of the eldest child node of the current node. The jump_to_child function takes the relative xpath of the current node as an input and returns the relative xpath of its eldest child node. The implementation code of the jump_to_child function is as follows:
[0144]
[0145]
[0146] The jump_to_neighbour function is used to calculate the relative XPath of the sibling node of the current node. The jump_to_neighbour function takes the relative XPath of the current node as input and returns the XPath of its sibling node. The implementation code of the jump_to_neighbour function is as follows:
[0147]
[0148] The jump_to_parent function is used to calculate the relative XPath of the parent node of the current node. The jump_to_parent takes the relative XPath of the current node as input and returns the relative XPath of the parent node of the current node. The implementation code of the jump_to_parent function is as follows:
[0149]
[0150] Call the above main function parse_user_structure_tree_locally to obtain the class tree content configured on the device page.
[0151] Test script 1 returns the tree data structure data in the form of a dictionary. The implementation code of test script 1 is as follows:
[0152]
[0153] Among them, the main function parse_user_structure_tree() takes 3 parameters. The first one is the name of the root directory " / ", the first False means not to print the process of traversing the tree, and the second True means to return the tree structure data in the form of a dictionary. As Figure 8 shown, it is a schematic diagram of the execution effect of test script 1.
[0154] Test script 2 returns the tree data structure data in the form of a list of absolute path strings. The implementation code of test script 2 is as follows:
[0155]
[0156] As Figure 9 shown, it is a schematic diagram of the execution effect of test script 2.
[0157] The steps of the methods described in the embodiments of this application can be directly embedded in hardware, software units executed by a processor, or a combination of the two. The software units can be stored in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium in the art. Exemplarily, the storage medium can be connected to the processor so that the processor can read information from the storage medium and write information to the storage medium. Optionally, the storage medium can also be integrated into the processor. The processor and the storage medium can be provided in an ASIC, and the ASIC can be provided in a UE. Optionally, the processor and the storage medium can also be provided in different components of the UE.
[0158] It should be understood that in various embodiments of this application, the magnitudes of the sequence numbers of the various processes do not imply the order of execution. The order of execution of the various processes should be determined according to their functions and internal logics, and should not constitute any limitation to the implementation process of the embodiments of this application.
[0159] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center to another website, a computer, a server, or a data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0160] For the same or similar parts among the various embodiments of this specification, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the description in the method embodiment part for the relevant parts.
[0161] Those skilled in the art can clearly understand that the technologies in the embodiments of the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or certain parts of the embodiments of the present application.
[0162] The above-described embodiments of the present application do not constitute a limitation on the protection scope of the present application.
Claims
1. A method for extracting page data, characterized in that, The described page data extraction method is used to extract the tree - shaped data structure data corresponding to the tree - shaped content displayed on a web page; the page data extraction method includes: Obtain the relative xpath of the root node in the tree - shaped data structure in the page to be tested; Use the relative xpath of the root node as the current relative xpath, and use a preset recursive call method to obtain a list of relative xpaths, where the list of relative xpaths includes the relative xpaths of each node arranged in depth - first order in the tree - shaped data structure; Convert the list of relative xpaths into a dictionary to obtain the tree - shaped data structure data in dictionary form; The preset recursive call method includes: Determine whether the current node exists in the corresponding record dictionary, where the current node is the node corresponding to the current relative xpath, and the record dictionary is used to record the traversed nodes and the corresponding first quantity and second quantity. The first quantity is the number of child nodes included in the target node, and the second quantity is the number of child nodes corresponding to the target node that have been recorded. The target node is any one of the traversed nodes; If the current node exists in the record dictionary, then determine whether the first quantity of the current node is equal to the second quantity; If the first quantity of the current node is not equal to the second quantity, then obtain the relative xpaths of the child nodes of the current node according to the current relative xpath; Record the relative xpaths of the child nodes of the current node into the list of relative xpaths, and increment the second quantity corresponding to the current node by 1; Use the relative xpaths of the child nodes of the current node as the current relative xpath and re - execute the preset recursive call method; If the first quantity of the current node is equal to the second quantity, then determine whether the current node has a younger brother node; If the current node has a younger brother node, then obtain the relative xpath of the younger brother node of the current node according to the current relative xpath; Record the relative xpath of the younger brother node of the current node into the list of relative xpaths, and increment the second quantity corresponding to the parent node of the current node by 1; Use the relative xpath of the younger brother node of the current node as the current relative xpath and re - execute the preset recursive call method; If the current node does not have a younger brother node, then obtain the relative xpath of the parent node of the current node according to the current relative xpath; Use the relative xpath of the parent node of the current node as the current relative xpath and re - execute the preset recursive call method.
2. The page data extraction method according to claim 1, wherein If the current node does not exist in the record dictionary, then obtain the first quantity of the current node according to the current relative xpath; Initialize the second quantity of the current node to 0; Record the current node, the corresponding first quantity, and the second quantity into the record dictionary, and continue to execute the step of judging whether the first quantity of the current node is equal to the second quantity.
3. The page data extraction method according to claim 1, wherein If the current node exists in the record dictionary, before judging whether the first quantity of the current node is equal to the second quantity, the page data extraction method further includes: Judge whether the current xpath relative path is the xpath relative path of the root node; If the current xpath relative path is the xpath relative path of the root node, terminate the execution of the recursive call method; If the current xpath relative path is not the xpath relative path of the root node, continue to execute the preset recursive call method.
4. The page data extraction method according to claim 1, wherein According to the current xpath relative path, obtain the xpath relative path of the child node of the current node, including: Add the xpath offset value of the child node to the current xpath relative path to obtain the xpath relative path of the child node of the current node.
5. The page data extraction method according to claim 1, wherein The obtaining the xpath relative path of the younger brother node of the current node according to the current xpath relative path includes: Delete the xpath offset value of the current node on the current xpath relative path to obtain the xpath relative path of the parent node of the current node; Add the xpath offset value of the younger brother node to the xpath relative path of the parent node of the current node to obtain the xpath relative path of the younger brother node of the current node.
6. The page data extraction method according to claim 1, wherein The obtaining the xpath relative path of the parent node of the current node according to the current xpath relative path includes: Delete the xpath offset value of the current node on the current xpath relative path to obtain the xpath relative path of the parent node of the current node.
7. The page data extraction method according to claim 1, wherein The converting the xpath relative path list into a dictionary to obtain the tree-shaped data structure data in dictionary form includes: Extract the absolute path string of each node from the xpath relative path list; Obtain the tree-shaped data structure data in dictionary form according to the absolute path string of each node.
8. A page automation testing method, characterized in that, Include: According to the page data extraction method according to any one of claims 1-7, obtain the tree-shaped data structure data corresponding to the tree-shaped content displayed on the page to be tested to obtain a test dictionary; Perform page testing according to the test dictionary.
9. A terminal device, characterized in that, Include: At least one processor and a memory; The memory is used to store program instructions; The processor is used to call and execute the program instructions stored in the memory, so that the terminal device executes the page data extraction method according to any one of claims 1-7.
Citation Information
Patent Citations
Method and device for assembling XML message
CN111628975A
Method for extracting webpage target information, electronic equipment and medium
CN112559929A