Path positioning method and device based on dictionary tree structure and weight learning and medium
Through the dictionary tree structure and weight learning path positioning method, the problems of accuracy and high maintenance cost of dynamic web page element positioning are solved, and more stable element positioning and more efficient test script maintenance are achieved.
Patent Information
- Application Number
- CN202511232799.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-29
AI Technical Summary
Traditional web automation testing tools rely on XPath or CSS selectors when locating dynamic web page elements, resulting in high technical coupling. Dynamically generated class names produce characteristic noise interference, affecting positioning accuracy. The maintenance workload and cost of automated test scripts are large and high.
A path positioning method based on dictionary tree structure and weight learning is adopted. By obtaining document object model elements, evaluating node weight scores, generating comprehensive feature encoding, and matching similar features with the preset dictionary tree, the target path is recommended.
It improves positioning accuracy, reduces technical coupling, reduces maintenance costs, and improves test case development efficiency.
Smart Images

Figure CN120743791A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer science, and in particular to a path positioning method, device and medium based on a dictionary tree structure and weight learning. Background Art
[0002] In the field of web automation testing, element location is a critical step in ensuring that test scripts can accurately identify and manipulate web page elements, which is crucial for test accuracy and reliability. Current mainstream web automation testing tools, such as Selenium, primarily rely on XPath (XML Path Language) or CSS (Cascading Style Sheets) selectors to locate page elements.
[0003] However, dynamic web pages are highly flexible and interactive, and their DOM (Document Object Model) structure changes frequently based on factors such as user operations and data loading. This makes element locators based on traditional positioning methods extremely prone to failure, posing a huge challenge to Web automated testing. Summary of the Invention
[0004] The present invention provides a path positioning method, device and medium based on a dictionary tree structure and weight learning, so as to at least solve the problems that traditional element positioning relies on paths such as array subscripts, has a high degree of technical coupling, dynamically generated class names in dynamic web pages generate characteristic noise interference, affecting positioning accuracy, and has a large workload and high maintenance cost for locators in automated test scripts.
[0005] The present invention provides a path positioning method based on a dictionary tree structure and weight learning, comprising the following steps: obtaining a document object model element; evaluating the weight scores of the nodes of the document object model element, and determining a target node that meets a preset importance condition based on the weight scores of the nodes of the document object model element; extracting feature information of the target node to generate a comprehensive feature code, performing similar feature matching based on the comprehensive feature code of the target node and a preset dictionary tree, and when the comprehensive feature code of the target node successfully matches the preset dictionary tree, generating a recommended target path based on the structure tree with successful similar feature matching.
[0006] The present invention also provides a path positioning system based on a dictionary tree structure and weight learning, comprising: an acquisition module for acquiring document object model elements; an evaluation module for evaluating the weight scores of the nodes of the document object model elements, and determining a target node that meets a preset importance condition based on the weight scores of the nodes of the document object model elements; a positioning module for extracting feature information of the target node to generate a comprehensive feature code, performing similar feature matching based on the comprehensive feature code of the target node and a preset dictionary tree, and generating a recommended target path based on the structure tree with successful similar feature matching when the comprehensive feature code of the target node successfully matches the preset dictionary tree.
[0007] The present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of the above-mentioned path positioning method based on a dictionary tree structure and weight learning when executing the computer program.
[0008] The present invention also provides a non-volatile computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the path positioning method based on the dictionary tree structure and weight learning are implemented.
[0009] The present invention also provides a computer program product, including a computer program, which implements the above-mentioned path positioning method based on dictionary tree structure and weight learning when executed by a processor.
[0010] The present invention obtains a document object model element, evaluates the weight scores of the nodes of the document object model element, determines a target node that meets a preset importance condition based on the weight scores of the nodes of the document object model element, extracts feature information of the target node to generate a comprehensive feature code, performs similar feature matching based on the comprehensive feature code of the target node and a preset dictionary tree, and, if the comprehensive feature code of the target node successfully matches the preset dictionary tree, generates a recommended target path based on the structure tree in which the similar feature matching succeeds. This solves the problems of traditional element positioning relying on paths such as array subscripts, high technical coupling, feature noise interference generated by dynamically generated class names in dynamic web pages, affecting positioning accuracy, and high maintenance workload and cost for locators in automated test scripts. Positioning accuracy is improved, technical coupling is reduced, and in test case development, XPath structures can be recommended based on user-selected DOM elements, saving development time. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0012] Figure 1 A flowchart of a path positioning method based on a dictionary tree structure and weight learning according to an embodiment of the present invention; Figure 2 A flowchart of a path positioning method based on a dictionary tree structure and weight learning according to an embodiment of the present invention; Figure 3 A schematic diagram of a node of a DOM structure according to an embodiment of the present invention; Figure 4 A schematic diagram of a path positioning system based on a dictionary tree structure and weight learning according to an embodiment of the present invention; Figure 5 FIG. 1 is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0013] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0014] It should be noted that, in the description of the present invention, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence.
[0015] In the field of web automation testing, element locator failure is a long-standing technical pain point. Current mainstream tools (such as Selenium) rely on XPath or CSS selectors to locate page elements. However, dynamic changes in the DOM structure of web pages (such as class name changes and node hierarchy adjustments) can cause widespread locator failure.
[0016] Traditional solutions have three limitations: High degree of technical coupling: The positioning path includes global xpath paths such as div[3] / button[2] that use array subscripts, and front-end fine-tuning will fail.
[0017] Feature noise interference: Dynamically generated class names (such as class="c7b3f") cannot be used as stable identifiers.
[0018] Maintenance costs are out of control: In automated test scripts, locator maintenance accounts for 67% of the overall workload. To address the above problem, an embodiment of the present invention provides a path positioning method based on a dictionary tree structure and weight learning.
[0019] XPath is a query language used to locate and select specific nodes in XML or HTML documents. Based on path expressions, it allows developers to precisely match target nodes based on criteria such as element name, attributes, hierarchical relationships, and text content. XPath is widely used in web automated testing (such as Selenium), data scraping (such as crawlers), and XML data processing. It provides both absolute and relative path locations to enhance the flexibility and accuracy of element locating.
[0020] like Figure 1 As shown, the path positioning method based on the dictionary tree structure and weight learning includes the following steps: Step S101: Obtain a document object model element.
[0021] Specifically, Figure 2 As shown, users install a specific browser plug-in and select the document object model element information they need to locate on the page. The plug-in traces the DOM tree until the body structure, and then transmits the document object model element to the program core.
[0022] It should be noted that there are a large number of similar open source plug-ins on the market. Regarding the development of plug-ins and DOM positioning, the following article will no longer describe information related to plug-ins and DOM collection.
[0023] Step S102 : evaluating the weight scores of the nodes of the document object model elements, and determining target nodes that meet a preset importance condition based on the weight scores of the nodes of the document object model elements.
[0024] Optionally, in some embodiments, evaluating the weight score of a node of a document object model element includes: determining the attribute weight score, semantic value weight score, and structural value of the node of the document object model element; calculating the sum of the attribute weight score and the semantic value weight score of the node of the document object model element, and obtaining the weight score of the node of the document object model element based on the product of the sum and the structural value.
[0025] After obtaining the document object model element, the program evaluates the weight score of each HTML DOM node from three dimensions: node attributes, semantic value, and structural value. It then selects the target node with the preset importance condition based on the weight score. The calculation formula for the weight score of the node of the document object model element is: (node attribute weight score + semantic value weight score) * structural value.
[0026] For example, if the node attribute weight score of a certain node of the document object model element is 40, the semantic value weight score is 20, and the structural value is 1.5, then the weight score of the certain node of the document object model element is 90 points.
[0027] Through the above technical solution, the weight scores of nodes are evaluated from a multi-dimensional and fine-grained perspective, which improves the accuracy of screening important nodes and the accuracy of path recommendations.
[0028] Optionally, in some embodiments, determining a target node that meets a preset importance condition based on the weight score of the node of the document object model element includes: determining whether there is a first node in the document object model element whose weight score is greater than a first threshold; if there is a first node, determining whether the number of the first nodes is greater than or equal to a first preset value; if the number of the first nodes is greater than or equal to the first preset value, determining that the first node is a target node that meets the preset importance condition.
[0029] The first threshold value and the first preset value may be threshold values pre-set by a user, may be threshold values obtained through a finite number of experiments, or may be threshold values obtained through a finite number of computer simulations.
[0030] In the embodiment of the present invention, taking the first threshold value as 80 points and the first preset value as 5 as an example: If there is a first node with a weight score greater than 80 in the document object model element, it is further determined whether the number of first nodes is greater than or equal to 5. If the number of first nodes is greater than or equal to 5, the first node is used as the target node that meets the preset importance conditions.
[0031] Through the above technical solution, through the dual constraints of "weight score threshold + quantity threshold", we can avoid misjudgments caused by abnormal weights of individual nodes or overly strict / loose settings of a single threshold, and improve the stability of screening important nodes.
[0032] Optionally, in some embodiments, after determining whether the number of first nodes is greater than or equal to a first preset value, it includes: if the number of first nodes is less than the first preset value and the number of first nodes is greater than or equal to a second preset value, then the second node in the first node that meets the preset number condition is used as the target node.
[0033] The second preset value may be a threshold value preset by a user, a threshold value obtained through a finite number of experiments, or a threshold value obtained through a finite number of computer simulations.
[0034] In the embodiment of the present invention, taking the second preset value as 3 as an example: It can be understood that if the number of first nodes is less than 5, it is determined whether the number of first nodes is greater than or equal to 3. If the number of first nodes is greater than or equal to 3, the first nodes are sorted from high to low according to the weight scores, and at least the top three nodes in the weight score ranking are taken as target nodes.
[0035] If multiple nodes have the same weight score (i.e., "duplicate score"), the node at the end of the sort (i.e., "post-node") is given priority.
[0036] For example, if nodes A (weight score = 90), B (weight score = 90), and C (weight score = 80) are sorted as [A, B, C], and if the last nodes are prioritized, the first three nodes in [B, A, C] are selected as the target nodes.
[0037] Through the above technical solution, by setting two thresholds, the first preset value and the second preset value, the processing strategy can be dynamically adjusted according to the actual number of nodes. When there are sufficient nodes, all nodes can be directly processed to avoid information loss caused by excessive screening. When there are insufficient nodes, secondary screening is initiated to ensure that at least a certain number of key nodes (such as 3) are retained to avoid the inability to support core functions due to too few nodes.
[0038] Optionally, in some embodiments, determining the attribute weight score of the node of the document object model element includes: identifying the current attribute of the node of the document object model element; if the current attribute is a unique identifier attribute, assigning a first score as the attribute weight score of the node of the document object model element; if the current attribute is a business attribute, assigning a second score as the attribute weight score of the node of the document object model element; if the current attribute is a role attribute, assigning a third score as the attribute weight score of the node of the document object model element; if the current attribute is a class name attribute, assigning a fourth score as the attribute weight score of the node of the document object model element, wherein the first score is greater than the second score, and the second score is greater than the third score.
[0039] The first score, the second score, the third score, and the fourth score may be thresholds preset by a user, may be thresholds obtained through a finite number of experiments, or may be thresholds obtained through a finite number of computer simulations.
[0040] Specifically, the program extracts the built-in attributes of each Document Object Model element node, including but not limited to: unique identifier attribute id, class name attribute class, business attributes data-*, role attribute role, value attribute value, title attribute title, etc. For different attribute pairs, the program calculates the attribute weight score of the corresponding node.
[0041] In the embodiment of the present invention, the unique identifier ID attribute has the highest weight score, which can be 40 points, that is, the first score is 40. Because ID is usually a unique identifier in HTML, the ID attribute is directly assigned a full 40 points when it exists, regardless of whether its value is randomly generated or semantically clear.
[0042] For example, if a node has id="submit-btn" or id="a8b3c2", the corresponding attribute weight score is 40 points.
[0043] Business data-* attributes (such as data-type and data-role) can be weighted to 30 points, giving them a secondary score of 30. These custom attributes are often used in business logic and receive a score of 30 simply for their presence. Their potential value is prioritized, but their weight may be adjusted based on semantic evaluation.
[0044] Role attribute: The weight score can be 20 points, that is, the third score is 20 points. role is an ARIA role attribute used for accessibility. Its presence gives 20 points, such as role="button".
[0045] Class name attribute: The weighted score can be 20 points. The presence of a class name grants a base score of 20 points. However, it should be noted that this score is based on the assessment of "existence." Additional points may be added or deducted based on the business semantic value of the class name (for example, a style class such as class="clearfix" may be downgraded). Therefore, the fourth score may be less than or equal to 20 points, or greater than or equal to 20 points.
[0046] Through the above technical solution, the attributes of the Document Object Model (DOM) elements are weighted and scored to accurately quantify the importance of the nodes, thus providing a reliable basis for subsequent path recommendations.
[0047] Optionally, in some embodiments, determining the semantic value weight score of a node of a document object model element includes: determining whether the node of the document object model element has a target business keyword; if the node of the document object model element has a target business keyword, assigning a fifth score to the semantic value weight score of the node of the document object model element, otherwise, determining whether the node of the document object model element meets a preset random generation feature; if the node of the document object model element meets the preset random generation feature, assigning a sixth score to the semantic value weight score of the node of the document object model element, otherwise, assigning a seventh score to the semantic value weight score of the node of the document object model element, wherein the fifth score is greater than the sixth score and the seventh score, and the sixth score is less than the seventh score.
[0048] The fifth score, the sixth score, and the seventh score may be thresholds preset by the user, may be thresholds obtained through a finite number of experiments, or may be thresholds obtained through a finite number of computer simulations.
[0049] It is understandable that if a node attribute has value based on its "existence", semantics will perform a secondary check on the "existence" information and the values of related attributes. This embodiment of the present invention divides the semantic value of the nodes of the Document Object Model elements into three cases based on the built-in language model: If there are meaningful business keywords in the node of the document object model element, the semantic value weight score is 20 points, that is, the fifth score is 20 points, such as id="btn-submit", class="add-panel", etc.
[0050] If the target business keyword exists in the node of the document object model element, it is determined whether the node of the document object model element meets the preset random generation feature. If the node of the document object model element contains a randomly generated string, the semantic value weight score of the node of the document object model element is -20 points, that is, the sixth score is -20 points, for example, id="a8b3c2"; If the node of the document object model element is a string or a value without a business keyword, the semantic value weight score is 0, that is, the seventh score is 0, for example, data-index="0" and data-type="a".
[0051] For the above technical solution, the highest semantic weight is given to business keywords to accurately identify business core nodes. Neutral weight is given when there are no business keywords, which can filter out irrelevant noise nodes. Random strings are given negative weights to reduce priority and reduce security risks.
[0052] Optionally, in some embodiments, determining the structural value of a node of a document object model element includes: determining whether the node of the document object model element is a main content area node; if the node of the document object model element is a main content area node, assigning a first preset value as the structural value of the node of the document object model element, otherwise, determining whether the node of the document object model element is a footer node or a sidebar node, or determining whether the node of the document object model element is a transition container node, or determining whether the node of the document object model element is a business carrying node; if the node of the document object model element is a footer node or a sidebar node, assigning a second preset value as the structural value of the node of the document object model element, if the node of the document object model element is a transition container node, assigning a third preset value as the structural value of the node of the document object model element, if the node of the document object model element is a business carrying node, assigning a fourth preset value as the structural value of the node of the document object model element, wherein the first preset value is greater than the fourth preset value, the fourth preset value is greater than the second preset value, and the second preset value is greater than the third preset value.
[0053] The first preset value, the second preset value, the third preset value and the fourth preset value may be thresholds preset by a user, may be thresholds obtained through a limited number of experiments, or may be thresholds obtained through a limited number of computer simulations.
[0054] It can be understood that the structural value of the node of the document object model element is obtained by determining the position coefficient or container role of the node of the document object model element. The position coefficient includes the main content area, footer or sidebar, transition container, etc., and the container role includes business-carrying nodes such as form container or list container.
[0055] In the embodiment of the present invention, if the position coefficient of the node of the document object model element is a main content area node, the structural value of the node of the document object model element is 1.5, that is, the first preset value is 1.5; If the node of the document object model element is a footer node or a sidebar node, the structural value of the node of the document object model element is 0.8, that is, the second preset value is 0.8; If the node of the document object model element is a transition container node, the structural value of the node of the document object model element is, that is, the third preset value is 0.5; If the container role of the node of the document object model element is a business carrying node, such as a form container or a list container, the structural value of the node of the document object model element is 1.3, that is, the fourth preset value is 1.3.
[0056] For example, the attributes, semantic values, and structural values of a node of a document object model element are described as follows: Figure 3As shown, the weight scores of nodes A, B, C, and D of the document object model element are evaluated.
[0057] Evaluation of node A: The weight score of node A = the existence of the unique identifier ID (i.e. the first score is 40 points) + the category name contains business words (i.e. the fourth score is 15 points) × the main content area node (i.e. the first preset value is 1.5) = 82.5 points.
[0058] Evaluation of node B: Node B's weight score = class name existence (i.e., attribute weight score is 20 points) × sidebar position (i.e., structural value is 0.8) = 16 points (not a target node: filtering is performed).
[0059] Evaluation of node C: The weight score of node C = existence of data-role (i.e. attribute weight score is 25 points) × content area coefficient (i.e. structural value is 1.5) = 37.5 points.
[0060] Node D evaluation: Node D's weight score = class name existence (i.e., attribute weight score is 20 points) × excessive container (i.e., structure value is 0.5) = 10 points, (not belonging to the target node: filtering is performed).
[0061] Node E evaluation: ID existence (i.e., attribute weight score is 40 points) + business text (i.e., semantic value weight score is 10 points) = 50 points. Based on the weight scores of nodes A, B, C, and D, the target node sequence is output: [A, C, E] → the target path XPath is generated: / / *[@id='app'] / main[@data-role='content-area'] / button[@id='submit-btn'].
[0062] Through the above technical solution, the main content area node is given the highest structural value, which can accurately locate the core area of the page; the business carrying node is given the second highest structural value, which can focus on monitoring the business node; the footer / sidebar node is given the second lowest structural value to distinguish the auxiliary function area; the transition container node is given the lowest structural value to clean up redundant structure.
[0063] Step S103: extract the feature information of the target node to generate a comprehensive feature code, perform similar feature matching based on the comprehensive feature code of the target node and the preset dictionary tree, and if the comprehensive feature code of the target node successfully matches the preset dictionary tree, generate a recommended target path based on the structure tree with successful similar feature matching.
[0064] It should be understood that the present invention is an embodiment that utilizes the data structure of the dictionary tree itself, and puts the characteristic information of each node of a set of Xpath into the node in sequence to form a corresponding dictionary tree. When it is confirmed that a certain node needs to be located, it is only necessary to read all the node information on the corresponding path in sequence starting from the root node, and after splicing them together, it is a complete Xpath.
[0065] Taking the example above as an example, assuming that the XPath has been added to the preset dictionary tree, its internal structure is as follows: body→[type=div:2, id=*app*]→[type=main:2, data-role=content-*]→[type=button:leaf, id=*btn, text=A3F9].
[0066] There is a new DOM structure. After Dom weight evaluation and feature extraction, its XPath information is as follows: body→ / *[@id='app-v2']→ / main[@data-role='content-area']→ / div[@id='submit-parent']→ / button[@id='submit-btn']. Compared with the existing information in the dictionary tree, it has an additional parent div wrapped on the button. After matching, the information already in the dictionary tree will be recommended first, followed by the current XPath information.
[0067] Optionally, in some embodiments, extracting feature information of the target node to generate a comprehensive feature code includes: extracting structural features, attribute features, and semantic features of the target node; and splicing the structural features, attribute features, and semantic features of the target node to generate a comprehensive feature code of the target node.
[0068] Specifically, after the DOM weight algorithm selects the target node, the feature extractor reads the target node's feature information and performs feature distillation. It should be noted that the feature extractor reads the original DOM node, not the XPath structure, and still extracts the target node's features from three dimensions, including the target node's structural features, attribute features, and semantic features. The specific extraction process is as follows: 1. Structural fingerprint features, judging the location of the element from the target node structure of the context, for example: Node label + direct child element type combination (such as div>h3+ul).
[0069] The position ratio in the parent container (e.g. first child node / 5 child nodes → "0.2").
[0070] 2. Stabilize attribute features and determine element features from some common attribute information, such as: Common prefix extraction of class names (e.g. btn-primary / btn-danger → feature "btn-*").
[0071] data attribute pattern abstraction (e.g. data-type="user-123" → feature "data-type=user-*").
[0072] 3. Semantic features, similar to weights, extract keywords from a semantic perspective, for example: Business text keyword fingerprint ("Submit" → code A3F9, "Cancel" → B2E4).
[0073] Role type flag (role="dialog" → encoding NAV).
[0074] For the previous example node E (<buttonid="submit-btn"> Submit), extraction process example: Structural features: button tag + no child elements → "button:leaf"; Attribute characteristics: ID contains "submit" → mode "id=*btn"; Semantic features: text "submit" → fingerprint "A3F9"; Finally, the structural features, attribute features and semantic features of the target node are spliced together to generate the comprehensive feature code of the target node: [type=button:leaf, id=*btn, text=A3F9].
[0075] Through the above technical solution, comprehensive feature coding is used to replace single attribute matching, which reduces mismatching and significantly improves the robustness of feature matching.
[0076] Optionally, in some embodiments, similarity matching is performed based on the comprehensive feature code of the target node and the preset dictionary tree. If the comprehensive feature code of the target node successfully matches the preset dictionary tree, a recommended target path is generated based on the structure tree in which the similarity matching is successful, including: sequentially passing the comprehensive feature code of the target node from the root node of the preset dictionary tree for matching; if there is a matching node in the target node whose comprehensive feature code matches the similarity feature of the preset dictionary tree more than a second threshold, it is determined that the comprehensive feature code of the target node successfully matches the preset dictionary tree; and the recommended target path is returned according to the weight value of the matching node.
[0077] The second threshold may be a threshold preset by a user, a threshold obtained through a finite number of experiments, or a threshold obtained through a finite number of computer simulations.
[0078] Specifically, the comprehensive feature code of the target node passed in cyclically is matched with the preset dictionary tree (it should be noted that: a maximum of 3 layers of children of the preset dictionary tree are matched with the comprehensive feature code of the target node). If there is a matching node in the target node whose comprehensive feature code has a similar feature matching degree with the preset dictionary tree greater than the second threshold, then it is determined that the comprehensive feature code of the target node is successfully matched with the preset dictionary tree. A maximum of 3 nodes are taken as matching nodes in order from high to low according to the weight value of the successfully matched nodes, and the recommended target path is returned based on the structure of the matching nodes.
[0079] It should be noted that if a matching node exists in the direct child of a node in the preset dictionary tree, the weight value of the child is increased.
[0080] Through the above technical solution, comprehensive feature coding is used to replace single attribute matching to reduce mismatches, and the matching scope is limited by hierarchical matching rules (root node → up to 3 layers of child nodes), taking into account both retrieval efficiency and accuracy.
[0081] Optionally, in some embodiments, after generating the target path, it includes: obtaining user feedback information based on the target path; if the feedback information is that the path is available, then increasing the weight value of the target path in the preset dictionary tree; if the feedback information is that the path is unavailable, then adding a new node in the preset dictionary tree.
[0082] From the above description, it can be seen that the XPath positioning method provided by this program is based on the initial XPath given by Dom weight evaluation, supplemented by comparison with the dictionary tree after feature extraction. In fact, the results given by the dictionary tree are more accurate than the Dom weight evaluation, and the reason is user feedback. After the program gives the recommended target path XPath, the user can provide feedback based on the actual situation, asking whether the target path XPath is usable. If it is, the target path's feature information will be stored in the dictionary library and the target path's weight value will be increased in the preset dictionary tree. The more a path is used, the higher its weight value will be, achieving an effect of increasing accuracy with use.
[0083] If the user feedback information is that the target path is unavailable, a new node is added to the preset dictionary tree, and the new node is recorded as the default weight value.
[0084] When the user's actual target path is used and a new path is selected on the original route, a new path will be given in the dictionary tree according to the node characteristics. At this time, there are two links in the dictionary tree for the button. Subsequently, according to the user's choice, the weight value of a link will become higher and higher. When similar paths are passed in, the path with a high weight value will always be recommended to the user first.
[0085] Through the above technical solution, the dynamic update of the dictionary tree is driven by user feedback information (path availability / unavailability), so that the dynamic adaptability of the dictionary tree is enhanced.
[0086] Optionally, in some embodiments, when the comprehensive feature code of the target node fails to match the preset dictionary tree, it includes: returning a target path generated based on the target node.
[0087] It is understandable that when the comprehensive feature coding of the target node cannot match the preset dictionary tree, the target path is dynamically generated based on the attribute or structural information of the target node, and the generated target path is returned to the user, and a new node is added to the preset dictionary tree, recorded as the default weight value.
[0088] Through the above technical solution, when the dictionary tree matching fails, the target path (such as XPath) is directly generated based on the original features of the target node, avoiding recommendation interruption due to matching failure and ensuring system availability.
[0089] Optionally, in some embodiments, the above-mentioned path positioning method based on dictionary tree structure and weight learning further includes: obtaining the number of times at least one path in a preset dictionary tree is used; and adjusting the weight value of at least one path according to the number of times at least one path is used.
[0090] Specifically, the dictionary tree is traversed periodically (such as every hour / day) and the weight is updated according to the number of times the path is used.
[0091] For example, for high-frequency paths that are used more times than the first time, the weight value of the corresponding path will be increased, and priority will be given to matching the path the next time the user uses it. For low-frequency paths that are used less than the second time, the weight value of the corresponding path will be reduced, and matching will be reduced when the user uses it.
[0092] Through the above technical solution, the weight of the dictionary tree node is dynamically adjusted according to the number of times the path is used, so that high-frequency paths can be given priority recommendation.
[0093] Optionally, in some embodiments, the above-mentioned path positioning method based on dictionary tree structure and weight learning further includes: cleaning nodes in the preset dictionary tree that have not been used for more than a preset time and whose corresponding weight values are less than a third threshold based on a preset period.
[0094] The preset period, the preset time and the third threshold may be thresholds preset by the user, may be thresholds obtained through a limited number of experiments, or may be thresholds obtained through a limited number of computer simulations.
[0095] It should be understood that if a node in a preset dictionary tree is not used for a period greater than a preset period of time and the weight value of the node is less than a third threshold value every week, the node may be cleaned up.
[0096] For example, taking the preset time of 30 days and the third threshold of 50 as an example, if it is detected that a node A has not been used for 40 days and the current weight of the node A is 40, the node A will be cleaned up.
[0097] Through the above technical solution, by regularly cleaning up nodes that have not been used for a long time and have low weight, the dictionary tree can be prevented from expanding infinitely due to sparse user behavior or dynamic changes in pages, thereby reducing storage costs and query delays.
[0098] In order to enable those skilled in the art to further understand the path positioning based on the dictionary tree structure and weight learning of the embodiment of the present application, the following is a detailed description in conjunction with specific embodiments, such as Figure 2 shown.
[0099] Step 1: The user installs a specific browser plug-in and selects the element information they need to locate on the page.
[0100] Step 2: The plug-in traces the DOM tree until it reaches the body structure and then passes the information to the program core.
[0101] Step 3: The Dom weight evaluation program parses the incoming HTML structure according to the preset algorithm and selects important nodes.
[0102] Step 4: The feature extractor traverses the nodes to extract feature information for each important node. It then searches the dictionary tree for similar structures. If not, the important node is directly returned. If similar structures exist, all similar structures are extracted from the dictionary and returned one by one according to their weight as the recommended XPath.
[0103] Step 5: The user uses the XPath provided by the system to perform actual verification and provide feedback. The feedback mechanism optimizes the dictionary tree based on the user's feedback.
[0104] In summary, the technical effects brought about by the embodiments of the present invention are as follows.
[0105] (1) Improve positioning stability by establishing a dictionary tree to filter noise and calculate node weights, making the locator more stable in front-end iterations.
[0106] (2) Reduce technical coupling, elevate positioning to the business logic layer, and reduce large-scale modifications to the locator due to front-end changes.
[0107] (3) Achieve an optimization closed loop, using user feedback to drive incremental learning, so that the locator becomes more accurate with use.
[0108] (4) Improve development efficiency. In test case development, it can recommend XPath structures based on the DOM elements selected by the user, saving development time.
[0109] Next, a path positioning method based on a dictionary tree structure and weight learning proposed in an embodiment of the present invention will be described with reference to the accompanying drawings.
[0110] Figure 4 Schematic diagram of a path positioning system based on a dictionary tree structure and weight learning according to an embodiment of the present invention.
[0111] like Figure 4 As shown, the path positioning system 10 based on the dictionary tree structure and weight learning includes: an acquisition module 100, an evaluation module 200 and a positioning module 300.
[0112] Among them, the acquisition module 100 is used to obtain the document object model element; the evaluation module 200 is used to evaluate the weight score of the node of the document object model element, and determine the target node that meets the preset importance condition based on the weight score of the node of the document object model element; the positioning module 300 is used to extract the feature information of the target node to generate a comprehensive feature code, perform similar feature matching based on the comprehensive feature code of the target node and the preset dictionary tree, and when the comprehensive feature code of the target node successfully matches the preset dictionary tree, generate a recommended target path based on the structure tree with successful similar feature matching.
[0113] Optionally, in some embodiments, after generating the target path, the positioning module 300 is further used to: obtain user feedback information based on the target path; if the feedback information is that the path is available, then increase the weight value of the target path in the preset dictionary tree; if the feedback information is that the path is unavailable, then add a new node in the preset dictionary tree.
[0114] Optionally, in some embodiments, the evaluation module 200 is further used to: determine the attribute weight score, semantic value weight score and structural value of the node of the document object model element; calculate the sum of the attribute weight score and the semantic value weight score of the node of the document object model element, and obtain the weight score of the node of the document object model element based on the product of the sum and the structural value.
[0115] Optionally, in some embodiments, the evaluation module 200 is also used to: determine whether there is a first node in the document object model element whose weight score is greater than a first threshold; if there is a first node, determine whether the number of first nodes is greater than or equal to a first preset value; if the number of first nodes is greater than or equal to the first preset value, determine that the first node is a target node that meets the preset importance condition.
[0116] Optionally, in some embodiments, after determining whether the number of first nodes is greater than or equal to a first preset value, the evaluation module 200 is further used to: if the number of first nodes is less than the first preset value and the number of first nodes is greater than or equal to a second preset value, then the second node in the first node that meets the preset quantity condition is used as the target node.
[0117] Optionally, in some embodiments, the evaluation module 200 is further used to: identify the current attribute of the node of the document object model element; if the current attribute is a unique identifier attribute, assigning a first score as the attribute weight score of the node of the document object model element; if the current attribute is a business attribute, assigning a second score as the attribute weight score of the node of the document object model element; if the current attribute is a role attribute, assigning a third score as the attribute weight score of the node of the document object model element; if the current attribute is a class name attribute, assigning a fourth score as the attribute weight score of the node of the document object model element, wherein the first score is greater than the second score, and the second score is greater than the third score.
[0118] Optionally, in some embodiments, the evaluation module 200 is also used to: determine whether the node of the document object model element has the target business keyword; if the node of the document object model element has the target business keyword, assign the fifth score to the semantic value weight score of the node of the document object model element; otherwise, determine whether the node of the document object model element meets the preset random generation feature; if the node of the document object model element meets the preset random generation feature, assign the sixth score to the semantic value weight score of the node of the document object model element; otherwise, assign the seventh score to the semantic value weight score of the node of the document object model element, wherein the fifth score is greater than the sixth score and the seventh score, and the sixth score is less than the seventh score.
[0119] Optionally, in some embodiments, the evaluation module 200 is also used to: determine whether the node of the document object model element is a main content area node; if the node of the document object model element is a main content area node, assign a first preset value to the structural value of the node of the document object model element; otherwise, determine whether the node of the document object model element is a footer node or a sidebar node, or determine whether the node of the document object model element is a transition container node, or determine whether the node of the document object model element is a business carrying node; if the node of the document object model element is a footer node or a sidebar node, assign a second preset value to the structural value of the node of the document object model element; if the node of the document object model element is a transition container node, assign a third preset value to the structural value of the node of the document object model element; if the node of the document object model element is a business carrying node, assign a fourth preset value to the structural value of the node of the document object model element, wherein the first preset value is greater than the fourth preset value, the fourth preset value is greater than the second preset value, and the second preset value is greater than the third preset value.
[0120] Optionally, in some embodiments, the positioning module 300 is further used to: extract the structural features, attribute features, and semantic features of the target node; and concatenate the structural features, attribute features, and semantic features of the target node to generate a comprehensive feature code of the target node.
[0121] Optionally, in some embodiments, the positioning module 300 is further used to: sequentially pass in the comprehensive feature code of the target node from the root node of the preset dictionary tree for matching; if there is a matching node in the target node whose comprehensive feature code matches the similar feature of the preset dictionary tree with a degree greater than a second threshold, then it is determined that the comprehensive feature code of the target node matches the preset dictionary tree successfully; and return the recommended target path according to the weight value of the matching node.
[0122] Optionally, in some embodiments, when the comprehensive feature code of the target node fails to match the preset dictionary tree, the positioning module 300 is further used to: return a target path generated based on the target node.
[0123] Optionally, in some embodiments, the above-mentioned path positioning system 10 based on dictionary tree structure and weight learning further includes: an adjustment module for obtaining the number of times at least one path in a preset dictionary tree is used, and adjusting the weight value of at least one path according to the number of times at least one path is used.
[0124] Optionally, in some embodiments, the above-mentioned path positioning system 10 based on dictionary tree structure and weight learning further includes: a cleaning module for cleaning nodes in a preset dictionary tree that have not been used for more than a preset time and whose corresponding weight values are less than a third threshold based on a preset period.
[0125] It should be noted that the description of the features in the embodiment corresponding to the path positioning system based on the dictionary tree structure and weight learning can be found in the relevant description of the embodiment corresponding to the path positioning method based on the dictionary tree structure and weight learning, and will not be repeated here.
[0126] Figure 5 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. The electronic device may include: Memory 501 , processor 502 , and computer programs stored in the memory 501 and executable on the processor 502 .
[0127] When the processor 502 executes the program, the path positioning method based on the dictionary tree structure and weight learning provided in the above embodiment is implemented.
[0128] Furthermore, the electronic device further includes: The communication interface 503 is used for communication between the memory 501 and the processor 502 .
[0129] The memory 501 is used to store computer programs that can be run on the processor 502 .
[0130] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0131] If the memory 501, processor 502, and communication interface 503 are implemented independently, the communication interface 503, memory 501, and processor 502 can be interconnected via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 5 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0132] Optionally, in a specific implementation, if the memory 501, the processor 502 and the communication interface 503 are integrated on a chip, the memory 501, the processor 502 and the communication interface 503 can communicate with each other through an internal interface.
[0133] The processor 502 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0134] An embodiment of the present invention also provides a non-volatile computer-readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned path positioning method embodiments based on dictionary tree structure and weight learning when running.
[0135] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0136] An embodiment of the present invention further provides a computer program product, including a computer program, which implements the above-mentioned path positioning method based on dictionary tree structure and weight learning when executed by a processor.
[0137] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0138] The above is a detailed introduction to the path positioning method, device and medium based on the dictionary tree structure and weight learning provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only applicable to help understand the method of the present invention and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, the present invention can also be improved and modified in several ways, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.
Claims
1. A path positioning method based on dictionary tree structure and weight learning, characterized in that: The following steps are involved: Get the document object model element; Evaluating the weight scores of the nodes of the document object model element, and determining a target node that meets a preset importance condition based on the weight scores of the nodes of the document object model element; Extract the feature information of the target node to generate a comprehensive feature code, perform similar feature matching based on the comprehensive feature code of the target node and a preset dictionary tree, and if the comprehensive feature code of the target node successfully matches the preset dictionary tree, generate a recommended target path based on the structure tree with successful similar feature matching.
2. The path positioning method based on dictionary tree structure and weight learning according to claim 1 is characterized in that: After generating the target path, the following steps are included: Obtaining user feedback information based on the target path; If the feedback information indicates that the path is available, the weight value of the target path is increased in the preset dictionary tree; if the feedback information indicates that the path is unavailable, a new node is added in the preset dictionary tree.
3. The path positioning method based on dictionary tree structure and weight learning according to claim 1 is characterized in that: Evaluating a weight score of a node of the document object model element, including: Determining an attribute weight score, a semantic value weight score, and a structural value of a node of the document object model element; The sum of the attribute weight score and the semantic value weight score of the node of the document object model element is calculated, and the weight score of the node of the document object model element is obtained based on the product of the sum and the structural value.
4. The path positioning method based on dictionary tree structure and weight learning according to claim 1 is characterized in that: The determining of the target node that meets a preset importance condition based on the weight score of the node of the document object model element includes: Determining whether there is a first node in the document object model element whose weight score is greater than a first threshold; If the first node exists, determining whether the number of the first nodes is greater than or equal to a first preset value; If the number of the first nodes is greater than or equal to a first preset value, the first nodes are determined to be target nodes that meet a preset importance condition.
5. The path positioning method based on dictionary tree structure and weight learning according to claim 4 is characterized in that: After determining whether the number of the first nodes is greater than or equal to a first preset value, the method includes: If the number of the first nodes is less than the first preset value and the number of the first nodes is greater than or equal to a second preset value, the second node among the first nodes that meets the preset number condition is used as the target node.
6. The path positioning method based on dictionary tree structure and weight learning according to claim 3 is characterized in that: Determines the attribute weight score of a node of a Document Object Model element, including: Identifying current attributes of a node of the document object model element; If the current attribute is a unique identifier attribute, the first score is assigned to the attribute weight score of the node of the document object model element; if the current attribute is a business attribute, the second score is assigned to the attribute weight score of the node of the document object model element; if the current attribute is a role attribute, the third score is assigned to the attribute weight score of the node of the document object model element; if the current attribute is a class name attribute, the fourth score is assigned to the attribute weight score of the node of the document object model element, wherein the first score is greater than the second score, and the second score is greater than the third score.
7. The path positioning method based on dictionary tree structure and weight learning according to claim 3 is characterized in that: Determine the semantic value weight score of the node of the document object model element, including: Determining whether a target business keyword exists in a node of the document object model element; If the target business keyword exists in the node of the document object model element, assigning the fifth score to the semantic value weight score of the node of the document object model element; otherwise, determining whether the node of the document object model element meets the preset randomly generated characteristics; If the node of the document object model element meets the preset randomly generated characteristics, the sixth score is assigned as the semantic value weight score of the node of the document object model element; otherwise, the seventh score is assigned as the semantic value weight score of the node of the document object model element, wherein the fifth score is greater than the sixth score and the seventh score, and the sixth score is less than the seventh score.
8. The path positioning method based on dictionary tree structure and weight learning according to claim 3 is characterized in that: Determines the structural value of a node for a Document Object Model element, including: Determine whether the node of the document object model element is a main content area node; If the node of the document object model element is a main content area node, assigning the first preset value to the structural value of the node of the document object model element; otherwise, determining whether the node of the document object model element is a footer node or a sidebar node, or determining whether the node of the document object model element is a transition container node, or determining whether the node of the document object model element is a service bearing node; If the node of the document object model element is a footer node or a sidebar node, the second preset value is assigned to the structural value of the node of the document object model element; if the node of the document object model element is a transition container node, the third preset value is assigned to the structural value of the node of the document object model element; if the node of the document object model element is a business carrying node, the fourth preset value is assigned to the structural value of the node of the document object model element, wherein the first preset value is greater than the fourth preset value, the fourth preset value is greater than the second preset value, and the second preset value is greater than the third preset value.
9. The path positioning method based on dictionary tree structure and weight learning according to claim 1 is characterized in that: The step of extracting the characteristic information of the target node and generating a comprehensive characteristic code includes: Extracting structural features, attribute features, and semantic features of the target node; The structural features, the attribute features and the semantic features of the target node are concatenated to generate a comprehensive feature code of the target node.
10. The path positioning method based on dictionary tree structure and weight learning according to claim 1, characterized in that: The similarity feature matching is performed based on the comprehensive feature code of the target node and a preset dictionary tree. If the comprehensive feature code of the target node successfully matches the preset dictionary tree, a recommended target path is generated based on the structure tree with successful similarity feature matching, including: The comprehensive feature codes of the target nodes are sequentially passed in from the root node of the preset dictionary tree for matching; If there is a matching node in the target node whose comprehensive feature code matches the similarity feature of the preset dictionary tree more than a second threshold, it is determined that the comprehensive feature code of the target node matches the preset dictionary tree successfully; Return a recommended target path based on the weight value of the matching node.
11. The path positioning method based on dictionary tree structure and weight learning according to claim 1, characterized in that: In the case where the comprehensive feature code of the target node fails to match the preset dictionary tree, the method includes: Returns the target path generated based on the target node.
12. The path positioning method based on dictionary tree structure and weight learning according to claim 1, characterized in that: Also includes: Obtaining the number of times at least one path in the preset dictionary tree is used; The weight value of the at least one path is adjusted according to the number of times the at least one path is used.
13. The path positioning method based on dictionary tree structure and weight learning according to claim 1, characterized in that: Also includes: Based on a preset period, nodes in the preset dictionary tree that have not been used for more than a preset time and whose corresponding weight values are less than a third threshold are cleared.
14. An electronic device, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the path positioning method based on the dictionary tree structure and weight learning as described in any one of claims 1 to 13.
15. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the path positioning method based on the dictionary tree structure and weight learning as described in any one of claims 1 to 13.
Citation Information
Patent Citations
Webpage node path information determination method and apparatus
CN106599280A
Dictionary tree construction method, statement search method and device, equipment and storage medium
CN109740165A
Data storage structure, retrieval method, data storage method and terminal equipment
CN113127692A
Application program testing method and device, electronic equipment and storage medium
CN115904969A
Text processing method and device, equipment, medium and program product
CN119337863A