Page element positioning method and device

By generating descriptive text information and combining multiple text recognition schemes and similarity calculations, the problem of inaccurate DOM element positioning is solved, and more efficient browser automation testing is achieved.

CN120632475AActive Publication Date: 2025-09-12THREE GORGES HI TECH INFORMATION TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510619068.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-09-12
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

The existing technology has low accuracy in DOM element positioning, which results in insufficient accuracy and efficiency in browser automation testing.

Method used

By receiving user instructions, it generates descriptive text information, uses multiple text recognition schemes to determine the similarity of page elements, and locates them based on the total similarity, including optical character recognition, target page code analysis and server-side multimodal visual model, combined with similarity weight calculation to finally locate the target page element.

Benefits of technology

Improves the robustness and accuracy of page element positioning, and enhances the efficiency and accuracy of browser automation testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632475A_ABST
    Figure CN120632475A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a page element positioning method and device. The method comprises the steps that a user instruction of a target page is received; the target page comprises at least one page element; generating description text information according to the user instruction; the description text information at least comprises a user intention and a descriptive target; determining the similarity between at least one piece of text information of the page element and the description text information; the text information is obtained through at least one text recognition scheme; determining the total similarity of the page elements according to the similarity of the page elements; and positioning a target page element from the page elements in the target page according to the total similarity. According to the embodiment of the invention, the robustness and accuracy of page element positioning are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of computer technology, and in particular to a page element positioning method, a page element positioning device, an electronic device, and a computer-readable storage medium. Background Art

[0002] A DOM (Document Object Model Element) element is a structured representation of a page, a collection of tree-like objects generated by a browser after parsing the code.

[0003] In specific implementations, DOM element positioning is a commonly used technology in the field of browser automation testing. However, the accuracy of DOM element positioning is currently not high. Summary of the Invention

[0004] In view of the above problems, a method and device for locating page elements is proposed to overcome or at least partially solve the above problems. The specific technical solution is as follows:

[0005] In a first aspect of the present invention, a method for locating a page element is provided, the method comprising:

[0006] Receiving a user instruction for a target page; the target page includes at least one page element;

[0007] generating descriptive text information according to the user instruction; the descriptive text information at least including the user intention and the descriptive target;

[0008] Determining a similarity between at least one text information of the page element and the description text information; the text information is obtained by at least one text recognition scheme;

[0009] Determining the total similarity of the page elements according to the similarity of the page elements;

[0010] A target page element is located from the page elements in the target page according to the total similarity.

[0011] In one embodiment of the present invention, determining the similarity between at least one text information of the page element and the descriptive target includes:

[0012] Filtering candidate page elements from page elements in the target page according to the user intention and / or the descriptive goal;

[0013] A similarity between at least one text information of the candidate page element and the user intention and / or the descriptive target is determined.

[0014] In one embodiment of the present invention, determining the similarity between at least one text information of the candidate page element and the user intention and / or the descriptive target includes:

[0015] Obtaining a page element screenshot of the candidate page element in the target page;

[0016] The text information of the page element screenshot is recognized by optical character recognition technology, and a first similarity between the text information and the user intention and / or the descriptive target is determined.

[0017] In one embodiment of the present invention, determining the similarity between at least one text information of the candidate page element and the user intention and / or the descriptive target includes:

[0018] Obtaining the page code of the candidate page element in the target page;

[0019] The text information corresponding to the candidate page element is extracted from the page code, and a second similarity between the text information and the user intention and / or the descriptive target is determined.

[0020] In one embodiment of the present invention, determining the similarity between at least one text information of the candidate page element and the user intention and / or the descriptive target includes:

[0021] Obtaining a page element screenshot of the candidate page element in the target page;

[0022] Sending the page element screenshot to a server; the server is configured to call a multimodal visual model to identify text information of the page element screenshot;

[0023] The text information sent by the server is received, and a third similarity between the text information and the user intention and / or the descriptive target is determined.

[0024] In one embodiment of the present invention, sending the screenshot of the page element to the server includes:

[0025] When the preset conditions are met, the screenshot of the page elements is sent to the server;

[0026] The preset conditions include at least: the description text information includes preset text information, the difference between the first similarity and the second similarity is greater than a preset similarity threshold, and / or the network connection status with the server is higher than a preset network connection status.

[0027] In one embodiment of the present invention, determining the total similarity of the page elements according to the similarity of the page elements includes:

[0028] Obtaining text recognition schemes corresponding to the similarities of the page elements; the text recognition schemes respectively have corresponding weight values;

[0029] The total similarity of the page elements is determined according to the similarity and weight values ​​of the page elements.

[0030] In one embodiment of the present invention, generating descriptive text information according to the user instruction includes:

[0031] Obtaining an action type corresponding to the user instruction, and using the action type as the user intention;

[0032] A descriptive target is generated according to the user instruction, and the user intention and the descriptive target are used as descriptive text information.

[0033] In a second aspect of the present invention, a device for locating page elements is provided, the device comprising:

[0034] A user instruction receiving module, configured to receive a user instruction for a target page; the target page includes at least one page element;

[0035] A description text information generating module, configured to generate description text information according to the user instruction; the description text information at least includes the user intention and a descriptive target;

[0036] a similarity determination module, configured to determine a similarity between at least one text information of the page element and the description text information; the text information being obtained by at least one text recognition scheme;

[0037] A total similarity determination module, configured to determine the total similarity of the page elements according to the similarities of the page elements;

[0038] The target page element locating module is configured to locate a target page element from the page elements in the target page according to the total similarity.

[0039] In another aspect of the present invention, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium. When the computer-readable storage medium is executed on a computer, the computer executes any of the above-mentioned methods for locating page elements.

[0040] In another aspect of the present invention, a computer program product comprising instructions is provided, which, when executed on a computer, enables the computer to execute any of the above-mentioned methods for locating page elements.

[0041] Compared with the related art, the embodiments of the present invention have at least the following advantages:

[0042] In an embodiment of the present invention, a user instruction for a target page is received, the target page may include at least one page element, descriptive text information is generated according to the user instruction, wherein the descriptive text information may include at least the user intention and a descriptive target, and then the similarity between at least one text information of the page element and the descriptive text information is determined, the text information is obtained through at least one text recognition scheme, and the total similarity of the page elements is determined according to the similarity of the page elements, so as to locate the target page element from the page elements in the target page according to the total similarity. In an embodiment of the present invention, the total similarity can be determined based on the similarity between the text information of at least one page element and the descriptive text information of the user instruction, and then based on at least one similarity, the target page element can be located from the page elements in the target page based on the total similarity. Since the page element is not located in the target page based on a single similarity, the robustness and accuracy of the page element location are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for describing the embodiments or the prior art.

[0044] Figure 1 A flowchart of a method for locating page elements provided in an embodiment of the present invention;

[0045] Figure 2 This is a structural block diagram of a page element positioning device provided in an embodiment of the present invention;

[0046] Figure 3 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0047] The technical solutions in the embodiments of the present invention will be described below with reference to the accompanying drawings in the embodiments of the present invention.

[0048] Reference Figure 1 , is a flow chart of the steps of a page element positioning method provided in an embodiment of the present invention, such as Figure 1 As shown, the method may specifically include the following steps:

[0049] Step 101: Receive a user instruction for a target page; the target page includes at least one page element.

[0050] Step 102: Generate descriptive text information according to the user instruction; the descriptive text information at least includes user intention and a descriptive goal.

[0051] In practice, automated testing is performed on the browser (client) to simulate user actions on one or more pages of the application. This ensures that the application pages can be displayed in the browser and can achieve normal human-computer interaction, thereby ensuring the user experience of the browser page. The page can include one or more DOM elements, i.e., page elements, such as a login button, search button, input box, or image.

[0052] In an embodiment of the present invention, during browser automation testing, a specific intent list may be given. The specific intent list may include generating one or more user instructions for the target page in the browser, such as "click the search button on the homepage", "click the login button on the homepage", "click the circular button on the details page", etc., and then locating the corresponding page element in the browser according to the user instruction, and then determining whether the page element correctly implements the corresponding human-computer interaction according to the user instruction. For example, when the user instruction is "click the login button on the homepage", it will locate the location of the login button on the homepage, and then verify whether the login button changes state (for example, whether it changes color, becomes brighter or darker, etc.), and whether it jumps to the login interface, etc. Among them, the target page can be a certain interface / page of the application, or it can be multiple interfaces of the application, or even all interfaces, and the embodiment of the present invention does not need to limit this.

[0053] In an embodiment of the present invention, a browser may receive one or more user instructions for a target page, and then generate descriptive text information based on the user instructions, wherein the descriptive text information may include at least user intent and a descriptive target, wherein the user intent may include at least an action type, such as "click", "slide", "input" and "select" and other action types, and the descriptive target refers to a semantic feature extracted from the user instruction for locating page elements, for example, from the user instruction "click the login button in the home page", the descriptive targets "login" and "home page" may be extracted.

[0054] Step 103: Determine the similarity between at least one text information of the page element and the description text information; the text information is obtained through at least one text recognition scheme.

[0055] In an embodiment of the present invention, different text information corresponding to each page element of the target page can be determined through different text recognition schemes, and then the similarity between at least one text information of each page element and the descriptive text information can be determined. For example, assuming there are three text recognition schemes, three text information corresponding to the page elements can be obtained through the three text recognition schemes. It can be understood that due to the different implementation principles of different text recognition schemes or the different computing capabilities of the devices, the accuracy of their text information recognition is also different. The text information of the page elements obtained based on different text recognition schemes is also different, and thus the similarity between the text information of the page elements and the descriptive text information is also different.

[0056] Step 104: Determine the total similarity of the page elements according to the similarity of the page elements.

[0057] Step 105: Locate a target page element from the page elements in the target page according to the total similarity.

[0058] In an embodiment of the present invention, after obtaining one or more similarities corresponding to each page element, the total similarity of the page element can be calculated based on the similarities of each page element. Then, the corresponding target page element can be located in the target page based on the total similarities of each page element. For example, the page element with the highest total similarity can be located as the target page element, and then it can be determined whether the target page element can correctly realize the corresponding human-computer interaction according to the user instruction.

[0059] In an embodiment of the present invention, a user instruction for a target page is received, the target page may include at least one page element, descriptive text information is generated according to the user instruction, wherein the descriptive text information may include at least the user intention and a descriptive target, and then the similarity between at least one text information of the page element and the descriptive text information is determined, the text information is obtained through at least one text recognition scheme, and the total similarity of the page elements is determined according to the similarity of the page elements, so as to locate the target page element from the page elements in the target page according to the total similarity. In an embodiment of the present invention, the total similarity can be determined based on the similarity between the text information of at least one page element and the descriptive text information of the user instruction, and then based on at least one similarity, the target page element can be located from the page elements in the target page based on the total similarity. Since the page element is not located in the target page based on a single similarity, the robustness and accuracy of the page element location are improved.

[0060] In one embodiment of the present invention, determining the similarity between the at least one text information of the page element and the descriptive target may include:

[0061] Filtering candidate page elements from page elements in the target page according to the user intention and / or the descriptive goal;

[0062] A similarity between at least one text information of the candidate page element and the user intention and / or the descriptive target is determined.

[0063] In specific implementations, the DOM tree is a structured object model that the browser parses the HTML (Hypertext Markup Language) document of a page into. It represents the hierarchical relationship of all elements in a web page (such as tags, text, attributes, etc.) in a tree structure and is the core foundation of browser automation and front-end development.

[0064] In an embodiment of the present invention, when parsing a DOM tree of a target page, the DOM tree of the target page can be optimized and reduced based on user intent, a descriptive target, or both, i.e., based on user intent and / or descriptive target, to retain only a list of candidate targets that meet the user intent and / or descriptive target, wherein the list of candidate targets includes one or more candidate page elements. Then, similarities between one or more textual information of the candidate page elements and the user intent, the descriptive target, or both, i.e., the similarities between one or more textual information of the candidate page elements and the user intent and / or descriptive target, can be evaluated.

[0065] For example, if the user's intention is "click" and the descriptive target is "button", the DOM tree of the target page is optimized and reduced, and the page elements in the target page that are buttons and allow clicks can be filtered as candidate page elements. In this way, when there are many page elements in the target page, data processing processes such as similarity calculation can be performed only on the candidate page elements in the target page, which greatly reduces the amount of data processing and thus improves the efficiency of browser automation testing.

[0066] In one embodiment of the present invention, step 102, generating descriptive text information according to the user instruction, may include:

[0067] Obtaining an action type corresponding to the user instruction, and using the action type as the user intention;

[0068] A descriptive target is generated according to the user instruction, and the user intention and the descriptive target are used as descriptive text information.

[0069] In a specific implementation, user instructions each have a corresponding action type, and the action type can reflect the user's intention. For example, when the user instruction is "click the blue login button on the homepage", the action type corresponding to the user instruction can be "click", and "click" can be determined as the user's intention. In an embodiment of the present invention, a descriptive target refers to the semantic feature of the user's positioning of the page element extracted from the user instruction. For example, from the user instruction "click the blue login button on the homepage", the descriptive targets "blue", "login" and "homepage" can be extracted. The user intention and the descriptive target can then be used as descriptive text information for data processing processes such as screening candidate page elements in the target page and similarity calculation.

[0070] In one embodiment of the present invention, determining the similarity between at least one text information of the candidate page element and the user intention and / or the descriptive target may include:

[0071] Obtaining a page element screenshot of the candidate page element in the target page;

[0072] The text information of the page element screenshot is recognized by optical character recognition technology, and a first similarity between the text information and the user intention and / or the descriptive target is determined.

[0073] The text recognition solution may include a solution based on OCR (Optical Character Recognition) technology, and the text recognition solution is implemented through a browser.

[0074] In an embodiment of the present invention, the browser can use OCR technology to identify the page element screenshot of the candidate page element, and identify the text information of the page element screenshot through COR, and then determine the first similarity between the text information and the user intention, descriptive target, or the user intention and the descriptive target, that is, the first similarity between the text information and the user intention and / or the descriptive target.

[0075] In one embodiment of the present invention, determining the similarity between at least one text information of the candidate page element and the user intention and / or the descriptive target may include:

[0076] Obtaining the page code of the candidate page element in the target page;

[0077] The text information corresponding to the candidate page element is extracted from the page code, and a second similarity between the text information and the user intention and / or the descriptive target is determined.

[0078] The text recognition solution is implemented based on the page code of the target page, and the text recognition solution is implemented through a browser.

[0079] In an embodiment of the present invention, the browser can obtain the page code of the candidate page elements of the target page, extract the text information corresponding to the candidate page elements from the page code, and then determine the second similarity between the text information and the user intention, the descriptive target, or the user intention and the descriptive target, that is, the second similarity between the text information and the user intention and / or the descriptive target.

[0080] In one embodiment of the present invention, determining the similarity between at least one text information of the candidate page element and the user intention and / or the descriptive target may include:

[0081] Obtaining a page element screenshot of the candidate page element in the target page;

[0082] Sending the page element screenshot to a server; the server is configured to call a multimodal visual model to identify text information of the page element screenshot;

[0083] The text information sent by the server is received, and a third similarity between the text information and the user intention and / or the descriptive target is determined.

[0084] Among them, the text recognition solution is a solution implemented based on a multimodal visual model deployed on the server, and the text recognition solution is implemented through the server.

[0085] In an embodiment of the present invention, the browser can take a screenshot of the page element of the candidate page element and send the screenshot of the page element to the server. The server can call the modal visual model to obtain the text information from the screenshot of the page element and send the text information to the browser. The browser can determine the third similarity between the text information and the user intention, the descriptive target, or the user intention and the descriptive target, that is, the third similarity between the text information and the user intention and / or the descriptive target.

[0086] In one embodiment of the present invention, sending the screenshot of the page element to the server includes:

[0087] When the preset conditions are met, the screenshot of the page elements is sent to the server;

[0088] The preset conditions include at least: the description text information includes preset text information, the difference between the first similarity and the second similarity is greater than a preset similarity threshold, and / or the network connection status with the server is higher than a preset network connection status.

[0089] In actual applications, the computing power of the server is much higher than that of the browser. Therefore, if high accuracy is required, a screenshot of the page element can be sent to the server to calculate the third similarity through the multimodal visual model of the server, and then the total similarity can be calculated based on multiple similarities such as the first similarity, the second similarity and the third similarity, thereby improving the positioning accuracy of the target page element of the target page. However, due to the interaction between the browser and the server, affected by factors such as the network connection status between the browser and the server and the server cost, if the third similarity is calculated through the multimodal visual model of the server in all cases, it will lead to increased costs and increased overall testing time. Therefore, in an embodiment of the present invention, the screenshot of the page element can be sent to the server only when the preset conditions are met, so as to calculate the third similarity through the multimodal visual model of the server.

[0090] In one embodiment of the present invention, the descriptive text information includes preset text information, and the preset text information reflects that the page element screenshot of the page element requires strong computing power to be identified. For example, the preset text information can be visual text information, such as "blue". For another example, the preset text information can be shape information, such as "circle", indicating that it is necessary to identify the color information or outline information in the page element screenshot of the page element. Then, it can be determined that the page element screenshot requires strong computing power to be identified, and the page element screenshot of the page element can be sent to the server to calculate the third similarity through the multimodal visual model of the server.

[0091] In another embodiment of the present invention, the difference between the first similarity and the second similarity is greater than a preset similarity threshold, indicating that the similarity results calculated by the text recognition scheme based on OCR technology and the text recognition scheme implemented based on the page code of the target page are quite different, indicating that the text information recognized by one of the text recognition schemes is unreliable / mismatched, and a screenshot of the page element of the page element can be sent to the server to calculate the third similarity through the multimodal visual model of the server, and then the similarities of multiple text recognition schemes including the text recognition scheme implemented based on the server can be combined to locate the target page element in the target page, thereby ensuring the accuracy of page element positioning.

[0092] In another embodiment of the present invention, a screenshot of the page element will be sent to the server only when the network connection status between the browser and the server is higher than the preset network connection status, so as to calculate the third similarity through the multimodal visual model of the server. This can avoid calling the multimodal visual model of the server when the network connection status between the browser and the server is not good, which will lead to the problem of increasing the overall test time.

[0093] In one embodiment of the present invention, determining the total similarity of the page elements according to the similarity of the page elements may include:

[0094] Obtaining text recognition schemes corresponding to the similarities of the page elements; the text recognition schemes respectively have corresponding weight values;

[0095] The total similarity of the page elements is determined according to the similarity and weight values ​​of the page elements.

[0096] In the embodiment of the present invention, different weight values ​​may be assigned to similarities obtained from different text recognition schemes, so that the total similarity of the page elements may be obtained based on the similarities and weight values ​​of multiple page elements.

[0097] In one embodiment, assuming there are three text recognition schemes, the calculation formula for the total similarity is as follows:

[0098] s=a*l+b*m+c*n;

[0099] a+b+c=1

[0100] Among them, s represents the total similarity; l, m, and n represent the similarities calculated based on the text information corresponding to the three text recognition schemes; a, b, and c represent the weight values ​​corresponding to the three text recognition schemes respectively.

[0101] In order to help those skilled in the art better understand the embodiments of the present invention, a specific example is used below to illustrate the method for matching DOM elements in a target page based on action type and descriptive target. The specific process can be as follows:

[0102] 1. Parse the DOM tree of the target page and optimize and reduce it based on the description text information, retaining only the candidate target list that meets the requirements. The description text information may include user intent (action type) and descriptive target.

[0103] 2. Use client-side OCR to identify the text information in the candidate page elements in the candidate target list and calculate the similarity l with the description text information;

[0104] 4. Directly calculate the similarity m between the text information contained in the page code of the candidate target list and the description text information;

[0105] 5. Use the server-side multimodal vision model to identify the similarity n between the candidate target list and the description text information;

[0106] 6. Calculate the total similarity s = a*l+b*m+c*n, a+b+c=1.

[0107] In the above embodiment, the context of the target page is compressed by pruning, and the final total similarity is calculated by weighted method to improve the robustness of page element recognition in the target page. At the same time, the multimodal visual model on the server side is used to better understand user instructions in the case of text mismatch, thereby improving the accuracy of locating target page elements in the browser based on user instructions.

[0108] It should be noted that for the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions involved are not necessarily required for the embodiments of the present invention.

[0109] Reference Figure 2 , is a structural block diagram of a page element positioning device provided in an embodiment of the present invention, such as Figure 2 As shown, the device may specifically include the following modules:

[0110] The user instruction receiving module 201 is used to receive a user instruction for a target page; the target page includes at least one page element;

[0111] A description text information generating module 202 is configured to generate description text information according to the user instruction; the description text information at least includes the user intention and a descriptive goal;

[0112] A similarity determination module 203 is configured to determine the similarity between at least one text information of the page element and the description text information; the text information is obtained by at least one text recognition scheme;

[0113] A total similarity determination module 204 is configured to determine the total similarity of the page elements according to the similarities of the page elements;

[0114] The target page element locating module 205 is configured to locate a target page element from the page elements in the target page according to the total similarity.

[0115] In one embodiment of the present invention, the similarity determination module 203 is configured to:

[0116] Filtering candidate page elements from page elements in the target page according to the user intention and / or the descriptive goal;

[0117] A similarity between at least one text information of the candidate page element and the user intention and / or the descriptive target is determined.

[0118] In one embodiment of the present invention, the similarity determination module 203 is configured to:

[0119] Obtaining a page element screenshot of the candidate page element in the target page;

[0120] The text information of the page element screenshot is recognized by optical character recognition technology, and a first similarity between the text information and the user intention and / or the descriptive target is determined.

[0121] In one embodiment of the present invention, the similarity determination module 203 is configured to:

[0122] Obtaining the page code of the candidate page element in the target page;

[0123] The text information corresponding to the candidate page element is extracted from the page code, and a second similarity between the text information and the user intention and / or the descriptive target is determined.

[0124] In one embodiment of the present invention, the similarity determination module 203 is configured to:

[0125] Obtaining a page element screenshot of the candidate page element in the target page;

[0126] Sending the page element screenshot to a server; the server is configured to call a multimodal visual model to identify text information of the page element screenshot;

[0127] The text information sent by the server is received, and a third similarity between the text information and the user intention and / or the descriptive target is determined.

[0128] In one embodiment of the present invention, the similarity determination module 203 is further configured to:

[0129] When the preset conditions are met, the screenshot of the page elements is sent to the server;

[0130] The preset conditions include at least: the description text information includes preset text information, the difference between the first similarity and the second similarity is greater than a preset similarity threshold, and / or the network connection status with the server is higher than a preset network connection status.

[0131] In one embodiment of the present invention, the total similarity determination module 204 is configured to:

[0132] Obtaining text recognition schemes corresponding to the similarities of the page elements; the text recognition schemes respectively have corresponding weight values;

[0133] The total similarity of the page elements is determined according to the similarity and weight values ​​of the page elements.

[0134] In one embodiment of the present invention, generating descriptive text information according to the user instruction includes:

[0135] Obtaining an action type corresponding to the user instruction, and using the action type as the user intention;

[0136] A descriptive target is generated according to the user instruction, and the user intention and the descriptive target are used as descriptive text information.

[0137] In an embodiment of the present invention, a user instruction for a target page is received, the target page may include at least one page element, descriptive text information is generated according to the user instruction, wherein the descriptive text information may include at least the user intention and a descriptive target, and then the similarity between at least one text information of the page element and the descriptive text information is determined, the text information is obtained through at least one text recognition scheme, and the total similarity of the page elements is determined according to the similarity of the page elements, so as to locate the target page element from the page elements in the target page according to the total similarity. In an embodiment of the present invention, the total similarity can be determined based on the similarity between the text information of at least one page element and the descriptive text information of the user instruction, and then based on at least one similarity, the target page element can be located from the page elements in the target page based on the total similarity. Since the page element is not located in the target page based on a single similarity, the robustness and accuracy of the page element location are improved.

[0138] As for the above-mentioned device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0139] It should be noted that the embodiments of the present invention may involve the use of user data. In actual applications, user-specific personal data may be used in the scheme described herein within the scope permitted by applicable laws and regulations of the country where the user is located (for example, with the user's explicit consent, effective notification to the user, etc.).

[0140] The embodiment of the present invention further provides an electronic device, such as Figure 3 As shown, it includes a processor 301, a communication interface 302, a memory 303 and a communication bus 304, wherein the processor 301, the communication interface 302, and the memory 303 communicate with each other through the communication bus 304.

[0141] Memory 303, for storing computer programs;

[0142] The processor 301 is configured to implement the page element positioning method described in any one of the above embodiments when executing the program stored in the memory 303:

[0143] Receiving a user instruction for a target page; the target page includes at least one page element;

[0144] generating descriptive text information according to the user instruction; the descriptive text information at least including the user intention and the descriptive target;

[0145] Determining a similarity between at least one text information of the page element and the description text information; the text information is obtained by at least one text recognition scheme;

[0146] Determining the total similarity of the page elements according to the similarity of the page elements;

[0147] A target page element is located from the page elements in the target page according to the total similarity.

[0148] The communication bus mentioned in the terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.

[0149] The communication interface is used for communication between the above terminal and other devices.

[0150] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage. Alternatively, the memory may be at least one storage device located away from the processor.

[0151] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.

[0152] In another embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores instructions that, when executed on a computer, enable the computer to execute the page element locating method described in any one of the above embodiments.

[0153] In another embodiment of the present invention, a computer program product including instructions is provided. When the computer program product is run on a computer, the computer executes the page element positioning method described in any one of the above embodiments.

[0154] In the above embodiments, they can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0155] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0156] Each embodiment in this specification is described in a related manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are described briefly because they are generally similar to the method embodiments. For related portions, reference can be made to the description of the method embodiments.

[0157] The above description is only a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the scope of protection of the present invention.

Claims

1. A page element positioning method, characterized in that: The method comprises: Receiving a user instruction for a target page; the target page includes at least one page element; generating descriptive text information according to the user instruction; the descriptive text information at least including the user intention and the descriptive target; Determining a similarity between at least one text information of the page element and the description text information; the text information is obtained by at least one text recognition scheme; Determining the total similarity of the page elements according to the similarity of the page elements; A target page element is located from the page elements in the target page according to the total similarity.

2. The method according to claim 1, characterized in that Determining the similarity between at least one text information of the page element and the descriptive target includes: Filtering candidate page elements from page elements in the target page according to the user intention and / or the descriptive goal; A similarity between at least one text information of the candidate page element and the user intention and / or the descriptive target is determined.

3. The method according to claim 2, characterized in that Determining the similarity between at least one text information of the candidate page element and the user intention and / or the descriptive target includes: Obtaining a page element screenshot of the candidate page element in the target page; The text information of the page element screenshot is recognized by optical character recognition technology, and a first similarity between the text information and the user intention and / or the descriptive target is determined.

4. The method according to claim 3, characterized in that Determining the similarity between at least one text information of the candidate page element and the user intention and / or the descriptive target includes: Obtaining the page code of the candidate page element in the target page; The text information corresponding to the candidate page element is extracted from the page code, and a second similarity between the text information and the user intention and / or the descriptive target is determined.

5. The method according to claim 4, characterized in that Determining the similarity between at least one text information of the candidate page element and the user intention and / or the descriptive target includes: Obtaining a page element screenshot of the candidate page element in the target page; Sending the page element screenshot to a server; the server is configured to call a multimodal visual model to identify text information of the page element screenshot; The text information sent by the server is received, and a third similarity between the text information and the user intention and / or the descriptive target is determined.

6. The method according to claim 5, characterized in that Sending a screenshot of the page elements to the server includes: When the preset conditions are met, the screenshot of the page elements is sent to the server; The preset conditions include at least: the description text information includes preset text information, the difference between the first similarity and the second similarity is greater than a preset similarity threshold, and / or the network connection status with the server is higher than a preset network connection status.

7. The method according to claim 1, characterized in that Determining the total similarity of the page elements according to the similarity of the page elements includes: Obtaining text recognition schemes corresponding to the similarities of the page elements; the text recognition schemes respectively have corresponding weight values; The total similarity of the page elements is determined according to the similarity and weight values ​​of the page elements.

8. The method according to claim 1, characterized in that Generating description text information according to the user instruction, including: Obtaining an action type corresponding to the user instruction, and using the action type as the user intention; A descriptive target is generated according to the user instruction, and the user intention and the descriptive target are used as descriptive text information.

9. A page element positioning device, characterized in that: The device comprises: A user instruction receiving module, configured to receive a user instruction for a target page; the target page includes at least one page element; A description text information generating module, configured to generate description text information according to the user instruction; the description text information at least includes the user intention and a descriptive target; a similarity determination module, configured to determine a similarity between at least one text information of the page element and the description text information; the text information being obtained by at least one text recognition scheme; A total similarity determination module, configured to determine the total similarity of the page elements according to the similarities of the page elements; The target page element locating module is configured to locate a target page element from the page elements in the target page according to the total similarity.

10. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory for storing computer programs; A processor, configured to implement the method according to any one of claims 1 to 8 when executing a program stored in a memory.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Page control method and device

    CN107919129A

  • Page element positioning method, electronic equipment and storage medium

    CN117472744A

  • Chinese text recognition method based on deep paradigm

    CN118247796A

  • Blurred image judgment method and device based on image OCR comparison result

    CN119671950A

  • Interface element control method and device, electronic equipment and storage medium

    CN119690288A