Web page operation method and device, equipment and medium

By receiving user instructions in a web page and determining the DOM node to generate operation instructions, the problem of users having difficulty locating operation elements in diverse web pages is solved, thereby improving operation efficiency and user experience.

CN120743151APending Publication Date: 2025-10-03CHONGQING JINKANG NEW ENERGY VEHICLE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510673878.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

When users use diverse and complex web pages, it is difficult for them to effectively locate action buttons or input boxes, resulting in low operational efficiency and poor user experience. This is especially true when filling out multi-field forms or navigating deep-level menus, where cognitive load is excessive.

Method used

By receiving the operation object description in the user instruction, determining the DOM node to be processed, and generating the corresponding operation instruction, the operation is performed directly in the DOM tree, avoiding the problem that the user cannot find the target element in a complex Web page.

Benefits of technology

It improves the efficiency of user interaction with web pages and reduces the problems of low operation efficiency and poor user experience caused by not being able to find operation buttons or input boxes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743151A_ABST
    Figure CN120743151A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a Web page operation method and device, equipment and a medium, which are used for improving the efficiency of operating a Web page by a user so as to improve the user experience. The method comprises the steps of receiving a user instruction; wherein the user instruction comprises operation object description; determining a to-be-processed DOM node based on the operation object description, and generating an operation instruction; and executing the operation instruction on the DOM node to be processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, apparatus, device, and medium for operating a web page. Background Art

[0002] To meet diverse functional needs, web page layouts are becoming increasingly rich and complex. Different web pages exhibit varying design styles due to their varying interaction logic. This requires users to devote more energy and attention to locating the desired button among numerous icons and buttons when using a web page.

[0003] Especially when it comes to filling out multi-field forms or navigating deep-level menus, users not only need to deal with the visual parsing pressure of high-density information flow, but also have to bear the cognitive load brought by multi-step operation processes, resulting in low processing efficiency. Summary of the Invention

[0004] Based on this, it is necessary to provide a web page operation method, device, equipment and medium to address the above technical problems, so as to improve the efficiency of user operation of web pages and thus enhance user experience.

[0005] In a first aspect, an embodiment of the present application provides a method for operating a web page, including:

[0006] Receiving a user instruction; wherein the user instruction includes an operation object description;

[0007] Based on the operation object description, determine the DOM node to be processed and generate an operation instruction;

[0008] Execute the operation instruction on the DOM node to be processed.

[0009] In some embodiments, determining a DOM node to be processed based on the operation object description and generating an operation instruction includes:

[0010] Determining, based on the operation object description, a visual element in the first web page and element information of the visual element;

[0011] In a first DOM tree corresponding to the first Web page, based on the element information, identifying a DOM node to be processed corresponding to the visual element;

[0012] Based on the operation object description, the operation instruction corresponding to the DOM node to be processed is generated.

[0013] In some embodiments, generating the operation instruction corresponding to the DOM node to be processed based on the operation object description includes:

[0014] Extracting the operation type from the operation object description and the operation parameter value corresponding to the operation type;

[0015] According to a preset order, the operation type and the operation parameter value are encapsulated into the operation instruction of a preset structure.

[0016] In some embodiments, in a first DOM tree corresponding to the first web page, identifying a to-be-processed DOM node corresponding to the visual element based on the element information includes:

[0017] Determining a positioning rule in the element information;

[0018] Accessing the first DOM tree corresponding to the first Web page;

[0019] In the first DOM tree, the DOM node to be processed is identified based on the positioning rule in the element information of the visual element.

[0020] In some embodiments, determining the positioning rule in the element information includes:

[0021] In response to the element information containing an XPath expression or an element index, the XPath expression or the element index is determined to be the positioning rule.

[0022] In some embodiments, determining a visual element in the first web page and element information of the visual element based on the operation object description includes:

[0023] Inputting the operation object description into a pre-trained text processing model to obtain feature information of the visualization element;

[0024] In response to the feature information including visual information and text information of the visual element, detecting, based on the visual information, a visual element to be verified that matches the visual information in the first web page;

[0025] Based on the text information in the feature information, the visual element to be verified is verified to obtain the visual element; and then the element information corresponding to the visual element is determined.

[0026] In some embodiments, determining a visual element in the first web page and element information of the visual element based on the operation object description includes:

[0027] Based on the characteristic information of the visual element in the description of the operation object, identifying the visual element in the first web page;

[0028] The element information corresponding to the visual element is determined.

[0029] In a second aspect, an embodiment of the present application provides a web page operation device, including:

[0030] A user module is configured to receive a user instruction; wherein the user instruction includes an operation object description;

[0031] An instruction module is used to determine the DOM node to be processed based on the operation object description and generate an operation instruction;

[0032] The node module is used to execute the operation instruction on the DOM node to be processed.

[0033] In some embodiments, the instruction module is specifically used to determine a visual element and element information of the visual element in a first web page based on the operation object description; identify a DOM node to be processed corresponding to the visual element in a first DOM tree corresponding to the first web page based on the element information; and generate the operation instruction corresponding to the DOM node to be processed based on the operation object description.

[0034] In some embodiments, the instruction module is specifically used to extract the operation type in the operation object description and the operation parameter value corresponding to the operation type; and encapsulate the operation type and the operation parameter value into the operation instruction of a preset structure according to a preset order.

[0035] In some embodiments, the instruction module is further used to extract the operation type in the operation object description and the operation parameter value corresponding to the operation type; and encapsulate the operation type and the operation parameter value into the operation instruction of a preset structure according to a preset order.

[0036] In some embodiments, the instruction module is further configured to determine a positioning rule in the element information; access the first DOM tree corresponding to the first web page; and identify the DOM node to be processed in the first DOM tree based on the positioning rule in the element information of the visual element.

[0037] In some embodiments, the instruction module is further configured to, in response to the element information containing an XPath expression or an element index, determine the XPath expression or the element index as the positioning rule.

[0038] In some embodiments, the instruction module is further used to input the operation object description into a pre-trained text processing model to obtain feature information of the visual element; in response to the feature information containing visual information and text information of the visual element, based on the visual information, a visual element to be verified that matches the visual information is detected in the first web page; based on the text information in the feature information, the visual element to be verified is verified to obtain the visual element; and then the element information corresponding to the visual element is determined.

[0039] In some embodiments, the instruction module is further configured to identify the visual element in the first web page based on feature information of the visual element in the operation object description; and determine the element information corresponding to the visual element.

[0040] In a third aspect, an embodiment of the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and for running on the processor, wherein when the processor executes the computer program, the method described in the first aspect and any embodiment is implemented.

[0041] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the method described in the first aspect and any one of the embodiments is implemented.

[0042] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the method described in the first aspect and any one of the embodiments.

[0043] In the web page operation method provided in the embodiments of the present application, after receiving a user instruction, the method extracts the operation object description of the user instruction, determines the pending DOM node that matches the user intent contained in the user instruction, and generates an operation instruction corresponding to the pending DOM node. In this way, by executing the operation instruction on the pending DOM node, a direct response to the user instruction is achieved, avoiding the problem of low operation efficiency and poor user experience caused by users being unable to find visual elements such as operation buttons or input boxes on web pages, especially when users are faced with diverse and rich web pages.

[0044] Other features and advantages of the present application will be described in the following description and, in part, will become apparent from the description or may be learned through practice of the present application. The objectives and other advantages of the present application may be achieved and obtained through the structures particularly pointed out in the written description, claims, and drawings. It should be understood that the above general description and the detailed description that follow are merely exemplary and explanatory and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0046] Figure 1 A flowchart of a method for operating a web page provided in an embodiment of the present application;

[0047] Figure 2 A flowchart of a method for generating an operation instruction provided in an embodiment of the present application;

[0048] Figure 3 A diagram illustrating the training of a text processing model provided in an embodiment of the present application;

[0049] Figure 4 A structural block diagram of a web page operation device provided in an embodiment of the present application;

[0050] Figure 5 This is a diagram of the internal structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0052] It should be noted that the diagrams provided in the present embodiment are only schematic illustrations of the basic concept of the present application. The diagrams only show the components related to the present application rather than the number, shape and size of the components when actually implemented. The type, quantity and ratio of each component can be changed at will during actual implementation, and the component layout pattern may also be more complicated. The structures, ratios, sizes, etc. illustrated in the drawings of this specification are only used to match the content disclosed in the specification for people familiar with this technology to understand and read. They are not used to limit the restrictive conditions that can be implemented in this application. Therefore, they have no technical significance. Any modification of the structure, change of the proportional relationship or adjustment of the size should still fall within the scope of the technical content disclosed in this application without affecting the effect and purpose that can be achieved by this application. At the same time, the terms such as "upper", "lower", "left", "right", "middle" and "one" quoted in this specification are only for the convenience of description and are not used to limit the scope of the implementation of this application. The change or adjustment of their relative relationship should also be considered as the scope of the implementation of this application without substantial change in the technical content.

[0053] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various places herein does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0054] As used herein, unless the context clearly indicates otherwise, the terms "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include additional steps or elements.

[0055] The definition of inclusion herein, such as the terms “having”, “may have”, “include” or “may include” as used herein, indicates the existence of the corresponding functions, operations, elements, etc. herein, and does not limit the existence of one or more other functions, operations, elements, etc. In addition, it should be understood that the terms “including” or “having” as used herein indicate the existence of the features, numbers, steps, operations, elements, components or their combination described in the specification, and do not exclude the existence or addition of one or more other features, numbers, steps, operations, elements, components or their combination.

[0056] In the embodiments of the present application, prefixes such as "first" and "second" are used only to distinguish different description objects and have no limiting effect on the position, order, priority, quantity or content of the described objects. In the embodiments of the present application, the use of prefixes such as ordinal numbers to distinguish description objects does not constitute a restriction on the described objects. For the statement of the described + object, please refer to the description in the context of the claims or embodiments, and no unnecessary restrictions should be constituted due to the use of such prefixes. In addition, in the description of this embodiment, unless otherwise specified, the meaning of "plurality" is two or more.

[0057] To facilitate understanding of the technical solutions provided by the embodiments of the present application, the following first introduces the design concepts of the embodiments of the present application:

[0058] A web page (World Wide Web) is the basic information unit of the World Wide Web. It is written in HTML (Hypertext Markup Language) and is an electronic document accessible through a browser. Web pages are used to present information to users and meet their interactive needs. However, as web page functionality has diversified, web page styles have become richer and more diverse. This has reduced the efficiency of users searching for input boxes or buttons that match their needs when using web pages.

[0059] To this end, an embodiment of the present application provides a web page operation method, which determines the DOM node to be processed through the operation object description in the user instruction, and generates an operation instruction for the DOM node to be processed, thereby converting the user instruction into an operation instruction for processing the DOM node to be processed in the DOM tree, achieving efficient response to user instructions and improving the interaction efficiency between the user and the web page.

[0060] The following is a detailed description of the web page operation method provided by the embodiment of the present application in conjunction with the accompanying drawings. Figure 1 , the method comprises the following implementation steps:

[0061] Step 101: Receive user instructions.

[0062] The user instruction includes an operation object description.

[0063] Specifically, the user instruction can be obtained by preprocessing the user's interaction demand information. The interaction demand information can be voice data input by the user in response to a web page operation via a virtual voice input button. Alternatively, the interaction demand information can be text data input by the user via a command box in a pre-specified location on the web page.

[0064] If the interaction demand information is voice data, the interaction demand information can be first translated into text data and then pre-processed, which includes: denoising to remove noise characters, redundant punctuation marks, etc. in the text data.

[0065] Furthermore, user instructions can be parsed based on NLU (Natural Language Understanding). During the parsing process, grammatical errors or incomplete instructions are automatically completed and corrected to obtain a description of the operation object in the user instruction.

[0066] The operation object description corresponds to the user intention in the user instruction. The operation object description includes at least feature information of a visual element in the first web page. The visual element refers to a visual element relative to the user.

[0067] For example, the operation object description in the user instruction may be: delete the data cell in the second row and third column of the table. Then the visual element in the operation object description is: the data cell in the second row and third column of the table.

[0068] Step 102: Based on the operation object description, determine the DOM node to be processed and generate an operation instruction.

[0069] Specifically, the DOM node to be processed refers to a node located in a first DOM (Document Object Model) tree.

[0070] In one embodiment, the DOM node to be processed may include at least one of a DOM element node, a DOM text node, and a DOM script and style node.

[0071] The DOM node to be processed corresponds to the visual element in the first web page, and the first DOM tree corresponds to the first web page.

[0072] In one embodiment, the first DOM tree corresponding to the first web page can be constructed when parsing the corresponding HTML file and rendering the first web page. Alternatively, it can be reconstructed after executing an operation instruction corresponding to the user operation object description. Therefore, in step 202, the DOM node to be processed can be determined by accessing the first DOM tree. Specifically, the first web page can be processed using a DOM parsing interface or a pre-specified DOM parsing library to obtain the pre-constructed first DOM tree.

[0073] Furthermore, the above-mentioned operation instruction may include an operation type, an operation parameter corresponding to the operation type, and location information of a DOM node to be processed.

[0074] For example, the operation type may be clicking, inputting text, selecting a drop-down option, submitting a form, and the like.

[0075] Exemplarily, the operation parameter corresponding to the operation type may be the coordinates of the click, the output text content, or the selected value in the drop-down option.

[0076] Step 103: Execute the operation instruction on the DOM node to be processed.

[0077] Specifically, the DOM node to be processed corresponds to the operation instruction.

[0078] In one embodiment, the operation instructions may be extracted from the operation instruction queue according to a queue order, wherein the operation instruction queue is ordered according to the generation time of the operation instructions.

[0079] Then, the operation instruction is parsed to obtain the position information of the DOM node to be processed in the first DOM tree, as well as the operation type and operation parameter values.

[0080] Finally, based on the location information in the first DOM tree, the processing corresponding to the operation type can be performed on the DOM node to be processed according to the operation parameter value, thereby executing the operation instruction. By reading the location information in the operation instruction, the DOM node to be processed can be located in the first DOM tree, and based on the operation type and parameter value in the operation instruction, the processing corresponding to the DOM node to be processed in the first DOM tree can be performed.

[0081] After the operation instruction is executed, the first DOM tree is updated to obtain a second DOM tree, and the first web page is also updated accordingly to obtain a second web page.

[0082] The web page operation method described in steps 101 to 103 converts the visual elements of the first web page into operation instructions for the DOM nodes to be processed in the first DOM tree based on the operation object description in the user instruction, and executes the operation instructions on the DOM nodes to be processed. This can avoid the problem of low processing efficiency caused by the user having difficulty finding the button or form to be processed in the web page due to the complexity of the web page.

[0083] In one embodiment, the operation instruction is generated in the following manner, please refer to Figure 2 .

[0084] Step 201: Determine a visual element and element information of the visual element in a first web page based on an operation object description.

[0085] The operation object description may include feature information describing the location, appearance, and other characteristics of the visual element. Based on the operation object description, the visual element can be located within the first web page, and element information corresponding to the visual element can be determined. Optionally, the element information can be generated when constructing the first DOM tree corresponding to the first web page, and pre-bound to the visual element to establish a corresponding relationship.

[0086] In one embodiment, the visual element may include a first type of visual element and a second type of visual element.

[0087] Exemplarily, the first type of visual element may be a form element, such as a text input box, a radio button, a check box, or the like.

[0088] The second type of visual element may be a multimedia element, such as an image element, an audio element, a video element, and the like.

[0089] Step 202 : In a first DOM tree corresponding to the first Web page, based on the element information, identify a DOM node to be processed corresponding to the visual element.

[0090] In one embodiment, a positioning rule in the element information may be determined first. Then, a first DOM tree corresponding to the first web page may be accessed. Finally, a DOM node to be processed may be quickly identified in the first DOM tree based on the positioning rule in the element information of the visual element.

[0091] In one embodiment, the positioning rule may also be composed of multiple attribute values, so as to locate the aforementioned DOM node to be processed in the first DOM tree based on the multiple attribute values. The multiple attribute values ​​may include, but are not limited to, button, select, option (drop-down list and options in the drop-down list), input, etc.

[0092] To further improve the efficiency of locating the DOM node to be processed, in one embodiment, in response to the element information containing an XPath expression or an element index, the XPath expression or element index is determined as the positioning rule. The element index can be an element ID.

[0093] Step 203: Generate an operation instruction corresponding to the DOM node to be processed based on the operation object description.

[0094] Specifically, a pre-trained text processing model may be used to extract the operation type in the operation object description and the operation parameter value corresponding to the operation type.

[0095] Then, according to a preset sequence, the operation type and the operation parameter value are encapsulated into an operation instruction of a preset structure.

[0096] Furthermore, after the aforementioned operation instruction is generated, the corresponding relationship between the operation instruction and the DOM node to be processed may be recorded.

[0097] Furthermore, the aforementioned user instructions may include descriptions of operation objects corresponding to multiple visual elements. Based on the aforementioned pre-trained text processing model, multiple descriptions of the operation objects can be obtained, arranged in the order of the operations, to form a first sequence. Then, operation instructions can be generated that correspond one-to-one to the multiple descriptions of the operation objects. Finally, the operation instructions can be arranged based on the order of the multiple descriptions of the operation objects in the first sequence to obtain an instruction sequence.

[0098] In one embodiment, the visualization element may be determined in the first web page by extracting feature information from the description of the operation object and based on the feature information.

[0099] The characteristic information of the visualization element may include at least one of visual information, text information, and position information of the visualization element.

[0100] The visual information is used to describe the visual features of the visual element in the first web page. The visual information may include shape information and / or color information. For example, it is used to describe a folder storing deleted data as a "gray trash can shape."

[0101] The text information is the text information carried by the visual element, such as the LOGO (LOGO type) carried on an icon. The position information can be used to describe the position of the visual element in the first web page.

[0102] The following provides an embodiment to illustrate the determination of a visual element: First, the description of the operation object can be input into a pre-trained text processing model to obtain feature information of the visual element. In response to the feature information including the visual information and text information of the visual element, based on the visual information, a visual element to be verified that matches the visual information is detected in the first web page. Specifically, this can be obtained by processing a screenshot of the first web page through target detection. Then, based on the text information in the feature information, the visual element to be verified can be verified to accurately obtain the visual element. That is, in response to the text data contained in the visual element to be verified being consistent with the text information, the visual element to be verified can be determined to be a visual element.

[0103] Alternatively, in response to the text data contained in the visual element to be verified being inconsistent with the text information, the first web page is again screenshotted and the target detection process is performed until a visual element containing text data consistent with the text information in the feature information is identified.

[0104] Furthermore, to identify visual elements, the first HTML tag corresponding to the first web page can be used for identification. This is because the first web page is obtained by parsing an HTML file, and during the process of parsing the HTML file to obtain the first web page, a corresponding first DOM tree is also constructed. Therefore, the visual elements in the first web page correspond to the DOM nodes in the first DOM tree. The HTML nodes in the corresponding HTML document also correspond to the DOM nodes in the first DOM tree. Based on this, in one embodiment, the characteristic information of the aforementioned visual elements may also include the first HTML tag. Visual elements can then be identified using the following method: First, based on the operation characteristic description, the first HTML tag information can be extracted and the first HTML tag generated. Within the HTML document corresponding to the first web page, HTML nodes whose HTML tags satisfy the first HTML tag are screened. Then, based on the location information in the characteristic information, the visual element in the first web page can be uniquely located. For example, if the operation object description is "a data cell in the second row and third column of a table," and the first HTML tag corresponding to the data cell in the operation object description is a tag, then based on this tag and the location information (second row and third column), the data cell (i.e., the visual element) in the first web page can be identified.

[0105] Furthermore, the characteristic information of the operation object description and the visual elements in the operation object description in the embodiment of the present application can be obtained by inputting the user instruction into a pre-trained text processing model, or inputting the operation object description into a pre-selected trained text processing model. The following describes the training of the pre-trained text processing model, please refer to Figure 3 .

[0106] Specifically, interaction data may be collected first. Then, the interaction data may be labeled and stored. Labeling the interaction data here indicates labeling the operation object description in the interaction data, the operation type in the operation object description, and the operation parameter value corresponding to the operation type.

[0107] Next, the labels obtained from the aforementioned annotated interaction data can be used together with the interaction data to conduct supervised training on the text processing model until the loss function in the text processing model converges, and the training is determined to be complete. The annotated interaction data is the labeled user instructions.

[0108] Finally, the trained text processing model can be used to extract descriptions of operational objects and / or feature information of visual elements. To avoid insufficient generalization or robustness of the trained text processing model, the trained text processing model can be optimized at predetermined intervals based on user instructions and user feedback received during its use.

[0109] The above text processing model can be set up based on the Transformer Architecture.

[0110] Furthermore, before executing an operation instruction on the DOM node to be processed, the loading status of the first web page can be verified to ensure that the DOM node to be processed is in an operable state, thereby avoiding failure in executing the operation instruction. For example, the loading status of the first web page is verified. If it is determined that the first web page is fully loaded, the DOM node to be processed can be determined to be in an operable state. Otherwise, the DOM node to be processed can be determined to be in an inoperable state. The first web page is then refreshed, or the loading status of the first web page is verified again after a preset period of time until the first web page is fully loaded, to ensure that the operation instruction can be successfully executed and improve the user experience.

[0111] Furthermore, after executing the operation instructions and generating the second web page and the second DOM tree, if response information is received from the server (for example, prompt information after filling in and submitting a form, etc.), it can be converted into natural language together with page performance indicators (such as operation execution time, network request time, etc.) and displayed to the user.

[0112] It should be understood that although Figure 1 、 Figure 2 、 Figure 3 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 1 、 Figure 2 、 Figure 3 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0113] Based on the same inventive concept, Figure 4 As shown, the embodiment of the present application provides a web page operation device, including: a user module 401, an instruction module 402 and a node module 403, wherein:

[0114] The user module 401 is used to receive user instructions.

[0115] Wherein, the user instruction includes an operation object description.

[0116] The instruction module 402 is used to determine the DOM node to be processed based on the operation object description and generate an operation instruction.

[0117] The node module 403 is configured to execute the operation instruction on the DOM node to be processed.

[0118] In one embodiment, the instruction module 402 is specifically configured to:

[0119] Based on the operation object description, a visual element in a first web page and element information of the visual element are determined; in a first DOM tree corresponding to the first web page, a DOM node to be processed corresponding to the visual element is identified based on the element information; and based on the operation object description, the operation instruction corresponding to the DOM node to be processed is generated.

[0120] In one embodiment, the instruction module 402 is specifically configured to:

[0121] Extracting the operation type and the operation parameter value corresponding to the operation type from the operation object description; and encapsulating the operation type and the operation parameter value into the operation instruction of a preset structure according to a preset sequence.

[0122] In one embodiment, the instruction module 402 is further configured to:

[0123] Extracting the operation type and the operation parameter value corresponding to the operation type from the operation object description; and encapsulating the operation type and the operation parameter value into the operation instruction of a preset structure according to a preset sequence.

[0124] In one embodiment, the instruction module 402 is further configured to:

[0125] Determine a positioning rule in the element information; access the first DOM tree corresponding to the first Web page; and identify the DOM node to be processed in the first DOM tree based on the positioning rule in the element information of the visual element.

[0126] In some embodiments, the instruction module is further configured to, in response to the element information containing an XPath expression or an element index, determine the XPath expression or the element index as the positioning rule.

[0127] In one embodiment, the instruction module 402 is further configured to:

[0128] The operation object description is input into a pre-trained text processing model to obtain feature information of the visualization element; in response to the feature information containing visual information and text information of the visualization element, based on the visual information, a visualization element to be verified that matches the visual information is detected in the first web page; based on the text information in the feature information, the visualization element to be verified is verified to obtain the visualization element; and the element information corresponding to the visualization element is determined.

[0129] In one embodiment, the instruction module 402 is further configured to:

[0130] Based on the characteristic information of the visual element in the operation object description, the visual element is identified in the first Web page; and the element information corresponding to the visual element is determined.

[0131] The specific definition of the web page operation device can be found in the definition of the web page operation method above and will not be repeated here. Each module in the above-mentioned web page operation device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor of the electronic device in hardware form, or can be stored in the memory of the electronic device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.

[0132] Based on the same inventive concept, see Figure 5 The present application also provides an electronic device. In one embodiment, the electronic device may include a memory 501, a communication module 503, and one or more processors 502 as shown in the figure.

[0133] The memory 501 is used to store computer programs executed by the processor 502. The memory 501 may mainly include a program storage area and a data storage area. The program storage area may store an operating system; the data storage area may store various operating instruction sets.

[0134] Memory 501 may be a volatile memory, such as random-access memory (RAM); a non-volatile memory, such as read-only memory, flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 501 may be a combination of the above memories.

[0135] The processor 502 may include one or more central processing units (CPUs) or digital processing units, etc. The processor 502 is configured to implement the above-mentioned Web page operation method when calling the computer program stored in the memory 501 .

[0136] The communication module 503 is used to communicate with terminal devices, site devices or other network devices.

[0137] The specific connection medium between the memory 501, the communication module 503 and the processor 502 is not limited in the embodiment of the present application. Figure 5 In the embodiment, the memory 501 and the processor 502 are connected via a bus 504. Figure 5 The connections between the other components are shown in bold lines, which are only for illustration and are not intended to be limiting. The bus 504 can be divided into an address bus, a data bus, a control bus, etc. For ease of description, Figure 5 The diagram shows a single thick line, but this does not indicate that there is only one bus or one type of bus.

[0138] The memory 501 stores a computer storage medium, which stores computer-executable instructions for implementing the method for determining a web page operation according to an embodiment of the present application. The processor 502 is configured to execute the web page operation method according to each embodiment of the computer-executable instructions.

[0139] In one embodiment, when the computer executable instructions are executed by the processor 502, the following steps are further implemented:

[0140] Receive a user instruction, wherein the user instruction includes an operation object description; determine a DOM node to be processed based on the operation object description, and generate an operation instruction; execute the operation instruction on the DOM node to be processed.

[0141] In one embodiment, when the computer executable instructions are executed by the processor 502, the following steps are further implemented:

[0142] Based on the operation object description, a visual element in a first web page and element information of the visual element are determined; in a first DOM tree corresponding to the first web page, a DOM node to be processed corresponding to the visual element is identified based on the element information; and based on the operation object description, the operation instruction corresponding to the DOM node to be processed is generated.

[0143] In one embodiment, when the computer executable instructions are executed by the processor 502, the following steps are further implemented:

[0144] Extracting the operation type and the operation parameter value corresponding to the operation type from the operation object description; and encapsulating the operation type and the operation parameter value into the operation instruction of a preset structure according to a preset sequence.

[0145] In one embodiment, when the computer executable instructions are executed by the processor 502, the following steps are further implemented:

[0146] Determine a positioning rule in the element information; access the first DOM tree corresponding to the first Web page; and identify the DOM node to be processed in the first DOM tree based on the positioning rule in the element information of the visual element.

[0147] In one embodiment, when the computer executable instructions are executed by the processor 502, the following steps are further implemented:

[0148] In response to the element information containing an XPath expression or an element index, the XPath expression or the element index is determined to be the positioning rule.

[0149] In one embodiment, when the computer executable instructions are executed by the processor 502, the following steps are further implemented:

[0150] The operation object description is input into a pre-trained text processing model to obtain feature information of the visualization element; in response to the feature information containing visual information and text information of the visualization element, based on the visual information, a visualization element to be verified that matches the visual information is detected in the first web page; based on the text information in the feature information, the visualization element to be verified is verified to obtain the visualization element; and the element information corresponding to the visualization element is determined.

[0151] In one embodiment, when the computer executable instructions are executed by the processor 502, the following steps are further implemented:

[0152] Based on the characteristic information of the visual element in the operation object description, the visual element is identified in the first Web page; and the element information corresponding to the visual element is determined.

[0153] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the electronic device to which the solution of the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0154] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0155] Receiving a user instruction; wherein the user instruction includes an operation object description;

[0156] Based on the operation object description, determine the DOM node to be processed and generate an operation instruction;

[0157] Execute the operation instruction on the DOM node to be processed.

[0158] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0159] Based on the operation object description, a visual element in a first web page and element information of the visual element are determined; in a first DOM tree corresponding to the first web page, a DOM node to be processed corresponding to the visual element is identified based on the element information; and based on the operation object description, the operation instruction corresponding to the DOM node to be processed is generated.

[0160] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0161] Extracting the operation type and the operation parameter value corresponding to the operation type from the operation object description; and encapsulating the operation type and the operation parameter value into the operation instruction of a preset structure according to a preset sequence.

[0162] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0163] Determine a positioning rule in the element information; access the first DOM tree corresponding to the first Web page; and identify the DOM node to be processed in the first DOM tree based on the positioning rule in the element information of the visual element.

[0164] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0165] In response to the element information containing an XPath expression or an element index, the XPath expression or the element index is determined to be the positioning rule.

[0166] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0167] The operation object description is input into a pre-trained text processing model to obtain feature information of the visualization element; in response to the feature information containing visual information and text information of the visualization element, based on the visual information, a visualization element to be verified that matches the visual information is detected in the first web page; based on the text information in the feature information, the visualization element to be verified is verified to obtain the visualization element; and the element information corresponding to the visualization element is determined.

[0168] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0169] Based on the characteristic information of the visual element in the operation object description, the visual element is identified in the first Web page; and the element information corresponding to the visual element is determined.

[0170] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0171] Based on the same inventive concept, an embodiment of the present application further provides a computer program product, including a computer program, which implements any of the above-mentioned Web page operation methods when executed by a processor.

[0172] The program code for executing the computer program product of the present application may be written in any combination of one or more programming languages, and the program code may be executed entirely on the user device, partially on the user device, as an independent software package, partially on the user device and partially on a remote device, or entirely on the remote device.

[0173] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0174] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0175] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0176] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of user-operated steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0177] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A method for operating a web page, characterized in that: include: Receiving a user instruction; wherein the user instruction includes an operation object description; Based on the operation object description, determine the DOM node to be processed and generate an operation instruction; Execute the operation instruction on the DOM node to be processed.

2. The method according to claim 1, wherein The step of determining a DOM node to be processed based on the operation object description and generating an operation instruction includes: Determining, based on the operation object description, a visual element in the first web page and element information of the visual element; In a first DOM tree corresponding to the first Web page, based on the element information, identifying a DOM node to be processed corresponding to the visual element; Based on the operation object description, the operation instruction corresponding to the DOM node to be processed is generated.

3. The method according to claim 2, wherein The generating, based on the operation object description, the operation instruction corresponding to the DOM node to be processed includes: Extracting the operation type from the operation object description and the operation parameter value corresponding to the operation type; According to a preset order, the operation type and the operation parameter value are encapsulated into the operation instruction of a preset structure.

4. The method according to claim 2, wherein Identifying, in a first DOM tree corresponding to the first web page, a DOM node to be processed corresponding to the visual element based on the element information, including: Determining a positioning rule in the element information; Accessing the first DOM tree corresponding to the first Web page; In the first DOM tree, the DOM node to be processed is identified based on the positioning rule in the element information of the visual element.

5. The method according to claim 4, wherein The determining of the positioning rule in the element information includes: In response to the element information containing an XPath expression or an element index, the XPath expression or the element index is determined to be the positioning rule.

6. The method according to any one of claims 2 to 5, wherein: The determining, based on the operation object description, a visual element in the first web page and element information of the visual element includes: Inputting the operation object description into a pre-trained text processing model to obtain feature information of the visualization element; In response to the feature information including visual information and text information of the visual element, detecting, based on the visual information, a visual element to be verified that matches the visual information in the first web page; Based on the text information in the feature information, the visual element to be verified is verified to obtain the visual element; and then the element information corresponding to the visual element is determined.

7. A web page operation device, characterized in that: include: A user module is configured to receive a user instruction; wherein the user instruction includes an operation object description; An instruction module is used to determine the DOM node to be processed based on the operation object description and generate an operation instruction; The node module is used to execute the operation instruction on the DOM node to be processed.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and configured to run on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.