Interaction method, apparatus, device, storage medium, and program product

By extracting information from interactive elements in the GUI page and combining it with DOM and CSS properties, high-confidence operation instructions are generated, solving the problem of low accuracy in GUI interaction in existing technologies and achieving more efficient interactive operations.

CN122431576APending Publication Date: 2026-07-21BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2026-04-29
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify interactive elements in automated testing and GUI interactions, resulting in low interaction accuracy. Furthermore, unstable visual perception and complex page designs lead to model prediction failures.

Method used

By acquiring interaction requirements and page screenshots, information about interactive elements is extracted. Combined with DOM node attributes and CSS styles, non-interactive elements are excluded, and high-confidence operation instructions are generated.

Benefits of technology

It improves the accuracy and efficiency of GUI interaction, reduces the error space, and ensures that operation instructions are obtained from a limited number of high-confidence elements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122431576A_ABST
    Figure CN122431576A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data processing, and discloses an interaction method, device, equipment, storage medium and program product, the method comprising the following steps: acquiring an interaction demand, and acquiring a first picture, the interaction demand and the first picture corresponding to a first interaction page; extracting information of a first element, the first element being used for representing an interactable element in the first interaction page, the first element being configured to be obtained according to the attributes of the elements in the first interaction page; generating a first operation instruction, the first operation instruction comprising the identification of a second element and first operation information, the first operation instruction being configured to be obtained according to the interaction demand, the first picture and the information of the first element; and executing the first operation instruction. The method improves the interaction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] It relates to the field of data processing technology, specifically to interaction methods, devices, equipment, storage media, and program products. Background Technology

[0002] In scenarios such as automated testing and process automation, devices typically need to automatically interact with a graphical user interface (GUI). Accurate identification of the elements requiring interaction is a prerequisite for accurate interaction. Therefore, an interaction method is needed to achieve accurate interaction with the GUI. Summary of the Invention

[0003] An interactive method, apparatus, device, storage medium, and program product for solving the problem of automatic interaction with a GUI.

[0004] Firstly, an interaction method includes: Obtain the interaction requirements and the first image, wherein the interaction requirements and the first image correspond to the first interactive page; Extract information from the first element, which is used to characterize the interactive element in the first interactive page, and the first element is configured to be obtained based on the attributes of the element in the first interactive page. Generate a first operation instruction, which includes the identifier of the second element and first operation information. The first operation instruction is configured to be obtained based on the interaction requirements, the first image, and the information of the first element. Execute the first operation instruction.

[0005] Secondly, an interactive device includes: The acquisition module is used to acquire interaction requirements and acquire a first image, wherein the interaction requirements and the first image correspond to a first interactive page; An extraction module is used to extract information of a first element, which represents an interactive element in the first interactive page. The first element is configured to be obtained based on the attributes of the elements in the first interactive page. A generation module is used to generate a first operation instruction, the first operation instruction including the identifier of a second element and first operation information, the first operation instruction being configured to be obtained based on the interaction requirements, the first image and the information of the first element; The execution module is used to execute the first operation instruction.

[0006] Thirdly, an electronic device includes: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the interaction method of the first aspect or any corresponding embodiment described above.

[0007] Fourthly, a computer-readable storage medium storing computer instructions for causing a computer to perform the interactive method of the first aspect or any corresponding embodiment thereof.

[0008] Fifthly, a computer program product includes computer instructions for causing a computer to perform the interactive method described in the first aspect or any corresponding embodiment thereof.

[0009] In some cases, the provided interaction method involves: obtaining the interaction requirement and a first image, which correspond to a first interactive page; extracting information about a first element, which represents an interactive element on the first interactive page and is configured to be obtained based on the attributes of elements on the first interactive page; generating a first operation instruction, which includes the identifier of a second element and first operation information, and is configured to be obtained based on the interaction requirement, the first image, and the information of the first element; and executing the first operation instruction. This method, by extracting information about interactive elements on the first interactive page, excludes non-interactive elements, compresses the error space, and ensures that the obtained second element is obtained from a limited set of high-confidence first elements, thus improving interaction accuracy. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in specific implementation methods or related technologies under certain circumstances, the accompanying drawings used in the description of specific implementation methods or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some implementation methods. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0011] Figure 1 These are schematic diagrams illustrating application scenarios in various situations; Figure 2 This is a flowchart illustrating the first type of interaction method in some scenarios; Figure 3 This is a flowchart illustrating the second type of interaction method in some situations; Figure 4 These are schematic diagrams of the second image in some scenarios; Figure 5 This is a flowchart illustrating the third type of interaction method in some situations; Figure 6 This is a flowchart illustrating the fourth type of interaction method in some situations; Figure 7 These are structural block diagrams of interactive devices in some scenarios; Figure 8 These are schematic diagrams of the hardware structure of electronic devices in some scenarios. Detailed Implementation

[0012] To make the objectives, technical solutions, and advantages clearer in some cases, the technical solutions in some cases will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments, not all embodiments. Based on the embodiments in some cases, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this solution.

[0013] It is understood that before using the technical solutions disclosed in the various embodiments in certain situations, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in certain situations and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.

[0014] For example, upon receiving a user's proactive request, a prompt message can be sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media that perform the operation based on the prompt message.

[0015] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0016] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the specific implementation method. Other methods that comply with relevant laws and regulations may also be applied to this implementation method.

[0017] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0018] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this document, "multiple" means two or more, unless otherwise explicitly specified.

[0019] In related technologies, a multimodal large language model is used to understand GUI screenshots and predict the positions of interactive elements, thereby generating and executing GUI operations. This method treats the interaction task as a screen-to-action transformation problem. The multimodal large language model locates interface elements and predicts interactive behaviors through pixel information in the GUI screenshot, and can achieve zero-shot generalization ability, adapting to any GUI.

[0020] However, the reliability of the above approach has certain limitations. Due to the inherent instability of visual perception, even minor changes in interface layout, different page themes, resolutions, or scaling ratios can cause model prediction failures, leading to lower interaction accuracy. Furthermore, much interaction state information, such as whether interactive elements are disabled or links are clickable, is determined by Document Object Model (DOM) properties or Cascading Style Sheets (CSS) styles, and this information is not visually visible. Relying solely on GUI screenshots may result in model outputs including non-interactive elements.

[0021] In other related technologies, page elements such as Hyper Text Markup Language (HTML) are injected into the model. However, because the raw HTML code contains a large amount of data, including a lot of noise information irrelevant to the current interaction task—such as invisible elements, style code, and recommendation scripts—this can lead to excessive occupation of the model's context window, resulting in an exponential increase in token consumption.

[0022] Furthermore, due to the injection of large amounts of data, the model often struggles to accurately focus on the key information that is truly relevant to the task. A large amount of redundant content distracts the model, preventing it from accurately extracting useful semantic and structural information from the data, which in turn affects the accuracy and efficiency of the model's output.

[0023] The complexity of page design further complicates model processing. For example, temporary overlays such as modal dialog boxes and dropdown menus can obscure underlying page elements, making it difficult for the model to determine hierarchical relationships. This can lead to the interception or incorrect location of predicted interactions. For instance, the event handling logic predicted by the model might be delegated to a visually dissimilar parent container. Therefore, models in related technologies struggle to identify the true event responder based solely on appearance.

[0024] If relying solely on visual information is insufficient to achieve highly reliable GUI operation, then in some cases, application-layer structural and semantic information is introduced. It should be understood that the application-layer structural and semantic information can be understood as the information of the first interactive element described below.

[0025] Based on this, in some cases, an interaction method is provided, which involves: obtaining an interaction requirement and a first image, wherein the interaction requirement and the first image correspond to a first interactive page; extracting information of a first element, wherein the first element is used to represent an interactive element in the first interactive page, and the first element is configured to be obtained based on the attributes of elements in the first interactive page; generating a first operation instruction, wherein the first operation instruction includes the identifier of a second element and first operation information, and the first operation instruction is configured to be obtained based on the interaction requirement, the first image, and the information of the first element; and executing the first operation instruction.

[0026] This method extracts information from interactive elements in the first interactive page, excludes non-interactive elements in the first interactive page, compresses the error space, and ensures that the obtained second element is obtained from the limited high-confidence first element, thereby improving the accuracy of interaction.

[0027] As an optional application scenario, such as Figure 1 As shown, application 101 is installed in terminal device 110, and user 130 can interact with application 101 through terminal device 110 and / or access device of terminal device 110.

[0028] For example, application 101 can be any application that provides question-and-answer related services. Figure 1 In the application scenario shown, if application 101 is active, the terminal device 110 can display the interface 102 of application 101. The interface 102 may include various pages that application 101 can provide, such as interactive pages, settings pages, query pages, etc.

[0029] In some embodiments, terminal device 110 is communicatively connected to server 120 to provide services to application 101. Terminal device 110 may be a mobile terminal, fixed terminal, or portable terminal, etc., including but not limited to mobile phones, desktop computers, laptop computers, multimedia tablets, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal device 110 may also support any type of interface, and server 120 may be various types of computing systems or servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, etc.

[0030] It should be noted that, Figure 1 This is merely an example of an application scenario and does not limit the scope of protection.

[0031] The embodiments will now be described with reference to the accompanying drawings. It should be understood that the pages shown in the drawings are merely examples, and various page designs are possible in practice. The various graphic elements on the page may have different arrangements and different visual representations; one or more elements may be omitted or replaced, and one or more other elements may also be present, without any limitation herein. Furthermore, the embodiments will be described below primarily with respect to the server side. It should be understood that the actions described relative to the terminal device 110 can be performed by the application 101 on the terminal device 110, or can be performed by the application 101 in conjunction with its server (e.g., server 120).

[0032] For example, an interactive application is installed on terminal device 110. The user inputs interactive commands via an interactive method, instructing terminal device 110 to automatically perform corresponding interactive operations on the interactive page. Correspondingly, the server obtains the user's interactive commands and a screenshot of the interactive page, processes them in conjunction with interactive methods provided in certain situations, obtains a first operation command, and sends it back to terminal device 110. Based on this, terminal device 110 executes the first operation command, automatically performing the operation on the interactive page.

[0033] In some cases, an embodiment of an interactive method is provided. It should be noted that the steps shown in the flowcharts in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0034] In some situations, this is a method of interaction that can be used on servers or cloud devices, etc. Figure 2 This is a flowchart of interaction methods in some situations, such as Figure 2 As shown, the process includes the following steps: Step S201: Obtain the interaction requirements and obtain the first image.

[0035] Among them, the interaction requirements and the first image correspond to the first interactive page.

[0036] Interaction requirements are used to characterize the interactive actions that need to be performed on the first interactive page. Interaction requirements can be input by the user through interactive methods, such as interacting with the terminal device through text or voice. Accordingly, the content obtained is called interaction requirements.

[0037] Interaction requirements can also be user-given instructions that require multiple interaction operations to complete. The next requirement after each interaction operation is the interaction requirement described above. For example, if a user-given instruction involves interaction operation 1, interaction operation 2, and interaction operation 3, then during automatic processing, the operation requirement for interaction operation 2 after interaction operation 1 is completed is the interaction requirement described above.

[0038] It should be understood that interaction requirements do not specifically refer to user-given requirements, but can also be requirements generated during the processing process, and no restrictions are placed on them here.

[0039] The first interactive page represents the page that requires interactive operation. The current state of the first interactive page can be active or inactive. In an active state, the first interactive page can be interacted with according to the subsequent first operation command; in an inactive state, the first interactive page needs to be woken up first before subsequent interaction can be performed.

[0040] The first image is a screenshot representing the first interactive page. The size of the first image can be the same as or smaller than the size of the first interactive page. For example, the first image can be a complete screenshot of the first interactive page, or it can be a screenshot of the page portion excluding the toolbar, etc.

[0041] It's important to understand that both the interaction requirements and the first image correspond to the first interactive page. That is, the interaction requirements are given in relation to the first interactive page, and the first image is a screenshot of that page.

[0042] Step S202: Extract the information of the first element.

[0043] The first element is used to represent the interactive element in the first interactive page, and the first element is configured to be obtained based on the attributes of the element in the first interactive page.

[0044] The first interactive page contains interactive elements, such as buttons, text boxes, and dropdown lists. It may also contain non-interactive elements, such as text descriptions and images. The interactivity of elements on the first interactive page can be obtained from element attributes. For example, for a browser page, this can be obtained from the DOM node attributes in the DOM tree. DOM nodes correspond to elements on the page. The DOM node attributes record the node's name, position, interactivity, parent node, and child nodes, etc. The specific content recorded is set according to actual needs and is not limited here.

[0045] Furthermore, by combining the element's attributes, we can also determine the occlusion status of elements on the first interactive page. For example, after clicking a dropdown list, the dropdown box will obscure other elements on the page. If those other elements are interactive, they will be uninterrupted due to being obscured. Therefore, the attributes of the first element include not only static attributes (name, position, etc.) but also visual attributes, such as occlusion status and whether it is hidden.

[0046] By analyzing the attributes of the elements on the first interactive page, we can determine which elements are actually interactive and which are not. Here, the elements that are actually interactive are referred to as the first elements.

[0047] It should be understood that the first element is used to represent interactive elements in the first interactive page. The interactivity here includes both the interactivity represented by the static attributes and the interactivity represented by the visual attributes. That is to say, the first element is configured to be both statically interactive and visually interactive.

[0048] The information of the first element includes, but is not limited to, the location information of the first element and the identifier of the first element. The location information of the first element is used to describe the location of the first element in the first interactive page, and the identifier of the first element can be the name of the first element or other information of the first element, as long as the other information can uniquely identify the first element.

[0049] Step S203: Generate the first operation instruction.

[0050] The first operation instruction includes the identifier of the second element and the first operation information. The first operation instruction is configured to be obtained based on the interaction requirements, the first image, and the information of the first element.

[0051] After the above steps, the interactive elements (static and visual interactive elements) in the first interactive page are obtained, namely, the first element.

[0052] Interaction requirements represent the purpose to be achieved through interaction, and these requirements pertain to the first interactive page. By analyzing the interaction requirements, the interaction intent can be determined. For example, if the interaction requirements describe "Operation 1" for "Control A" on the first interactive page, then by combining the description of "Control A" with the information of the first element, the second element corresponding to "Control A" can be obtained. Furthermore, by combining this with the first image, the identifier of the second element and the first operation information can be determined.

[0053] The second element represents a specific element within the first element, and the first operation information includes the interactive operation required for the second element. The identifier of the second element and the first operation information are combined to generate a first operation instruction for the first interactive page.

[0054] The first operation instruction is obtained based on the interaction requirements, the first image, and the information of the first element. This can be achieved by analyzing the text of the interaction requirements to extract the identifier of the element to be operated, and then combining this with the information of the first element and the first image to obtain the second element that needs to be interacted with. The information of the second element is then obtained from the information of the first element, and based on this, the first operation instruction is generated.

[0055] In some alternative implementations, a model can also be used for processing. For example, the interaction requirements, the first image, and the information of the first element can be used as input to the model, and the model can output the identifier of the second element and the first operation information. Based on this, a first operation command that the device can understand is generated by combining the identifier of the second element and the first operation information. Of course, the model can also directly output the first operation command without offline generation.

[0056] It should be understood that the above are only some optional ways to generate the first operation instruction, and other methods can also be used. No restrictions are placed on them here.

[0057] For example, the element's identifier is used to represent the element's name. That is, the element's name is directly used as the element's identifier. Since the element's name is readily available, using the element's name to represent the identifier can improve interaction efficiency.

[0058] Step S204: Execute the first operation instruction.

[0059] The first operation instruction includes first operation information containing a second element, and the first operation is processed by calling the corresponding control interface. For example, the first interactive page is a browser page; the second element in the first operation information is mapped to the browser page, and the browser control interface is called to complete the final interactive operations such as clicking, inputting, and scrolling.

[0060] It should be understood that after the first operation instruction is executed, the page state of the first interactive page changes, and the process of steps S201 to S204 can continue to be executed, entering the next cycle of perception, decision-making and execution, thus forming a complete interactive closed loop.

[0061] In some interaction methods, by extracting information from interactive elements in the first interactive page, non-interactive elements in the first interactive page are excluded, the error space is compressed, and the obtained second element is obtained from the limited high-confidence first element, thus improving the accuracy of the interaction.

[0062] In some situations, this is a method of interaction that can be used on servers or cloud devices, etc. Figure 3 This is a flowchart of interaction methods in some situations, such as Figure 3 As shown, the process includes the following steps: Step S301: Obtain the interaction requirements and the first image. The interaction requirements and the first image correspond to the first interactive page. See details below. Figure 2 Step S201 of the illustrated embodiment will not be described again here.

[0063] Step S302: Extract the information of the first element.

[0064] The first element represents an interactive element in the first interactive page, and is configured to be obtained based on the attributes of elements in the first interactive page. See details... Figure 2 Step S202 of the illustrated embodiment will not be described again here.

[0065] Step S303: Generate the first operation instruction.

[0066] The first operation instruction includes the identifier of the second element and the first operation information. The first operation instruction is configured to be obtained based on the interaction requirements, the first image, and the information of the first element.

[0067] For example, step S303 above includes: Step S3031: Generate the first prompt message.

[0068] The first prompt message is configured to be obtained based on the interaction requirements, the first image, and the information of the first element.

[0069] The initial prompt information can be injected into the first model based on context engineering. For example, a template with prompt information is set up, which includes fixed information and variable information. The fixed information includes, but is not limited to, indicating the role of the model, the task that the model needs to complete, etc.; the variable information can be represented in the template in the form of placeholders, including placeholder 1 corresponding to the interaction requirement, placeholder 2 corresponding to the first image, and placeholder 3 corresponding to the information of the first element, etc.

[0070] After obtaining the interaction information, the first image, and the information of the first element, replace the corresponding placeholders 1 to 3 to generate the first prompt information.

[0071] In some alternative implementations, step S3031 includes: Step a1: Label the first image to obtain the second image. The second image is configured to be obtained by labeling the first element based on the information of the first element.

[0072] Step a2: Using the interaction requirements, the second image, and the information from the first element, obtain the first prompt information.

[0073] For example, Figure 4 This illustrates the impact of operating dropdown button 402 on the page display content. Figure 4 The first image shown includes a toolbar 401, a drop-down button 402, and an area 403 corresponding to button 3 and text box 4. Before interaction with drop-down button 402, area 403 is not obscured on the page, and its corresponding element is an interactive element.

[0074] See also Figure 4 After interacting with the drop-down button 402, the drop-down box 404 is expanded. Due to the position of the drop-down box 404, part of area 403 becomes invisible. Consequently, the interactive elements in area 403 become actual non-interactive elements.

[0075] By combining the number of elements in the first interactive page, we can obtain a list of the actual interactive elements in the first interactive page.

[0076] The first image is a screenshot representing the first interactive page and does not include any other annotation information. To improve the accuracy of subsequent processing of the first model, the first element is annotated in the first image.

[0077] After extracting the information of the first element, the first operation instruction needs to be generated. Combining the first image, the interaction requirements, and the information of the first element, since the first image is used to represent a screenshot of the first interactive page and does not have any annotation information, in order to improve the accuracy of the subsequent processing of the first model, the first image is annotated to obtain the second image.

[0078] The annotation for the first image is used to identify the first element within it. In other words, by combining the information of the extracted first element, the first element is located within the first image and then annotated. Since the first element is annotated in the first image, to distinguish it from the original first image, the annotated first image is referred to as the second image.

[0079] The annotation methods for the first image include, but are not limited to, using a selection box to annotate the first element in the first image, for example, such as... Figure 4 As shown, button 1 is labeled using rectangle 405.

[0080] After obtaining the second image, the first prompt message is obtained by using the interaction requirements, the second image, and the information of the first element.

[0081] The information of the first element is marked in the first image to obtain the second image. The second image is used to provide additional visual auxiliary information for the first model, which further improves the reliability of the output results of the first model and thus improves the accuracy of the interaction.

[0082] By annotating the first image to obtain the second image, the first model is guided to simultaneously perceive the visual image of the second image and read a precise, reliable, and structured actionable object map. Based on this, the decision-making process of the first model is constrained from an open visual coordinate regression problem to a text matching and selection problem within a finite, high-confidence set.

[0083] Step S3032: Obtain the first operation instruction using the first model and the first prompt information.

[0084] The input of the first model includes a first prompt message, and the output may include a first operation instruction; or it may be the identifier of the second element and the first operation message, and then the identifier of the second element and the first operation message.

[0085] In some alternative implementations, the first model is configured as follows: Understand the interaction requirements and obtain the identifier of the second element; query the information of the first element to obtain the information of the second element. The information of the second element is configured to be obtained by querying the information of the first element based on the identifier of the second element. The information of the first element includes the identifier and position information of the first element; output the first operation instruction. The first operation information includes the position information of the second element and the first operation.

[0086] After receiving the initial prompt, the first model extracts and understands the interaction requirements, obtaining the identifier of the second element involved in the interaction requirements. Then, it uses the identifier of the second element to query the information of the first element to obtain the information of the second element.

[0087] The information of the first element includes its identifier and location information. For example, the information of the first element can be represented in the form of entries, with each entry corresponding to one first element, and each entry including the name of the first element and its location information. This location information can be represented by the coordinates of the first element within the first interactive page.

[0088] Using the identifier of the second element obtained from the interaction request, the information of the first element is queried to obtain the location information of the second element. That is, the query is performed by identifier matching. If there is an identifier in the information of the first element that matches the identifier of the second element, it means that the location information of the second element can be obtained; if there is no identifier in the information of the first element that matches the identifier of the second element, a prompt message can be issued to remind the user that the interaction request needs to be adjusted.

[0089] By transforming the task of the first model from coordinate regression to label matching and selection, and ensuring the accuracy of coordinate calculation by the application layer, the success rate of interaction is effectively improved.

[0090] Step S304: Execute the first operation instruction. See details below. Figure 2 Step S204 of the illustrated embodiment will not be described again here.

[0091] In some interaction methods, the interaction requirements, the first image, and the information of the first element are injected into the first model as the model context, instructing the first model to obtain the first operation instruction according to the prompt information. For different interaction requirements and interaction pages, only the first prompt information needs to be adjusted, without modifying the first model, thus improving interaction efficiency.

[0092] In some situations, this is a method of interaction that can be used on servers or cloud devices, etc. Figure 5 This is a flowchart of interaction methods in some situations, such as Figure 5 As shown, the process includes the following steps: Step S501: Obtain the interaction requirements and obtain the first image.

[0093] The interaction requirements and the first image correspond to the first interactive page. See details below. Figure 2 The detailed description of step S201 in the illustrated embodiment will not be repeated here.

[0094] Step S502: Extract the information of the first element.

[0095] The first element is used to represent the interactive element in the first interactive page, and the first element is configured to be obtained based on the attributes of the element in the first interactive page.

[0096] For example, step S502 above includes: Step S5021: Obtain the attributes of the elements in the first interactive page.

[0097] The attributes of elements in the first interactive page can include the attributes of DOM nodes, as well as other information, without any restrictions on them.

[0098] Step S5022: Extract optional interactive elements.

[0099] The optional interactive elements are configured to be obtained based on the attributes of the elements in the first interactive page.

[0100] Based on the attributes of the elements in the first interactive page, selectable interactive elements are filtered out. That is, what is obtained here are static interactive elements.

[0101] Step S5023: Filter the selectable interactive elements to obtain the information of the first element.

[0102] The selection of optional interactive elements is configured to be based on the visual characteristics of the optional interactive elements.

[0103] The visual characteristics of optional interactive elements can be obtained by combining DOM node properties with CSS styles. Visual characteristics are used to represent features reflected in the visual effects of the page, such as whether they are hidden, obscured, or disabled.

[0104] Based on the visual characteristics of the selectable interactive elements, the first element is obtained. Correspondingly, by combining the attributes of the DOM node, the identifier and position information of the first element can be obtained, that is, the information of the first element can be obtained.

[0105] For example, the information of the first element includes the identifier of the first element and its location information, and the method for obtaining the location information of the first element includes: Step b1: Obtain the first viewport information, which corresponds to the first interactive page.

[0106] Step b2: Obtain the first position information. The attributes of the first element include the first position information.

[0107] Step b3: Normalize the first position information to obtain the position information of the first element. The normalization is configured to be based on the first viewport information.

[0108] First viewport information is used to characterize the page size of the first interactive page, for example, such as Figure 4 As shown, the first viewport information is used to represent the size of the area in the first image excluding the toolbar. The first position information of the first element is its position in the coordinate system corresponding to the first interactive page. Since the page size corresponding to the first viewport information is smaller than the page size corresponding to the first interactive page, the first position information is normalized based on the first viewport information, transforming it to the position information in the coordinate system that the first model can process, thus obtaining the position information of the first element. The coordinate system that the first model can process is the coordinate system corresponding to the first viewport information.

[0109] Since the position of the element in the first interactive page and the position information understood by the first model are obtained in different coordinate systems, the first position information is normalized by combining the first viewport information of the first interactive page, which can obtain position information suitable for the processing of the first model, thereby improving the accuracy of the output results of the first model and providing accurate operation instructions for automatic interaction on the first interactive page.

[0110] Step S503: Generate the first operation instruction.

[0111] The first operation instruction includes the identifier of the second element and first operation information. The first operation instruction is configured to be obtained based on the interaction requirements, the first image, and the information of the first element. See details. Figure 2 Step S203 of the illustrated embodiment, or, Figure 3 The detailed description of step S303 in the illustrated embodiment will not be repeated here.

[0112] Step S504: Execute the first operation instruction. See details below. Figure 3 The detailed description of step S304 in the illustrated embodiment will not be repeated here.

[0113] In some cases, the interaction method combines element attributes and visual features, which is equivalent to integrating application-layer semantics into the interaction, thus improving the reliability of the interaction.

[0114] As a specific application example in some situations, such as Figure 6As shown, the sandbox browser displays a first interactive page. After obtaining the interaction requirements, a screenshot of the first interactive page is taken to obtain the first image, and the browser viewport information is also obtained. Based on the first image and the browser viewport information, interactive elements are extracted, and invisible, occluded, and disabled elements are filtered out to obtain the information of the first element. Combining this information with the first element's information, clickable elements in the first image are identified, and the first image is annotated to obtain the second image. Then, based on the second image and clickable elements, prompt information is constructed, a first operation command is generated, and the operation action is invoked. On this basis, the sandbox environment executes the action corresponding to the interaction requirements, such as clicking, scrolling, and inputting. After the action is executed, the page state of the first interactive page changes, and then the cycle of perception, decision-making, and execution begins again.

[0115] This method employs semantic information, CSS tags, front-end framework internal event handler detection, and visibility and non-disabled state verification to construct a high-confidence first element set. For the input of the first model, the information of the first element and the labeled second image are injected into the model, changing the model task from open coordinate regression constraint to limited candidate selection. This process combines the visual understanding of the first model with precise semantic analysis at the application layer, forming a closed loop from perception to execution, improving the accuracy and robustness of interactive operations.

[0116] This interaction method is used to provide enhanced GUI operation capabilities. It can be encapsulated in pluggable middleware and tools and can be integrated with different types of upper-layer agents or multimodal models without changing the business logic of the agent itself.

[0117] In some cases, an interactive device is also provided for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, hardware implementations, or a combination of software and hardware, are also possible and contemplated.

[0118] In some cases, an interactive device is provided, such as Figure 7 As shown, it includes: The acquisition module 701 is used to acquire interaction requirements and the first image, and to obtain the correspondence between the interaction requirements, the first image, and the first interactive page.

[0119] Extraction module 702 is used to extract information of the first element, which represents the interactive element in the first interactive page. The first element is configured to be obtained based on the attributes of the elements in the first interactive page.

[0120] The generation module 703 is used to generate a first operation instruction. The first operation instruction includes the identifier of the second element and the first operation information. The first operation instruction is configured to be obtained based on the interaction requirements, the first image, and the information of the first element.

[0121] Execution module 704 is used to execute the first operation instruction.

[0122] In some alternative implementations, the generation module 703 includes: The information generation unit is used to generate the first prompt information, which is configured to be obtained based on the interaction requirements, the first image, and the information of the first element.

[0123] The instruction acquisition unit is used to acquire a first operation instruction using the first model and the first prompt information.

[0124] In some optional implementations, the information generation unit includes: The annotation sub-unit is used to annotate the first image to obtain the second image, which is configured to be obtained by annotating the first element based on the information of the first element.

[0125] Obtain the sub-unit, which is used to obtain the first prompt information by utilizing the interaction requirements, the second image, and the information of the first element.

[0126] In some alternative implementations, the first model is configured as follows: Understand the interaction requirements to obtain the identifier of the second element.

[0127] Query the information of the first element to obtain the information of the second element. The information of the second element is configured to be obtained by querying the information of the first element based on the identifier of the second element. The information of the first element includes the identifier of the first element and its position information.

[0128] Output the first operation instruction. The first operation information includes the position information of the second element and the first operation.

[0129] In some alternative implementations, the extraction module 702 includes: The attribute retrieval unit is used to retrieve the attributes of elements in the first interactive page.

[0130] The element extraction unit is used to extract optional interactive elements, which are configured to be obtained based on the attributes of elements in the first interactive page.

[0131] The element filtering unit is used to filter optional interactive elements to obtain information about the first element. The filtering of optional interactive elements is configured to be based on the visual characteristics of the optional interactive elements.

[0132] In some optional implementations, the information of the first element includes the identifier and location information of the first element, and the module for obtaining the location information of the first element includes: The first acquisition unit is used to acquire first viewport information, which corresponds to the first interactive page.

[0133] The second acquisition unit is used to acquire the first position information, and the attributes of the first element include the first position information.

[0134] The normalization unit normalizes the first position information to obtain the position information of the first element. The normalization is configured to be obtained based on the first viewport information.

[0135] In some alternative implementations, a name is used to identify the element.

[0136] In some cases, the provided interactive device can execute the interactive methods provided in the above embodiments, possessing the corresponding functional modules and beneficial effects for executing the methods. Further functional descriptions of the various modules and units described above are the same as in the corresponding embodiments described above, and will not be repeated here.

[0137] Figure 8 This is a schematic diagram of the structure of an electronic device provided in certain situations.

[0138] The following is a detailed reference. Figure 8 This diagram illustrates a structural schematic suitable for implementing an electronic device in certain situations. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 801, which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) 802 or a program loaded from memory 808 into random access memory (RAM) 803. RAM 803 also stores various programs and data required for the operation of the electronic device. The processor 801, ROM 802, and RAM 803 are interconnected via bus 804. An input / output interface 805 is also connected to bus 804.

[0139] Typically, the following devices can be connected to the input / output interface 805: input devices 806 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 807 including, for example, a liquid crystal display, speaker, vibrator, etc.; memory devices 808 including, for example, magnetic tape, hard disk, etc.; and communication devices 809. Communication device 809 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.

[0140] Specifically, the processes described in the flowchart above can be implemented as computer software programs. For example, in some cases, a computer program product is also provided, comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication device 809, or installed from memory 808, or installed from ROM 802. When the computer program is executed by processor 801, it performs the functions defined in the interaction methods in some cases.

[0141] Figure 8 The electronic devices shown are merely examples and should not be construed as limiting their functionality or scope of use in any situation.

[0142] In some cases, a computer-readable storage medium is also provided, in which the above-described methods can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that a computer, processor, microprocessor controller, or programmable hardware includes storage components capable of storing or receiving software or computer code that, when accessed and executed by the computer, processor, or hardware, implements the interactive methods shown in the above embodiments.

[0143] Some of the above solutions can be applied as computer program products, such as computer program instructions. When executed by a computer, these instructions, through the operation of the computer, can invoke or provide the aforementioned methods and / or technical solutions. Those skilled in the art should understand that the forms in which computer program instructions exist in computer-readable media include, but are not limited to, source files, executable files, and installation package files. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instruction; the computer compiling the instruction and then executing the corresponding compiled program; the computer reading and executing the instruction; or the computer reading and installing the instruction and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0144] Although embodiments in some cases have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the above description, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. An interaction method, comprising: Obtain the interaction requirements and the first image, wherein the interaction requirements and the first image correspond to the first interactive page; Extract information from the first element, which is used to characterize the interactive element in the first interactive page, and the first element is configured to be obtained based on the attributes of the element in the first interactive page. Generate a first operation instruction, which includes the identifier of the second element and first operation information. The first operation instruction is configured to be obtained based on the interaction requirements, the first image, and the information of the first element. Execute the first operation instruction.

2. The method according to claim 1, wherein generating the first operation instruction comprises: Generate a first prompt message, which is configured to be obtained based on the interaction requirements, the first image, and the information of the first element; The first operation instruction is obtained using the first model and the first prompt information.

3. The method according to claim 2, wherein generating the first prompt information includes: The first image is labeled to obtain a second image, wherein the second image is configured to be obtained by labeling the first element based on the information of the first element. The first prompt information is obtained by using the interaction requirements, the second image, and the information of the first element.

4. The method according to claim 2, wherein the first model is configured as follows: Understand the interaction requirements and obtain the identifier of the second element; Query the information of the first element to obtain the information of the second element. The information of the second element is configured to be obtained by querying the information of the first element based on the identifier of the second element. The information of the first element includes the identifier of the first element and its location information. Output the first operation instruction, wherein the first operation information includes the position information of the second element and the first operation.

5. The method according to claim 1, wherein extracting the information of the first element includes: Get the attributes of the elements in the first interactive page; Extract optional interactive elements, which are configured to be obtained based on the attributes of elements in the first interactive page; The optional interactive elements are filtered to obtain information about the first element. The filtering of the optional interactive elements is configured to be based on the visual characteristics of the optional interactive elements.

6. The method according to claim 5, wherein the information of the first element includes the identifier and location information of the first element, and the method for obtaining the location information of the first element includes: Obtain first viewport information, which corresponds to the first interactive page; Obtain first location information, wherein the attributes of the first element include the first location information; The first position information is normalized to obtain the position information of the first element, wherein the normalization is configured to be based on the first viewport information.

7. The method according to any one of claims 1 to 6, wherein the identifier is used to characterize the name of the element.

8. An interactive device, comprising: The acquisition module is used to acquire interaction requirements and acquire a first image, wherein the interaction requirements and the first image correspond to a first interactive page; An extraction module is used to extract information of a first element, which represents an interactive element in the first interactive page. The first element is configured to be obtained based on the attributes of the elements in the first interactive page. A generation module is used to generate a first operation instruction, the first operation instruction including the identifier of a second element and first operation information, the first operation instruction being configured to be obtained based on the interaction requirements, the first image and the information of the first element; The execution module is used to execute the first operation instruction.

9. An electronic device, characterized in that, include: A memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, the processor executing the computer instructions to perform the interactive method of any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the interactive method of any one of claims 1 to 7.

11. A computer program product, characterized in that, Includes computer instructions for causing a computer to perform the interactive method of any one of claims 1 to 7.