A browser control method, device, electronic device and medium

By annotating screenshots of remote browser pages and generating target instructions, the problems of low work efficiency and poor user experience in remote browser technology are solved, and more efficient and accurate browser operations and user interaction are achieved.

CN119828937BActive Publication Date: 2025-07-18SHANGHAI DOUXIANG INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510307725.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-18
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

The existing remote browser technology has problems such as low work efficiency and poor user experience.

Method used

By annotating the screenshot of the browser page of the second electronic device on the first electronic device, the coordinates of the operable element are obtained, and the browser is controlled to operate in combination with the user input instructions, the target instructions based on the preset data structure are generated, and the corresponding tools are called for the browser operation.

Benefits of technology

It improves operation accuracy and response speed, reduces the use of computing resources and network bandwidth, and improves user experience and work efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119828937B_ABST
    Figure CN119828937B_ABST
Patent Text Reader

Abstract

The present invention relates to a browser control method, device, electronic device and medium, belonging to the field of computer technology. The method is applied to a first electronic device, and the first electronic device is connected to a second electronic device, and the second electronic device is used to install a browser. The method includes: when receiving an instruction related to the browser input by the user, controlling the browser in the second electronic device to perform a startup operation, and controlling the browser to start screen recording; displaying in real time on the first electronic device a page screenshot obtained by the browser during screen recording, wherein the page screenshot includes a first page screenshot after the browser performs the startup operation; marking operable elements in the first page screenshot to obtain a first picture; performing a first operation on the browser according to the first picture and the first step in the instruction. This application combines the marked browser page screenshot with the user instruction to guide the operation of the browser, which can simplify the task and improve work efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, and particularly relates to a browser control method, apparatus, electronic device, and computer-readable storage medium. Background Art

[0002] Remote browser technology is a technology that transfers the browser's running environment from a local device to a remote server or virtual environment. Current remote browser technology is mainly applied to achieve remote cross-device access, remote browser isolation (Remote Browser Isolation, RBI), etc. Among them, remote cross-device access means that users can access remote browser services through different clients to perform operations such as web browsing; remote browser isolation creates an isolation layer between the user device and Internet content to prevent malicious codes such as cross-site scripting (Cross-Site Scripting, XSS) from directly affecting the user's device and data. Remote browser technology can enhance security and prevent data loss, but it also has the following problems: low work efficiency, poor user experience, etc. Summary of the Invention

[0003] In view of this, the purpose of this application is to provide a browser control method, apparatus, electronic device, and computer-readable storage medium to improve the problems of low work efficiency and poor user experience existing in current remote browser technology.

[0004] The embodiments of this application are implemented as follows:

[0005] In a first aspect, an embodiment of this application provides a browser control method, which is applied to a first electronic device. The first electronic device is connected to a second electronic device, and the second electronic device is used to install a browser. The method includes: when receiving an instruction related to the browser input by a user, controlling the browser in the second electronic device to perform a startup operation, and controlling the browser to start screen recording; displaying a page screenshot obtained by the browser's screen recording in real time on the first electronic device, where the page screenshot includes a first page screenshot after the browser performs the startup operation; annotating operable elements in the first page screenshot to obtain a first picture; and performing a first operation on the browser according to the first picture and the first step in the instruction.

[0006] In the above embodiments, by annotating the screenshot of the browser page of the second electronic device, the first electronic device can accurately identify the operable elements in the browser page, thereby assisting the first electronic device to more precisely understand the user input instruction and convert it into a specific browser operation, reducing the execution failure caused by ambiguous or incorrect instructions, improving the accuracy of the operation, and thus enhancing the user experience; moreover, by combining the annotated picture (the first picture) with the user input instruction to operate the browser, complex tasks can be efficiently completed, and the occupation of computing resources and network bandwidth can also be reduced, thereby improving the response speed of the operation and increasing the work efficiency.

[0007] Combined with a possible implementation manner of the embodiment of the first aspect, the annotating the operable elements in the first page screenshot to obtain a first picture includes: obtaining the coordinates of the operable elements in the first page screenshot; and annotating the operable elements in the first page screenshot according to the coordinates of the operable elements to obtain a first picture.

[0008] In the above embodiments, by obtaining the coordinates of the operable elements in the first page screenshot and then using the coordinates of the operable elements to annotate the operable elements in the first page screenshot, the first electronic device can combine the coordinates of the operable elements with the user input instruction to guide the first electronic device to operate the browser, enhancing the instruction parsing ability of the first electronic device, providing clear operation guidance for the first electronic device, reducing the exception handling caused by ambiguous or incorrect instructions, improving the accuracy of the operation, and thus enhancing the user experience.

[0009] Combined with a possible implementation manner of the embodiment of the first aspect, before performing the first operation on the browser according to the first step in the first picture and the instruction, the method further includes: generating a prompt word according to the instruction; and decomposing the instruction into multiple steps according to the prompt word, where the multiple steps include the first step.

[0010] In the above embodiments, by generating a prompt word and decomposing the user input instruction according to the prompt word, complex tasks can be disassembled into multiple simple steps and then each step is executed step by step, thereby reducing the resource requirements for each step execution, reducing the complexity of resource calculation, and also dynamically adjusting the execution strategy of each step, and thus increasing the work efficiency; moreover, the process of decomposing the user input instruction can more accurately identify the true intention of the user, enhancing the instruction parsing ability, thereby improving the accuracy of the operation, enabling the user input instruction to be correctly executed, and thus enhancing the user experience.

[0011] In a possible implementation manner combining with the embodiments of the first aspect, the first operation on the browser according to the first picture and the first step in the instruction includes: when the first picture represents that the browser has been started, converting the first step into a target instruction based on a preset data structure; when the instruction type of the target instruction is a page element operation type, controlling the browser to traverse the operable elements in the browser page according to the coordinates of the operable elements in the target instruction, and controlling the browser to perform a first operation on the operable element corresponding to the target instruction, wherein the coordinates of the operable elements in the target instruction are obtained through the first picture.

[0012] In the above embodiment, converting the user input instruction into a target instruction based on a preset data structure can convert the user's vague, non-standard or colloquial expression into an instruction that the electronic device can understand. This can not only more accurately express the user's intention, reduce misunderstandings or ambiguities, but also increase the probability that the user input instruction is correctly processed, thereby improving the user experience. And when the instruction type of the target instruction is a page element operation type, the coordinates of the operable elements in the target instruction can accurately locate the page elements that need to be operated, which not only improves the accuracy of the operation, enables the user input instruction to be correctly executed, and thus improves the user experience, but also improves the response speed of the operation, thereby improving the work efficiency.

[0013] In a possible implementation manner combining with the embodiments of the first aspect, the method further includes: when the instruction type of the target instruction is a tool type, calling a tool corresponding to the target instruction to perform a first operation on the browser.

[0014] In the above embodiment, when the instruction type of the target instruction is a tool type, by calling a tool corresponding to the target instruction to perform a first operation on the browser, the operation steps can be reduced, time can be saved, and thus the work efficiency is improved. And by calling a tool corresponding to the target instruction to perform a first operation on the browser, the operation process can be made more fluent, and the target instruction can be easily completed, thereby improving the user experience.

[0015] In a possible implementation manner combining with the embodiments of the first aspect, after performing the first operation on the browser according to the first picture and the first step in the instruction, the method further includes: obtaining a second page screenshot after performing the first operation on the browser; annotating the operable elements in the second page screenshot to obtain a second picture; performing a second operation on the browser according to the second picture and the second step in the instruction until all steps in the instruction are completed.

[0016] In the above embodiments, by taking screenshots and making annotations for the page after each operation on the browser, it can be determined whether each operation is completed, and potential abnormal situations can be detected in a timely manner to guide the execution of subsequent operations, ensuring that the user input instructions are correctly executed, thereby improving the user experience. Moreover, by combining the annotated pictures after each operation with the user input instructions and operating the browser, complex tasks can be efficiently completed, thus improving the operation response speed and work efficiency.

[0017] Combined with a possible implementation manner of the first aspect of the embodiments, the method further includes: listening for browser events performed by the user on the page screenshots obtained by screen recording the browser, and controlling the browser to perform operations according to the browser events.

[0018] In the above embodiments, by listening for browser events performed by the user on the page screenshots obtained by screen recording the browser, the user's operations can be captured in real time, and the browser can be controlled to perform operations according to the browser events, which can reduce the time for the user to wait for the page to load or respond, thereby improving work efficiency. Moreover, by listening for the user's browser events and responding in real time, a smooth experience close to that of a local browser can be provided to the user, providing smooth interaction, thereby improving the user experience.

[0019] In a second aspect, an embodiment of the present application provides a browser control device, which is included in a first electronic device. The first electronic device is connected to a second electronic device, and the second electronic device is used to install a browser. The device includes: a function module, configured to control the browser in the second electronic device to perform a startup operation and control the browser to start screen recording when receiving an instruction related to the browser input by the user; a display module, configured to display in real time on the first electronic device the page screenshots obtained by the browser performing screen recording, where the page screenshots include the first page screenshot after the browser performs the startup operation; an annotation module, configured to annotate the operable elements in the first page screenshot to obtain a first picture; the function module is further configured to perform a first operation on the browser according to the first picture and the first step in the instruction.

[0020] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory and a processor, the processor is connected to the memory; the memory is used to store a program; the processor is used to call the program stored in the memory to execute the method provided in the first aspect of the embodiments and / or any possible implementation manner combined with the first aspect of the embodiments.

[0021] Fourthly, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it is configured to execute the method provided by any possible implementation manner of the above first aspect embodiment and / or in combination with the first aspect embodiment.

[0022] Other features and advantages of the present application will be described in the subsequent specification. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the written specification and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings. As shown in the drawings, the above and other objectives, features, and advantages of the present application will become clearer.

[0024] Figure 1 FIG. shows a schematic connection structure diagram of a first electronic device and a second electronic device provided by an embodiment of the present application.

[0025] Figure 2 FIG. shows a schematic flowchart of a browser control method provided by an embodiment of the present application.

[0026] Figure 3 FIG. shows a schematic diagram of labeling an operable element in a first page screenshot according to the coordinates of the operable element provided by an embodiment of the present application.

[0027] Figure 4 FIG. shows a schematic module diagram of a browser control device provided by an embodiment of the present application.

[0028] Figure 5 FIG. shows a schematic structure diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] The following will describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. The following embodiments can be used as examples to more clearly illustrate the technical solutions of the present application, but cannot be used to limit the protection scope of the present application. Those skilled in the art can understand that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0030] It should be noted that similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. At the same time, in the description of the present application, relational terms such as "first", "second", etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus.

[0031] Furthermore, the term "and / or" in the present application is merely a description of the relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent three cases: A exists alone, A and B exist simultaneously, and B exists alone.

[0032] In the description of the embodiments of the present application, unless otherwise clearly specified and limited, the technical term "connection" may be a direct connection or an indirect connection through an intermediate medium.

[0033] The embodiments of the present application provide a browser control method. By combining the screenshot of the browser page of the second electronic device after annotation with the user input instruction to guide the operation of the browser, the first electronic device can more accurately understand the user input instruction, simplify complex tasks, improve the accuracy and response speed of the operation, thereby improving work efficiency and enhancing the user experience.

[0034] The connection structure between the first electronic device and the second electronic device involved in the browser control method provided by the embodiments of the present application is as Figure 1As shown in the figure. Among them, the connection methods between the first electronic device and the second electronic device include, but are not limited to, network connection (communicating using the TCP / IP protocol), cloud service relay connection, Secure Shell Tunnel (SSH tunnel) connection, etc. The first electronic device controls the browser in the second electronic device to operate by executing the browser control method provided in the embodiments of the present application. In a possible implementation manner, the first electronic device may include an agent, and the agent executes the browser control method provided in the embodiments of the present application to control the browser in the second electronic device to operate. The second electronic device is used to install a browser for the first electronic device to use. The browsers in the second electronic device include, but are not limited to, Google Chrome, Mozilla Firefox, Brave, Tor Browser, etc. The above-mentioned first electronic device and second electronic device may be devices such as a computer, a mobile phone, and a server.

[0035] The browser control method provided in the embodiments of the present application can be applied to the above-mentioned first electronic device. The following will be combined with Figure 2 The flowchart shown will illustrate the browser control method from the perspective of the first electronic device. The browser control method may include the following steps:

[0036] Step S10: When receiving an instruction related to the browser input by the user, control the browser in the second electronic device to perform a startup operation, and control the browser to start screen recording.

[0037] In the above step, when the first electronic device receives an instruction related to the browser input by the user (in the embodiments of the present application, the user input instruction and the instruction both refer to the instruction related to the browser input by the user), it can control the browser in the second electronic device to perform a startup operation through the command line or the Application Programming Interface (API interface), and at the same time open the remote debugging protocol port of the browser in the second electronic device. The first electronic device uses the remote debugging protocol to control the browser to start screen recording through this remote debugging protocol port.

[0038] In a possible implementation manner, the first electronic device controls the browser in the second electronic device to perform a startup operation through the command line, which may be to send an instruction to the second electronic device through the command line to start the executable file of the browser in the second electronic device and make the browser run in headless mode. Among them, headless mode means that the browser runs without a Graphical User Interface (GUI).

[0039] Among them, when the first electronic device controls the browser in the second electronic device to perform a startup operation through an API interface, it can be through the API interfaces provided by tools such as Playwright and Puppeteer to control the browser in the second electronic device to perform a startup operation. Tools such as Playwright and Puppeteer are installed on the second electronic device. For example, the first electronic device sends an HTTP request to the second electronic device. After receiving the HTTP request, the second electronic device will trigger the Playwright script to run and start the browser installed in the second electronic device through the Playwright script. Among them, the content, format, structure, etc. of the HTTP request need to follow the specifications of the remote debugging protocol supported by the browser installed on the second electronic device.

[0040] The purpose of opening the remote debugging protocol port of the browser in the second electronic device above is to enable the first electronic device to control the browser in the second electronic device through this remote debugging protocol port using the remote debugging protocol. For example, when the browser in the second electronic device is a Chrome browser, by opening the remote debugging protocol (Chrome DevTools Protocol, CDP) port of the Chrome browser, the first electronic device can control the Chrome browser through the CDP protocol.

[0041] The specific process of the first electronic device controlling the browser to start screen recording through this remote debugging protocol port using the remote debugging protocol may include: The first electronic device sends a screen recording instruction to the browser through the remote debugging protocol. When the browser receives the screen recording instruction, it starts screen recording and returns the page screenshots obtained from the screen recording to the first electronic device in real time asynchronously through the remote debugging protocol.

[0042] In step S10, the instruction related to the browser input by the user can be an instruction based on natural language. An instruction based on natural language refers to a command expressed by the user in the form of daily language. For example, the user inputs "query the latest news through Baidu" on the first electronic device, and "query the latest news through Baidu" is an instruction based on natural language.

[0043] Step S20: Display the page screenshots obtained from the browser's screen recording in real time on the first electronic device, where the page screenshots include the first page screenshots after the browser performs the startup operation.

[0044] In step S20, when the first electronic device receives the page screenshot obtained by the browser during screen recording, it can use the Canvas Application Programming Interface (Canvas API) to render the page screenshot onto the page canvas of the first electronic device in real time, enabling the first electronic device to display the page screenshot obtained by the browser during screen recording in real time. This allows the user to intuitively understand the content of the browser on the second electronic device, providing the user with an experience similar to using a local browser and enhancing the user experience. Moreover, during the user's use of the browser on the second electronic device, the first electronic device will continuously display the page screenshot obtained by the browser during screen recording in real time until the user finishes using the browser on the second electronic device.

[0045] Among them, the Canvas API is a newly added tag in HTML5, and its core is a <canvas>The element provides a bitmap that can be manipulated with JavaScript. The Canvas API has functions such as drawing graphics, real-time rendering, image processing, and animation production.

[0046] Step S30: Label the operable elements in the first page screenshot to obtain the first picture.

[0047] In a possible implementation, the specific process for the first electronic device to label the operable elements in the first page screenshot to obtain the first picture may include: The first electronic device sends an instruction to the browser through the remote debugging protocol. The browser obtains the coordinates of the operable elements in the first page screenshot according to this instruction and returns the coordinates of the operable elements in the first page screenshot to the first electronic device. After the first electronic device obtains the coordinates of the operable elements in the first page screenshot, it labels the operable elements in the first page screenshot according to the coordinates of the operable elements to obtain the first picture. Among them, the process of labeling the operable elements in the first page screenshot according to the coordinates of the operable elements can be as Figure 3 shown. Among them, Figure 3 shows the process of labeling the operable element 1, operable element 2, operable element 3, and operable element 4 in the first page screenshot according to the coordinates of the operable elements. Among them, 1, 2, 3, and 4 are the labels of the operable elements, indicating the order of the operable elements.

[0048] In a possible implementation, when the first electronic device includes a large language model, the specific process for the first electronic device to label the operable elements in the first page screenshot to obtain the first picture may include: The first electronic device sends an instruction to the browser through the remote debugging protocol. The browser obtains the coordinates of the operable elements in the first page screenshot according to this instruction and returns the coordinates of the operable elements in the first page screenshot to the first electronic device. After the first electronic device obtains the coordinates of the operable elements in the first page screenshot, it numbers the coordinates of the operable elements to obtain the numbered coordinates. Then, it labels the operable elements in the first page screenshot according to the numbered coordinates to obtain the first picture.

[0049] Alternatively, the first electronic device sends an instruction to the browser through the remote debugging protocol. The browser obtains the coordinates of the operable elements in the first page screenshot according to this instruction and returns the coordinates of the operable elements in the first page screenshot to the first electronic device. After the first electronic device obtains the coordinates of the operable elements in the first page screenshot, it labels the operable elements in the first page screenshot according to the coordinates of the operable elements and numbers the labeled operable elements to obtain the first picture.

[0050] Wherein, when the first electronic device includes a large language model, numbering the coordinates of the actionable elements or numbering the annotated actionable elements enables the large language model to more quickly locate and identify the actionable elements in the first picture; it can also help the large language model extract text information from the first picture. For example, through numbering, the large language model can identify the text content of each element and associate the text content of each element with its number, thereby better understanding the page content. When converting the steps in the user input instruction into instructions based on a preset data structure, the numbering can help the large language model more clearly understand the page elements corresponding to each step, and thus more accurately execute the user input instruction. Moreover, in the embodiments of the present application, the large language model can be used to decompose the user input instruction and convert the decomposed instruction into an instruction based on a preset data structure. The large language model can be a multimodal model such as LLaMA, Qwen, VisCPM, etc.

[0051] In the above embodiments, the actionable elements include, but are not limited to, buttons, input boxes, selection boxes, links, etc. The coordinates of the actionable elements are square boxes, and the coordinates of the actionable elements include, but are not limited to, information such as the length, height, and highest point on the left of the box. In addition, the first electronic device can also annotate the actionable elements in the first page screenshot through an image vision library such as OpenCV.

[0052] By obtaining the coordinates of the actionable elements in the first page screenshot and then using the coordinates of the actionable elements to annotate the actionable elements in the first page screenshot, the first electronic device can combine the coordinates of the actionable elements and the user input instruction to guide the first electronic device to operate the browser, enhancing the instruction parsing ability of the first electronic device, providing clear operation guidance for the first electronic device, reducing exception handling caused by fuzzy or incorrect instructions, improving the accuracy of operations, and thus improving the user experience.

[0053] Step S40: Perform a first operation on the browser according to the first picture and the first step in the instruction.

[0054] In a possible implementation, according to the first image and the first step in the instruction, the specific process of performing the first operation on the browser may include: when the first image indicates that the browser has been started, converting the first step into a target instruction based on a preset data structure; and judging the instruction type of the target instruction, when the instruction type of the target instruction is a page element operation type, controlling the browser to traverse the operable elements in the browser page according to the coordinates of the operable elements in the target instruction, and controlling the browser to perform the first operation on the operable elements corresponding to the target instruction, wherein the coordinates of the operable elements in the target instruction are obtained through the first image. That is to say, when the first step is to operate on the page elements, when the first step is converted into a target instruction based on a preset data structure, the target instruction will contain the coordinates of the page element corresponding to the first step, and the coordinates are obtained through the first image.

[0055] In one possible implementation, the first electronic device may include an intelligent agent. Then, according to the first image and the first step in the instruction, the specific process of performing the first operation on the browser may include: when the first image indicates that the browser has been started, the intelligent agent converts the first step into a target instruction based on a preset data structure; and determines the instruction type of the target instruction. When the instruction type of the target instruction is a page element operation type, the intelligent agent controls the browser to traverse the operable elements in the browser page according to the coordinates of the operable elements in the target instruction, and controls the browser to perform the first operation on the operable elements corresponding to the target instruction, wherein the coordinates of the operable elements in the target instruction are obtained through the first image.

[0056] In one possible implementation, the first electronic device may include not only an intelligent agent but also a large language model; then, according to the first picture and the first step in the instruction, the specific process of performing a first operation on the browser may include: when the first picture indicates that the browser has been started, the large language model converts the first step into a target instruction based on a preset data structure, and sends the target instruction to the intelligent agent; when the intelligent agent receives the target instruction, it determines the instruction type of the target instruction, and when the instruction type of the target instruction is a page element operation type, it controls the browser to traverse the operable elements in the browser page according to the coordinates of the operable elements in the target instruction, and controls the browser to perform a first operation on the operable elements corresponding to the target instruction, wherein the coordinates of the operable elements in the target instruction are obtained through the first picture.

[0057] The large language model may be included in the intelligent agent, or may be a component of the first electronic device parallel to the intelligent agent.

[0058] In the above embodiment, the first step in the instruction refers to the first step of a plurality of steps having an execution order obtained after decomposing the instruction.

[0059] Among them, the specific process of converting the first step into a target instruction based on a preset data structure may include: converting the first step in the user input instruction based on natural language into a JSON object (target instruction based on the preset data structure) that conforms to the JSON_SCHEMA definition through a preset BrowserAction data model. In some embodiments, when the first electronic device includes an agent, the preset BrowserAction data model may be included in the agent; in some embodiments, when the first electronic device includes a large language model, the preset BrowserAction data model may be included in the large language model. The data dictionary of the above BrowserAction data model is shown in Table 1:

[0060] Table 1

[0061]

[0062] In the BrowserAction data model, the Element_Handler field represents the page element operation field. Among them, the operations performed on the page element include but are not limited to clicking, double-clicking, filling, hovering, pressing keys, etc.; the field description table of the Element_Handler field can be shown in Table 2:

[0063] Table 2

[0064]

[0065] The new_page field represents opening a new page, and it will return the page that needs to be opened by the browser in the second electronic device. The task_complete field represents the completion of the task, that is, the user input instruction is completed. Unable indicates that the operation cannot be performed, and generally this operation requires manual intervention, such as entering the user password. The tool field represents calling an automation tool, and the field description table of the tool field can be shown in Table 3:

[0066] Table 3

[0067]

[0068] Among them, the instruction type of the above target instruction being the page element operation type means that the instruction type of the target instruction is Element_Handler.

[0069] In the above embodiment, when the instruction type of the target instruction is the page element operation type, the page element to be operated can be accurately located through the coordinates of the operable element in the target instruction, which not only improves the accuracy of the operation, enables the user input instruction to be correctly executed, and thus improves the user experience; but also improves the response speed of the operation, and thus improves the work efficiency.

[0070] In a possible implementation, when the instruction type of the target instruction is a tool type, a tool corresponding to the target instruction is called to perform a first operation on the browser.

[0071] In the above embodiment, that the instruction type of the target instruction is a tool type means that the instruction type of the target instruction is "tool". When the instruction type of the target instruction is "tool", the first electronic device will call a tool corresponding to the target instruction to perform a first operation on the browser. For example, when the browser page in the second electronic device contains a verification code, the target instruction will select to call a verification code bypass tool to operate on the browser. The verification code bypass tool can identify the verification code on the browser page and return an instruction operation for automatic filling of the verification code.

[0072] In this implementation, by calling a tool corresponding to the target instruction to perform a first operation on the browser, the operation steps can be reduced, time can be saved, and thus the work efficiency is improved; and by calling a tool corresponding to the target instruction to perform a first operation on the browser, the operation process can be made more fluent, and the target instruction can be easily completed, thereby improving the user experience.

[0073] In a possible implementation, before performing a first operation on the browser according to the first picture and the first step in the instruction, the browser control method provided by the embodiments of the present application further includes: generating a prompt word according to the instruction; decomposing the instruction into multiple steps according to the prompt word, where the multiple steps include the first step.

[0074] In this implementation, the first electronic device may include an agent, and the agent can generate a prompt word according to the instruction and decompose the instruction into multiple steps according to the prompt word.

[0075] In this implementation, the first electronic device may not only include an agent, but also include a large language model. The agent can generate a prompt word according to the instruction and send the prompt word to the large language model, and the large language model decomposes the instruction into multiple steps according to the prompt word.

[0076] Among them, an agent refers to an entity that can perceive the environment and make decisions based on the perceived information to achieve a specific goal. An agent can be a software program, a robot, or a system. The functions of an agent include but are not limited to the following: perception function, decision-making function, action function, learning function, etc.

[0077] The prompt words in the above embodiments include but are not limited to the following: System prompt word: "You are a browser robot that supports using browser commnad to control the browser. You can choose your corresponding operations through screenshots and instructions provided by the user. Note that you only need to follow the steps.(You are a browser robot that supports using browser operation instructions to control the browser. You can choose your corresponding operations through screenshots and instructions provided by the user. Note that you only need to follow the steps)”、 Represents the user's natural language instructions, Represents the execution action descriptions of each stage, Represents a JSON object with a return type of "BrowserAction" according to the definition of JSON_SCHEMA, Represents the automated tool called.

[0078] The generated prompt words are illustrated by an example below. If the first step in the user input instruction requires calling a verification code bypass tool, the generated prompt word is: You can choose the tools to use according to your needs. The list of supported tools is as follows: (You can choose the appropriate tools according to your own needs. The following is the list of supported tools: )

[0079] In a possible implementation manner, after performing a first operation on the browser according to the first picture and the first step in the instruction, the browser control method provided by the embodiments of the present application further includes: obtaining a second page screenshot after performing the first operation on the browser; marking the operable elements in the second page screenshot to obtain a second picture; performing a second operation on the browser according to the second picture and the second step in the instruction until all steps in the instruction are completed. That is to say, after decomposing the instruction into multiple steps with an execution order, each time an operation is performed on the browser according to the steps, it is necessary to take a screenshot of the browser page after the operation and mark the screenshot until all steps in the instruction are completed. For example, if the instruction input by the user is "Baidu search for the latest news", it can be decomposed into three steps: the first step is to open the Baidu page, the second step is to enter "search for the latest news" in the search box on the Baidu page, and the third step is to click the search button on the Baidu page; after performing operations on the browser according to the first step, the second step, and the third step, it is necessary to take a screenshot of the page after the operation and mark the screenshot.

[0080] Among them, the second step in the above embodiment refers to the second step among the multiple steps with an execution order obtained after decomposing the instruction.

[0081] In this implementation manner, the specific process of performing a second operation on the browser according to the second picture and the second step in the instruction may include: when the second picture indicates that the first step has been completed, converting the second step into a first target instruction based on a preset data structure, and judging the instruction type of the first target instruction. When the instruction type of the first target instruction is a page element operation type, controlling the browser to traverse the operable elements in the browser page according to the coordinates of the operable elements in the first target instruction, and controlling the browser to perform a second operation on the operable element corresponding to the first target instruction, where the coordinates of the operable elements in the first target instruction are obtained through the second picture.

[0082] If the second picture indicates that the first step has not been completed or executed incorrectly, the first electronic device judges whether the first step cannot be executed according to the information in the second picture; if the information in the second picture indicates that the first step cannot be executed, the first electronic device sends a message that the instruction cannot be executed to the user, and saves the information related to the inability to execute the first step for subsequent analysis and improvement; if the information in the second picture indicates that the first step can be executed, the first electronic device controls the browser in the second electronic device to perform the first operation again.

[0083] In a possible implementation, before performing a second operation on the browser according to the second picture and the second step in the instruction, the browser control method provided by the embodiments of the present application further includes: adding a description related to the execution of the first step to the prompt word generated according to the instruction. The description related to the execution of the first step refers to the step content of the first step. For example, the description of the first step: opening the Baidu page is Currently executed to: step 1. {{ Open the Baidu page}} (Currently executed to: the first step {{ Open the Baidu page}}).

[0084] In a possible implementation, the browser control method provided by the embodiments of the present application further includes: listening for browser events performed by the user on the page screenshot obtained by screen recording the browser, and controlling the browser to perform operations according to the browser events.

[0085] In this implementation, the first electronic device can listen for browser events performed by the user on the page screenshot obtained by screen recording the browser through the Canvas canvas, and when a browser event is detected, send an instruction related to the browser event to the browser through the remote debugging protocol. When the browser receives the instruction related to the browser event, it performs the same operation on the browser page as the browser event.

[0086] Among them, browser events include but are not limited to mouse events, keyboard events, window and document events, form events, etc. Mouse events can include events such as mouse click, mouse double click, and mouse press. Keyboard events can include events such as keyboard press, keyboard release, and change of input box content. Window and document events can include events such as window size change, window scroll, and document loading completion. Form events can include events such as form submission, form reset, and text selection.

[0087] The browser control method provided by the embodiments of the present application will be described below through an example. Taking the first electronic device including an agent and a large language model as an example. If the user input instruction is "Baidu search for the latest news", when the agent receives the user input instruction, it controls the browser in the second electronic device to perform a startup operation and controls the browser to start screen recording; the page screenshot obtained by screen recording the browser is displayed in real time on the first electronic device, where the page screenshot includes the first page screenshot after the browser performs the startup operation; the operable elements in the first page screenshot are marked to obtain the first picture.

[0088] The agent generates the first prompt word according to the user input command "Baidu search for recent news", which is: You are a browser robot that supports using browser commands to control the browser. You can choose your corresponding operations through screenshots and instructions provided by the user. Note that you only need to follow the steps.

[0089] The user goals for the current task are: Search Baidu for recent news.

[0090] You need to return the results to me in JSON objects JSON objects oftype "BrowserAction" according to the following JSON Schema definitions:

[0091]

[0092] (You need to return the result to me as a JSON object of type "BrowserAction" according to the following JSON Schema definition: )

[0093] Return JSON object with 2 spaces of indentation.

[0094] The intelligent agent sends the above-mentioned first prompt word to the large language model. When receiving the above-mentioned first prompt word, the large language model decomposes the user input instruction "Baidu search recent news" into three steps according to the above-mentioned first prompt word: the first step is to open the Baidu page, the second step is to enter the search for recent news in the search box of the Baidu page, and the third step is to click the search button of the Baidu page.

[0095] In the case where the first picture represents that the browser has been started, the large language model converts the first step into a first instruction based on a preset data structure, the instruction type of the first instruction is new page, and the operation value is https: / / www.baidu.com , that is, the first instruction is to open a new page, the page URL is https: / / www.baidu.com The large language model sends the first instruction to the agent, and the agent then sends the first instruction to the browser. When the browser receives the first instruction, it opens the Baidu page. After the browser performs the first operation according to the first instruction, the agent obtains a screenshot of the second page after the first operation, and annotates the operable elements in the screenshot of the second page to obtain a second picture.

[0096] The agent adds a description of the execution of the first step to the generated first prompt word, and obtains the second prompt word, which is: You are a browser robot that supports using browser commands to control the browser. You can choose your corresponding operations through screenshots and instructions provided by the user. Note that you only need to follow the steps.

[0097] The user goals for the current task are: Search Baidu for recent news.

[0098] Currently executed to: step 1. {{ Open Baidu page}}.

[0099] You need to return the results to me in JSON objects JSON objects oftype "BrowserAction" according to the following JSON Schema definitions:

[0100]

[0101] (You need to return the result to me in the form of a JSON object of type "BrowserAction" according to the following JSON Schema definition: )

[0102] Return JSON object with 2 spaces of indentation.

[0103] The agent sends the second prompt word and the second picture to the large language model. The large language model can judge whether the first step is completed according to the second prompt word and the second picture. When the second picture represents that the first step is completed, the second step is converted into a second instruction based on a preset data structure, and the second instruction is sent to the agent. According to the second instruction, the agent performs a second operation on the browser and takes a screenshot of the browser page after the second operation to obtain a third page screenshot; the operable elements in the third page screenshot are marked to obtain a third picture. The agent then adds a description of the execution of the second step to the second prompt word to obtain a third prompt word; then the third prompt word and the third picture are sent to the large language model. The large language model judges whether the second step is completed according to the third prompt word and the third picture. When the third picture represents that the second step is completed, the third step is converted into a third instruction based on a preset data structure, and the third instruction is sent to the agent. According to the third instruction, the agent performs a third operation on the browser and takes a screenshot of the browser page after the third operation to obtain a fourth page screenshot; the operable elements in the fourth page screenshot are marked to obtain a fourth picture; the agent then adds a description of the execution of the third step to the third prompt word to obtain a fourth prompt word, and then the fourth prompt word and the fourth picture are sent to the large language model. The large language model can judge whether the third step is completed according to the fourth prompt word and the fourth picture. If the third step is completed, it represents that the user input instruction is completed.

[0104] As Figure 4 shown, Figure 4 shows a schematic structural diagram of a browser control device 10 provided by an embodiment of the present application. The browser control device 10 is included in the first electronic device, and the first electronic device is connected to the second electronic device, and the second electronic device is used to install a browser. The browser control device 10 includes a function module 11, a display module 12, and a marking module 13.

[0105] The function module 11 is used to control the browser in the second electronic device to perform a startup operation and control the browser to start screen recording when receiving a browser-related instruction input by a user.

[0106] A display module 12, configured to display in real time on a first electronic device a page screenshot obtained by screen recording of a browser, wherein the page screenshot includes a first page screenshot after the browser performs a startup operation.

[0107] A marking module 13, configured to mark the operable elements in the first page screenshot to obtain a first picture.

[0108] A function module 11 is further configured to perform a first operation on the browser according to the first picture and the first step in the instruction.

[0109] In a possible implementation manner, the marking module 13 is specifically configured to obtain the coordinates of the operable elements in the first page screenshot; mark the operable elements in the first page screenshot according to the coordinates of the operable elements to obtain a first picture.

[0110] In a possible implementation manner, the function module 11 is further configured to generate a prompt word according to the instruction; decompose the instruction into multiple steps according to the prompt word, wherein the multiple steps include a first step.

[0111] In a possible implementation manner, the function module 11 is specifically configured to, when the first picture indicates that the browser has been started, convert the first step into a target instruction based on a preset data structure; when the instruction type of the target instruction is a page element operation type, control the browser to traverse the operable elements in the browser page according to the coordinates of the operable elements in the target instruction, and control the browser to perform a first operation on the operable element corresponding to the target instruction, wherein the coordinates of the operable elements in the target instruction are obtained through the first picture.

[0112] In a possible implementation manner, the function module 11 is further configured to, when the instruction type of the target instruction is a tool type, call a tool corresponding to the target instruction to perform a first operation on the browser.

[0113] In a possible implementation manner, the marking module 13 is further configured to obtain a second page screenshot after the first operation is performed on the browser; mark the operable elements in the second page screenshot to obtain a second picture.

[0114] In a possible implementation manner, the function module 11 is further configured to perform a second operation on the browser according to the second picture and the second step in the instruction until all steps in the instruction are completed.

[0115] In a possible implementation manner, the function module 11 is further configured to monitor browser events performed by the user on the page screenshot obtained by screen recording of the browser, and control the browser to perform operations according to the browser events.

[0116] The browser control device 10 provided in the embodiments of the present application has the same implementation principle and technical effects as those in the foregoing method embodiments. For a brief description, for parts not mentioned in the device embodiments, reference may be made to the corresponding content in the foregoing method embodiments.

[0117] As Figure 5 shown, Figure 5 The structural block diagram of an electronic device 20 provided in the embodiments of the present application is shown. The electronic device 20 is the first electronic device in the foregoing embodiments. The electronic device 20 includes: a processor 21 and a memory 22.

[0118] Figure 5 The components and structure of the electronic device 20 shown are exemplary rather than restrictive. According to needs, the electronic device 20 may also have other components and structures.

[0119] The processor 21, the memory 22, and other components that may appear in the electronic device 20 are electrically connected directly or indirectly to each other to realize data transmission or interaction. For example, the processor 21, the memory 22, and other possible components may be electrically connected to each other through one or more communication buses or signal lines.

[0120] Among them, the memory 22 is used to store programs, such as the programs corresponding to the browser control method mentioned above.

[0121] The processor 21 is used to call the program stored in the memory to execute the browser control method mentioned above.

[0122] Among them, the memory 22 may be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.

[0123] The processor 21 may be an integrated circuit chip with the ability to process signals. The above-mentioned processor 21 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), an Accelerated Processing Unit, a Multimedia Application Processor (MAP), a microprocessor, etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. Or the processor 21 may also be any conventional processor, etc.

[0124] The embodiments of the present application also provide a non-volatile computer-readable storage medium (hereinafter referred to as the storage medium). A computer program is stored on the storage medium. When the computer program is run by the processor 21 as described above, it executes the method disclosed in any of the above embodiments.

[0125] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.

[0126] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0127] In addition, each functional module in various embodiments of this application may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part.

[0128] The aforementioned computer-readable storage media include: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0129] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.< / canvas>

Claims

1. A browser control method, characterized in that, Applied to a first electronic device, the first electronic device includes an agent and a large language model, the first electronic device is connected to a second electronic device, and the second electronic device is used to install a browser. The method includes: When receiving an instruction related to the browser input by the user, controlling the browser in the second electronic device to perform a startup operation, and controlling the browser to start screen recording; Realtime display on the first electronic device the page screenshot obtained by the browser during screen recording, wherein the page screenshot includes the first page screenshot after the browser performs the startup operation; Labeling the operable elements in the first page screenshot according to the coordinates of the operable elements in the first page screenshot to obtain a first picture; wherein the coordinates of the operable elements include the length, height of the box, and the highest point on the left; The agent generates a first prompt word according to the instruction; The large language model decomposes the instruction into multiple steps with an execution order according to the first prompt word; Performing a first operation on the browser according to the first picture and the first step in the instruction; The agent obtains a second page screenshot after performing the first operation on the browser; The agent labels the operable elements in the second page screenshot to obtain a second picture; The agent adds a description of the execution of the first step to the first prompt word to obtain a second prompt word, and sends the second prompt word and the second picture to the large language model; The large language model determines that the first step has been completed according to the second prompt word and the second picture; Performing a second operation on the browser according to the second picture and the second step in the instruction until all steps in the instruction are completed.

2. The method according to claim 1, wherein The performing a first operation on the browser according to the first picture and the first step in the instruction includes: When the first picture indicates that the browser has been started, converting the first step into a target instruction based on a preset data structure; When the instruction type of the target instruction is a page element operation type, controlling the browser to traverse the operable elements in the browser page according to the coordinates of the operable elements in the target instruction, and controlling the browser to perform a first operation on the operable element corresponding to the target instruction, wherein the coordinates of the operable elements in the target instruction are obtained through the first picture.

3. The method according to claim 2, wherein The method further includes: When the instruction type of the target instruction is a tool type, calling a tool corresponding to the target instruction to perform a first operation on the browser.

4. The method according to claim 1, wherein The method further includes: Listening for browser events of the user on the page screenshot obtained by the browser during screen recording, and controlling the browser to perform operations according to the browser events.

5. A browser control device, characterized in that, Included in a first electronic device, the first electronic device includes an agent and a large language model, the first electronic device is connected to a second electronic device, and the second electronic device is used to install a browser. The device includes: A function module, which is configured to control the browser in the second electronic device to start up and control the browser to start screen recording when receiving an instruction related to the browser input by the user; A display module, which is configured to display in real time on the first electronic device a page screenshot obtained by the browser during screen recording, wherein the page screenshot includes a first page screenshot after the browser starts up; A marking module, which is configured to mark the operable elements in the first page screenshot according to the coordinates of the operable elements in the first page screenshot to obtain a first picture; wherein the coordinates of the operable elements include the length, height and the highest point on the left side of the square; The function module is further configured to generate a first prompt word by using the agent according to the instruction; and decompose the instruction into multiple steps with an execution order by using the large language model according to the first prompt word; The function module is further configured to perform a first operation on the browser according to the first picture and the first step in the instruction; The marking module is further configured to obtain a second page screenshot after the first operation on the browser by using the agent; and mark the operable elements in the second page screenshot by using the agent to obtain a second picture; The function module is further configured to add a description of the execution of the first step to the first prompt word by using the agent to obtain a second prompt word, and send the second prompt word and the second picture to the large language model; and use the large language model to determine that the first step is completed according to the second prompt word and the second picture; The function module is further configured to perform a second operation on the browser according to the second picture and the second step in the instruction until all steps in the instruction are completed.

6. An electronic device, characterized in that, It includes: A memory and a processor, and the processor is connected to the memory; The memory is used to store programs; The processor is configured to call the program stored in the memory to execute the method according to any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is run by the processor, it executes the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Webpage end communication method and device, electronic equipment and storage medium

    CN111538601A

  • Webpage access method and device based on remote browser, medium and electronic equipment

    CN117520006A