Android Mobile Phone Intelligent Control Method and System Based on Large Model Agents
The large model-based intelligent agent addresses limitations of existing smartphone assistants by identifying and executing tasks on Android devices, ensuring efficient and adaptable operation execution with low costs and high accuracy.
Patent Information
- Application Number
- CN202411437367.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-15
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-10-15
AI Technical Summary
The existing technology has problems such as limited scope of function application, high initial deployment and maintenance costs, difficulty in understanding user complex language instructions and context association, page information acquisition and non-standard page processing defects.
The intelligent control method of Android mobile phones based on large model agents is adopted. By obtaining user input tasks, identifying and judging the status of the Android mobile phone page, combining accessibility services for operation planning and execution, including processing of page XML and screenshots, non-standard state judgment, compression simplification and completion processing, and the operation type and location are output using the Agent model.
It realizes intelligent control with wide applicability and low cost, can accurately understand complex language instructions and context, adapt to non-standard page states, and improves operational efficiency and flexibility.
Smart Images

Figure CN119342136B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and relates to an intelligent control method and system for Android mobile phones, and particularly to an intelligent control method and system for Android mobile phones based on large model agents. Background Art
[0002] Intelligent terminals such as mobile phones are excellent carriers for the implementation of large model technologies. The combination of intelligent terminals and large models can bring subversive product experiences and highly intelligent services to users. Currently, the mobile phone intelligent assistants and RPA tools on the market still have multiple limitations in assisting users to complete operation executions. The scope of application of their functions is extremely limited, not suitable for highly variable tasks, the initial deployment and maintenance costs are relatively high, and there are also many difficulties in understanding complex user language instructions and context associations.
[0003] Among them, mobile phone intelligent assistants such as Siri execute tasks related to application operations in various ways. These ways include using deep integration and API access, starting and operating third-party applications through the Intent mechanism provided by the operating system (such as in the Android system), using App Shortcut and Siri Shortcuts to activate in-app shortcut commands. In addition, mobile phone intelligent assistants first understand user intentions through natural language processing (NLP), then convert these intentions into specific tasks, and interact with application programs through protocols and APIs. Sometimes, mobile phone intelligent assistants also support dynamic frameworks and plugin systems, allowing developers to create specific plugins to achieve customized communication. For example, users can send text messages, play music, or open applications through Siri, and these tasks are implemented through APIs such as SiriKit or MusicKit. Such multiple ways ensure that intelligent assistants can execute a wide range of tasks, thereby improving user convenience and system flexibility. However, for the related tasks of operation execution proposed by users, mobile phone intelligent assistants generally implement task execution by calling APIs or SDKs. And not all applications are willing to open APIs to intelligent assistants, which makes intelligent assistants restricted in certain tasks, and the complex function calls of some applications require specific permissions, and intelligent assistants may not be able to directly access them, and can only assist users to complete relatively simple tasks.
[0004] A method for implementing automatic task execution by an RPA (Robotic Process Automation) tool on Android phones developed by third parties mainly includes UI automation through technologies such as screen coordinate clicking, image recognition, and OCR; control recognition by using control IDs, class names, or XPath, etc.; writing scripts for complex operations; recording and repeating user operations through the recording and playback function; directly calling Android system APIs; and analyzing and intercepting network requests, etc. The whole process involves preparing the device, defining and recording tasks, debugging and testing, and finally deploying and running to achieve task automation. Although using RPA technology to automate tasks can significantly improve efficiency, there are also some limitations, such as script invalidation caused by UI changes, limited complex task processing capabilities, compatibility issues brought by application updates and system updates, and performance and resource consumption during operation. In addition, RPA also faces challenges in anti-interference ability and multitasking processing. And RPA usually requires users to build their own workflows to complete, and the relatively high usage threshold is also one of the problems faced by the current implementation and popularization of related technologies.
[0005] An AI Agent (Artificial Intelligence Agent) is a software system based on artificial intelligence technology, which can be designed to autonomously execute specific tasks or solve problems. It can run in a specific environment and has certain autonomous decision-making and task planning capabilities. The Agent analyzes tasks based on the 'brain' composed of large models, designs and plans solutions and paths to problems, and at the same time has the ability to call tools and execute operations. As a new technical solution, it can be used to solve many problems faced by mobile phone intelligent control tasks. However, there is currently no Android phone intelligent control method based on large model agents. Even if some mobile phone intelligent control methods use large model technology, they still have certain defects in aspects such as page information acquisition, non-standard page judgment and processing, and page XML processing.
[0006] Therefore, in view of the defects existing in the above-mentioned prior art, it is necessary to develop a new Android phone intelligent control method based on large model agents. Summary of the Invention
[0007] In order to overcome the defects of the prior art, the present invention proposes an Android phone intelligent control method and system based on large model agents, which can identify and judge the state of the mobile phone page and plan and execute operations in combination with the user's intention.
[0008] To achieve the above object, the present invention provides the following technical solutions:
[0009] An Android phone intelligent control method based on large model agents, characterized by including the following steps:
[0010] 1) Obtain the task input by the user;
[0011] 2) Determine whether the task input by the user is within the processing scope of the Agent model. If so, proceed to step 3); otherwise, end the execution of the task input by the user.
[0012] 3) Obtain the page information of the Android phone, where the page information includes page XML and page screenshots.
[0013] 4) Perform judgment processing on the non-standard state of the page based on the page information to ensure that the obtained page information is for a standard page.
[0014] 5) Compress, simplify, and complete the page for the obtained page XML to obtain the processed page XML.
[0015] 6) Input the processed page XML and page screenshots into the Agent model, and the Agent model outputs the operation type and location for the next step of the Android phone to execute.
[0016] 7) Based on the output of the Agent model, perform operations on the Android phone based on the AccessibilityService of the accessibility service.
[0017] Preferably, step 3) specifically includes:
[0018] 31) Obtain the accessibility service node information AccessibilityServiceNodeInfo and page screenshot Screenshot of the page of the Android phone based on the AccessibilityService of the accessibility service.
[0019] 32) Parse the nodes Node in the accessibility service node information AccessibilityServiceNodeInfo and convert the parsing result into a mutable character sequence StringBuilder.
[0020] 33) Adjust the structure of the mutable character sequence StringBuilder and perform parallel parsing on the child nodes ChildNod of its corresponding node Node.
[0021] 34) Obtain the parsing results of each child node ChildNod and assemble them into the mutable character sequence StringBuilder, and then adjust the structure of the assembled mutable character sequence StringBuilder again.
[0022] 35) Obtain the final page XML based on the mutable character sequence StringBuilder after the second adjustment.
[0023] 36) Perform quality compression on the page screenshot Screenshot without changing the original resolution of the page screenshot Screenshot to obtain the final page screenshot.
[0024] Preferably, step 4) specifically includes:
[0025] 41) Page waiting processing, which is used to process pages that need to wait before becoming standard pages to ensure obtaining page information of a page that does not need to wait before becoming a standard page;
[0026] 42) Pop-up window processing, which is used to process the situation where there are pop-up windows on the page to ensure obtaining page information of a page without pop-up windows.
[0027] Preferably, step 41) specifically includes:
[0028] 411) Obtain the page XML of the Android phone again;
[0029] 412) Determine whether the difference in the total number of nodes in the page XMLs obtained in two adjacent acquisitions is less than D, and determine whether the position attributes of the first k nodes in the page XMLs obtained in two adjacent acquisitions are exactly the same. If the difference is less than D and the position attributes are exactly the same, then use the page XML obtained again as the page XML of the page that does not need to wait before becoming a standard page; if the difference is not less than D and / or the position attributes are not exactly the same, then determine whether the total elapsed time is greater than or equal to T seconds. If it is greater than or equal to T seconds, then use the page XML obtained again as the page XML of the page that does not need to wait before becoming a standard page. If it is less than T seconds, then return to step 411);
[0030] 413) Obtain the page screenshot of the Android phone again;
[0031] 414) Input the page screenshot obtained again into the single-image Wait model and input the page screenshots obtained in two adjacent acquisitions into the double-image Wait model to respectively determine whether the loading is completed. If the output results of both the single-image Wait model and the double-image Wait model are that the loading is completed, then use the page screenshot obtained again as the page screenshot of the page that does not need to wait before becoming a standard page; if the output result of the single-image Wait model and / or the double-image Wait model is that the loading is not completed, then determine whether the number of loops is greater than or equal to n times. If it is greater than or equal to n times, then use the page screenshot obtained again as the page screenshot of the page that does not need to wait before becoming a standard page. If it is less than n times, then return to step 413).
[0032] Preferably, step 42) specifically includes:
[0033] 421) Obtain a page screenshot of an Android phone;
[0034] 422) Detect whether there is a pop-up window in the page screenshot through a pop-up window detection model. If not, use the page screenshot as the page screenshot of the page without a pop-up window. If so, determine whether it has cycled once. If it has cycled once, hand it over to the user for operation. If it has not cycled once, hand it over to the pop-up window closing model to attempt to close the pop-up window by the pop-up window closing model. If the pop-up window closing model can return the closing position, automatically close the pop-up window according to the closing position and return to step 421). If the pop-up window closing model cannot return the closing position, hand it over to the user for operation.
[0035] Preferably, step 5) specifically includes:
[0036] 51) Determine whether it is necessary to retain the off-screen nodes in the page XML. If not, proceed to step 52). If so, proceed to step 53);
[0037] 52) Delete the off-screen nodes in the page XML;
[0038] 53) Delete the redundant nodes in the page XML;
[0039] 54) Determine whether the page of the Android phone is a special page. If so, proceed to step 55). If not, proceed to step 56);
[0040] 55) Perform deletion information or supplementary information processing on the special page;
[0041] 56) Use the optical character recognition method to identify all the text and the positions where the text is located in the page of the Android phone, compare the text recognized by the optical character recognition method with the text in the page XML, remove the text in the text recognized by the optical character recognition method that is repeated with the text in the page XML, and insert the remaining text into the page XML;
[0042] 57) Simplify the attributes in the inserted page XML.
[0043] Preferably, the intelligent control method for Android phones based on a large model agent further includes:
[0044] 8) Display the results of the operations performed on the Android phone.
[0045] In addition, the present invention also provides an intelligent control system for Android phones based on a large model agent, which is characterized by including:
[0046] A task input module, which is used to obtain the tasks input by the user;
[0047] A task discrimination model, which is used to determine whether the task input by the user is within the processing scope of the Agent model to determine whether to continue executing the task input by the user;
[0048] A data acquisition module, which is used to acquire the page information of the Android mobile phone, and the page information includes page XML and page screenshots;
[0049] A page status judgment and processing module, which is used to judge and process the non-standard status of the page based on the page information to ensure that the obtained page information is of a standard page;
[0050] A data processing module, which is used to compress, simplify and complete the page for the obtained page XML to obtain the processed page XML;
[0051] An Agent model, which is used to output the operation type and position to be executed by the Android mobile phone next based on the processed page XML and page screenshots;
[0052] A task execution module, which is used to operate and execute the Android mobile phone based on the AccessibilityService according to the output of the Agent model.
[0053] Moreover, the present invention also provides an Android mobile phone intelligent control device based on a large model agent, which is characterized by including:
[0054] One or more processors;
[0055] A memory, which is used to store one or more programs;
[0056] When the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned Android mobile phone intelligent control method based on a large model agent.
[0057] Finally, the present invention provides a computer-readable storage medium, on which a computer program is stored, and is characterized in that when the program is executed by a processor, the steps of the above-mentioned Android mobile phone intelligent control method based on a large model agent are implemented.
[0058] Compared with the prior art, the Android mobile phone intelligent control method and system based on a large model agent of the present invention have one or more of the following beneficial technical effects:
[0059] 1. The present invention can identify and judge the mobile phone page status, and plan and execute operations in combination with the user's intention.
[0060] 2. The present invention has a wide range of applications, low initial deployment and maintenance costs, and there is no difficulty in accurately understanding complex language instructions and context associations of users. Brief Description of the Drawings
[0061] Figure 1 is a flowchart of the Android phone intelligent control method based on the large model agent of the present invention.
[0062] Figure 2 is a flowchart of obtaining the page information of the Android phone in the present invention.
[0063] Figure 3 is a flowchart of the judgment and processing of the non-standard state of the page in the present invention.
[0064] Figure 4 is a flowchart of the page waiting processing in the present invention.
[0065] Figure 5 is a flowchart of the pop-up window processing in the present invention.
[0066] Figure 6 is a flowchart of the compression, simplification and page completion processing of the obtained page XML in the present invention.
[0067] Figure 7 is a schematic diagram of the composition of the Android phone intelligent control system based on the large model agent of the present invention. Detailed Description of the Invention
[0068] Before detailing any embodiment of the present invention, it should be understood that in its application, the present invention is not limited to the construction and arrangement details of the components described in the following description or illustrated in the following drawings. The present invention is capable of other embodiments and can be practiced or carried out in various ways. Additionally, it should be understood that the wording and terminology used herein are for the purpose of description and should not be regarded as restrictive. As used herein, "including" or "having" and their variants are intended to cover the items listed hereinafter and their equivalents as well as additional items. Unless otherwise specified or limited, the terms "installed", "connected", "supported" and "coupled" and their variants are used broadly and cover direct installation and indirect installation, connection, support and coupling. Further, "connection" and "coupling" are not limited to physical or mechanical connection or coupling.
[0069] Moreover, on the first aspect, in the disclosure of the present invention, the orientation or positional relationships indicated by terms such as "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the above terms should not be construed as limiting the present invention. On the second aspect, the term "one" should be understood as "at least one" or "one or more". That is, in one embodiment, the number of an element can be one, while in other embodiments, the number of the element can be multiple. The term "one" should not be construed as a limitation on the quantity.
[0070] Before introducing the specific content of the present invention, some technical terms used in the present invention will be briefly introduced to facilitate those skilled in the art to better understand the present invention.
[0071] 1. RPA, the full name is Robotic Process Automation, which is a technology that uses software robots to imitate the regular and repetitive tasks performed by humans on a computer. RPA robots usually perform tasks such as data input, data processing, and data transmission automatically through the interaction of the user interface, thereby greatly improving efficiency and reducing human errors.
[0072] 2. Android's AccessibilityService, which was originally designed to help disabled users with visual, hearing impairments, etc. operate their mobile phones more conveniently. This special permission allows applications to perform some operations that are usually only possible for users, such as clicking buttons, swiping the screen, entering text, etc. Therefore, it is widely used in fields such as automated testing and task automation.
[0073] 3. ASR is the abbreviation of Automatic Speech Recognition. This technology can convert speech signals into text. Important applications of ASR include voice assistants (such as Siri, Google Assistant), voice input methods, intelligent devices with voice control functions, etc.
[0074] 4. XML (Extensible Markup Language) is a widely used format in Android development for defining user interface layouts, resource files, and various configuration files. It is mainly used to define the UI layout of an application. Developers can describe the positions, attributes, and values of UI components through a hierarchical XML structure. These layout files usually exist in the form of files, which can intuitively display the interface design, thereby improving the readability and maintainability of the code.
[0075] 5. Bounding Box (the bounding box of a node), abbreviated as bbox. In the accessibility service of Android development, the AccessibilityNodeInfo class is used to describe elements in the view hierarchy, and each element is called a node. The Bounding Box defines its position and size on the screen. It is a rectangular area used to accurately describe the boundary position of the node on the screen, specifically including the left boundary, upper boundary, right boundary, and lower boundary.
[0076] Figure 1 The flowchart of the Android phone intelligent control method based on the large model agent of the present invention is shown. As Figure 1 shown, the Android phone intelligent control method based on the large model agent of the present invention includes the following steps:
[0077] I. Obtain the task input by the user.
[0078] In the present invention, the user can input tasks in two forms: voice and text. If the user inputs a task in voice form, the client of the Android phone calls the ASR tool to convert the input voice into text form.
[0079] II. Determine whether the task input by the user is within the processing scope of the Agent model. If so, continue to process and enter step III; otherwise, it means that the Agent model cannot process the task input by the user, and the execution of the task input by the user is ended.
[0080] In the present invention, a task discrimination model can be used to determine whether the task input by the user is within the processing scope of the Agent model. The task discrimination model can be a binary classifier fine-tuned and trained based on the GLM-4-9B model. Before use, some data with the input being the task input by the user and the output being the operation types that the Agent model can process can be collected as the training data set, and the GLM-4-9B model can be fine-tuned and trained using this training data set to obtain the task discrimination model. This part of the content is technical that those skilled in the field of large language models can understand and master. For the sake of simplicity, it will not be described in detail here.
[0081] III. Obtain the page information of the Android phone, where the page information includes the page XML and the page screenshot.
[0082] As Figure 2 shown, in the present invention, obtaining the page information of the Android phone specifically includes:
[0083] 1. Obtain the accessibility service node information AccessibilityServiceNodeInfo and the page screenshot Screenshot of the Android phone based on the accessibility service AccessibilityService.
[0084] 2. Parse the nodes Node in the accessibility service node information AccessibilityServiceNodeInfo and convert the parsing result into a mutable character sequence StringBuilder.
[0085] In the present invention, the accessibility service node information AccessibilityServiceNodeInfo is passed into RecursiveTask (a subclass of ForkJoinPool, specifically used for recursive tasks with return values) to execute the node Node parsing. RecursiveTask parses the passed-in node Node to obtain the key fields and then converts them into StringBuilder. Moreover, when parsing the node Node, the parsing starts from the root node RootNode.
[0086] The code logic of RecursiveTask is as follows:
[0087] (1) Pass parameters. The parameters passed include: Node, that is, the node Node object to be parsed; index, that is, the order of the current node Node in the parent node FatherNode. When assembling the result, it is spliced according to the index order to prevent the XML structure order from being disordered; depth, that is, the depth of the current node Node in the node tree NodeTree, which is used to supplement the number of spaces when formatting indentation.
[0088] (2) Method. The method is compute(), which is the core method of RecursiveTask. If the current task needs to be split, the task is split into more subtasks, and the fork() method is used to process these subtasks in parallel. After the parsing is completed, the join() method can be called to obtain the result after the corresponding task is completed. The node Node parsing logic is executed asynchronously in this method.
[0089] In the `compute()` method, the key fields of the current `Node` are first obtained, then the child nodes `ChildNode` of the current `Node` are obtained, and then split into a fixed number of subtasks according to the number of child nodes `ChildNode`. Each child node `ChildNode` and the sequence index `index` where the current `Node` is located are passed into a new `RecursiveTask`, and the `fork()` method is called to execute the parsing task of each child node `ChildNode`, and the parsing results of the child nodes `ChildNode` are obtained through `join()`.
[0090] In the present invention, parsing the `Node` in the `AccessibilityServiceNodeInfo` of the accessibility service node information and converting the parsing result into a mutable character sequence `StringBuilder` specifically includes:
[0091] (1) Create an `XmlSerializer` object (the present invention uses the `XmlSerializer` object to ensure that the data conforms to the XML specification), and set the `StringWriter` object and the start document `startDocument` attribute and the start tag `startTag` attribute. The specific example code for setting is as follows:
[0092] val stringWriter = StringWriter()
[0093] val serializer = Xml.newSerializer()
[0094] serializer.setOutput(stringWriter)
[0095] serializer.startDocument("UTF-8", true)
[0096] serializer.startTag("", "node")
[0097] (2) Obtain the key fields of the `Node` in the `AccessibilityServiceNodeInfo` of the accessibility service node information.
[0098] The `Node` itself carries some information fields, and the key fields of the `Node` can be directly obtained through method calls. An example of loading the key fields of the `Node` into the `XmlSerializer` object is as follows:
[0099] serializer.attribute("", "class", node.className.toString());
[0100] serializer.attribute("", "text", node.text.toString());
[0101] Among them, node.className can obtain the information of the className field of the node Node, and node.text can obtain the information of the text field of the node Node.
[0102] In the present invention, the key fields of the node Node include: className, contentDescription, text, viewIdResourceName, isClickable, isEnabled, isChecked, isFocused, isFocusable, isScrollable, isSelected, isLongClickable, and boundsInScreen.
[0103] (3) Set the endTag attribute and endDocument attribute of the XmlSerialize object, insert the key fields of the obtained node Node into the XmlSerializer object, and convert it into a mutable character sequence StringBuilder using the StringWriter object. The specific example code for setting is as follows:
[0104] serializer.endTag("", "node")
[0105] serializer.endDocument()
[0106] val stringBuilder = StringBuilder()
[0107] stringBuilder.append(stringWriter.toString());
[0108] Among them, stringWriter is the StringWriter object set in step (1).
[0109] Thus, an exemplary StringBuilder obtained is as follows:
[0110] <? xml version='1.0'encoding='UTF-8'standalone='yes'? > <node NAF="true"index="0"text=""resource-id=""class="android.view.ViewGroup"package="com.zhipu.agent"content-desc=""checkable="false"checked="false"clickable="false"enabled="true"focusable="false"focused="false"scrollable="false"long-clickable="false"password="false"selected="false"
[0111] bounds="[0,148][1224,2776]" / >.
[0112] 3. Adjust the structure of the variable character sequence StringBuilder, and perform parallel parsing on the child nodes ChildNod of the corresponding node Node. Specifically, it includes:
[0113] (1) Remove the version, encoding, and standalone information of the XmlSerializer object in the mutable character sequence StringBuilder, and only keep the version, encoding, and standalone information of the outermost node Node to avoid duplication. The sample code is as follows:
[0114] StringBuilder.toString().replace("<?xml version='1.0'encoding='UTF-8'
[0115] standalone='yes'? >","").
[0116] (2) Determine whether the node Node corresponding to the variable character sequence StringBuilder contains a child node ChildNode. If the node Node corresponding to the variable character sequence StringBuilder contains a child node ChildNode, remove the end tag endTag of the node Node in the variable character sequence StringBuilder, otherwise a structural error will occur when assembling the child node ChildNode. The sample code is as follows:
[0117] StringBuilder.toString().replace(" / >",">").
[0118] (3) Parallelly obtain the key fields of each child node ChildNod and set a depth parameter for each child node ChildNod to record the depth of each child node ChildNod. Save the key fields and depth parameter of each obtained child node ChildNod into RecursiveTaskList.
[0119] Among them, the way to obtain the key fields of each child node ChildNod is the same as the way to obtain the key fields of node Node in step 2. For the sake of simplicity, it will not be described again. However, the key fields of multiple child nodes ChildNod of a node Node can be obtained in parallel, so as to improve the parsing efficiency.
[0120] Of course, for the convenience of subsequent assembly, it is also necessary to save the index of each child node ChildNod into RecursiveTaskList.
[0121] 4. Obtain the parsing results of each child node ChildNod and assemble them into a mutable character sequence StringBuilder, and then adjust the structure of the assembled mutable character sequence StringBuilder again. Specifically, it includes:
[0122] (1) Traverse RecursiveTaskList to obtain the key fields and depth parameter of each child node ChildNod. Of course, if there are multiple child nodes ChildNod, it is also necessary to obtain the index of each child node ChildNod.
[0123] (2) Concatenate the key fields of each obtained child node ChildNod to the mutable character sequence StringBuilder, and set the number of indented spaces according to the depth parameter to complete format indentation. The example code for format indentation is as follows:
[0124] StringBuilder.append("\r\n").append("".repeat(depth)).
[0125] Among them, "\r\n" is the line feed character, and depth is the depth of the current child node ChildNod.
[0126] Example code for splicing the key fields of each obtained child node ChildNod into a mutable character sequence StringBuilder is as follows:
[0127] StringBuilder.append(childNodeStr).
[0128] Among them, childNodeStr is the parsing result of the child node ChildNode. If there are multiple child nodes ChildNode, childNodeStr needs to be assembled according to the index order to prevent the overall structure from being chaotic.
[0129] (3) After all the key fields of the child nodes ChildNod are spliced into the mutable character sequence StringBuilder, complete the end tag endTag of the mutable character sequence StringBuilder. The example code is as follows:
[0130] StringBuilder.append("\r\n").append("".repeat(depth - 1)).append("").
[0131] Thus, starting from the parsing of the root node RootNode, through the gradual parsing and splicing of its child nodes ChildNod and the child nodes ChildNod of the child nodes ChildNod, all the nodes Node in the accessibility service node information AccessibilityServiceNodeInfo of the page will be assembled into a StringBuilder.
[0132] 5. Obtain the final page XML based on the mutable character sequence StringBuilder after re - adjustment.
[0133] The final StringBuilder, that is, the final XML of the current page, can be obtained through the ForkJoinPool.invoke() method.
[0134] An exemplary final XML is as follows:
[0135] <?xml version='1.0' encoding='UTF - 8' standalone='yes'?
[0136] <hierarchy rotation="0">
[0137] <node index="0" text="" resource - id=""
[0138] class="android.widget.FrameLayout" package="com.autonavi.minimap" content-desc="" checkable="false" checked="false" clickable="false" enabled="true" focusable="true" focused="false" scrollable="false" long-clickable="false" password="false" selected="false"
[0139] bounds="[0,0][1224,2776]">
[0140] <node index="0" text="" resource-id="">
[0141] class="android.widget.LinearLayout" package="com.autonavi.minimap" content-desc="" checkable="false" checked="false" clickable="false" enabled="true" focusable="false" focused="false" scrollable="false" long-clickable="false" password="false" selected="false"
[0142] bounds="[0,0][1224,2776]">
[0143] <node index="0" text="" resource-id="android:id / content"
[0144] class="android.widget.FrameLayout" package="com.autonavi.minimap" content-desc="" checkable="false" checked="false" clickable="false" enabled="true" focusable="false" focused="false" scrollable="false" long-clickable="false" password="false" selected="false"
[0145] bounds="[0,0][1224,2776]">
[0146] <node index="0" text=""
[0147] resource-id="com.autonavi.minimap:id / root_view"
[0148] class="android.widget.LinearLayout" package="com.autonavi.minimap" content-desc="" checkable="false" checked="false" clickable="false" enabled="true" focusable="false" focused="false" scrollable="false" long-clickable="false" password="false" selected="false"
[0149] bounds="[0,0][1224,2776]">
[0150] <node index="0" text=""
[0151] resource-id="com.autonavi.minimap:id / fl_content_view"
[0152] class="android.widget.FrameLayout" package="com.autonavi.minimap" content-desc="" checkable="false" checked="false" clickable="false" enabled="true" focusable="false" focused="false" scrollable="false" long-clickable="false" password="false" selected="false"
[0153] bounds="[0,0][1224,2776]">
[0154] <node index="0" text=""
[0155] resource-id="com.autonavi.minimap:id / map_container"
[0156] class="android.widget.FrameLayout" package="com.autonavi.minimap" content-desc="" checkable="false" checked="false" clickable="false" enabled="true" focusable="false" focused="false" scrollable="false" long-clickable="false" password="false" selected="false"
[0157] bounds="[0,0][1224,2776]">
[0158] <node index="0" text=""
[0159] resource-id="com.autonavi.minimap:id / atmapsView"
[0160] class="android.widget.FrameLayout" package="com.autonavi.minimap" content-desc="" checkable="false" checked="false" clickable="false" enabled="true" focusable="false" focused="false" scrollable="false" long-clickable="false" password="false" selected="false"
[0161] bounds="[0,0][1224,2776]">
[0162] <node index="0" text="" resource-id="" class="android.view.View" package="com.autonavi.minimap" content-desc="" checkable="false" checked="false" clickable="false" enabled="true" focusable="false" focused="false" scrollable="false" long-clickable="false" password="false" selected="false" bounds="[0,0][1224,2776]" / >
[0163]
[0164]
[0165] <node index="1" text=""
[0166] resource-id="com.autonavi.minimap:id / temporary_layer"
[0167] class="android.widget.FrameLayout" package="com.autonavi.minimap" content-desc="" checkable="false" checked="false" clickable="false" enabled="true" focusable="false" focused="false" scrollable="false" long-clickable="false" password="false" selected="false"
[0168] bounds="[0,0][1224,2776]" /
[0169] <node index="2" text="" resource-id="" /
[0170] class="android.widget.RelativeLayout" package="com.autonavi.minimap" content-desc="" checkable="false" checked="false" clickable="false" enabled="true" focusable="false" focused="false" scrollable="false" long-clickable="false" password="false" selected="false"
[0171] bounds="[0,0][1224,2776]">
[0172] <node index="0" text=""
[0173] resource-id="com.autonavi.minimap:id / notification_layer"
[0174] class="android.widget.FrameLayout" package="com.autonavi.minimap" content-desc="" checkable="false" checked="false" clickable="false" enabled="true" focusable="false" focused="false" scrollable="false" long-clickable="false" password="false" selected="false"
[0175] bounds="[0,0][1224,2776]" / >
[0176]
[0177]
[0178]
[0179]
[0180]
[0181]
[0182] 。
[0183] 6. Perform quality compression on the page screenshot Screenshot without changing the original resolution of the page screenshot Screenshot to obtain the final page screenshot.
[0184] It is possible to perform quality compression on the screenshot Screenshot without changing the original resolution of the screenshot Screenshot to accelerate the picture transmission efficiency and reduce the traffic consumption, thereby obtaining the final screenshot of the page.
[0185] In the present invention, the Bitmap.compress() method can be used to compress the Bitmap image data after screenshotting. The method signature is: public boolean compress(Bitmap.CompressFormat format, int quality, OutputStream stream). Among them, the parameter format is the format of the compressed image, and the Bitmap.CompressFormat enumeration is used to specify the format, such as JPEG, PNG, and WEBP; the parameter quality is the quality of the compressed image, and the integer range is from 0 to 100, where 0 represents the lowest quality (minimum compression), and 100 represents the highest quality (minimum compression). If the format is PNG, this parameter will be ignored because PNG is lossless compression; the parameter stream is the output stream to which the compressed image data will be written.
[0186] Based on the AccessibilityServiceNodeInfo of the accessibility service node information, the present invention parses in parallel to obtain the ChildNode information of the child nodes, and at the same time obtains the information of multiple child nodes ChildNode, and then combines the obtained ChildNode information of the child nodes in sequence to obtain the complete XML content, which has been greatly improved in terms of time. Moreover, the present invention retains the information of the corresponding node Node at the end of the parsing of each node Node, and then reversely complements the node tree NodeTree to obtain the complete XML, which not only improves the efficiency but also ensures the integrity of the page information. Through a large number of experiments, it is found that under the same mobile phone interface, the speed of obtaining page information by the present invention saves nearly half of the time compared with recursive parsing and recursive parsing, and the integrity of the XML is also well guaranteed.
[0187] 4. Based on the page information, judge and process the non-standard state of the page to ensure that the obtained page information is of a standard page.
[0188] The prior art for the parsing and processing of Android GUI (Graphical User Interface) pages is all based on standard pages (fully loaded and no pop-up windows appear). However, during the actual operation of the Agent model on an Android mobile phone, non-standard page states (such as waiting for loading, page jumping, pop-up windows, etc.) often occur, and the prior art cannot make judgments and process these non-standard pages, so it will fail. The problem to be solved in this step is to distinguish these non-standard page states from the standard page states and provide a method for processing non-standard pages.
[0189] As Figure 3 shown, in the present invention, the judgment and processing of the non-standard state of the page specifically include:
[0190] 1. Page waiting processing.
[0191] The page waiting processing is used to process pages such as "page loading" and "page jumping" that need to wait for a while before becoming standard pages, so as to ensure obtaining page information of a page that is not a page that needs to wait before becoming a standard page (i.e., a standard page).
[0192] Figure 4 The flowchart of the page waiting processing in the present invention is shown. This process can take effect at any time during the operation of an application on an Android phone. After this process, the obtained page XML and page screenshot can be guaranteed not to be in non-standard page states such as "page loading" and "page jumping".
[0193] In the present invention, the page waiting processing is divided into two parallel processing processes: the XML (Extensible Markup Language) side and the screenshot side. Its core idea is to cyclically obtain the page XML and page screenshot of the page of the Android phone, and respectively judge whether waiting is still needed based on the page XML and page screenshot. If the judgment result is that waiting is still needed, enter the next cycle until the judgment result is that waiting is no longer needed or the number of cycles or the time consumption exceeds a preset threshold.
[0194] Specifically, as Figure 4 shown, the page waiting processing specifically includes:
[0195] (1) Obtain the page XML and page screenshot of the page of the Android phone.
[0196] In the present invention, since the page XML and page screenshot of the page of the Android phone have been obtained in step three, therefore, when performing the page waiting processing, this step can be omitted.
[0197] After obtaining the page XML and page screenshot of the page of the Android phone, it is divided into two parallel processing processes: the XML (Extensible Markup Language) side and the screenshot side.
[0198] Among them, the processing process on the XML side is as follows:
[0199] (2) Obtain the page XML of the page of the Android phone again.
[0200] The method in step three can still be used to obtain the page XML of the page of the Android phone again.
[0201] (3) Determine whether the difference in the total number of nodes (Node class nodes in the present invention) in the page XML obtained twice in succession is less than D, and determine whether the position attributes of the first k nodes (Node class nodes in the present invention) in the page XML obtained twice in succession are exactly the same. If the difference is less than D and the position attributes are exactly the same, it indicates that it is not in non-standard page states such as "page loading" or "page jumping", and the page XML obtained again can be used as the page XML of the page that does not need to wait to become a standard page; if the difference is not less than D and / or the position attributes are not exactly the same, determine whether the total time consumed on the XML side is greater than or equal to T seconds. If it is greater than or equal to T seconds, the page XML obtained again can be used as the page XML of the page that does not need to wait to become a standard page. If it is less than T seconds, return to step (2) to obtain the page XML of the Android phone page again and start the next loop.
[0202] In the present invention, by traversing the obtained XML file tree, the statistics and difference calculation of the total number of node class nodes can be realized. By pre-order traversing the XML file tree, the first k node class nodes and their bounds attributes (the bounds attribute represents the pixel value of the position of the element on the screen and is the position attribute) can be obtained. And if the total time consumed on this side is greater than or equal to T seconds, the loop is exited to avoid infinite looping.
[0203] Preferably, D is between 10 and 20, k is between 5 and 10, and T is between 1.5 and 2.5.
[0204] The processing process on the screenshot side is as follows:
[0205] (4) Obtain the page screenshot of the Android phone again.
[0206] The method in step three can still be used to obtain the page screenshot of the Android phone again.
[0207] (5) Input the page screenshot obtained again into the single-image Wait model and input the page screenshots obtained twice in succession into the double-image Wait model to respectively determine whether the loading is completed. If the output results of both the single-image Wait model and the double-image Wait model are that the loading is completed, the page screenshot obtained again can be used as the page screenshot of the page that does not need to wait to become a standard page; if the output result of the single-image Wait model and / or the double-image Wait model is that the loading is not completed, determine whether the number of loops is greater than or equal to n times (preferably, n is between 5 and 10). If it is greater than or equal to n times, the page screenshot obtained again can be used as the page screenshot of the page that does not need to wait to become a standard page. If it is less than n times, return to step (4) to start the next loop.
[0208] Among them, the situations mainly concerned by the single - picture Wait model include: (1) There are obvious prompts on the page indicating that it is in the loading state; (2) The page is obviously in the transition state between two - page jumps. The situations mainly concerned by the double - picture Wait model include those that the single - picture Wait model cannot handle: (1) There is no page jump, but some changes within the page have not ended; (2) There is a page jump, but due to various reasons (such as good mobile phone performance and network conditions), the screenshots before and after are relatively complete. Thus, through the single - picture Wait model and the double - picture Wait model, it is possible to judge various pages such as "page loading" and "page jump" that need to wait for a while before becoming standard pages.
[0209] In the present invention, both the single - picture Wait model and the double - picture Wait model are trained end - to - end prediction models based on convolutional neural networks. When specifically used, CNN (convolutional neural network) can be used as the single - picture Wait model and the double - picture Wait model, and training data is collected and manually labeled. Then, the constructed single - picture Wait model and double - picture Wait model are trained with the labeled training data to obtain the single - picture Wait model and the double - picture Wait model.
[0210] Among them, the training data of the single - picture Wait model are the pages with obvious prompts on the page indicating that it is in the loading state after manual annotation and the pages that are obviously in the transition state between two - page jumps after manual annotation. The training data of the double - picture Wait model are the pages where there is no page jump but some changes within the page have not ended after manual annotation and the pages where there is a page jump but due to various reasons the screenshots before and after are relatively complete after manual annotation.
[0211] 2. Pop - up window processing.
[0212] Pop - up window processing is used to handle the situation where there are pop - up windows on the page to ensure obtaining the page information of a page without pop - up windows (i.e., a standard page). That is to say, pop - up window processing is responsible for handling the situation where there are "pop - up windows" on the page. Here, the "pop - up window" mainly refers to an unexpected interactive window that appears in the operation path. The appearance of this window is uncertain and interferes with and hinders the normal completion of the original task.
[0213] Figure 5 The flowchart of pop - up window processing in the present invention is shown. As Figure 5 shown, the pop - up window processing specifically includes:
[0214] (1) Obtain the page screenshot of the Android mobile phone page.
[0215] Since the page screenshot of the Android phone has been obtained in Step 3, when starting to handle pop-ups, this step can be omitted first and only executed in the subsequent loop.
[0216] Of course, this step can still use the method in Step 3 to obtain the page screenshot of the Android phone.
[0217] (2) Use the pop-up window detection model to detect whether there is a pop-up window in the page screenshot. If not, use the page screenshot as the page screenshot of the page without a pop-up window. If so, determine whether it has been looped once (that is, whether it has been processed by the pop-up window closing model once). If it has been looped once, hand it over to the user for operation. If it has not been looped once, hand it over to the pop-up window closing model, and the pop-up window closing model attempts to close the pop-up window. If the pop-up window closing model can return the closing position, automatically close the pop-up window according to the closing position and return to step (1). If the pop-up window closing model cannot return the closing position, hand it over to the user for operation.
[0218] Both the pop-up window detection model and the pop-up window closing model are also end-to-end prediction models based on convolutional neural networks that have been trained. The pop-up window detection model is used to determine whether there is a pop-up window in the page screenshot, and the pop-up window closing model is used to determine whether there is a click position for closing the pop-up window in the page screenshot. If so, predict its position.
[0219] In specific use, CNN (Convolutional Neural Network) can be used as the pop-up window detection model and the pop-up window closing model, and training data is collected and manually labeled. The constructed pop-up window detection model and pop-up window closing model are trained with the labeled training data to obtain the pop-up window detection model and the pop-up window closing model.
[0220] Among them, the training data of the pop-up window detection model is the page with a pop-up window on the page after manual annotation, and its label is "there is a pop-up window". The training data of the pop-up window closing model is the page with a pop-up window on the page after manual annotation, and its label is "the pixel coordinates of the close button in the pop-up window".
[0221] In addition, in the present invention, during user operation, the user can choose to manually close the pop-up window existing in the page or manually confirm that there is no pop-up window on the current page, thereby ending the process. For the case where it is manually confirmed that there is no pop-up window on the current page, the obtained page screenshot can be directly used as the page screenshot of the page without a pop-up window. For the case where the user chooses to manually close the pop-up window existing in the page, return to step (1) again to obtain the page screenshot again, and use the page screenshot obtained again as the page screenshot of the page without a pop-up window.
[0222] When performing the judgment and processing of non-standard pages, the present invention takes into account the detection of various non-standard page states such as "page loading", "page jumping", "pop-up windows", etc., and can give corresponding processing methods for different non-standard page states. Therefore, the present invention is more comprehensive, flexible, and more adaptable to the needs of the Agent model.
[0223] V. Compress, simplify, and complete the obtained page XML to obtain the processed page XML.
[0224] As Figure 6 shown, in the present invention, the compression, simplification, and completion processing of the obtained page XML includes the following steps:
[0225] 1. Determine whether it is necessary to retain the off-screen nodes in the page XML. If not, go to step 2; if so, go to step 3.
[0226] Since the page XML is used to define the layout and elements of the page of an Android mobile phone and contains all the components in a page, there are a large number of nodes in the page XML that exist only for structural and layout needs. These nodes do not contain actual useful page information, which is also the main reason for the excessive length of the page XML. In addition, the number of nodes contained in a page is often more than the part displayed on the screen. For example, for some pages that can be scrolled, there will also be many off-screen nodes in the page XML.
[0227] Determining whether to retain the off-screen nodes is controlled by an input parameter remain_nodes (retain nodes when remain_nodes = True), and it can be selected according to one's own needs. In the present invention, if it is necessary to summarize the information of the entire page, in order to save the number of operation steps (such as scrolling the screen to view the complete text), the off-screen nodes can be retained, so that the text information of the entire page can be directly obtained without additional operations. If it is a requirement related to operation simulation, such as the requirement of simulating clicks or swipes, the off-screen nodes can be selected to be deleted to prevent interference with the Agent model.
[0228] 2. Delete the off-screen nodes in the page XML.
[0229] In the page XML, the bounds attributes of all on-screen nodes must be within the range of [0,0][Window_Height,Window_Width] and must be contained by their parent nodes. Therefore, it is only necessary to determine whether the bounds of the current node are contained by its parent node to determine all the nodes within the screen range. For the requirement of deleting off-screen nodes, after determining all the nodes within the screen range, delete other nodes not within the screen range.
[0230] 3. Delete redundant nodes in the page XML.
[0231] The page XML also contains a large number of nodes that only exist for structural and layout needs and do not contain useful page information. Therefore, the present invention will delete these redundant nodes. The present invention will determine whether a node is redundant based on the attribute information of the node. If at least one of the attributes such as "checkable", "checked", "clickable", "focusable", "scrollable", "long - clickable", "password", and "selected" of a node is True, or the text and content - desc attributes are not empty, then this node is considered a functional node, and those that do not meet this requirement are redundant nodes. The present invention will delete all nodes that do not meet this requirement.
[0232] 4. Determine whether the page of the Android phone is a special page. If it is, go to step 5; if not, go to step 6.
[0233] Due to different ideas of APP developers during development or on different models, for some pages, page XMLs with different structures or even information may be obtained, which can easily confuse the Agent model. At the same time, for components such as pop - up windows and dropdown boxes, there will be an overlapping phenomenon between different components. However, in the page XML, these components are displayed flat, which also means that it is impossible to tell which component is on top and which is at the bottom only through the page XML (in the actual operation process, the component at the bottom of the overlap will not be triggered by any operation), which can also easily confuse the Agent model. Therefore, the present invention designs a separate processing logic for special pages. The present invention will determine whether it is a special page based on the text content that exists fixedly on each page, and then execute the processing logic for special pages.
[0234] In the present invention, it will be determined whether it is a special page based on the page content contained in the page XML. Specifically, corresponding text information can be preset for all special pages in advance. For example, for the pages containing dropdown boxes in each APP, the preset text information can be selected according to the specific information contained in the dropdown boxes on the page. In this way, when the text or content - desc attributes of some nodes in the page XML contain the text information preset for the special page, the current page is considered a special page.
[0235] 5. Process the special page by deleting information or supplementing information;
[0236] Special pages can be divided into two categories in total: special pages for deleting information and special pages for supplementing information.
[0237] Special pages for deleting information mainly target the situation where multiple components overlap on some pages. Among them, the component on the upper layer is operable, but the component on the lower layer cannot be operated (covered by the upper component). At this time, the nodes of the lower component still exist in the page XML, which will cause confusion to the Agent model. For example, the pop-up window when selecting specifications in the food delivery APP, and various pages containing dropdown boxes. When the pop-up window and dropdown box appear, they will cover the components originally in this position, resulting in an overlapping phenomenon. When processing these special pages for deleting information, the present invention will first select a text node that remains fixed on the current page as a marker node (information that is not affected by time and network changes), and then find the top position of the node tree to be deleted based on the position of this node and delete it (XML is shown in a tree structure, so only the root node of the subtree where all the nodes to be deleted are located needs to be deleted).
[0238] Special pages for supplementing information mainly target some pages. Due to the implementation logic, some information (some attributes of a certain node are missing) or the nodes corresponding to some operable components do not exist in the page XML. For example, the like button for a WeChat Moments post does not have a specific node corresponding in the page XML and needs to be added through special processing. Similar to the processing method in special pages for deleting information, for special pages for supplementing information, the present invention will also select a marker node and find the node that needs to supplement information based on its position and supplement the information (supplement the missing attributes or add nodes under this node).
[0239] 6. Use the optical character recognition method to recognize all the text and the positions where the text is located in the page of an Android phone, compare the text recognized by the optical character recognition method with the text in the page XML, remove the text that repeats with the text in the page XML from the text recognized by the optical character recognition method, and insert the remaining text into the page XML.
[0240] In many pages, the page XML does not contain all the text information in the page, but often this text information is crucial for the Agent model to understand the page information. The present invention will introduce OCR (optical character recognition) technology to supplement the missing text information in the page XML. Specifically, first, it will use OCR to recognize all the text and the positions (bounds) where the text is located in the page of an Android phone. Subsequently, compare the text recognized by OCR with the text (text and content-desc) contained in the page XML, remove the repeated text and insert the remaining text into the page XML.
[0241] Specifically, we can use the MinHash-based algorithm to compare text similarity and try to determine whether the given OCR text contains the same or most of the information as the text in the existing page XML (sometimes due to the accuracy of OCR, the text recognized by OCR may not be completely consistent with the page XML). We use technologies such as MinHash and Levenshtein edit distance to process text matching and determine the matching coverage under a certain threshold. First, we create a MinHashLSH object for locality sensitive hash (LSH) query, and generate MinHash for all page XML descriptions and insert LSH, while removing punctuation and spaces and saving them in the dictionary. Then we generate MinHash for each OCR text and perform LSH query. If it reaches a certain threshold, it is considered a successful match (parameter of the MinHash algorithm). If the match fails, we further process it according to the length of the text: for texts with a length of 2 to 4 characters, we use Levenshtein edit distance to compare, and if the distance is less than or equal to 1, it is considered a match; for single-character text, we directly perform character matching. Finally, we get all non-duplicate texts. For these non-repeated texts, they will be inserted into the page XML as a node according to their corresponding position information (bounds). According to the aforementioned fact that the bounds of a parent node must contain child nodes, the present invention will select the smallest bounds of all OCR-recognized texts as its parent node for insertion.
[0242] The following is an example. Assume that all page XML descriptions correspond to a text list ["How!", "document."], and the text list recognized by OCR is ["HoW!", "丶document."]. First, remove all punctuation and special symbols to get the text list of the page XML description ["How", "document"] and the text list recognized by OCR ["HoW", "document"]. Then generate a MinHash value for each element in the text list of the page XML description and insert it into the LSH. After that, for each element in the text list recognized by OCR, also generate a MinHash value and match it with the element in the LSH (the MinHash corresponding to the XML list). Finally, document can be matched successfully, which means that "document." and "丶document." correspond to the same content, so this text does not need to be added to the page XML. However, HoW and How do not match, and the next step is to judge. Since the Levenshtein edit distance between How and HoW is 1, it means that the two match, and it is also considered that this text recognized by OCR is repeated with the text in the page XML.
[0243] 7. Simplify the attributes in the page XML after insertion.
[0244] The description of each attribute in the page XML is too redundant and will take up a lot of tokens. In the end, the present invention will simplify the description of these attributes. For the functional attributes of "checkable", "checked", "clickable", "focusable", "scrollable", "long-clickable", "password" and "selected", since most of them are False in most cases, only the attributes with the value of True are displayed. The three attributes of "index", "resource-id" and "package" are not helpful for the Agent model to understand the page and will be deleted directly. The "class" attribute will also represent the main function of the node to a certain extent, so the last part will be retained (the class must be composed of the format of xxxx, and the number of points is variable. The present invention only retains the content after the last point, for example, in android.widget.FrameLayout, only FrameLayout is retained). The two attributes "text" and "content-desc" represent the text information of the node, and the two are combined and displayed separately. The "bounds" attribute represents the position of the node in the page, which is one of the most critical attributes and is also displayed separately.
[0245] Finally, for the following nodes:
[0246] <node index="0"text="炷 steam"resource-id=""class="android.view.View"
[0247] package="com.autonavi.minimap" content-desc="checkable="false" checked="false"
[0248] clickable="false"enabled="true" focusable="false" focused="false" scrollable="false"
[0249] long-clickable="false" password="false" selected="false" bounds="[290,844][346,885]" / >
[0250] Will be simplified to:
[0251] [n42]View;;; Steaming; [290, 844][346, 885]
[0252] Through node deletion, the present invention deletes redundant nodes and nodes not within the screen, simplifies the attribute representation of nodes, and rewrites the page XML into a new format to obtain a more concise page XML. At the same time, through special processing of the page and OCR technology, information is deleted or supplemented for specific pages, thereby improving the processing ability of the Agent model.
[0253] Sixth, the processed page XML and page screenshots are input into the Agent model, and the Agent model outputs the operation type and location for the next step of the Android phone to execute.
[0254] In the present invention, the Agent model can be obtained by fine-tuning and training based on the GLM-4 model. Before use, some data with the processed page XML and page screenshots as input and the operation type and operation location as output can be collected as the training dataset, and the GLM-4 model can be fine-tuned and trained with this training dataset to obtain the Agent model. This part of the content is a technology that those skilled in the field of large language models can understand and master. For the sake of simplicity, it will not be described in detail here.
[0255] In the present invention, after fine-tuning and training, the input of the Agent model is mainly the processed page XML and page screenshots. Table 1 below shows the operation types that the Agent can output and their corresponding output contents. element is the bbox corresponding to the node Node in the page XML. In addition to some conventional operation types such as click, long press, slide, back, etc., the present invention can also handle some special types of operations. For example, call api, which is for some tasks that require text content understanding and generation (such as: Help me write an abstract for this WeChat official account article / Help me send a 300-word apology letter to xx). The Agent model will combine the task input by the user to generate an instruction, and call the api of GLM-4-public in combination with the page XML to complete the task and generate a message.
[0256] In addition, for some special operation types, it is necessary to process them in combination with the page screenshots. Among them, if the actions predicted by the Agent model are Note and interact operation types, it is necessary to further perform inference and prediction by using the fine-tuned GLM-4v vision model.
[0257] Among them, the Note operation exists in the form of annotations. For some tasks that require using the information viewed in the middle (for example: Compare which is closer to me, the Summer Palace or the Forbidden City? - This task requires remembering the distance between the Summer Palace and the current location, and then comparing it with the current distance of the Forbidden City to draw a conclusion), the Agent model will predict the steps that need to be memorized, add "#Note:True" before the output, and extract all the steps predicted by the Agent model that need to be noted in this trace and the screenshot of the current staying page before finish, splice them into a large picture, call the visual model to generate the final information finish message according to the information in it, and feedback it to the client.
[0258] Interact mainly generates interaction information based on the page screenshot. During the process of the task, multiple results that meet the user's needs may be found (for example: Help me buy a high-speed rail ticket to Shanghai tomorrow morning - For specific requirements such as which train to choose and whether it is a second-class seat or a first-class seat, the user needs to make further selections). In response to such situations, the Agent model will predict the steps that need to be interacted with, and call the fine-tuned GLM-4v model to generate interaction information (for example: Multiple results that meet the conditions, such as xx and xx, have been retrieved for you. Which one do you want to choose?), and continue to execute the task after the user supplements it.
[0259] Table 1 Operation types processed by the Agent model
[0260]
[0261]
[0262] VII. Based on the output of the Agent model, perform operations on the Android phone based on the AccessibilityService for accessibility services.
[0263] The output of the Agent model will be sent to the client of the Android phone, and the client will simulate operations such as clicking, swiping, long-pressing, and inputting based on the AccessibilityService of Android to execute.
[0264] VIII. Display the results of the operations performed on the Android phone.
[0265] On the client side, the information messages of operations such as call api and interact can be displayed.
[0266] Figure 7 Shows the schematic diagram of the composition of the Android phone intelligent control system based on the large model agent of the present invention. As Figure 7As shown in the figure, the Android mobile phone intelligent control system based on the large model intelligent agent of the present invention includes:
[0267] 1. Task input module
[0268] The task input module is used to obtain the tasks input by the user.
[0269] 2. Task discrimination model.
[0270] The task discrimination model is used to judge whether the tasks input by the user are within the processing scope of the Agent model to determine whether to continue to execute the tasks input by the user.
[0271] 3. Data acquisition module.
[0272] The data acquisition module is used to acquire the page information of the Android mobile phone, and the page information includes page XML and page screenshots.
[0273] 4. Page status judgment and processing module.
[0274] The page status judgment and processing module is used to judge and process the non-standard status of the page based on the page information to ensure that the obtained page information is of a standard page.
[0275] 5. Data processing module.
[0276] The data processing module is used to compress, simplify and complete the page XML obtained to obtain the processed page XML.
[0277] 6. Agent model.
[0278] The Agent model is used to output the operation type and position for the next execution of the Android mobile phone based on the processed page XML and page screenshots.
[0279] 7. Task execution module.
[0280] The task execution module is used to operate and execute the Android mobile phone based on the output of the Agent model and the AccessibilityService.
[0281] Of course, the Android mobile phone intelligent control system based on the large model intelligent agent of the present invention may further include:
[0282] 8. Display module.
[0283] The display module is used to display the results of the operation execution of the Android mobile phone.
[0284] In addition, the present invention further provides an Android mobile phone intelligent control device based on a large model agent, which includes: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the Android mobile phone intelligent control method based on a large model agent as described above.
[0285] Finally, the present invention provides a computer-readable storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the Android mobile phone intelligent control method based on a large model agent as described above are implemented.
[0286] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting the protection scope of the present invention. Those skilled in the art can modify or equivalently replace the technical solutions of the present invention according to the idea of the present invention, without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. An intelligent control method for Android mobile phones based on large model agents, characterized in that, Including the following steps: 1) Obtain the task input by the user; 2) Determine whether the task input by the user is within the processing scope of the agent model. If so, proceed to step 3); otherwise, end the execution of the task input by the user. Among them, collect the data with the input being the task input by the user and the output being the operation types that the agent model can process as the training dataset, and use this training dataset to fine-tune and train the model to obtain the task discrimination model, and use the task discrimination model to determine whether the task input by the user is within the processing scope of the agent model; 3) Obtain the page information of the Android phone, where the page information includes page XML and page screenshots. Among them, based on the accessibility service node information AccessibilityServiceNodeInfo, parallelly parse to obtain the child node ChildNode information, and at the same time obtain the information of multiple child nodes ChildNode, and then sequentially combine the obtained child node ChildNode information to obtain the complete XML content. Retain the information of the corresponding node Node when each node Node parsing ends, and then reverse-complement the node tree NodeTree to obtain the complete XML; perform quality compression on the page screenshot Screenshot without changing the original resolution of the page screenshot Screenshot to obtain the final page screenshot; 4) Based on the page information, perform judgment and processing on the non-standard state of the page to ensure that the obtained page information is of a standard page. The specific steps of step 4) include: 41) Page waiting processing, which is used to process the page that needs to wait before it can become a standard page, so as to ensure that the page information of a page that does not need to wait before becoming a standard page is obtained; 42) Pop-up window processing, which is used to process the situation where there are pop-up windows on the page, so as to ensure that the page information of a page without pop-up windows is obtained. The specific steps of step 42) include: 421) Obtain the page screenshot of the Android phone; 422) Detect whether there is a pop-up window in the page screenshot through the pop-up window detection model. If not, use the page screenshot as the page screenshot of the page without pop-up windows; if so, determine whether it has cycled once. If it has cycled once, hand it over to the user for operation. If it has not cycled once, hand it over to the pop-up window closing model, and the pop-up window closing model attempts to close the pop-up window. If the pop-up window closing model can return the closing position, automatically close the pop-up window according to the closing position and return to step 421). If the pop-up window closing model cannot return the closing position, hand it over to the user for operation; 5) Compress, simplify and page complete the obtained page XML to obtain the processed page XML; 6) Input the processed page XML and page screenshot into the agent model, and the agent model outputs the operation type and position for the next execution of the Android phone; 7) Based on the output of the agent model, perform operations on the Android phone based on the accessibility service AccessibilityService.
2. The intelligent control method for Android mobile phones based on large model agents according to claim 1, wherein The specific steps of step 3) include: 31) Obtain the accessibility service node information AccessibilityServiceNodeInfo and the page screenshot Screenshot of the Android phone's page based on the accessibility service AccessibilityService; 32) Parse the nodes Node in the accessibility service node information AccessibilityServiceNodeInfo, and convert the parsing result into a mutable character sequence StringBuilder; 33) Adjust the structure of the mutable character sequence StringBuilder, and perform parallel parsing on the child nodes ChildNod of the corresponding node Node; 34) Obtain the parsing results of each child node ChildNod and assemble them into the mutable character sequence StringBuilder, and adjust the structure of the assembled mutable character sequence StringBuilder again; 35) Obtain the final page XML based on the mutable character sequence StringBuilder after the second adjustment; 36) Perform quality compression on the page screenshot Screenshot without changing the original resolution of the page screenshot Screenshot to obtain the final page screenshot.
3. The intelligent control method for Android phones based on large model agents according to claim 2, wherein, The specific steps of step 41) include: 411) Obtain the page XML of the Android phone again; 412) Determine whether the difference in the total number of nodes in the page XML obtained twice in succession is less than D, and determine whether the position attributes of the first k nodes in the page XML obtained twice in succession are exactly the same. If the difference is less than D and the position attributes are exactly the same, then use the page XML obtained again as the page XML of the page that does not need to wait to become a standard page; if the difference is not less than D and / or the position attributes are not exactly the same, then determine whether the total elapsed time is greater than or equal to T seconds. If it is greater than or equal to T seconds, then use the page XML obtained again as the page XML of the page that does not need to wait to become a standard page. If it is less than T seconds, then return to step 411); 413) Obtain the page screenshot of the Android phone again; 414) Input the page screenshot obtained again into the single-image Wait model and input the page screenshots obtained twice in succession into the double-image Wait model to respectively determine whether the loading is completed. If the output results of both the single-image Wait model and the double-image Wait model are that the loading is completed, then use the page screenshot obtained again as the page screenshot of the page that does not need to wait to become a standard page; if the output result of the single-image Wait model and / or the double-image Wait model is that the loading is not completed, then determine whether the number of loops is greater than or equal to n times. If it is greater than or equal to n times, then use the page screenshot obtained again as the page screenshot of the page that does not need to wait to become a standard page. If it is less than n times, then return to step 413).
4. The Android phone intelligent control method based on the large model agent according to any one of claims 1-3, characterized in that, The specific steps of step 5) include: 51) Determine whether it is necessary to retain the off-screen nodes in the page XML. If not, proceed to step 52). If so, proceed to step 53). 52) Delete the off-screen nodes in the page XML. 53) Delete the redundant nodes in the page XML. 54) Determine whether the page of the Android phone is a special page. If so, proceed to step 55). If not, proceed to step 56). 55) Perform information deletion or information supplementation processing on the special page. 56) Use the optical character recognition method to identify all the text and the positions where the text is located in the page of the Android phone. Compare the text recognized by the optical character recognition method with the text in the page XML, remove the text that is repeated in the text recognized by the optical character recognition method and the text in the page XML, and insert the remaining text into the page XML. 57) Simplify the attributes in the inserted page XML.
5. The Android mobile phone intelligent control method based on the large model intelligent agent according to claim 4, wherein Further includes: 8) Display the results of the operations performed on the Android phone.
6. An Android mobile phone intelligent control system based on a large model intelligent agent, characterized in that, Includes: A task input module, which is used to obtain the task input by the user. A task discrimination model, which is used to determine whether the task input by the user is within the processing scope of the agent model to determine whether to continue to execute the task input by the user; among them, collect the data with the input being the task input by the user and the output being the operation type that the agent model can process as the training data set, and use this training data set to fine-tune the model to obtain the task discrimination model, and use the task discrimination model to determine whether the task input by the user is within the processing scope of the agent model. A data acquisition module, which is used to obtain the page information of the Android phone, and the page information includes page XML and page screenshots; among them, based on the AccessibilityServiceNodeInfo of the accessibility service node information, parallelly parse to obtain the ChildNode information of the sub-nodes, and at the same time obtain the information of multiple sub-nodes ChildNode, and then sequentially combine the obtained ChildNode information of the sub-nodes to obtain the complete XML content. Retain the information of the corresponding node Node when each node Node parsing ends, and then reverse-complement the node tree NodeTree to obtain the complete XML; perform quality compression on the page screenshot Screenshot without changing the original resolution of the page screenshot Screenshot to obtain the final page screenshot. A page status judgment and processing module, which is used to judge and process the non-standard status of the page based on the page information to ensure that the obtained page information is for a standard page. Specifically, it includes: page waiting processing, which is used to process pages that need to wait before becoming standard pages to ensure that the obtained page information is for a page that does not need to wait before becoming a standard page; pop-up window processing, which is used to process the situation where there are pop-up windows on the page to ensure that the obtained page information is for a page without pop-up windows. The pop-up window processing specifically includes: 1) obtaining a page screenshot of the Android phone; 2) detecting whether there is a pop-up window in the page screenshot through a pop-up window detection model. If not, using the page screenshot as the page screenshot of a page without a pop-up window; if so, judging whether it has cycled once. If it has cycled once, handing it over to the user for operation. If it has not cycled once, handing it over to a pop-up window closing model, and the pop-up window closing model attempts to close the pop-up window. If the pop-up window closing model can return the closing position, automatically close the pop-up window according to the closing position and return to step 1). If the pop-up window closing model cannot return the closing position, handing it over to the user for operation. A data processing module, which is used to compress, simplify, and complete the page XML obtained to obtain the processed page XML. An agent model, which is used to output the operation type and position for the next step of the Android phone based on the processed page XML and the page screenshot. A task execution module, which is used to perform operations on the Android phone based on the output of the agent model and the AccessibilityService.
7. An intelligent control device for Android phones based on large model agents, characterized in that, including: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the Android phone intelligent control method based on the large model agent as described in any one of claims 1-5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the Android phone intelligent control method based on the large model agent as described in any one of claims 1-5.
Citation Information
Patent Citations
An automated, non-invasive, barrier-free support detection method for Android applications
CN109359029A
Interaction method and device based on multi-agent large model, equipment and medium
CN118363695A