Large and small model end cloud collaboration method oriented to agent interface information leakage reduction

Through the local-cloud collaborative decision-making mechanism, semantic blocks are divided based on the user interface XML tree structure, and the collaborative decision-making of local and cloud LLM is combined to solve the problem of user interface information leakage in mobile task automation and achieve efficient and secure task execution.

CN120743271AActive Publication Date: 2025-10-03SHANGHAI JIAOTONG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510869762.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-10-03
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

In the existing technology of mobile task automation, user interface information leakage is serious, resulting in unnecessary exposure of user privacy information, and it is difficult to balance task execution efficiency and success rate.

Method used

A local-cloud collaborative method is adopted to divide semantically related structured blocks through the interface structure perception preprocessing stage. Combined with the collaborative decision-making mechanism of local LLM and cloud LLM, only key area information is uploaded for task planning and decision-making, ensuring task success rate and privacy protection.

Benefits of technology

The user interface information exposure was significantly reduced by 55.60%, the task success rate was increased by 36.36%, and the cloud inference time was reduced by 19.16%, achieving a balance between task execution efficiency and privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743271A_ABST
    Figure CN120743271A_ABST
Patent Text Reader

Abstract

A large and small model end cloud collaboration method oriented to agent interface information leakage reduction comprises the steps that in the preprocessing stage of interface structure perception, a user interface is divided into semantic-related structured blocks by analyzing an XML hierarchical structure; in a local-cloud collaborative planning stage, a local LLM proposes a plurality of sub-task candidates according to divided blocks, and a cloud LLM selects or generates a more accurate current sub-task based on global understanding of all candidate items; in the local-cloud collaborative decision-making stage, the local LLM performs preliminary screening and sorting on the interface blocks according to the subtasks, and the cloud LLM further performs fine-grained decision-making in the high-correlation blocks; according to the method, the local-cloud collaborative mobile terminal task automatic processing combining the strong reasoning capability of the cloud LLM and the privacy protection advantage of the local LLM enables the intelligent agent to only interact with the key area on the user interface information, and unnecessary user interface information exposure is greatly reduced while the task execution efficiency and the success rate are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of information security, specifically a large-scale model-end-cloud collaboration method for reducing information leakage in intelligent body interfaces. Background Art

[0002] With the popularity of smartphones, users have an increasing demand for automatically completing operational tasks on mobile devices. Task automation technology aims to enable intelligent agents to autonomously complete a series of operational tasks based on the user's natural language instructions, thereby improving the user experience and reducing the user's operational burden. In recent years, with the development of large language model (LLM) technology, there are many existing LLM-based mobile operation agents that can simulate human operation processes, gradually perceive the interface status, plan subtasks, and interact with the graphical user interface (GUI) until the task is completed. In practical applications, it is usually necessary to upload complete interface information (such as XML structure or screenshots) to the cloud model for decision-making in each step of the operation. This process leads to unnecessary leakage of user privacy information. Summary of the Invention

[0003] To address the aforementioned shortcomings of the existing technology, the present invention proposes a large-scale, small-scale, and cloud-based collaborative method for reducing information leakage in intelligent agent interfaces. This method combines the strong reasoning capabilities of cloud-side LLMs with the privacy protection advantages of local LLMs for local-cloud collaborative mobile task automation. This allows intelligent agents to interact only with key areas of user interface information, significantly reducing unnecessary exposure to user interface information while ensuring task execution efficiency and success rate. Experiments have shown that the present invention reduces user interface information exposure by up to 55.60%, approaches task success rates (within 5%), reduces cloud-side reasoning time by 19.16%, and increases task success rates by up to 36.36%.

[0004] The present invention is achieved through the following technical solutions:

[0005] The present invention relates to a large-scale model-end-cloud collaboration method for reducing information leakage in an intelligent agent interface, comprising:

[0006] 1) In the preprocessing stage of interface structure perception, the user interface is divided into semantically related structured blocks by analyzing the XML hierarchical structure;

[0007] 2) In the local-cloud collaborative planning phase, the local LLM proposes multiple subtask candidates based on the divided blocks, and the cloud LLM selects or generates a more accurate current subtask based on a global understanding of all candidates;

[0008] 3) In the local-cloud collaborative decision-making stage, the local LLM preliminarily screens and sorts the interface blocks according to the subtasks, and the cloud LLM further makes fine-grained decisions within the highly relevant blocks.

[0009] When the cloud-based LLM determines that the current information is insufficient to support reliable decision-making, it requests more context from the local LLM, achieving a trade-off between ensuring mission success and reducing information exposure.

[0010] The present invention relates to a system for realizing the above-mentioned collaborative method, comprising: a user interface segmentation unit, a local subtask generation unit, a cloud subtask screening and correction unit, a local user interface block sorting and screening unit, a multi-round information accumulation control unit, a cloud decision unit and an execution unit, wherein: the user interface segmentation unit obtains the XML tree structure of the current interface, extracts important interactive elements with semantic information therefrom, and groups these elements by analyzing their common ancestor nodes in the tree structure to obtain at least three semantically coherent UI blocks; the local subtask generation unit uses the local LLM to independently generate candidate subtasks that may be executed in the current block and do not involve specific interface elements for each UI block, in combination with the user task description and historical operations; the cloud subtask screening and correction unit receives all candidate subtasks generated by the local LLM through the cloud LLM, and in combination with the global task description and historical operations The cloud-based LLM determines the most suitable subtask for the current task; the local user interface block sorting and screening unit uses the local LLM to score and sort the importance of all UI blocks according to their relevance to the subtask based on the most suitable subtask for the current task; the multi-round information accumulation control unit uses the cloud-based LLM to determine whether the UI block information uploaded in the first round and ranked first in importance is sufficient to complete the current subtask; if the information is insufficient, the next round of upload mechanism is triggered, and the suboptimal block is selected from the remaining blocks to continue uploading until the cloud-based LLM determines that there is enough information to support reliable decision-making; the cloud-based decision-making unit performs fine-grained analysis and decision-making on the received UI block set through the cloud-based LLM to determine the specific target UI elements and their interactive actions; the execution unit receives the final decision and automatically executes it on the user's mobile device to complete the current step; after execution, the system records the operation and prepares to enter the next round of interaction until all subtasks are completed.

[0011] Technical Effects

[0012] The present invention summarizes block information into subtask descriptions that can be used to guide operations without exposing specific content to the cloud; a local-cloud collaborative decision-making mechanism: the local LLM condenses UI block information into subtasks and uploads them. After the cloud LLM confirms the optimal subtask, the local LLM ranks each block according to its importance to completing the subtask, allowing the cloud LLM to perform fine-grained decisions only in the most relevant blocks. Compared with the existing technology, the XML-based interface partitioning method of the present invention helps LLM to deeply understand the partition structure of the UI interface, abandon traditional visual perception, and realize block partitioning according to the tree-like hierarchical logic of the UI design itself, laying the foundation for LLM to accurately perceive key blocks and eliminate irrelevant information in the future; the UI block semantic condensation method based on local LLM effectively protects the privacy of the specific content of the user interface and avoids direct uploading of sensitive information, but transmits key information to the cloud-side LLM through the condensed sub-task description, allowing it to have a better perception of the current page status and thus make better planning, making up for the lack of local model planning capabilities; the local-cloud collaborative task execution mechanism takes into account the privacy protection capabilities of the local LLM and the reasoning advantages of the cloud-side LLM. Through block screening, only the interface information most relevant to the current task is uploaded, enabling the cloud-side LLM to focus on key areas and carry out high-quality reasoning. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 Flowchart of the present invention;

[0014] Figure 2 This is a specific example diagram in an actual application scenario. DETAILED DESCRIPTION

[0015] like Figure 1 As shown, this embodiment relates to a large-scale model-end-cloud collaboration method for reducing information leakage in an intelligent agent interface, including:

[0016] Step 1: Preprocessing of interface structure perception: By analyzing the XML hierarchical structure, the user interface is divided into semantically related structured blocks, including:

[0017] 1.1. Use the uiautomator2 tool to extract the XML file containing the attributes and hierarchical structure of each UI element in the current UI state from the mobile device;

[0018] 1.2. Parse the XML file to obtain the corresponding XML tree structure and extract the key nodes that meet the following conditions through depth-first traversal: a. Clickable and containing description information; b. Editable; c. Containing text, description, and hint semantic information fields;

[0019] 1.3. During the traversal process of step 1.2, assign unique numbers to all nodes and record the ancestral paths of important nodes. The path format is ,in: is the root node number, For this important node The direct parent node number of

[0020] 1.4. Set of ancestor paths for all important nodes , from depth To maximum path depth Perform traversal;

[0021] 1.5. At every depth , according to the first Layer ancestor number Divide and divide The important nodes of the same depth should be grouped into the same block, while the important nodes with different ancestors at the same depth should be grouped into different blocks.

[0022] 1.6. If the number of blocks divided at a certain depth is not less than three, the division result at that depth is used as the final block division result;

[0023] 1.7. Arrange the important nodes in each block in depth-first order in the XML tree, and retain their attributes in the original XML tree structure, specifically: text content, description, class name, resource identifier, and bounding box position coordinate information;

[0024] 1.8. Number all the divided user interface blocks in sequence to get , each Represents a block, which consists of corresponding important nodes Composition, that is .

[0025] Step 2: Local-cloud collaborative planning phase, the local LLM will plan the plan based on each block. Multiple subtask candidates are proposed. Cloud-based LLM selects or generates the most accurate current subtask based on a global understanding of all candidates, including:

[0026] 2.1. For each UI block obtained by the block division method , which is compared to the task description provided by the user , Completed historical operation sequence Combined as input Send to local LLM;

[0027] 2.2. The local LLM determines the task progress and outputs a logical subtask that may be executed within the current block , specifically, the following conditions must be met at the same time: a. Based only on the current block information b. It does not rely on information from other blocks; c. It does not leak any specific UI element information; and c. It expresses logical intent in natural language, such as "open settings" and "enter query content."

[0028] 2.3. Repeat steps 2.1-2.2 to obtain the candidate subtask set corresponding to all blocks ;

[0029] 2.4. Collect candidate subtasks Together with the user task description and historical operation sequences Combined as input Send to the cloud LLM for planning, including:

[0030] i) The cloud-based LLM infers the functional distribution of the overall user interface based on the candidate subtasks on each block, ensuring that task planning has strong contextual consistency and practical executability when specific UI information is unavailable.

[0031] ii) Cloud-based LLM uses its strong planning capabilities to comprehensively analyze the candidate subtask set , select the most reasonable subtask from it As the subtask that should be executed currently;

[0032] iii) If the set If the candidate subtasks in the task list are all unreasonable, the cloud-based LLM can generate a new subtask based on its own reasoning and analysis. .

[0033] Step 3: Local-cloud collaborative decision-making phase: The local LLM performs preliminary screening and sorting of interface blocks based on subtasks, and the cloud LLM further makes fine-grained decisions within highly relevant blocks, specifically including:

[0034] 3.1. Local LLM is based on the user interface block set obtained by previous division and the current subtasks derived from collaborative planning , for each block Perform importance assessment and generate corresponding importance scores , specifically: , that is, the sum of the scores generated by the large model is 1.

[0035] 3.2. According to the scoring results, the importance of all blocks is ranked, and the block sequence after ranking is ,in: The block with the highest score is most likely to contain the key information required to solve the subtask;

[0036] 3.3. The local LLM will block the highest score Upload to the cloud LLM and combine it with the complete task description , historical interaction records As well as the current block content, it first evaluates whether there is enough information to make this decision. If the cloud LLM determines that the current block information is sufficient, it makes a decision and jumps directly to step 3.6.

[0037] 3.4. If the cloud LLM determines that the current block information is insufficient to support a reliable decision, it will request the next block with the second highest score but not yet uploaded locally. , continue uploading to the cloud and merging with the received block information;

[0038] 3.5. Each time a new block is uploaded, the cloud LLM re-makes a decision based on the currently accumulated information of multiple blocks. This process uses a progressive information accumulation mechanism until the cloud LLM is confident that the current information is sufficient to make an accurate and executable decision, or until all blocks have been uploaded. This mechanism effectively achieves a dynamic trade-off between reducing the amount of uploaded information and ensuring decision accuracy. Even if the local LLM's initial sorting and screening errors or the task itself is highly complex, the cloud LLM can still obtain sufficient and necessary user interface information through multiple rounds of iterations to make the right decision.

[0039] 3.6. The final decision result of Cloud LLM is ,in: The type of action to interact with the user interface (click, long press, or text input). Number the specific UI element that needs to be interacted with. For input content (when Valid when it is text input, otherwise );

[0040] If the cloud-based LLM still cannot make a valid decision after all blocks have been uploaded, it indicates that it believes the information contained in the current user interface is insufficient to complete the task. At this point, the intelligent agent will automatically perform a downward swipe operation to obtain the new user interface state and upload this state for the cloud-based LLM to continue its decision-making attempt.

[0041] The sliding, uploading, and decision-making process can be repeated up to five times to avoid infinite sliding. If no valid decision result is obtained after the sliding operation reaches the upper limit, the cloud LLM is determined to be unable to complete the current task, and the system terminates the task process.

[0042] like Figure 1 As shown, the information security protection system for implementing the above method in this embodiment includes: a user interface segmentation unit, a local subtask generation unit, a cloud subtask screening and correction unit, a local user interface block sorting and screening unit, a multi-round information accumulation control unit, a cloud decision unit and an execution unit.

[0043] The user interface segmentation unit includes: a structure extraction module and a block division module, wherein: the structure extraction module parses the hierarchical structure and attributes of the user interface elements based on the XML data of the current page to obtain a processed UI tree structure; the block division module traverses the UI tree obtained by the structure extraction module, identifies important interactive nodes, and groups them based on their ancestor paths in the tree, outputting at least three semantically coherent UI blocks.

[0044] The local subtask generation unit includes: a context fusion module and a semantic condensation module, wherein: the context fusion module takes the user's current task description and historical operation sequence as input, and combines the content of each block as the input context of the local model; the semantic condensation module generates executable subtask descriptions in the corresponding block through the local LLM, without involving specific UI element information, to ensure that the user interface privacy is not uploaded.

[0045] The cloud-based subtask screening and correction unit includes: a context fusion module and a subtask screening and correction module, wherein the context fusion module is responsible for receiving all subtask candidate descriptions generated by the local LLM, and combining the complete task description and historical operation records as the input context of the cloud-based LLM; the subtask screening and correction module evaluates the rationality and feasibility of each subtask candidate based on the input context, and selects the most suitable one as the final subtask. If no candidate meets the task requirements, the cloud-based LLM will correct or regenerate the subtask to ensure that the subtask has clear operational guidance and meets the current task advancement goals.

[0046] The local user interface block sorting and screening unit includes: a task association analysis module and a block sorting module, wherein: the task association analysis module analyzes the importance of each block content to the completion of the subtask based on the current subtask confirmed by the cloud LLM, and the local LLM analyzes the importance of the content of each block to the completion of the subtask, and scores them in turn; the block sorting module sorts all UI blocks according to the scores of the local LLM, and determines the importance ranking and upload priority of each block.

[0047] The multi-round information accumulation control unit includes: an information sufficiency judgment module and a block progressive upload module, wherein: the information sufficiency judgment module uses the cloud LLM to analyze the currently uploaded high-priority UI block ranked first to determine whether the content of the block is sufficient to support the cloud LLM to complete the subtask decision; if the information is insufficient, the block progressive upload module selects the suboptimal blocks from the remaining unuploaded blocks in the locally generated priority order and uploads them in sequence until the information sufficiency judgment module believes that the current content has met the cloud decision-making needs.

[0048] The cloud-based decision-making unit includes: an information aggregation module and an action decision module, wherein: the information aggregation module integrates the UI block content uploaded to the cloud and the current subtask to construct contextual information for decision-making; the action decision module completes fine-grained interaction decisions through the cloud-based LLM, that is, it clarifies the target UI elements and the type of interaction to be performed.

[0049] The execution unit includes: an operation execution module and a status update module, wherein: the operation execution module receives the cloud decision output and calls tools such as ADB to control the mobile device to perform corresponding operations (click, input, long press); the status update module records the current operation to the interaction history and automatically enters the next round of task processing until the task is completed.

[0050] This embodiment is specifically experimentally verified under the following environmental configuration: the local end uses a consumer-grade personal computer equipped with an NVIDIA GeForce RTX 4090D graphics card, deploys a quantized local large language model (Gemma2-9B-Instruct, Qwen 2.5-7B-Instruct, LLaMA 3.1-8B-Instruct), and runs it through the Ollama tool. The cloud model uses GPT-4o (version 2024-11-20) and is called through the OpenAI official API interface. The test terminals include a real mobile phone running Android 9 (Honor Play3, suitable for DroidTask dataset) and a Pixel 7 Pro emulator running Android 13 (suitable for AndroidLab dataset). The user interface state is extracted in the form of an XML file and obtained using the uiautomator2 tool. The interactive operation is implemented through Android Debug Bridge (ADB).

[0051] Under the above settings, we conducted comparative experiments using different methods, and calculated the task success rate and UI information upload reduction rate. The experimental data is shown in Table 1.

[0052] Table 1

[0053] The complete workflow of the mobile intelligent agent system of this embodiment is as follows:

[0054] Step 1. First, connect your smartphone to a local personal computer with basic computing power. Deploy a local LLM on the computer, such as running the Gemma2-9B model using the Ollama tool. If graphics memory resources are insufficient, use a quantized version to reduce resource requirements. For closed-source cloud LLMs deployed on a cloud server, call the corresponding paid API, such as the GPT-4o API provided by OpenAI.

[0055] Step 2. The user enters a specific task described in natural language on the computer and specifies the target application. The mobile intelligent agent system uses the Android Debug Bridge (ADB) tool to start the target application and execute the command adbshell am start <mainactivity>Enter the main interface of the application;

[0056] Step 3. After the application is launched, the mobile intelligent agent extracts the XML file of the current user interface and divides the UI into blocks as input for subsequent collaborative processing. The specific steps have been detailed above;

[0057] Step 4. Conduct local-cloud collaborative planning and decision-making. The local LLM is responsible for performing block sorting and preliminary screening, uploading the highest-scoring blocks to the cloud LLM, which makes fine-grained decisions based on the task description, interaction history, and block content. The specific steps have been detailed above.

[0058] Step 5. After receiving the decision from the cloud LLM, the mobile intelligent agent uses the ADB tool to perform the corresponding interactive operation: a. If the decision is a click operation, it is executed by clicking the center coordinates of the target element's bounding box: adb shell input tap <x> <y>b. If it is a long press operation, extend the time of clicking the center coordinates of the target element's bounding box: adb shell inputtouchscreen swipe <x> <y> <x> <y> <duration>c. If you want to input text, you need to clear the input box first before inputting. You can use the following command combination to complete it: adb shell am broadcast -a ADB_CLEAR_TEXT, adb shell am broadcast -a ADB_INPUT_TEXT –es msg<input_text> This method requires ADBKeyBoard to be pre-installed on the mobile smartphone, which supports Chinese input.

[0059] Step 6. After each round of interaction, the mobile intelligent agent records the interaction and updates the interaction history. The user interface state is also updated due to the interaction. Steps 3, 4, and 5 are repeated in the next round. This process terminates after a round of local-cloud LLM collaborative planning determines that there are no remaining subtasks.

[0060] Compared to existing technologies, this invention utilizes a UI block partitioning method based on the user interface XML tree structure. During the preprocessing phase, UI blocks with high semantic cohesion are partitioned according to the UI structure. This allows subsequent processing to focus on block-level semantic information, avoiding complex element-level filtering. This significantly reduces the amount of interface information uploaded while ensuring task completion rates. Field measurements show that compared to pure cloud-based LLM methods, this invention can reduce interface content uploads by 55.60%, significantly enhancing user data privacy protection. Secondly, this invention uses a local large language model to condense the semantics of UI blocks locally, summarizing them into subtask descriptions with operational guidance. These descriptions are then sent to the cloud instead of specific UI content, ensuring privacy while enabling information transfer. This mechanism enables the cloud-based LLM to determine the task progress reflected in the current UI and identify appropriate subtasks without directly accessing the original UI content. Thirdly, this invention leverages a local-cloud collaborative decision-making mechanism. The local model ranks and filters blocks based on subtasks confirmed by the cloud, uploading only the most relevant blocks to the cloud. This reduces redundant information interference and significantly improves the inference efficiency of the cloud-based model. Actual measured data shows that cloud-based decision-making time is reduced by 19.16%.

[0061] The above-mentioned specific implementation can be partially adjusted in different ways by those skilled in the art without departing from the principles and purpose of the present invention. The scope of protection of the present invention shall be based on the claims and shall not be limited by the above-mentioned specific implementation. All implementation schemes within its scope shall be subject to the constraints of the present invention.< / duration> < / y> < / x> < / y> < / x> < / y> < / x> < / mainactivity>

Claims

1. A large-scale model-end-cloud collaboration method for reducing information leakage in intelligent agent interfaces, characterized in that: In the preprocessing phase of interface structure perception, the user interface is divided into semantically related structured blocks by analyzing the XML hierarchical structure. In the local-cloud collaborative planning phase, the local LLM proposes multiple subtask candidates based on the divided blocks, and the cloud LLM selects or generates the most accurate current subtask based on a global understanding of all candidates. In the local-cloud collaborative decision-making phase, the local LLM preliminarily screens and sorts the interface blocks according to the subtasks, and the cloud LLM further makes fine-grained decisions within highly relevant blocks. When the cloud-based LLM determines that the current information is insufficient to support reliable decision-making, it requests more context from the local LLM, achieving a trade-off between ensuring mission success and reducing information exposure.

2. The large-scale model-end-cloud collaboration method for reducing information leakage in intelligent agent interfaces according to claim 1 is characterized in that: The pre-processing stage includes: 1.

1. Use the uiautomator2 tool to extract the XML file containing the attributes and hierarchical structure of each UI element in the current UI state from the mobile device; 1.

2. Parse the XML file to obtain the corresponding XML tree structure and extract the key nodes that meet the following conditions through depth-first traversal: a. Clickable and containing description information; b. Editable; c. Containing text, description, and hint semantic information fields; 1.

3. During the traversal process of step 1.2, assign unique numbers to all nodes and record the ancestral paths of important nodes. The path format is ,in: is the root node number, For this important node The direct parent node number of 1.

4. Set of ancestor paths for all important nodes , from depth To maximum path depth Perform traversal; 1.

5. At every depth , according to the first Layer ancestor number Divide and divide The important nodes of the same depth should be divided into the same block, and the important nodes with different ancestors at the same depth should be divided into different blocks; 1.

6. If the number of blocks divided at a certain depth is not less than three, the division result at that depth is used as the final block division result; 1.

7. Arrange the important nodes in each block in depth-first order in the XML tree, and retain their attributes in the original XML tree structure, specifically: text content, description, class name, resource identifier, and bounding box position coordinate information; 1.

8. Number all the divided user interface blocks in sequence to get , each Represents a block, which consists of corresponding important nodes Composition, that is .

3. The large-scale model-end-cloud collaboration method for reducing information leakage in intelligent agent interfaces according to claim 1 is characterized in that: The local-cloud collaborative planning phase specifically includes: 2.

1. For each UI block obtained by the block division method , which is compared to the task description provided by the user , Completed historical operation sequence Combined as input Send to local LLM; 2.

2. The local LLM determines the task progress and outputs a logical subtask that may be executed within the current block , specifically, the following conditions must be met at the same time: a. Based only on the current block information It is concluded that it does not rely on information from other blocks. b. It does not leak any specific UI element information. c. It expresses logical intent in natural language. 2.

3. Repeat steps 2.1-2.2 to obtain the candidate subtask set corresponding to all blocks ; 2.

4. Collect candidate subtasks Together with the user task description and historical operation sequences Combined as input Send to the cloud LLM for planning, including: i) The cloud-based LLM infers the functional distribution of the overall user interface based on the candidate subtasks on each block, ensuring that task planning has strong contextual consistency and practical executability when specific UI information is not available; ii) Cloud-based LLM uses its strong planning capabilities to comprehensively analyze the candidate subtask set , select the most reasonable subtask from it As the subtask that should be executed currently; iii) If the set If the candidate subtasks in the task list are all unreasonable, the cloud-based LLM can generate a new subtask based on its own reasoning and analysis. .

4. The large-scale model-end-cloud collaboration method for reducing information leakage in intelligent agent interfaces according to claim 1 is characterized in that: The local-cloud collaborative decision-making stage specifically includes: 3.

1. Local LLM is based on the user interface block set obtained by previous division and the current subtasks derived from collaborative planning , for each block Perform importance assessment and generate corresponding importance scores , specifically: , that is, the sum of the scores generated by the large model is 1; 3.

2. According to the scoring results, the importance of all blocks is ranked, and the block sequence after ranking is ,in: The block with the highest score is most likely to contain the key information required to solve the subtask; 3.

3. The local LLM will block the highest score Upload to the cloud LLM and combine it with the complete task description , historical interaction records As well as the current block content, first evaluate whether there is enough information to complete this decision. If the cloud LLM determines that the current block information is sufficient, it makes a decision and jumps directly to step 3.6; 3.

4. If the cloud LLM determines that the current block information is insufficient to support a reliable decision, it will request the next block with the second highest score but not yet uploaded locally. , continue uploading to the cloud and merging with the received block information; 3.

5. After each new block is uploaded, the cloud-based LLM re-makes a decision based on the currently accumulated information from multiple blocks. This process uses a progressive information accumulation mechanism until the cloud-based LLM is confident that the current information is sufficient to make an accurate and executable decision, or until all blocks have been uploaded. This mechanism effectively achieves a dynamic trade-off between reducing the amount of uploaded information and ensuring decision accuracy. Even if the local LLM's initial sorting and screening errors or the task itself is complex, the cloud-based LLM can still obtain sufficient and necessary user interface information through multiple rounds of iterations to make the right decision. 3.

6. The final decision result of Cloud LLM is ,in: The type of action to interact with the user interface (click, long press, or text input). Number the specific UI element that needs to be interacted with. For input content (when Valid when it is text input, otherwise ).

5. A system for implementing the collaborative method according to any one of claims 1 to 4, characterized in that: include: The user interface segmentation unit, the local subtask generation unit, the cloud subtask screening and correction unit, the local user interface block sorting and screening unit, the multi-round information accumulation control unit, the cloud decision unit and the execution unit, wherein: the user interface segmentation unit obtains the XML tree structure of the current interface, extracts important interactive elements with semantic information from it, and groups these elements by analyzing their common ancestor nodes in the tree structure to obtain at least three semantically coherent UI blocks; the local subtask generation unit uses the local LLM to independently generate candidate subtasks that may be executed in the current block and do not involve specific interface elements for each UI block, combining the user task description and historical operations; the cloud subtask screening and correction unit receives all candidate subtasks generated by the local LLM through the cloud LLM, and combines the global task description and historical operations to determine the most suitable subtask at the moment. Task; the local user interface block sorting and screening unit uses the local LLM to score and sort the importance of all UI blocks according to their relevance to the subtask based on the current most suitable subtask; the multi-round information accumulation control unit uses the cloud LLM to determine whether the UI block information uploaded in the first round and ranked first in importance is sufficient to complete the current subtask; if the information is insufficient, the next round of upload mechanism is triggered, and the suboptimal block is selected from the remaining blocks to continue uploading until the cloud LLM determines that there is enough information to support reliable decision-making; the cloud decision-making unit performs fine-grained analysis and decision-making on the received UI block set through the cloud LLM to determine the specific target UI elements and their interactive actions; the execution unit receives the final decision and automatically executes it on the user's mobile device to complete the current step; after execution, the system records the operation and prepares to enter the next round of interaction until all subtasks are completed.

6. The system according to claim 5, characterized in that: The user interface segmentation unit includes: a structure extraction module and a block division module, wherein: the structure extraction module parses the hierarchical structure and attributes of the user interface elements based on the XML data of the current page to obtain a processed UI tree structure; the block division module traverses the UI tree obtained by the structure extraction module, identifies important interactive nodes, and groups them based on their ancestor paths in the tree, outputting at least three semantically coherent UI blocks.

7. The system according to claim 5, wherein: The local subtask generation unit includes: a context fusion module and a semantic condensation module, wherein: the context fusion module takes the user's current task description and historical operation sequence as input, and combines the content of each block as the input context of the local model; the semantic condensation module generates executable subtask descriptions within the corresponding block through the local LLM, without involving specific UI element information, to ensure that user interface privacy is not uploaded; The cloud-based subtask screening and correction unit includes: a context fusion module and a subtask screening and correction module, wherein the context fusion module is responsible for receiving all subtask candidate descriptions generated by the local LLM, and combining the complete task description and historical operation records as the input context of the cloud-based LLM; the subtask screening and correction module evaluates the rationality and feasibility of each subtask candidate based on the input context, and selects the most suitable one as the final subtask. If no candidate meets the task requirements, the cloud-based LLM will correct or regenerate the subtask to ensure that the subtask has clear operational guidance and meets the current task advancement goals.

8. The system according to claim 5, wherein: The local user interface block sorting and screening unit includes: a task association analysis module and a block sorting module, wherein: the task association analysis module analyzes the importance of each block content to the completion of the subtask based on the current subtask confirmed by the cloud LLM and the local LLM, and scores them in turn; the block sorting module sorts all UI blocks according to the scores of the local LLM to determine the importance ranking and upload priority of each block; The multi-round information accumulation control unit includes: an information sufficiency judgment module and a block progressive upload module, wherein: the information sufficiency judgment module uses the cloud LLM to analyze the currently uploaded high-priority UI block ranked first to determine whether the content of the block is sufficient to support the cloud LLM to complete the subtask decision; if the information is insufficient, the block progressive upload module selects the suboptimal blocks from the remaining unuploaded blocks in the locally generated priority order and uploads them in sequence until the information sufficiency judgment module believes that the current content has met the cloud decision-making needs.

9. The system according to claim 5, wherein: The cloud-based decision-making unit includes: an information aggregation module and an action decision module, wherein: the information aggregation module integrates the UI block content uploaded to the cloud and the current subtask to construct contextual information for decision-making; the action decision module completes fine-grained interaction decisions through the cloud-based LLM, that is, it clarifies the target UI elements and the type of interaction to be performed.

10. The system according to claim 5, wherein: The execution unit includes: an operation execution module and a status update module, wherein: the operation execution module receives the cloud decision output and calls tools such as ADB to control the mobile device to perform corresponding operations (click, input, long press); the status update module records the current operation to the interaction history and automatically enters the next round of task processing until the task is completed.

Citation Information

Patent Citations

  • Cloud-edge collaborative big language model intelligent customer service deployment optimization method

    CN117808481A

  • Optimization method for lower end side LLM of end cloud LLM hybrid service framework

    CN118747166A

  • Transform large model training method based on cloud edge collaboration

    CN119294444A

  • Cloud-edge collaborative inference method and inference system

    CN119783823A

  • System and method using intelligent privacy assistant model for large language model operation

    US20250131122A1