Mobile terminal agent information processing system based on prior rule and multi-modal large model
By combining prior rules and multimodal large models in the mobile agent information processing system, the precise positioning and functional recognition of icons are achieved, and the task execution path is optimized, which solves the accuracy and efficiency of the multimodal large model when processing the graphical user interface, and improves the system's response speed and execution efficiency.
Patent Information
- Application Number
- CN202510266979.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-06
AI Technical Summary
When processing graphical user interfaces, the current multimodal model cannot accurately generate accurate coordinates and detailed functional descriptions for each icon, button or other interactive element, resulting in insufficient understanding and operation accuracy of the system's interface and inability to effectively perform complex and high-frequency automation tasks.
Using a mobile agent information processing system based on prior rules and multimodal large models, the positioning and identification module, the execution efficiency improvement module, the operation optimization module for specific applications, and the navigation and task data collection and update module are adopted to realize the precise positioning and function recognition of icons, optimize the task execution path, and improve the system's response speed and execution efficiency.
It realizes accurate positioning and functional recognition of icons, improves the stability and accuracy of task execution, reduces computing resource consumption and response delay, is suitable for complex and high-frequency interaction scenarios, and improves the operating efficiency and accuracy of the system in specific applications.
Smart Images

Figure CN120161973A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a mobile intelligent agent information processing system based on prior rules and multimodal large models. Background Art
[0002] Multimodal large models are a rapidly developing direction in the field of artificial intelligence, and have made remarkable progress particularly in cross-modal understanding and generation in fields such as vision, language, and audio. For example, Tongyi Qianwen - a large visual language model, Shusheng Wanxiang multimodal large model, GPT-4 multimodal version, etc. These models can process and generate cross-modal data through deep learning techniques, such as generating images based on text, generating descriptions from images, or even reasoning based on multimodal inputs, greatly enhancing the flexibility and application breadth of the models, being able to handle multiple tasks simultaneously and being applied in multiple industries.
[0003] However, when current multimodal large models process graphical user interfaces, they cannot accurately generate precise coordinates and detailed function descriptions for each icon, button, or other interactive element. This problem directly affects the system's understanding and operation accuracy of the interface, resulting in a series of problems during the execution of automated tasks. For example, the system may make incorrect clicks due to the lack of accurate coordinates of interactive elements, or may not recognize certain specific interactive elements, thereby causing the task to fail. Or the model can recognize certain interactive elements, but the descriptions they generate are usually too simple (such as "button" or "input box"), lacking sufficient semantic information to effectively distinguish the different functions of the same type of elements, thus reducing the accuracy of task execution. More complicatedly, when the page contains multiple interactive elements or dynamically loaded content, existing models may not be able to correctly understand the relative relationships between elements or update the states of interactive elements in a timely manner, resulting in incorrect operations.
[0004] In addition, when dealing with complex pages and multiple tasks, it is necessary to comprehensively consider multiple factors such as the visual information of the current page, user instructions, and task progress to generate prompt words and corresponding operation strategies. However, the process of generating these prompt words not only takes time, but also due to the large amount of context information involved, the prompt words and image data occupy a large amount of resources. In particular, when the page content is more complex or multiple tasks are executed in parallel, the required amount of resources increases sharply, and the processing time and resource consumption increase accordingly, resulting in a more obvious delay in system processing and a lower execution efficiency, which is not suitable for high-frequency interaction scenarios that require real-time feedback.
[0005] On the other hand, each application usually has its unique interface design, interaction logic, and operation process. When the current intelligent agent processes different applications, it lacks in-depth optimization and customization support for the operation paths of these applications. Although it has a certain degree of generality and can identify and execute some standardized operations, when faced with complex interactions in specific applications, the flexibility and efficiency of the system are often not high enough. Specifically, most applications have fixed operation processes. For example, in social applications, the operation processes such as users browsing messages, replying to comments, sending private messages, and modifying settings also have certain rules. When the current intelligent agent executes these operations, it does not optimize or pre-learn according to the inherent path of the application, but relies on general model reasoning and multi-modal information generation, which results in its performance in specific applications being inferior to that of a specially optimized automated system. Due to the inability to deeply understand and adapt to the operation process of a specific application, the current intelligent agent may encounter unnecessary complexity and inefficiency problems when executing tasks. For example, the system cannot shorten the execution path through experience accumulation in repetitive tasks. Instead, due to the lack of optimization for the application, it may cause more calculation and judgment processes. This makes the current intelligent agent perform poorly in high-frequency, low-latency, and high-precision operations, especially in certain specific scenarios (such as some operations that need to be completed accurately and quickly), and it cannot be as efficient as human-computer interaction.
[0006] Therefore, it is necessary to provide a mobile intelligent agent information processing system that combines deep learning and interface data parsing to provide a reference basis for accurate icon positioning and recognition, making it suitable for high-frequency interaction scenarios that require real-time feedback. Summary of the Invention
[0007] The object of the present invention is to provide a mobile intelligent agent information processing system based on prior rules and multi-modal large models, which can achieve accurate icon positioning and function recognition, shorten the task execution time, improve the response speed of the system in high-frequency interaction scenarios, deeply customize the operation path, and continuously update through self-learning ability.
[0008] To achieve the above object, the present invention provides a mobile intelligent agent information processing system based on prior rules and multi-modal large models, including: An icon accurate positioning and recognition module, which accurately locates and recognizes the icons in the interface and infers the functions of the icons by combining the position information, visual features, and semantic reasoning of the icons through different levels of recognition methods; An execution efficiency improvement module, which improves the execution speed by task decomposition and optimized path design according to the changes in the application program and the content of the user instruction; An operation optimization module for specific applications, which is used to record the position information, size, and corresponding icon functions of each interface element, build a navigation path library for specific applications, and optimize the execution path of periodic tasks; The navigation and task data collection and update module is responsible for collecting and updating the dynamic data of navigation paths and task execution, ensuring the continuous optimization and learning of the system.
[0009] Preferably, the recognition methods at different levels include: First, parse the interface source data of the application, extract the position information of the icons, and initially locate the icon area as the candidate area; Next, use a deep learning model to combine image features and language information to perform visual content analysis on each candidate area to confirm whether it is an icon and classify it; Then, infer the specific function of the icon by analyzing the position of the icon and the semantic context of the surrounding elements.
[0010] Preferably, the task decomposition and optimization path design include: First, split the user instructions according to the changes in the application to avoid interference caused by complex instructions and ensure that each task can be efficiently executed one by one; Then, divide the task execution process into three stages: "select and open the application, page navigation, and execute specific tasks". By using a fixed and easy-to-collect operation path, it avoids repeatedly calling the multi-modal large model for complex reasoning, thereby significantly reducing the consumption of computing resources and execution latency.
[0011] Preferably, the stage of executing specific tasks includes caching the operation paths and interface states of common tasks to avoid starting reasoning and calculation from scratch each time.
[0012] Preferably, build a navigation path library for specific applications, including learning common page navigation paths by monitoring and recording the operation processes of users in specific applications.
[0013] Preferably, optimizing the execution path of periodic tasks includes presetting task execution strategies to automatically complete tasks, reducing user participation, and further optimizing the path according to the success rate and time consumption of task execution to ensure task execution efficiency.
[0014] Preferably, collecting and updating the dynamic data of navigation paths and task execution includes, for unrecorded tasks or page navigations, calling an agent based on a multi-modal large model to perform step-by-step reasoning, execute the task, and at the same time, record the execution path of the task in real time and store it in the operation path library.
[0015] Preferably, the real-time recording of the task execution path includes the operation process, timestamp, and task status; When the task status shows that the path execution is successful, mark the corresponding path as valid and update the operation path of the multi-modal large model in the path library; When the task status shows that the path execution fails, the agent based on the multimodal large model submits the record of the failed execution to the human reviewer for verification, and then modifies or optimizes it according to the operation path and execution status.
[0016] Preferably, when updating the navigation path also includes changes in interface elements or adjustments to application logic, the operation path library is automatically updated to adapt to the new task execution mode.
[0017] Therefore, the mobile intelligent agent information processing system based on the prior rules and multimodal large model of the present invention has the following technical effects: (1) It can accurately locate the coordinates of icons and interactive elements, and combine deep learning algorithms to infer their specific functions, thus effectively avoiding misoperations and recognition errors, and improving the stability and accuracy of task execution, especially in the case of complex interfaces and dynamic interfaces.
[0018] (2) By optimizing task decomposition and path design, it effectively reduces the consumption of computing resources and response latency. At the same time, by caching common operation paths, it avoids starting from scratch for reasoning every time a task is executed, thereby improving the execution efficiency, especially suitable for complex and high-frequency interaction scenarios.
[0019] (3) Through the unique operation path and interaction logic of the deep learning application, it provides highly customized optimization support for specific applications, and can self-learn and continuously optimize the path library, and adjust the execution strategy in a timely manner according to interface changes to ensure continuous improvement of task execution efficiency and accuracy.
[0020] (4) It realizes real-time recording and dynamic update of task paths and operation processes, can continuously optimize the operation path library according to interface element changes and application logic adjustments to adapt to changing application requirements, and ensure the efficiency and accuracy of the path.
[0021] The technical solution of the present invention will be further described in detail below through the drawings and embodiments. Description of the Drawings
[0022] Figure 1 It is a schematic diagram of a mobile intelligent agent information processing system based on prior rules and multimodal large models. Detailed Embodiments
[0023] The present invention can be more specifically explained through the following embodiments. The purpose of disclosing the present invention is to protect all changes and improvements within the scope of the present invention. The present invention is not limited to the following embodiments.
[0024] As Figure 1 shown, the present invention provides a mobile intelligent agent information processing system based on prior rules and multimodal large models, including: The figure standard positioning and recognition module aims to accurately position and recognize icons in the interface and infer their functions. Specifically: First, by parsing the interface source data of the application (such as an Extensible Markup Language file), the position information of the icons (such as the coordinates and dimensions of the rectangular box) is extracted, thereby initially positioning potential icon areas and forming a preliminary list of candidate areas; Second, for the interface pictures of the application, a deep learning model is used for image content analysis to enrich the scope of the candidate areas and avoid omissions; Then, for the image content input into each candidate area, it is judged whether the area is an icon and classified; After that, for the areas that have been recognized as icons, the system further combines the coordinate information, relative size, and image content of the icons to obtain the meaning of the icons.
[0025] The execution efficiency improvement module, whose core goal is to reduce unnecessary calculations and improve the execution speed through task decomposition and optimized path design. The specific implementation is as follows: First, by splitting the user instructions according to the changes in the application, the task is split into multiple executable subtasks. Each subtask only involves a small range of operations, avoiding the interference caused by complex instructions and ensuring that each task can be executed efficiently one by one. For example, a complex task (such as "Modify the notification settings in the settings page and save") is split into four subtasks: "Open the settings page", "Navigate to the notification settings", "Modify the notification settings", and "Save the settings".
[0026] Next, the task execution process is divided into three stages: "Select and open the application, page navigation, execute specific tasks". The common operation paths of the application are recorded in advance, and the paths are solidified into an optimized execution process, avoiding repeated calls to the multi-modal large model for complex reasoning, thereby greatly reducing the consumption of computing resources and execution latency. For example, (1) the execution path from the home page to the settings page of the user is fixed; (2) the "Confirm Search" button on Twitter is hidden in the lower right corner of a specific input method, and a specific input method switch needs to be executed to wake up the "Confirm Search" button, etc. By recording the execution path and caching it as a rule, when encountering similar tasks, these rules can be retrieved as a reference or directly reused, without having to reason step by step each time, reducing the amount of calculation.
[0027] For the stage of executing specific tasks, it is also possible to cache the operation paths and interface states of common tasks, avoiding starting from scratch for reasoning and calculation each time. This approach enables the system to quickly load and execute existing operation paths, reducing execution time and computational burden. For parallel tasks, the system can also execute multiple subtasks simultaneously through multi-threading or asynchronous mechanisms based on task priorities and resource limitations, effectively reducing the overall execution time and improving the system's response speed. In addition, the system can also make dynamic adjustments based on factors such as the urgency of tasks, resource usage, and operation dependencies, optimizing task scheduling to ensure the efficiency of the system when dealing with complex tasks and parallel execution of multiple tasks. In particular, when dealing with complex pages and parallel execution of multiple tasks, the optimized execution process and dynamic task scheduling mechanism can significantly shorten the task execution time and improve the response speed in high-frequency interaction scenarios.
[0028] The operation optimization module for specific applications aims to provide precise optimization for specific application programs, improving the accuracy and efficiency of task execution. The specific implementation is as follows: For each application, the system accurately records the position information, size, and corresponding icon functions of each interface element by analyzing the page structure, layout, and common interaction patterns of the application. For example, the "send message" icon in a social application is usually located at the bottom of the interface, and the system records this position and binds it to the icon function. The system also learns common page navigation paths by monitoring and recording the user's operation process in a specific application.
[0029] For each application program, the system constructs a specific navigation path library. For example, the path from the main interface to the settings page is fixed and can be directly reused; the path from the settings page to a certain function page can also be automatically recognized and executed.
[0030] For periodic tasks (such as daily logins, timed message checks, etc.), the system records and optimizes the execution paths of these tasks. By setting task execution strategies in advance, these tasks can be automatically completed, reducing user participation. Moreover, the system will gradually optimize the path based on the success rate and time consumption of task execution to ensure the improvement of task execution efficiency.
[0031] The navigation and task data collection and update module is responsible for collecting and updating the dynamic data of navigation paths and task execution to ensure the continuous optimization and learning of the system. The specific implementation is as follows: For each unrecorded task or page navigation, the system will call an intelligent agent based on a multimodal large model for step-by-step reasoning and execute the task. At the same time, the system will automatically record the execution path of the task in real time, including key information such as the operation, timestamp, and task status (success, failure) of each step, and store it in the operation path library.
[0032] If a certain operation path is executed successfully, the system will mark this path as valid and update the task and its corresponding operation path to the path library (i.e., the prior rule library).
[0033] For the records of the failure of the intelligent agent based on the multi-modal large model, the system will submit them to human reviewers for verification. The reviewers can view the operation path, execution status, and modify or optimize them. After the review is passed, the operation path will be automatically added to the system path library. When the intelligent agent executes tasks subsequently, it can search from the prior rule library according to the task content for reference or reuse the existing operation path. The system realizes the continuous update of the path library by continuously collecting, reviewing, and optimizing the task paths. In addition, whenever the system detects changes in interface elements or adjustments to application logic, the path library will be automatically updated to adapt to the new task execution mode.
[0034] Therefore, the present invention adopts the above-mentioned mobile intelligent agent information processing system based on prior rules and multi-modal large models, combines deep learning with interface data parsing to achieve accurate icon positioning and function recognition; significantly improves the execution efficiency and reduces the consumption of computing resources through task decomposition and path optimization design; provides customized operation optimization for specific applications, and continuously improves the accuracy and adaptability of the path library through automatic learning; collects and updates task paths in real time to ensure that the system can handle interface changes and continuously optimize.
[0035] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A mobile agent information processing system based on prior rules and multimodal large models, characterized by: include: Icon accurate positioning and recognition module, through different levels of recognition methods, combined with the icon's location information, visual features and semantic reasoning, accurately locates and recognizes the icons in the interface, and infers the icon's function; The execution efficiency improvement module improves execution speed by decomposing tasks and optimizing path design according to changes in the application and the content of user instructions; The operation optimization module for specific applications is used to record the location information, size and corresponding icon function of each interface element, build a navigation path library for specific applications, and optimize the execution path of periodic tasks; The navigation and task data collection and update module is responsible for collecting and updating the dynamic data of navigation paths and task execution, and continuously performing self-learning and optimization.
2. The mobile terminal agent information processing system based on prior rules and multimodal large model according to claim 1 is characterized in that: Different levels of identification methods, including: First, parse the interface source data of the application, extract the location information of the icon, and preliminarily locate the icon area as the candidate area; Next, we use a deep learning model to combine image features and language information to perform visual content analysis on each candidate area to confirm whether it is an icon and classify it; Then, the specific function of the icon is inferred by analyzing the location of the icon and the semantic context of the surrounding elements.
3. The mobile terminal intelligent agent information processing system based on prior rules and multimodal large model according to claim 1 is characterized in that: Task decomposition and optimized path design, including: First, split the user instructions according to the application changes; Then, the task execution process is divided into three stages: "select and open the application, page navigation, and perform specific tasks", forming a fixed and easy-to-collect operation path.
4. The mobile terminal intelligent agent information processing system based on prior rules and multimodal large model according to claim 3 is characterized in that: The specific task execution phase includes caching the operation paths and interface states of common tasks.
5. The mobile terminal intelligent agent information processing system based on prior rules and multimodal large model according to claim 1 is characterized in that: Build a navigation path library for a specific application, including learning common page navigation paths by monitoring and recording the user's operation flow in a specific application.
6. The mobile terminal intelligent agent information processing system based on prior rules and multimodal large model according to claim 1 is characterized in that: Optimizing the execution path of periodic tasks includes pre-setting task execution strategies, completing tasks automatically, reducing user involvement, and further optimizing the path based on the success rate and time consumption of task execution.
7. The mobile terminal intelligent agent information processing system based on prior rules and multimodal large model according to claim 1 is characterized in that: Collect and update dynamic data of navigation paths and task execution, including calling the intelligent agent based on the multimodal large model to perform step-by-step reasoning and execute tasks for unrecorded tasks or page navigations, while recording the execution path of the task in real time and storing it in the operation path library.
8. The mobile terminal intelligent agent information processing system based on prior rules and multimodal large model according to claim 7 is characterized in that: Real-time recording of the task execution path including operation process, timestamp, and task status; When the task status shows that the path is executed successfully, the corresponding path is marked as valid, and the path library is updated for the operation path executed for the multimodal large model; When the task status shows that the path execution has failed, the intelligent agent based on the multimodal large model will submit the record of the execution failure to the manual reviewer for verification, and then modify or optimize it according to the operation path and execution status.
9. The mobile terminal intelligent agent information processing system based on prior rules and multimodal large model according to claim 1 is characterized in that: Updating the navigation path also includes automatically updating the operation path library when interface elements change or application logic is adjusted to adapt to the new task execution mode.
Citation Information
Patent Citations
Method for intelligent operation based on image recognition
CN109062474A
UI element analysis method and system of human-computer interaction interface, terminal and medium
CN118587713A
Graphical user interface navigation method and apparatus
US20060010402A1
Cited By
Intelligent right steward implementation method and device based on AI capability
CN121144621A
Browser user behavior recording intelligent workflow generation and execution method and system
CN121412135A