Mobile agent information processing system based on prior rules and multimodal large model

By combining prior rules and multimodal large models with a mobile intelligent agent information processing system, the problem of low recognition and operation efficiency of multimodal large models in graphical user interfaces is solved, precise positioning and function recognition of icons are achieved, task paths are optimized, customized operations are adapted to specific applications, and response speed and execution efficiency are improved in high-frequency interaction scenarios.

CN120161973BActive Publication Date: 2025-09-23HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510266979.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-09-23
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

When processing graphical user interfaces, existing large multimodal models are unable to accurately identify and generate the coordinates and functional descriptions of interactive elements, resulting in increased misoperations and resource consumption. They are unable to adapt to complex pages and multi-tasking scenarios, and lack deep optimization for specific applications, resulting in low execution efficiency.

Method used

A mobile intelligent agent information processing system based on prior rules and multimodal large models is adopted. Through the icon positioning and recognition module, execution efficiency improvement module, specific application operation optimization module and navigation task data collection and update module, combined with deep learning and interface data analysis, it can achieve precise icon positioning and function recognition, optimize task paths, and self-learn to adapt to interface changes.

Benefits of technology

It achieves accurate positioning and function identification of icons, reduces computing resource consumption, improves the stability and efficiency of task execution, adapts to customized operations for specific applications, and ensures fast response and accuracy in high-frequency interaction scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120161973B_ABST
    Figure CN120161973B_ABST
Patent Text Reader

Abstract

The present invention discloses a mobile terminal intelligent body information processing system based on prior rules and multimodal large models, which relates to the field of artificial intelligence technology, including an icon accurate positioning and recognition module, which accurately positions and recognizes icons in the interface through different levels of recognition methods, and infers the functions of the icons; an execution efficiency improvement module, which improves the execution speed by task disassembly and optimization path design; an operation optimization module for specific applications, which is used to record the position information, size and corresponding icon function of each interface element, build a navigation path library for specific applications, and optimize the execution path of periodic tasks; a navigation and task data collection and update module, which is responsible for collecting and updating the dynamic data of navigation paths and task execution, and continuously performing self-learning and optimization. Therefore, the above method can achieve accurate icon positioning and function recognition, improve execution efficiency and reduce computing resource consumption, and at the same time be able to cope with interface changes and continuously optimize.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a mobile terminal intelligent agent information processing system based on priori rules and multimodal large models. Background Art

[0002] Large multimodal models are a rapidly developing area in the field of artificial intelligence, with significant progress in cross-modal understanding and generation across fields such as vision, language, and audio. Examples include the Tongyi Qianwen (Tongyi Qianwen) large-scale visual-language model, the Shusheng Wanxiang (Shusheng Wanxiang) large multimodal model, and the multimodal version of GPT-4. These models leverage deep learning techniques to process and generate cross-modal data, such as generating images from text, generating descriptions from images, and even reasoning based on multimodal input. This significantly enhances the models' flexibility and applicability, enabling them to simultaneously handle multiple tasks and find application across multiple industries.

[0003] However, current large multimodal models are unable to accurately generate precise coordinates and detailed functional descriptions for each icon, button, or other interactive element when processing graphical user interfaces. This problem directly affects the system's understanding of the interface and operational accuracy, leading to a series of problems in the execution of automated tasks. For example, the system may make erroneous clicks due to the lack of accurate coordinates of interactive elements, or fail to identify certain specific interactive elements, resulting in the inability to complete the task. Alternatively, the model can identify certain interactive elements, but the descriptions they generate are often too simple (such as "button" or "input box"), lacking sufficient semantic information, and unable to effectively distinguish between different functions of elements of the same type, thereby reducing the accuracy of task execution. To make matters more complicated, when a page contains multiple interactive elements or dynamically loads content, existing models may not be able to correctly understand the relative relationships between elements or update the status of interactive elements in a timely manner, leading to incorrect operations.

[0004] Furthermore, when processing complex web pages and multiple tasks, it's necessary to integrate multiple factors, including the current page's visual information, user instructions, and task progress, to generate prompts and corresponding action strategies. However, generating these prompts is not only time-consuming, but also consumes significant resources due to the extensive contextual information involved, as the prompts and image data consume significant resources. In particular, when web page content is complex or multiple tasks are executed in parallel, the number of resources required increases dramatically, leading to increased processing time and resource consumption. This results in more noticeable system processing delays, lower execution efficiency, and makes the system unsuitable for high-frequency interactive scenarios requiring real-time feedback.

[0005] On the other hand, each application typically has its own unique interface design, interaction logic, and operational processes. Current intelligent agents lack deep optimization and customized support for these application-specific operational paths when handling diverse applications. While they possess a certain degree of versatility and can recognize and execute some standardized operations, they often lack flexibility and efficiency when faced with the complex interactions of specific applications. Specifically, most applications have fixed operational processes. For example, in social applications, users' operational processes from browsing messages, replying to comments, sending private messages, to modifying settings follow a certain pattern. Current intelligent agents do not optimize or pre-learn based on the application's inherent operational processes when executing these operations. Instead, they rely on general model reasoning and multimodal information generation, resulting in inferior performance compared to specifically optimized automated systems for specific applications. Due to their inability to deeply understand and adapt to the operational processes of a specific application, current intelligent agents may encounter unnecessary complexity and inefficiency when performing tasks. For example, in repetitive tasks, the system cannot shorten the execution path through experience accumulation. Instead, the lack of application-specific optimization may lead to additional calculations and judgments. This results in current intelligent agents performing poorly in high-frequency, low-latency, and high-precision operations, especially in certain specific scenarios (such as some operations that need to be completed accurately and quickly), and they cannot be as efficient as human interactions.

[0006] Therefore, it is necessary to provide a mobile intelligent agent information processing system that combines deep learning and interface data analysis to provide a reference basis for accurate icon positioning and identification, making it suitable for high-frequency interaction scenarios that require real-time feedback. Summary of the Invention

[0007] The purpose of the present invention is to provide a mobile intelligent agent information processing system based on prior rules and multimodal large models, which can achieve precise positioning and function identification of icons, shorten task execution time, improve the response speed of the system in high-frequency interaction scenarios, deeply customize operation paths and continuously update through self-learning capabilities.

[0008] To achieve the above objectives, the present invention provides a mobile agent information processing system based on prior rules and a multimodal large model, comprising:

[0009] The icon positioning and recognition module uses different levels of recognition methods, combining the icon's location information, visual features, and semantic reasoning to accurately locate and identify icons in the interface and infer their functions;

[0010] The execution efficiency improvement module improves execution speed by decomposing tasks and optimizing path design based on changes in the application and the content of user instructions;

[0011] An application-specific operation optimization module, which records the location information, size, and corresponding icon function of each interface element, builds a navigation path library for specific applications, and optimizes the execution path of periodic tasks;

[0012] The navigation and task data collection and update module is responsible for collecting and updating dynamic data of navigation paths and task execution to ensure continuous optimization and learning of the system.

[0013] Preferably, different levels of identification methods include:

[0014] First, parse the application interface source data, extract the icon location information, and preliminarily locate the icon area as the candidate area;

[0015] Next, using a deep learning model, combined with image features and language information, a visual content analysis is performed on each candidate region to confirm whether it is an icon and classify it;

[0016] Then, the specific function of the icon is inferred by analyzing the location of the icon and the semantic context of the surrounding elements.

[0017] Preferably, task decomposition and optimized path design include:

[0018] First, user instructions are split according to application changes to avoid interference caused by complex instructions and ensure that each task can be executed efficiently one by one;

[0019] Then, the task execution process is divided into three stages: "selecting and opening the application, page navigation, and executing specific tasks". Through a fixed and easy-to-collect operation path, repeated calls to multimodal large models for complex reasoning are avoided, thereby greatly reducing computing resource consumption and execution delays.

[0020] Preferably, executing the specific task stage includes caching the operation paths and interface states of common tasks to avoid reasoning and calculation from scratch each time the task is executed.

[0021] Preferably, building a navigation path library for a specific application includes learning common page navigation paths by monitoring and recording user operation processes in the specific application.

[0022] Preferably, optimizing the execution path of periodic tasks includes pre-setting task execution strategies, automatically completing tasks, reducing user involvement, and further optimizing the path based on the success rate and time consumption of task execution to ensure task execution efficiency.

[0023] Preferably, dynamic data of navigation paths and task execution are collected and updated, including calling an intelligent agent based on a multimodal large model to perform step-by-step reasoning and execute tasks for unrecorded tasks or page navigations, while recording the execution path of the task in real time and storing it in an operation path library.

[0024] Preferably, the execution path of the task is recorded in real time, including the operation process, timestamp, and task status;

[0025] When the task status shows that the path is successfully executed, the corresponding path is marked as valid, and the path library is updated to indicate the operation path executed for the multimodal large model;

[0026] When the task status shows that the path execution has failed, the intelligent agent based on the multimodal large model will submit the execution failure record to the manual reviewer for verification, and then modify or optimize it according to the operation path and execution status.

[0027] Preferably, updating the navigation path also includes automatically updating the operation path library when interface elements change or application logic is adjusted, so as to adapt to the new task execution mode.

[0028] Therefore, the present invention adopts the above-mentioned mobile terminal intelligent agent information processing system based on prior rules and multimodal large model, which has the following technical effects:

[0029] (1) It can accurately locate the coordinates of icons and interactive elements, and combine deep learning algorithms to infer their specific functions, thereby effectively avoiding misoperations and recognition errors, and improving the stability and accuracy of task execution, especially in the case of complex and dynamic interfaces.

[0030] (2) By optimizing task decomposition and path design, the consumption of computing resources and response delay are effectively reduced. At the same time, by caching common operation paths, reasoning from scratch is avoided every time a task is executed, thereby improving execution efficiency. It is particularly suitable for complex and high-frequency interaction scenarios.

[0031] (3) Through the unique operation paths and interaction logic of deep learning applications, it provides highly customized optimization support for specific applications, and is able to self-learn and continuously optimize the path library, and adjust the execution strategy in time according to interface changes to ensure that the efficiency and accuracy of task execution continue to improve.

[0032] (4) Realize real-time recording and dynamic updating of task paths and operation processes, and continuously optimize the operation path library according to changes in interface elements and application logic adjustments to adapt to changing application needs and ensure the efficiency and accuracy of the path.

[0033] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 This is a schematic diagram of a mobile intelligent agent information processing system based on prior rules and multimodal large models. DETAILED DESCRIPTION

[0035] The present invention can be explained in more detail by the following examples. The purpose of disclosing the present invention is to protect all changes and improvements within the scope of the present invention. The present invention is not limited to the following examples.

[0036] like Figure 1 As shown, the present invention provides a mobile terminal intelligent agent information processing system based on prior rules and multimodal large models, including:

[0037] The icon positioning and recognition module aims to accurately locate and identify icons in the interface and infer their functions. Specifically: First, by parsing the application's interface source data (such as an extensible markup language file), the icon's location information (such as the coordinates and dimensions of the rectangular box) is extracted to preliminarily locate potential icon areas and form a preliminary list of candidate areas. Second, a deep learning model is used to analyze the image content of the application's interface images to enrich the scope of candidate areas and avoid omissions. Then, for each candidate area input, the image content is determined to determine whether the area is an icon and classify it. After that, for areas identified as icons, the system further combines the icon's coordinate information, relative size, and image content to obtain the icon's meaning.

[0038] The core goal of the execution efficiency improvement module is to reduce unnecessary calculations and improve execution speed by decomposing tasks and optimizing path design. The specific implementation is as follows:

[0039] First, tasks are broken down into multiple executable subtasks by splitting user instructions based on application changes. Each subtask only involves a narrow range of operations, avoiding the interference caused by complex instructions and ensuring that each task can be executed efficiently one by one. For example, a complex task (such as "Modify and save notification settings in the settings page") is broken down into four subtasks: "Open the settings page," "Navigate to notification settings," "Modify notification settings," and "Save settings."

[0040] Next, the task execution process is divided into three stages: "selecting and opening an application, navigating the page, and executing a specific task." Common application operation paths are recorded in advance and solidified into an optimized execution process, avoiding repeated calls to multimodal large models for complex reasoning, thereby significantly reducing computing resource consumption and execution delays. For example, (1) the execution path from the homepage to the settings page is fixed; (2) Twitter's "Confirm Search" button is hidden in the lower right corner of a specific input method, and a specific input method must be switched to activate the "Confirm Search" button. By recording the execution path and caching it as rules, these rules can be retrieved when encountering similar tasks, used as a reference, or directly reused, without the need to re-infer step by step each time, thus reducing the amount of computation.

[0041] During the execution of specific tasks, the operation paths and interface states of common tasks can also be cached to avoid reasoning and calculations from scratch each time the task is executed. This method enables the system to quickly load and execute existing operation paths, reducing execution time and computational burden. For parallel tasks, the system can also execute multiple subtasks simultaneously through multi-threading or asynchronous mechanisms based on task priority and resource constraints, effectively reducing overall execution time and improving system response speed. In addition, the system can also dynamically adjust based on factors such as the urgency of the task, resource usage, and operation dependencies to optimize task scheduling and ensure the system's efficiency when executing complex tasks and multiple tasks in parallel. In particular, when processing complex pages and multiple tasks in parallel, the optimized execution process and dynamic task scheduling mechanism can significantly shorten task execution time and improve response speed in high-frequency interaction scenarios.

[0042] The application-specific operation optimization module aims to provide precise optimization for specific applications and improve the accuracy and efficiency of task execution. The specific implementation is as follows:

[0043] For each app, the system analyzes the app's page structure, layout, and common interaction patterns, accurately recording the location, size, and corresponding icon function of each interface element. For example, the "Send Message" icon in a social app is typically located at the bottom of the interface; the system records this location and associates it with the icon's function. The system also learns common page navigation paths by monitoring and recording user actions within specific apps.

[0044] For each application, the system builds a specific navigation path library. For example, the path from the main interface to the settings page is fixed and can be directly reused; the path from the settings page to a specific function page can also be automatically identified and executed.

[0045] For periodic tasks (such as daily logins and scheduled message checks), the system records and optimizes their execution paths. By pre-setting task execution strategies, these tasks can be completed automatically, reducing user involvement. Furthermore, the system gradually optimizes the paths based on task success rates and time consumption, ensuring improved task execution efficiency.

[0046] The navigation and task data collection and update module is responsible for collecting and updating dynamic data of navigation paths and task execution to ensure continuous optimization and learning of the system. The specific implementation is as follows:

[0047] For each unrecorded task or page navigation, the system invokes an agent based on a large multimodal model to perform step-by-step reasoning and execute the task. Simultaneously, the system automatically records the execution path of the task in real time, including key information such as the operation, timestamp, and task status (success, failure) for each step, and stores it in the operation path library.

[0048] If an operation path is executed and succeeds, the system will mark this path as valid and update the task and its corresponding operation path to the path library (i.e., the prior rule library).

[0049] For records of failed executions of agents based on multimodal large models, the system will submit them to manual reviewers for verification. Reviewers can view the operation path and execution status and modify or optimize them. After the review is passed, the operation path will be automatically added to the system path library. When the agent performs tasks subsequently, it can search from the prior rule library for reference or reuse existing operation paths based on the task content. The system continuously updates the path library by continuously collecting, reviewing and optimizing task paths. In addition, whenever the system detects changes to interface elements or adjustments to application logic, the path library will be automatically updated to adapt to the new task execution mode.

[0050] Therefore, the present invention adopts the above-mentioned mobile intelligent agent information processing system based on prior rules and multimodal large models, and realizes accurate icon positioning and function identification through the combination of deep learning and interface data analysis; significantly improves execution efficiency and reduces computing resource consumption through task decomposition and path optimization design; provides customized operation optimization for specific applications, and continuously improves the accuracy and adaptability of the path library through automatic learning; collects and updates task paths in real time to ensure that the system can respond to interface changes and continuously optimize.

[0051] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A mobile agent information processing system based on prior rules and multimodal large models, characterized by: include: The icon positioning and recognition module uses different levels of recognition methods, combining the icon's location information, visual features, and semantic reasoning to accurately locate and identify icons in the interface and infer their functions; Among them, different levels of identification methods include: First, parse the application interface source data, extract the icon location information, and preliminarily locate the icon area as the candidate area; Next, using a deep learning model, combined with image features and language information, a visual content analysis is performed on each candidate region to confirm whether it is an icon and classify it; Then, by analyzing the icon’s location and the semantic context of surrounding elements, the specific function of the icon is inferred; The execution efficiency improvement module improves execution speed by breaking down tasks and optimizing path design based on application changes and user instructions, including: First, split the user instructions according to the application changes; Then, the task execution process is divided into three stages: "select and open the application, page navigation, and perform specific tasks", forming a fixed and easy-to-collect operation path; An application-specific operation optimization module, which records the location information, size, and corresponding icon function of each interface element, builds a navigation path library for specific applications, and optimizes the execution path of periodic tasks; The navigation and task data collection and update module is responsible for collecting and updating dynamic data of navigation paths and task execution, and continuously performing self-learning and optimization.

2. The mobile terminal intelligent agent information processing system based on prior rules and multimodal large model according to claim 1 is characterized in that: The specific task execution phase includes caching the operation paths and interface states of common tasks.

3. The mobile agent information processing system based on prior rules and multimodal large models according to claim 1 is characterized in that: Build a navigation path library for a specific application, including learning common page navigation paths by monitoring and recording user operation processes in a specific application.

4. The mobile agent information processing system based on prior rules and multimodal large models according to claim 1 is characterized in that: Optimizing the execution path of periodic tasks involves pre-setting task execution strategies, completing tasks automatically, reducing user involvement, and further optimizing the path based on the success rate and time consumption of task execution.

5. The mobile terminal intelligent agent information processing system based on prior rules and multimodal large model according to claim 1 is characterized in that: Collect and update dynamic data of navigation paths and task execution, including calling the intelligent agent based on the multimodal large model to perform step-by-step reasoning and execute tasks for unrecorded tasks or page navigations, while recording the execution path of the task in real time and storing it in the operation path library.

6. The mobile terminal intelligent agent information processing system based on prior rules and multimodal large model according to claim 5 is characterized in that: Real-time recording of the task execution path including operation process, timestamp, and task status; When the task status shows that the path is successfully executed, the corresponding path is marked as valid, and the path library is updated to indicate the operation path executed for the multimodal large model; When the task status shows that the path execution has failed, the intelligent agent based on the multimodal large model will submit the execution failure record to the manual reviewer for verification, and then modify or optimize it according to the operation path and execution status.

7. The mobile terminal intelligent agent information processing system based on prior rules and multimodal large model according to claim 1 is characterized in that: Updating the navigation path also includes automatically updating the operation path library when interface elements change or application logic is adjusted to adapt to the new task execution mode.

Citation Information

Patent Citations

  • Method for intelligent operation based on image recognition

    CN109062474A

  • UI element analysis method and system of human-computer interaction interface, terminal and medium

    CN118587713A