Terminal control system and control method based on large model and interface combined driving

By building a terminal control system based on the joint drive of large models and interfaces, the problem of efficient control of smart terminals in complex scenarios is solved, multimodal interaction and automatic execution of cross-application tasks are realized, and the user experience and terminal automation capabilities are improved.

CN120704834APending Publication Date: 2025-09-26BEIJING QIANGUAN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510830541.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The interaction methods of existing smart terminals are difficult to meet the needs of efficient control in complex scenarios. They have problems such as limited command comprehension, limited operation capabilities, lack of context awareness and limited computing resources.

Method used

Build a terminal control system based on the joint drive of large models and interfaces, monitor user input through peripheral intelligent carriers, combine the terminal interface status, and use large language models to generate operation instructions to realize the automatic execution of multi-step tasks. The peripherals dominate the control process and reduce terminal resource usage.

Benefits of technology

It supports natural interaction methods for multimodal input, improves user experience, realizes flexible and intelligent cross-application task planning, reduces terminal computing burden, and improves the automation capabilities of mobile terminals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704834A_ABST
    Figure CN120704834A_ABST
Patent Text Reader

Abstract

The invention discloses a terminal control system and control method based on large model and interface combined driving, and the control system comprises a peripheral agent carrier, an instruction analysis and task generation module, a terminal interface state collection and feedback module, and a task execution control module. The peripheral agent carrier comprises an input sensing unit, a connecting module and a processing chip; the terminal control system and the control method support a natural interaction mode of multi-modal input, improve user experience, utilize peripherals to decouple control logic and terminal resources, reduce local calculation burden, realize flexible and intelligent task planning in combination with a large language model and an interface state, can execute a complex task chain in a cross-application manner, and improve the user experience. The automation capability of the mobile terminal is improved; the terminal control system is open in structure and suitable for various forms such as earphones, watches, back clips and pens.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence terminal control and human-computer interaction technology, and in particular to a terminal control system and a control method based on the joint drive of a large model and an interface. Background Art

[0002] With the widespread adoption of smart devices (such as mobile phones, tablets, and wearable devices), users' demands for operational efficiency and interactive intelligence are constantly increasing. Traditional graphical user interfaces (GUIs), which rely on manual clicks and swipes, are unable to meet the demand for efficient control in complex scenarios.

[0003] Although voice assistant products such as Siri, Google Assistant, and Xiao Ai exist, they generally suffer from the following technical bottlenecks:

[0004] 1. Limited command comprehension: Most assistants rely on keyword matching or limited intent templates and are unable to interpret complex logic or multi-step tasks in natural language.

[0005] 2. Limited operational capabilities: Unable to make dynamic decisions based on the current interface status, only preset commands can be executed;

[0006] 3. Lack of context awareness: The current interface structure and status of the terminal cannot be obtained, resulting in a lack of intelligence and adaptability in task execution;

[0007] 4. Limited computing resources: Running complex models locally on the terminal is expensive and has high latency.

[0008] Therefore, existing technologies are unable to realize multi-modal input-driven, large-model reasoning and control processes based on interface state perception without relying on the permanent residence of terminal resources. There is an urgent need for an intelligent task execution system that integrates software and hardware, decouples control, and has contextual decision-making capabilities. Summary of the Invention

[0009] The purpose of this paper is to build a terminal control system and method jointly driven by a large model and an interface, supporting users to issue task requests through any form of natural instructions, and realizing the automatic execution of multi-step, cross-application tasks through joint reasoning of external collaborative devices and a large language model, combined with the current interface status of the terminal.

[0010] In order to achieve the above object, the present invention adopts the following technical solutions:

[0011] The terminal control system based on the joint drive of large model and interface includes peripheral intelligent body carrier, instruction parsing and task generation module, terminal interface status collection and feedback module, and task execution control module.

[0012] The peripheral intelligent body carrier includes an input sensing unit, a connection module, and a processing chip;

[0013] The peripheral intelligent agent carrier can continuously monitor user input or perceive task signals, actively wake up the terminal device, start the debugging communication channel, and serve as the "entrance" and "executor" of the large model assistant.

[0014] The instruction parsing and task generation module converts the natural task instructions input by the user into a semantic structure; at the same time, it combines the current interface state of the terminal; and at the same time calls the large language model to generate an operation instruction sequence or step-by-step reasoning to generate the next operation.

[0015] The terminal interface status acquisition and feedback module can obtain the current interface structure through the terminal supporting software; at the same time, it can feedback the operation execution results to the model or user.

[0016] The task execution control module can perform simulated operations on the terminal according to the operation instructions generated by the model; it can support functions such as status feedback, exception recovery, and task rollback; and it can provide feedback in the form of voice, text, images, etc.

[0017] Preferably, the natural instructions in any form include but are not limited to voice, text, images, gestures, etc.

[0018] Preferably, the input sensing unit includes but is not limited to a microphone, a camera, and a touch panel;

[0019] Preferably, the connection module includes but is not limited to USB, Bluetooth, and Wi-Fi;

[0020] Preferably, the natural task instructions include but are not limited to voice, text, and images;

[0021] Preferably, the front interface state includes but is not limited to the interface control hierarchical structure, foreground App information, etc.;

[0022] Preferably, the simulation operation includes but is not limited to clicking, sliding, and inputting.

[0023] A terminal control method based on a combined drive of a large model and an interface comprises the following steps:

[0024] S1: Task instruction acquisition stage

[0025] Users input task instructions through voice, text, images, touch, etc.; peripherals continuously sense and identify task triggering conditions, such as keywords, image features or gestures, and support multimodal interaction.

[0026] S2: Terminal status awareness stage

[0027] The peripheral activates the terminal control agent through the communication interface to obtain the current interface status, including the foreground app, interface structure hierarchy, focus elements, etc. The system is controlled by the peripheral and has the ability to stay online, so it can wake up the terminal and initiate task execution at any time.

[0028] S3: Intent parsing and action generation stage

[0029] The peripheral packages the command data and terminal status and sends them to a remote or local large language model. The model combines contextual reasoning to generate the next operation instruction. Task reasoning and scheduling are controlled by the peripheral, and the terminal is only responsible for status collection and action execution, reducing resource usage.

[0030] S4: Operation execution and loop judgment stage

[0031] The peripheral or terminal agent executes the operation instructions generated by the model, and after execution, obtains the interface status again as the input for the next round of reasoning, forming a loop until the task is completed; the terminal agent cooperates to extract interface structure information to provide contextual decision-making basis for large model reasoning.

[0032] S5: Task completion and feedback stage

[0033] After the task is completed, the system will feedback the results to the user through voice synthesis, interface display or other means; after the task is completed, the results can be fed back to the user through voice, graphics, vibration, etc., forming a closed-loop intelligent experience.

[0034] The present invention is based on a terminal control system and control method jointly driven by a large model and an interface, supports natural interaction with multimodal input, improves user experience, utilizes peripherals to decouple control logic and terminal resources, reduces local computing burden, and combines a large language model with interface status to achieve flexible and intelligent task planning. It can execute complex task chains across applications and enhance the automation capabilities of mobile terminals. The terminal control system of the present invention has an open structure and is applicable to various forms such as headphones, watches, back clips, and pens. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 This is a system structure diagram of the present invention.

[0036] Figure 2 A logic diagram of the connection between the peripheral device and the terminal of the present invention; DETAILED DESCRIPTION

[0037] This section will describe in detail the specific embodiments of the present invention. The preferred embodiments of the present invention are shown in the accompanying drawings. The purpose of the accompanying drawings is to supplement the description of the text part of the specification with graphics, so that people can intuitively and vividly understand each technical feature and the overall technical solution of the present invention, but it should not be understood as a limitation on the scope of protection of the present invention.

[0038] In the description of the present invention, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.

[0039] In the description of this invention, terms such as "greater than," "less than," and "exceed" are understood to exclude the number itself, while terms such as "above," "below," and "within" are understood to include the number itself. The use of terms such as "first" and "second" is solely for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance, or implicitly specifying the number of the indicated technical features, or implicitly specifying the order of the indicated technical features.

[0040] like Figure 1 As shown in the figure, the terminal control system based on the joint drive of large model and interface includes peripheral intelligent body carrier, instruction parsing and task generation module, terminal interface status collection and feedback module, and task execution control module.

[0041] The peripheral intelligent body carrier includes an input sensing unit, a connection module, and a processing chip;

[0042] The peripheral intelligent agent carrier can continuously monitor user input or perceive task signals, actively wake up the terminal device, start the debugging communication channel, and serve as the "entrance" and "executor" of the large model assistant.

[0043] The instruction parsing and task generation module converts the natural task instructions input by the user into a semantic structure; at the same time, it combines the current interface state of the terminal; and at the same time calls the large language model to generate an operation instruction sequence or step-by-step reasoning to generate the next operation.

[0044] The terminal interface status acquisition and feedback module can obtain the current interface structure through the terminal supporting software; at the same time, it can feedback the operation execution results to the model or user.

[0045] The task execution control module can perform simulated operations on the terminal according to the operation instructions generated by the model; it can support functions such as status feedback, exception recovery, and task rollback; and it can provide feedback in the form of voice, text, images, etc.

[0046] Furthermore, the natural instructions in any form include but are not limited to voice, text, image, gesture, etc. Furthermore, the input sensing unit includes but is not limited to a microphone, a camera, and a touchpad;

[0047] Furthermore, the connection module includes but is not limited to USB, Bluetooth, and Wi-Fi;

[0048] Furthermore, the natural task instructions include but are not limited to voice, text, and images;

[0049] Furthermore, the front interface state includes but is not limited to the interface control hierarchical structure, foreground App information, etc.; further, the simulated operation includes but is not limited to clicking, sliding, and input.

[0050] like Figure 2 As shown, the terminal control method based on the combined drive of the large model and the interface includes the following steps:

[0051] S1: Task instruction acquisition stage

[0052] Users input task instructions through voice, text, images, touch, etc.; peripherals continuously sense and identify task triggering conditions, such as keywords, image features or gestures, and support multimodal interaction.

[0053] S2: Terminal status awareness stage

[0054] The peripheral activates the terminal control agent through the communication interface to obtain the current interface status, including the foreground app, interface structure hierarchy, focus elements, etc. The system is controlled by the peripheral and has the ability to stay online, so it can wake up the terminal and initiate task execution at any time.

[0055] S3: Intent parsing and action generation stage

[0056] The peripheral packages the command data and terminal status and sends them to a remote or local large language model. The model combines contextual reasoning to generate the next operation instruction. Task reasoning and scheduling are controlled by the peripheral, and the terminal is only responsible for status collection and action execution, reducing resource usage.

[0057] S4: Operation execution and loop judgment stage

[0058] The peripheral or terminal agent executes the operation instructions generated by the model, and after execution, obtains the interface status again as the input for the next round of reasoning, forming a loop until the task is completed; the terminal agent cooperates to extract interface structure information to provide contextual decision-making basis for large model reasoning.

[0059] S5: Task completion and feedback stage

[0060] After the task is completed, the system will feedback the results to the user through voice synthesis, interface display or other means; after the task is completed, the results can be fed back to the user through voice, graphics, vibration, etc., forming a closed-loop intelligent experience.

[0061] The above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, but these descriptions should not be understood as limiting the scope of the present invention. The scope of protection of the present invention is defined by the appended claims, and any changes based on the claims of the present invention are within the scope of protection of the present invention.

Claims

1. The terminal control system based on the joint drive of large model and interface is characterized by: include: Peripheral intelligent body carrier, instruction parsing and task generation module, terminal interface status collection and feedback module, task execution control module, the peripheral intelligent body carrier includes an input perception unit, a connection module, and a processing chip; the peripheral intelligent body carrier can continuously monitor user input or perceive task signals, can actively wake up the terminal device, start the debugging communication channel, and serve as the "entrance" and "executor" of the large model assistant; the instruction parsing and task generation module converts the natural task instructions input by the user into a semantic structure; at the same time, it combines the current interface status of the terminal; at the same time, it calls the large language model to generate an operation instruction sequence or step-by-step reasoning to generate the next operation; the terminal interface status collection and feedback module can obtain the current interface structure through the terminal supporting software; at the same time, it can feedback the operation execution results to the model or user; the task execution control module can perform simulated operations on the terminal according to the operation instructions generated by the model; it can support functions such as status feedback, exception recovery, and task rollback; and it can provide feedback in the form of voice, text, images, etc.

2. The terminal control system based on large model and interface joint drive according to claim 1 is characterized in that: The instruction parsing includes user task instructions, which can be input through various methods such as voice, text, image, action, etc., supporting multimodal interaction.

3. The terminal control system based on large model and interface joint drive according to claim 1 is characterized in that: The peripheral intelligent body carrier includes a peripheral dominant control process, and the peripheral dominant control process has a permanent online capability and can wake up the terminal and initiate task execution at any time.

4. The terminal control system based on large model and interface joint drive according to claim 1 is characterized in that: The task generation module includes task reasoning and scheduling, which are controlled by the peripheral intelligent body carrier. The terminal is only responsible for state collection and action execution, reducing resource usage.

5. The terminal control system based on large model and interface joint drive according to claim 1 is characterized in that: The terminal interface status collection includes a terminal agent program, and the terminal agent program cooperates to extract interface structure information to provide context decision basis for large model reasoning.

6. The terminal control system based on large model and interface joint drive according to claim 1 is characterized in that: The feedback module includes a model that updates judgments in real time based on the results of each round, generates subsequent operation instructions as needed, and supports the decomposition of complex tasks.

7. The terminal control system based on large model and interface joint drive according to claim 1 is characterized in that: The task execution control module consists of lightweight peripherals and a large model, and has good cross-platform adaptability and deployment flexibility.

8. A terminal control method based on a combined drive of a large model and an interface, the method comprising the following steps: S1: Task instruction acquisition stage Users input task instructions through voice, text, images, touch, etc. The peripheral device continuously senses and identifies task triggering conditions, such as keywords, image features, or gestures, and supports multimodal interaction. S2: Terminal status awareness stage The peripheral activates the terminal control agent through the communication interface to obtain the current interface status, including the foreground app, interface structure hierarchy, focus elements, etc. The system is controlled by the peripheral and has the ability to stay online, so it can wake up the terminal and initiate task execution at any time. S3: Intent parsing and action generation stage The peripheral packages the command data and terminal status and sends them to a remote or local large language model. The model combines contextual reasoning to generate the next operation instruction. Task reasoning and scheduling are controlled by the peripheral, and the terminal is only responsible for status collection and action execution, reducing resource usage. S4: Operation execution and loop judgment stage The peripheral or terminal agent executes the operation instructions generated by the model, and after execution, obtains the interface status again as the input for the next round of reasoning, forming a loop until the task is completed; the terminal agent cooperates to extract interface structure information to provide contextual decision-making basis for large model reasoning. S5: Task completion and feedback stage After the task is completed, the system will feedback the results to the user through voice synthesis, interface display or other means; after the task is completed, the results can be fed back to the user through voice, graphics, vibration, etc., forming a closed-loop intelligent experience.