Auxiliary operation method, electronic equipment, storage medium and computer program product
By performing multiple rounds of auxiliary operations on the terminal device and inputting the information of each round of operations into the AI large model to obtain output instructions, the problem of difficulty in executing complex tasks of auxiliary operation functions in the existing technology is solved, and more efficient auxiliary operation logical coherence and reduced reasoning costs are achieved.
Patent Information
- Application Number
- CN202510889854.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-19
AI Technical Summary
In existing technologies, the auxiliary operation functions of terminal devices are difficult to effectively perform complex tasks due to low intelligence levels and insufficient computing power, and the proxy models face high inference costs and computing burdens when deployed on the terminal side.
In response to receiving auxiliary operation instructions for the target task, at least one round of auxiliary operations is performed for the target task, and screen display information is obtained in each round of operation. The input information is input into the AI big model, and output information is obtained, including operation action descriptions and instructions. The operation is executed and the action description is written into the operation trajectory information until the task is completed.
Ensure the coherence and consistency of auxiliary operation logic, reduce the storage and computing requirements of historical information irrelevant to auxiliary operation decisions, thereby reducing the inference cost of large AI models and improving their efficiency in terminal deployment.
Smart Images

Figure CN120670078A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of terminal technology, and in particular to an auxiliary operation method, electronic device, storage medium, and computer program product. Background Art
[0002] With the continuous development and widespread adoption of intelligent terminal devices, assisted operation functions have also attracted widespread attention. Limited by the relatively low intelligence level and insufficient computing power of terminal devices, assisted operation functions are often limited to simple command parsing. Although proxy models can achieve end-to-end task execution, user commands are often complex, requiring proxy models to perform multiple inferences to complete the task. This increases the computational burden of model inference and poses significant challenges to device-side deployment. Summary of the Invention
[0003] The embodiments of the present application provide an auxiliary operation method, an electronic device, a storage medium, and a computer program product to at least solve the problems of high inference cost and difficult terminal-side deployment of the proxy model during the auxiliary operation process.
[0004] In order to solve the above technical problems, this application is implemented as follows: In the first aspect, an embodiment of the present application provides an auxiliary operation method, comprising: in response to receiving an auxiliary operation instruction corresponding to a target task, performing at least one round of auxiliary operation for the target task; wherein, in each round of auxiliary operation, screen display information is obtained; by inputting input information into an AI big model, output information of the AI big model is obtained, the input information includes the auxiliary operation instruction, the screen display information and the operation trajectory information corresponding to the previous round of auxiliary operation, and the output information includes the operation action description and operation instruction of the current round of auxiliary operation; executing the operation action corresponding to the operation instruction, and writing the operation action description into the operation trajectory information, obtaining the operation trajectory information corresponding to the current round of auxiliary operation, until the current round of auxiliary operation meets the termination condition and the auxiliary operation is completed; the termination condition includes at least one of the following: the number of rounds of the auxiliary operation exceeds a preset number threshold, and the task status of the target task is determined to be a completed state.
[0005] In a second aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method described in the first aspect above are implemented.
[0006] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect above are implemented.
[0007] In a fourth aspect, an embodiment of the present application provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer performs the steps of the method described in the first aspect above.
[0008] In an embodiment of the present application, in response to receiving an auxiliary operation instruction corresponding to a target task, at least one round of auxiliary operation is performed for the target task; wherein, in each round of auxiliary operation, screen display information is obtained; by inputting input information into the AI big model, output information of the AI big model is obtained, the input information includes the auxiliary operation instruction, the screen display information and the operation trajectory information corresponding to the previous round of auxiliary operation, and the output information includes the operation action description and operation instruction of the current round of auxiliary operation; the operation action corresponding to the operation instruction is executed, and the operation action description is written into the operation trajectory information to obtain the operation trajectory information corresponding to the current round of auxiliary operation, until the current round of auxiliary operation meets the termination condition and the auxiliary operation is completed. In this way, by writing the operation action description in each round of auxiliary operation into the operation trajectory information, in the current round of auxiliary operation, the operation trajectory information input by the AI big model includes the operation action description of all rounds of auxiliary operations before the current round of auxiliary operation, which can ensure the coherence and consistency of the auxiliary operation logic, while reducing the storage requirements and computing burden of historical information irrelevant to the auxiliary operation decision, thereby reducing the inference cost of the AI big model and improving the efficiency of the AI big model on the end side.
[0009] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0011] Figure 1 A schematic diagram showing a flow chart of an auxiliary operation method provided in some embodiments of the present application is shown; Figure 2 A schematic diagram showing input information and output information of a multimodal large model provided in some embodiments of the present application; Figure 3 A schematic diagram showing an interactive operation interface provided by some embodiments of the present application is shown; Figure 4 Schematic diagrams showing interactive operation interfaces provided by other embodiments of the present application; Figure 5 A schematic diagram showing a flow chart of auxiliary operation methods provided in other embodiments of the present application is shown; Figure 6 One of the interactive methods provided by some embodiments of the present application is shown; Figure 7 The second schematic diagram of the interaction method provided in some embodiments of the present application is shown; Figure 8 A schematic diagram of an auxiliary operation process based on an agent framework provided in some embodiments of the present application is shown; Figure 9 A schematic diagram showing the structure of the proxy framework provided by some embodiments of the present application is shown; Figure 10 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0012] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0013] Terminal device assisted operation, also known as GUI Agent or OS Agent. There are generally two approaches to the workflow of agent architectures. One is the agent framework, which leverages the understanding and reasoning capabilities of high-level foundational models to enhance the flexibility of task execution. Multiple models are combined into action agents through pre-defined roles. Although agent frameworks offer greater adaptability than rule-based systems, they still rely on manually defined workflows to build action processes. Whenever tasks, interfaces, or usage scenarios change, developers need to redesign or extend the workflow's manual rules or prompts, an error-prone and labor-intensive process. Furthermore, complex tasks often require multiple modules (such as visual parsing, memory storage, and long-term planning) to coordinate through prompts or bridge code. Inconsistencies or errors in any module can derail the entire process, and diagnosing these issues often requires domain experts to debug the process. These factors often lead to agent frameworks facing drawbacks in practical applications, such as system fragility, high maintenance costs, high risk of module incompatibility, and poor generalization.
[0014] The second type is the native agent model. Specifically, this involves embedding workflow knowledge directly into the policy model through targeted learning. In this paradigm, tasks are learned and executed in an end-to-end manner. By unifying capabilities such as perception, reasoning, memory, and action in a continuously evolving model through model training, the model maintains the ability to continuously self-improve through data-driven learning. Compared to the first approach, the native agent framework is fundamentally data-driven, allowing the agent to seamlessly adapt to new tasks, interfaces, or user needs. It is more suitable for models to continuously improve themselves through training, thereby seamlessly adapting to changing tasks, interfaces, or user needs. At the same time, there is no need to rely on manually written prompts or predefined rules, significantly reducing manual engineering.
[0015] However, due to the complexity of user commands, native proxy models often require multiple operations to complete tasks. In a common implementation, the proxy framework interacts with the large model through multi-round dialogues. Within this multi-round interaction paradigm, the large model must obtain input information such as user commands and screenshots during each round of operation, and output the user's thought process, decision, and operational instructions. As the number of interactions increases, this historical operation information grows, occupying a large number of tokens. This is particularly true for image data, where the number of tokens occupied can range from hundreds to thousands depending on the resolution, and this phenomenon is particularly pronounced in the case of multi-round dialogues. To maintain decision coherence and consistency, the large model must recalculate this historical information during each inference. This places a significant computational burden on the model's inference, resulting in high computational load and inference costs, and complicating its on-device deployment.
[0016] In response to the problems existing in the auxiliary operation process of the above-mentioned large model, an embodiment of the present application provides an auxiliary operation method. This method writes the operation action description in each round of auxiliary operation into the operation trajectory information. In the current round of auxiliary operation, the operation trajectory information input by the AI large model includes the operation action description of all rounds of auxiliary operations before the current round of auxiliary operation, which can ensure the coherence and consistency of the auxiliary operation logic, while reducing the storage requirements and computing burden of historical information irrelevant to the auxiliary operation decision, thereby reducing the inference cost of the AI large model and improving the efficiency of the AI large model on the end side.
[0017] See also Figure 1 , Figure 1The following is a flow chart illustrating an auxiliary operation method provided by some embodiments of the present application. The method may be performed by a terminal device equipped with an operating system, including a mobile phone, tablet computer, television, wearable device, vehicle-mounted device, augmented reality / virtual reality device, laptop computer, ultra-mobile personal computer, netbook, personal digital assistant, etc. As shown in the figure, the auxiliary operation method 100 may include the following steps: Step 101: in response to receiving an auxiliary operation instruction corresponding to a target task, performing at least one round of auxiliary operation on the target task; Among them, in each round of auxiliary operation, screen display information is obtained; by inputting input information into the AI big model, the output information of the AI big model is obtained, the input information includes auxiliary operation instructions, screen display information and operation trajectory information corresponding to the previous round of auxiliary operations, and the output information includes the operation action description and operation instructions of the current round of auxiliary operations; the operation action corresponding to the operation instruction is executed, and the operation action description is written into the operation trajectory information to obtain the operation trajectory information corresponding to the current round of auxiliary operations, until the current round of auxiliary operations meets the termination conditions and the auxiliary operations are completed; the termination conditions include at least one of the following: the number of rounds of auxiliary operations exceeds the preset number threshold, and the task status of the target task is determined to be completed.
[0018] Exemplarily, the screen display information in the above step 101 can be a screenshot or information obtained from a screen layout file; the AI big model in the above step 101 includes but is not limited to a multimodal big model, a language big model, and a combination model of a speech big model and a multimodal model; the number threshold in the above step 101 can be set to 15 times, 20 times, 50 times, etc., and can be set according to actual needs.
[0019] In one exemplary embodiment, a user can trigger a terminal device to enter an auxiliary operation process using a preset wake-up command. After receiving the wake-up command, the terminal device obtains auxiliary operation instructions corresponding to the target task based on the wake-up command. Specifically, the terminal device can determine the wake-up command by: determining the wake-up command based on detected biometric information such as voice or gestures; or determining the wake-up command based on the user's touch operation on a physical key or virtual key.
[0020] After receiving the wake-up instruction, the terminal device obtains the auxiliary operation instruction of the target task corresponding to the wake-up instruction. For example, if the user has set an auxiliary operation shortcut key in advance, the terminal device directly enters the auxiliary operation process of the target task after the user clicks the shortcut key; for another example, the terminal device can generate an auxiliary operation interface, and obtain instruction information based on the auxiliary operation interface, and determine the auxiliary operation instruction corresponding to the target task according to the instruction information.
[0021] After the terminal device receives the auxiliary operation instruction corresponding to the target task, it performs at least one round of auxiliary operation for the target task, and obtains screen display information in each round of auxiliary operation; inputs the obtained auxiliary operation instruction, screen display information, and operation trajectory information corresponding to the previous round of auxiliary operation into the AI big model, and obtains the operation action description and operation instruction of the current round of auxiliary operation output by the AI big model; executes the operation action corresponding to the operation instruction, and writes the operation action description output by the AI model into the operation trajectory information to obtain the operation trajectory information corresponding to the current round of auxiliary operation.
[0022] It can be understood that since the operation action description output by the AI big model is written into the operation trajectory information in each round of auxiliary operation, in the current round of auxiliary operation, the operation trajectory information input by the AI big model includes the operation action descriptions of all rounds of auxiliary operations before the current round of auxiliary operation. In this way, the AI big model can obtain coherent auxiliary operation logic from the operation trajectory information without the need to store and calculate historical information that is not related to the decision of this round of auxiliary operation. For example, in the 5th round of auxiliary operation, the AI big model can obtain the operation action descriptions corresponding to the previous 1st, 2nd, 3rd, and 4th rounds of auxiliary operations from the operation trajectory information corresponding to the 4th round of auxiliary operation, without the need to store and infer historical screenshots, historical operation instructions, historical thinking processes and other information corresponding to the previous 4 rounds of auxiliary operations, thereby reducing the number of tokens used in AI model reasoning and reducing the reasoning cost of the AI big model.
[0023] When the current round of auxiliary operations meets the termination conditions, the auxiliary operations are completed. The termination condition may be that the number of rounds of the auxiliary operations exceeds a preset threshold. For example, the preset threshold is set to 50 times. If the number of rounds of the auxiliary operations is 50 times, the auxiliary operations are completed. In this way, unnecessary repeated operations or falling into an infinite loop can be avoided. The termination condition may be that the task status of the target task is determined to be a completed state. For example, when the operation action is to indicate that the task has been completed, the task status of the target task is determined to be a completed state, and the auxiliary operation is completed at this time.
[0024] In some embodiments, the output information further includes memory information associated with the auxiliary operation instruction determined based on the screen display information; in the above step 101, after inputting the input information into the AI big model to obtain the output information of the AI big model, the step further includes: Write the memory information into the operation track information.
[0025] It can be understood that since in each round of auxiliary operation, the memory information associated with the auxiliary operation instructions determined by the AI big model based on the screen display information is written into the operation trajectory information, in the current round of auxiliary operation, the operation trajectory information input by the AI big model contains the memory information of all rounds of auxiliary operations before the current round of auxiliary operation. In this way, the reasoning information of the AI big model for the target task can be enriched, and the coherence and consistency of the auxiliary action decision can be further improved.
[0026] In some possible implementations, in order to improve the accuracy of AI large model reasoning, the output information of the AI large model also includes the thinking process of obtaining output information based on input information.
[0027] In an exemplary embodiment, taking the AI big model as a multimodal big model as an example, Figure 2 As shown, the input information of the multimodal large model includes: system prompts, user instructions (i.e., the auxiliary operation instructions mentioned above), screenshots (i.e., the screen display information mentioned above), historical operation records, historical memory, and model operation instructions. The output information of the multimodal large model includes: thought process, current round memory, current round operation action description, and current round operation instructions. The current round memory refers to the memory information corresponding to the auxiliary operation in the current round, and the historical memory refers to all memory information before the current round auxiliary operation written into the operation trajectory information. After each round of reasoning, the current round operation instructions are executed on the terminal device, thereby obtaining an updated screenshot. In the next round of the iterative single-round dialogue, the screenshots, historical records, and historical memory in the input information are updated. Specifically, the updated screenshot is the new screenshot obtained after the current round of actions are executed. The current round operation action description is added to the historical record, and the current round memory (if any) is added to the historical memory. Using this method, iterative reasoning operations are performed to ultimately complete the user task. The historical operation instructions specifically refer to the record of the operation action description output by the model in each round, and the historical memory refers to the record of the current round memory output by the model in each round. It should be noted that during each round of reasoning, the model will decide whether to output the memory of this round associated with the user's instructions based on the target task and the current screen display information.
[0028] It should be emphasized that, unlike single-round dialogue in the general sense, the iterative single-round dialogue proposed in the embodiment of the present application is also a calling method of single-round dialogue in terms of its expression, but it also contains information on the interactions of multiple iterative operations, such as historical operation records and historical memories. This information is continuously updated with the iterative operations, while ensuring the consistency and coherence of the multimodal large model decision-making, effectively reducing the length of the text description used, reducing the reasoning cost and delay.
[0029] The following uses the user instruction "Order me a latte" as an example to illustrate the input and output information of the multimodal model. The input and output examples are shown in the following table: Table 1. Examples of inputs and outputs of a large multimodal model
[0030] Among them, CLICK: (456,789) represents the screen coordinate position corresponding to the click operation.
[0031] In some embodiments, in step 101 above, executing an operation corresponding to the operation instruction includes: An operation type of the operation instruction is determined, where the operation type includes an execution operation and an interaction operation; and an operation action corresponding to the operation instruction is performed according to the operation type.
[0032] In an exemplary embodiment, in each round of auxiliary operation, the AI big model can output an execution-type action such as CLICK: (456,789) in Table 1. For this execution-type action, the operation action corresponding to the operation instruction can be executed on the terminal device; the AI big model can also output an interactive operation such as "Interaction: Would you like to choose a large cup or a small cup?" For this operation-type action, an interactive interface for selecting "large cup" or "small cup" can be displayed on the interactive terminal device, and the user's supplementary information for the target task is determined based on the user's selection operation on the interactive interface.
[0033] In some possible implementations, the above-mentioned execution of an operation action corresponding to the operation instruction according to the operation type includes: In response to the operation type being an execution-type operation, an operation action corresponding to the operation instruction is executed by calling an interface of the terminal device.
[0034] In an exemplary embodiment, when the operation instruction output by the AI large model is an execution-type operation, for example, [Operation instruction for this round]: CLICK: (456, 789), the interface of the terminal device is called to execute the operation action corresponding to the operation instruction, for example, performing a click operation at the screen coordinate position (456, 789).
[0035] In some other possible implementations, the above-mentioned execution of an operation action corresponding to the operation instruction according to the operation type includes: In response to the operation type being an interactive operation, an operation action corresponding to the operation instruction is executed. The operation action includes a display action of annotation information. The annotation information is determined based on the operation attribute information of all target controls in the screen layout file corresponding to the screen display information. The operation attribute information is used to characterize the operations that can be performed on the target controls.
[0036] In an exemplary embodiment, when the operation instruction output by the AI large model is an interactive operation, for example, [Operation Instructions for This Round]: Interaction: Would you like a large or small cup?, annotation information corresponding to the operation instruction is displayed. This annotation information can be determined based on the operation attribute information of all target controls in the screen layout file corresponding to the screen display information. The operation attribute information is used to represent the operations that can be performed on the target control, such as clicking, swiping up, swiping down, etc.
[0037] In some possible implementations, the operational attribute information of all the target controls described above is determined in the following manner: By traversing all controls in the screen layout file corresponding to the screen display information, the operation attributes of each control are obtained; candidate controls whose attribute values of the operation attributes are preset values are selected from all controls in the screen layout file; and the operation attribute information of all target controls is determined based on the candidate controls and the operation attributes of the candidate controls.
[0038] In an exemplary embodiment, all elements in a screen layout file of a terminal device may be sequentially traversed, where each element corresponds to a control on the screen, and operational attributes of each control may be obtained. These operational attributes may include "focusable" and "clickable" attributes. During the traversal process, if the attribute value of an element's "focusable" or "clickable" attribute is a preset value, such as "true," the element is recorded in a target list. The element attributes that need to be recorded include: "clickable," "scrollable," "long-clickable," "text," "bounds," and "class." Based on all candidate controls in the target list and the operational attributes of each candidate control, the operational attribute information of all target controls is determined.
[0039] In some possible implementations, the above-mentioned determination of the operational attribute information of all target controls based on the candidate controls and the operational attributes of the candidate controls includes: The first control and the second control in the candidate controls are merged to obtain a third control, wherein the overlapping surface between the first control and the second control meets preset conditions, and the preset conditions include that the ratio of the overlapping surface to the area of the first control is greater than a preset ratio threshold, and the ratio of the overlapping surface to the area of the second control is greater than the preset ratio threshold; according to the operation properties of the first control and the operation properties of the second control, the operation attribute information of the third control is determined; according to all third controls and the operation attribute information of the third control, the operation attribute information of all target controls is determined.
[0040] The above-mentioned preset ratio threshold value can be set to 90%, 85%, 70%, etc., and the preset ratio threshold value can be set according to actual needs.
[0041] In an exemplary embodiment, after completing the traversal of all controls in the screen layout file, the elements in the target list are compared pairwise. If the overlapping area of two controls exceeds 90% of the area of each control at the same time, the two controls are merged into one control. The position of the third control after the merger can adopt the position of the overlapping part, and the text is the text merge of the text attributes of the two original controls, and the various control attributes mentioned above are merged in an "or" logical manner. In particular, if the class value of a control is android.widget.EditText4, the control is editable, and the control remains editable after the merger. The position of the border after the merger is the border position of the overlapping area of the two controls before the merger.
[0042] In some possible implementations, the above-mentioned determination of the annotation information based on the operation attribute information of all target controls includes: Obtain the operation identifier corresponding to the operation attribute information of each target control, and determine the annotation information according to each target control and the operation identifier corresponding to each target control; or generate candidate operation items according to the operation attribute information of all target controls, and determine the annotation information according to the candidate operation items.
[0043] In an exemplary embodiment, an operation identifier corresponding to the operation attribute information of the target control is obtained, for example, an up arrow or a down arrow is used to indicate sliding, and a “ " indicates clickable; determine the labeling information based on each target control and the operation identifier corresponding to each target control. For example, based on the target control "big cup" and the operation identifier corresponding to the target control " ", determine the annotation information, in addition, you can also use a dotted box to indicate the clickable area, by executing the display action of the annotation information, you can get the following Figure 3 The interactive operation interface shown.
[0044] In another exemplary embodiment, candidate operation items are generated based on the operation attribute information of all target controls. For example, the operation attribute information of the target control "large cup" is "clickable", and the operation attribute information of the target control "medium cup" is "clickable". The generated candidate operation items are "1. Click the large cup specification option; 2. Click the medium cup specification option", etc., and the annotation information is determined based on the candidate operation items. By executing the display action of the annotation information, the following is obtained: Figure 4 The interactive operation interface shown.
[0045] In this way, during the execution of the task, situations such as further communication with the user are detected, and the positions of all target controls on the current screen and their allowed operations are obtained by parsing the screen layout file of the terminal device (such as an XML file). While interacting with the user, the candidate operation set is displayed to the user on the screen using operation identifiers or candidate operation items, thereby assisting the user to express his or her intention quickly and accurately, helping the user to narrow the set of optional candidate operations, and guiding the user to express his or her demands in more clear language, thereby improving the interaction efficiency and the success rate of task execution.
[0046] In some possible implementations, after executing the operation corresponding to the operation instruction, the following steps are further included: Based on the annotation information, a candidate operation set is determined; based on the candidate operation set and the operation information returned based on the annotation information, supplementary information for the target task is determined; the supplementary information is written into the operation trajectory information to obtain the operation trajectory information corresponding to the auxiliary operation of the current round.
[0047] In one exemplary embodiment, a candidate operation set is determined based on the annotation information. The user then returns corresponding operation information based on the annotation information displayed on the screen, for example, clicking or selecting "Large." Based on the candidate operation set and the operation information, supplementary information specific to the target task is determined. For example, an interaction is initiated by asking the user: "Would you like a large or medium cup?" The user responds: "Large." This supplementary information is written into the operation trajectory information, resulting in the operation trajectory information corresponding to the current round of auxiliary operations.
[0048] Figure 5 Schematic diagrams of the auxiliary operation methods provided in other embodiments of the present application are shown. Figure 5 As shown, the auxiliary operation method may include the following steps: Step 501: Acquire auxiliary operation instructions and screen display information in the current state, where the screen display information includes a screenshot and screen layout information, which can be acquired from a screen layout file; Step 502: Determine the current state of the operation instruction based on the auxiliary operation instruction, screenshot, and screen layout information; if the operation type of the operation instruction is an execution operation, proceed to step 303; if the operation type of the operation instruction is an interaction operation, proceed to step 304; Step 503: Execute the operation corresponding to the operation instruction on the terminal device and update the terminal information; Step 504: interact with the user by displaying the annotation information corresponding to the operation instruction to obtain the user's supplementary information on the target task; Step 505: If the number of rounds of auxiliary operations currently being performed does not reach the preset threshold and the operation action is not to end the task, the relevant information is updated, including the current screenshot, screen layout information, and the user's supplementary information on the target task, and the next round of auxiliary operations is performed; otherwise, the task ends.
[0049] In this way, the user's true expression intention can be accurately obtained, the success rate of task execution can be improved, and thus user satisfaction can be improved.
[0050] In some embodiments, in step 101 above, before performing at least one round of auxiliary operations on the target task in response to receiving the auxiliary operation instruction corresponding to the target task, the process further includes: Receive an initial operation instruction for a target task; obtain information to be supplemented for the target task by inputting the initial operation instruction into an AI big model; wherein the information to be supplemented is determined by the AI big model based on prior knowledge in training data, and the training data is determined based on multiple operation steps for operating the terminal device for the target task; generate a query statement based on the information to be supplemented; obtain an answer statement returned based on the query statement; determine the auxiliary operation instruction corresponding to the target task based on the initial operation instruction and the answer statement.
[0051] In an exemplary embodiment, Figure 6 and Figure 7 As shown, an initial operation instruction for a target task is received, for example, the initial operation instruction is "order me a cup of milk tea"; by inputting the initial operation instruction into the AI big model, the AI big model determines the information to be supplemented for the target task based on the initial operation instruction and the prior knowledge in the training data. For example, the AI big model determines based on the prior knowledge in the training data that to complete the task of "ordering a cup of milk tea", the user is also required to supplement key information such as "milk tea shop", "full sugar or half sugar", "large cup or small cup", and uses these key information as the information to be supplemented for the task; based on the information to be supplemented, a query statement is generated, for example, "Which store's milk tea would you like to order?", "Would you like full sugar or half sugar for your milk tea?"; an answer statement returned by the user based on the query statement is obtained; based on the initial operation instruction and the answer statement, the auxiliary operation instruction corresponding to the target task is determined, for example, the auxiliary operation instruction is "Please help me order a cup of milk tea at XXX milk tea shop, choose full sugar for milk tea and a large cup for the specification."
[0052] In this way, by constructing the expected training data to train the AI big model, the AI big model's understanding of the target task is strengthened, so that the AI big model can determine the key information of the target task, and then independently decide the timing and method of interaction with the user, guiding the user to provide complete supplementary information at one time, avoiding the phenomenon of asking the user a large number of questions during the task execution process, reducing the number of interactions, and improving interaction efficiency.
[0053] In some possible implementations, the training data may be determined in the following manner: Get multiple steps for operating the terminal device for the target task and describe the process in the format of a JSON file. The annotation data example is shown in the following table: Table 2. Annotated data examples
[0054] The annotation content in Table 2 above includes the user's natural language description of the task, as well as screenshots of each round of operation, thinking process, current round memory, current round operation action description, current round operation instructions, etc. The format of the input information and output information of the multimodal large model is shown in Table 1. By reorganizing the key fields of the annotation data, the annotation data provided in Table 2 is converted into the format of the training data in Table 1. Specifically, each step will be converted into a piece of training data. In each piece of training data, the input of the multimodal large model includes: system prompts, user instructions, screenshots, historical operation records, historical memories, operation instructions and other information, among which the historical operation records are the sequential merger of the "current round operation action description" fields in the previous operation steps, and the historical memory is the sequential merger of the "current round memory" fields existing in the previous operation steps. The output of the multimodal large model includes: thinking process, current round memory, current round operation action description, and current round operation instructions.
[0055] QwenVL2.5-7B is used as the basic multimodal large model for training. During the training process, the multimodal large model reads the image and text information in the input information, uses the output information as the target value (label), and uses the cross-entropy loss function (Loss) to supervise the model's output results, resulting in the trained AI large model.
[0056] In an exemplary embodiment, the present application provides an agent framework, such as Figure 8 As shown in the figure, the auxiliary operation process based on the agent framework includes the following steps: Step 801: The agent framework obtains the terminal device status, user instructions, and user interaction information. The terminal device status includes screen display information and screen layout information. The screen display information, such as a screenshot, can be obtained from a screen layout file. The user instructions are auxiliary operation instructions. The user interaction information is operation track information, specifically including historical operation records, historical memory, supplementary information, etc. Step 802: The agent framework constructs input information for the AI large model based on the acquired information. The input information includes: system prompts, user instructions, screenshots, historical operation records, historical memory, and model operation instructions. Step 803: The agent framework calls the AI big model based on the input information; Step 804: The agent framework obtains the output information of the AI large model, which includes: the thinking process, the memory information of this round, the description of the operation action of this round, and the operation instruction of this round; Step 805: The proxy framework parses the output information of the AI large model and records the operation trajectory information; Step 806: The agent framework calls the interface of the terminal device according to the operation instruction in the output information, and performs the operation action corresponding to the operation instruction on the terminal device.
[0057] like Figure 9 As shown, the above-mentioned agent framework includes a large model input construction module 910, a large model output parsing module 920, a terminal device operation module 930 and a process control module 940; The large model input construction module 910 is used to obtain terminal device status, user instructions and user interaction information; and construct input information of the AI large model based on the obtained information; Large model output parsing module 920, used to parse the output information of the AI large model; The terminal device operation module 930 is used to call the terminal device interface according to the operation instruction in the output information and perform the operation action corresponding to the operation instruction on the terminal device; The process control module 940 is used for module scheduling, data organization and transmission of the large model input construction module, the large model output analysis module, and the terminal device operation module.
[0058] In an exemplary embodiment, taking the above-mentioned agent framework as an auxiliary operation assistant for a smartphone as an example, the auxiliary operation steps based on the agent framework are described as follows. The system structure is: a smartphone with AI model reasoning capabilities, agent framework software running on the smartphone, and a multimodal large model deployed and running on the smartphone platform. The method includes the following steps: Step 9101: Auxiliary operation assistant wake-up: The user can wake up the interactive interface through voice, and then input auxiliary operation instructions through voice or text, such as Figure 6 or Figure 7 As shown; Step 9102: Obtaining the current state information of the mobile phone: The agent framework obtains auxiliary operation instructions, the current screenshot and screen layout information, and the user's supplementary information for the target task (if any). The screenshot is the screen display information, and the screen layout information can be obtained from the screen layout file. Step 9103: Constructing the multimodal large model input: The agent framework organizes the auxiliary operation instructions, the current screenshot and screen layout information, the user's supplementary information on the target task (if any), and the operation history (if any) into a specific multimodal large model input format; Step 9104: Call the multimodal large model to obtain output: The agent framework calls the multimodal large model to perform reasoning in the form of a single-round dialogue. The input information is: screenshot, screen layout description, auxiliary operation instructions, supplementary information, historical operations, and output requirements. The output information is: thinking process, current operation action description, and operation command parameters corresponding to the current operation instruction. Step 9105: Analyze and record the output information of the multimodal large model: Analyze the various contents in the output information of the multimodal large model, including the thinking process, the description of the current operation action, the operation command parameters corresponding to the current operation instruction, and record the operation trajectory information; Step 9106: Execute the operation on the terminal device: The agent framework calls the terminal device interface based on the parsed operation command parameters, executes the operation (including executable operations and operations that interact with the user) on the terminal device, re-acquires the terminal device information after the operation is executed, and enters the next iteration; Step 9107: Execute execution operations on the terminal device: Execution operations are actions that simulate human gestures, including clicks, slides, long presses, text input, dragging, home buttons, back buttons, task buttons, opening apps, volume buttons, power buttons, confirmation buttons, and waiting. These actions simulate real user execution. The proxy framework can implement these operations by calling the interfaces provided by the terminal device. Step 9108: Perform interactive operations on the terminal device: Interactive operations include task status description (completed / incompleted) and interaction with the user (obtaining supplementary information from the user regarding ambiguous instructions). If, during the execution of the task, it is detected that the user's instructions are unclear, the user needs to provide additional information, the auxiliary operation status is fed back to the user, or further communication with the user is required during the auxiliary operation, such as Figure 6 or Figure 7As shown, the system interacts with the user through text or voice to supplement the information of the auxiliary operation instructions; Step 9109: Generating a candidate action set: When the action is identified as an interactive action, the agent framework obtains the screen layout description XML file from the operating system. The specific process is as follows: 1. Traverse all nodes in the XML file using a pre-order traversal method, where each node corresponds to a control on the screen; 2. During the pre-order traversal process, if the value of a node's "focusable" attribute or "clickable" attribute is "true", the node is recorded in a target list. The recorded attributes include: "clickable", "scrollable", "long-clickable", "text", "bounds", and "class". 3. After completing the above traversal, compare the nodes in the above target list one by one. If the overlapping area of two controls exceeds 90% of their respective areas, the two controls are merged into a single control. The position of the merged control adopts the position of the overlapping part, and the text is the merged text of the two original controls. The aforementioned control attributes are merged in an "or" logical manner. In particular, if the class value of a control is android.widget.EditText4, then the control is editable; 4. Finally, summarize all candidate operations of the page based on all controls and their attributes in the target list; 5. Figure 3 As shown, it can be displayed on the screen during interaction in the form of graphic annotations, or as Figure 4 As shown, it is presented to the user in the form of text description; Step 9110: Interaction of candidate class operations: Figure 3 As shown, while interacting with the user by voice, candidate execution operations are displayed to the user by using graphic annotations on the screen, and the user can select the candidate execution operations by voice or text; Figure 4 As shown, by displaying natural language descriptions of all execution operations allowed on the current screen to the user for selection, the user can select the execution operation through voice or text; Step 9111: Task end condition: If the total number of auxiliary operations currently executed does not reach the preset number threshold and the operation action is not to end the task, then update the current screenshot, screen layout information, and user's supplementary information for the target task, and then iterate. Otherwise, the task ends.
[0059] In an exemplary embodiment, taking the above-mentioned agent framework as an example of a smart TV operation assistant, the above-mentioned auxiliary operation steps based on the agent framework are described as follows. The system structure is: a smart TV, an agent framework software running on the smart TV, and a multimodal large model deployed and running on a server. The method includes the following steps: Step 9201: Auxiliary operation assistant wake-up: The user can use the wake-up button on the smart TV remote control to wake up the smart TV operation assistant and input auxiliary operation instructions through voice; Step 9202: Obtain the current state of the smart TV: The agent framework obtains auxiliary operation instructions, the current screenshot and screen layout information, and the user's supplementary information on the target task (if there is interaction). The screenshot is the screen display information, and the screen layout information can be obtained from the screen layout file. Step 9203: Generate candidate operation set: The agent framework reads the screen layout file (e.g., an XML file), identifies the operable elements on the current screen and the operations allowed for each element, summarizes the candidate operations corresponding to all elements, and adds interactive operations, including task status descriptions (completed / incompleted) and user interactions (obtaining supplementary information from the user regarding fuzzy instructions), to generate a candidate operation set. Step 9204: Constructing the multimodal large model input: The agent framework organizes the auxiliary operation instructions, the current screenshot and screen layout information, the user's supplementary information on the target task (if any), the operation history (if any), and the candidate operation set into a specific multimodal large model input format; Step 9205: Calling the multimodal large model to obtain output information: The agent framework calls the multimodal large model to perform reasoning in a single-round dialogue. The input information includes: screenshot, screen layout description, auxiliary operation instructions, supplementary information, historical operations, candidate operation sets, and output requirements. The output information includes: the thinking process and the current operation instructions. For example, the operation instruction is the sequence number of the selected candidate operation. Step 9206: Parse and record the output information of the multimodal large model: The agent framework parses the selected candidate operation sequence number in the large model output information and parses out the operation command parameters corresponding to the candidate operation; Step 9207: Generating a candidate action set: When the action is identified as an interactive action, the proxy framework obtains the screen layout description XML file from the operating system. The specific process is as follows: 1. Traverse all nodes in the XML file using a pre-order traversal method, where each node corresponds to a control on the screen; 2. During the pre-order traversal process, if the value of a node's "focusable" attribute or "clickable" attribute is "true", the node is recorded in a target list. The recorded attributes include: "clickable", "scrollable", "long-clickable", "text", "bounds", and "class". 3. After completing the above traversal, compare the nodes in the above target list one by one. If the overlapping area of two controls exceeds 90% of their respective areas, the two controls are merged into a single control. The position of the merged control adopts the position of the overlapping part, and the text is the merged text of the two original controls. The aforementioned control attributes are merged using the "OR" logic. In particular, if the class value of a control is android.widget.EditText4, then the control is editable; 4. Finally, summarize all candidate operations of the page based on all controls and their attributes in the target list; 5. Figure 3 As shown, it can be displayed on the screen during interaction in the form of graphic annotations; Step 9208: Perform interactive operations on the smart TV: If it is detected during the execution of the task that the user's instructions are unclear, the user needs to supplement relevant information, the auxiliary operation status is fed back to the user, or further communication with the user is required during the auxiliary operation, the user is guided to interact and complete the information supplement through text or voice, such as Figure 6 or Figure 7 At the same time, by using graphic annotations on the screen to show the user the operations allowed for each element on the screen, the user can provide feedback on the execution type of the operation selected by voice or text, such as Figure 3 As shown; Step 9209: Task end condition: If the number of auxiliary operations currently executed has not reached the preset threshold and the operation action is not to end the task, then update the current screenshot, screen layout information, and user's supplementary information for the target task, and then perform the iterative operation. Otherwise, the task ends.
[0060] In an exemplary embodiment, the auxiliary operation method provided by the embodiment of the present application is also applicable to multiple application scenarios such as the elderly, the visually impaired, office workers commuting, and driving. For the elderly, they often face difficulties in operating mobile phones, such as setting alarms and adjusting volume. Through the auxiliary operation method provided by the embodiment of the present application, users only need to issue instructions to automatically complete the corresponding operations, effectively lowering the threshold for using mobile phones and improving the user experience. For the visually impaired, the auxiliary operation method provided by the embodiment of the present application helps them easily operate their mobile phones, such as opening reading software and adjusting font size, ensuring the smoothness of their information acquisition and communication. For office workers commuting, their hands are occupied during the commute, making it difficult for them to manually operate their mobile phones. Through voice commands, they can complete tasks such as checking schedules and replying to simple messages, making full use of fragmented time. For driving scenarios, there are safety risks when manually operating a mobile phone while driving. The auxiliary operation method provided by the embodiment of the present application can be used to answer calls, switch navigation routes, and other operations to ensure driving safety.
[0061] Figure 10 A schematic diagram of the hardware structure of an electronic device that implements the embodiments of the present application is shown. Referring to this diagram, at the hardware level, electronic device 1000 includes a processor 1010, and optionally, an internal bus 1020, a network interface 1030, and a memory. The memory may include a memory 1041, such as a high-speed random-access memory (RAM), and may also include a non-volatile memory 1042, such as at least one disk storage device. Of course, electronic device 1000 may also include hardware required for other services.
[0062] The processor 1010, the network interface 1030, and the memory can be interconnected via an internal bus 1020. This internal bus 1020 can be an Advanced Microcontroller Bus Architecture (AMDBA) bus, a Wishbone bus, an Open Core Protocol (OCP) bus, an Avalon bus, or the like. Such buses can be categorized as address buses, data buses, and control buses. For ease of illustration, this figure uses only one bidirectional arrow, but this does not imply that there is only one bus or only one type of bus.
[0063] The memory stores programs. Specifically, the programs may include program codes, which include computer operating instructions. The memory may include internal memory 1041 and non-volatile memory 1042, and provides instructions and data to the processor 1010.
[0064] The processor 1010 reads the corresponding computer program from the non-volatile memory 1042 into the memory and then runs it, forming a device for locating the target user at the logical level. The processor 1010 executes the program stored in the memory and specifically performs the following: Figure 1 or Figure 5 The methods disclosed in the illustrated embodiments implement the functions and beneficial effects of the various methods described in the foregoing method embodiments, which will not be described in detail here.
[0065] The above application Figure 1 or Figure 5 The methods disclosed in the illustrated embodiments can be applied to or implemented by processor 1010. Processor 1010 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be performed by hardware integrated logic circuits or software instructions in processor 1010. The processor 1010 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0066] The computer device can also execute the methods described in the above method embodiments and realize the functions and beneficial effects of the methods described in the above method embodiments, which will not be repeated here.
[0067] Of course, in addition to software implementation, the electronic device 1000 of the present application does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0068] The embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores one or more programs, which, when executed by an electronic device including multiple application programs, enables the electronic device to execute Figure 1 or Figure 5 The methods disclosed in the illustrated embodiments implement the functions and beneficial effects of the various methods described in the foregoing method embodiments, which will not be described in detail here.
[0069] The computer-readable storage medium includes a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0070] Furthermore, an embodiment of the present application provides a computer program product, comprising a computer program stored on a non-transitory computer-readable storage medium, wherein the computer program comprises program instructions. When the program instructions are executed by a computer, the following process is implemented: Figure 1 or Figure 5 The methods disclosed in the illustrated embodiments implement the functions and beneficial effects of the various methods described in the foregoing method embodiments, which will not be described in detail here.
[0071] The embodiments of the present application can be applied to various electronic device collaboration or interconnection scenarios, including: collaboration and interconnection between mobile phones and laptops / tablets; collaboration and interconnection between mobile terminals and smart TVs / displays; collaboration and interconnection between mobile phones or tablets and in-car entertainment systems; collaboration and interconnection between mobile terminals and smart conference systems, etc., thereby meeting the diverse needs of users in scenarios such as smart homes, smart offices, and smart travel.
[0072] In short, the above description is only a preferred embodiment of the present application and does not limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
[0073] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0074] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0075] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0076] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
Claims
1. An auxiliary operation method, characterized in that: include: In response to receiving an auxiliary operation instruction corresponding to a target task, performing at least one round of auxiliary operation for the target task; In each round of auxiliary operation, screen display information is obtained; by inputting input information into the AI big model, the output information of the AI big model is obtained, the input information includes the auxiliary operation instructions, the screen display information and the operation trajectory information corresponding to the previous round of auxiliary operation, and the output information includes the operation action description and operation instructions of the current round of auxiliary operation; Execute the operation action corresponding to the operation instruction, and write the operation action description into the operation trajectory information to obtain the operation trajectory information corresponding to the current round of auxiliary operation, until the current round of auxiliary operation meets the termination condition and the auxiliary operation is completed; the termination condition includes at least one of the following: the number of rounds of the auxiliary operation exceeds the preset number threshold, and the task status of the target task is determined to be completed.
2. The method according to claim 1, characterized in that The output information further includes memory information associated with the auxiliary operation instruction determined based on the screen display information; after inputting the input information into the AI big model to obtain the output information of the AI big model, the method further includes: The memory information is written into the operation track information.
3. The method according to claim 1, characterized in that The executing of the operation action corresponding to the operation instruction includes: Determining an operation type of the operation instruction, where the operation type includes an execution operation and an interaction operation; According to the operation type, an operation action corresponding to the operation instruction is executed.
4. The method according to claim 3, characterized in that The performing of an operation action corresponding to the operation instruction according to the operation type includes: In response to the operation type being an execution-type operation, an operation action corresponding to the operation instruction is executed by calling an interface of the terminal device.
5. The method according to claim 3, characterized in that The performing of an operation action corresponding to the operation instruction according to the operation type includes: In response to the operation type being an interactive operation, an operation action corresponding to the operation instruction is performed, and the operation action includes a display action of annotation information. The annotation information is determined based on the operation attribute information of all target controls in the screen layout file corresponding to the screen display information. The operation attribute information is used to characterize the operations that can be performed on the target controls.
6. The method according to claim 5, characterized in that The operation attribute information of all target controls is determined in the following way: Obtaining an operation attribute of each control by traversing all controls in the screen layout file corresponding to the screen display information; Selecting a candidate control whose attribute value of an operation attribute is a preset value from all controls in the screen layout file; Determine the operational attribute information of all target controls according to the candidate controls and the operational attributes of the candidate controls.
7. The method according to claim 6, characterized in that The determining the operation attribute information of all target controls according to the candidate controls and the operation attributes of the candidate controls includes: Merging a first control and a second control in the candidate controls to obtain a third control, wherein the overlapping surface between the first control and the second control meets a preset condition, wherein the preset condition includes that a ratio of the overlapping surface to an area of the first control is greater than a preset ratio threshold, and a ratio of the overlapping surface to an area of the second control is greater than the preset ratio threshold; Determining the operation attribute information of the third control according to the operation attribute of the first control and the operation attribute of the second control; The operation attribute information of all target controls is determined according to all third controls and the operation attribute information of the third controls.
8. The method according to claim 5, characterized in that Determine the annotation information based on the operation attribute information of all target controls, including: Acquire the operation identifier corresponding to the operation attribute information of each target control, and determine the annotation information according to each target control and the operation identifier corresponding to each target control; or Candidate operation items are generated according to the operation attribute information of all target controls, and annotation information is determined according to the candidate operation items.
9. The method according to any one of claims 5 to 8, characterized in that After executing the operation action corresponding to the operation instruction, the method further includes: Determining a candidate operation set based on the annotation information; Determining supplementary information for the target task based on the candidate operation set and the operation information returned based on the annotation information; The supplementary information is written into the operation track information to obtain the operation track information corresponding to the auxiliary operation of the current round.
10. The method according to any one of claims 1 to 8, characterized in that Before performing at least one round of auxiliary operations on the target task in response to receiving the auxiliary operation instruction corresponding to the target task, the method further includes: Receive initial operation instructions for target tasks; By inputting the initial operation instruction into the AI big model, information to be supplemented for the target task is obtained; wherein the information to be supplemented is determined by the AI big model based on prior knowledge in training data, and the training data is determined based on multiple operation steps of operating the terminal device for the target task; Generate a query statement based on the information to be supplemented; Obtaining an answer statement returned based on the query statement; According to the initial operation instruction and the answer statement, the auxiliary operation instruction corresponding to the target task is determined.
11. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 10 are implemented.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program or instruction, and when the program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
13. A computer program product, characterized in that The computer program product comprises a computer program stored on a non-transitory computer-readable storage medium, wherein the computer program comprises program instructions, which, when executed by a computer, cause the computer to perform the steps of the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Mobile phone automatic operation method, device and equipment and computer storage medium
CN119363872A
Cited By
Closed-loop reasoning method, device and equipment based on multi-modal large model
CN120996209A