Cloud mobile phone automatic operation method and device

The cloud phone automation operation method that combines a visual language model with a multimodal large language model solves the problems of high threshold and poor adaptability of automation operation in existing technologies, and realizes low-threshold, intelligent and robust cross-application operations, especially improving recognition and execution efficiency in the Chinese environment. It is suitable for large-scale deployment in the cloud and efficient resource utilization.

CN120729980APending Publication Date: 2025-09-30GUANGZHOU DULING TECH CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511040207.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing UI automation operation technology has problems such as high automation threshold, poor adaptability, and difficulty in cross-application operation when handling complex mobile applications and web page operations, especially low recognition and processing efficiency in the Chinese environment.

Method used

By combining a visual language model with a multimodal large language model, the system obtains task sequences, uses the visual language model to infer current tasks and interface information, generates operation instructions, and executes them on the cloud phone, forming a closed-loop control system of perception-decision-action, supporting the automated operation of natural language instructions.

Benefits of technology

It achieves low-threshold, intelligent and robust automated operations, can adapt to interface changes, and operate across applications, improves recognition accuracy and operational efficiency in Chinese environments, reduces maintenance costs, and supports large-scale cloud deployment and efficient resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120729980A_ABST
    Figure CN120729980A_ABST
Patent Text Reader

Abstract

The invention provides an automatic operation method and device for a cloud mobile phone, and relates to the technical field of cloud computing, in particular to the technical field of cloud mobile phones. A specific embodiment of the method comprises the following steps: acquiring a task sequence of a cloud mobile phone; for a current task in the task sequence, executing the following operation steps: reasoning the current task and current interface information of the cloud mobile phone by utilizing a visual language model to generate a current operation instruction; after the current operation instruction is executed on the cloud mobile phone, next interface information is loaded; in response to the fact that the current task is the last task of the task sequence, it is determined that automatic operation of the cloud mobile phone is completed; and responding to the situation that the current task is not the last task of the task sequence, taking the next task in the task sequence as a new current task, taking the next interface information as new current interface information, and continuing to execute the operation steps.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of cloud computing technology, specifically the field of cloud phone technology. Background Art

[0002] Modern mobile applications and web pages are becoming increasingly versatile, and user scenarios on smartphones are becoming increasingly complex. This has led to a huge demand for UI (User Interface) automation, including automated use case execution in software testing, automated application operations, and cross-application information integration.

[0003] Currently, there have been some explorations and practices in UI automation and intelligent agent operation interfaces, which can be mainly divided into the following categories: First, traditional UI automated testing tools: Currently, the most widely used method is to use professional automated testing frameworks and scripting tools. Examples include Appium and Airtest for mobile devices, and Selenium and Puppeteer for web applications. These tools operate on UI elements through low-level interfaces, requiring test engineers to write scripts to locate elements and execute actions such as clicks and inputs.

[0004] Second, manual rules and OCR (Optical Character Recognition) solutions: In scenarios without direct programmable interfaces, some tools use image recognition and predefined rules to automate user interface operations. For example, computer vision-based tools like Sikuli take screenshots and match them to pre-provided control images to determine click locations. Some RPA (Robotic Process Automation) software also uses OCR to read text on the screen and then determines the next action based on pre-set rules.

[0005] Third, voice assistants and shortcuts: Another related technology for end consumers is voice assistants and some system-level automation on mobile phones. Users can use voice commands to ask assistants to perform simple tasks, such as sending text messages, setting reminders, and opening applications. Summary of the Invention

[0006] The embodiments of the present disclosure provide a cloud phone automated operation method, apparatus, device, storage medium, and program product.

[0007] In the first aspect, an embodiment of the present disclosure proposes a method for automated operation of a cloud phone, including: obtaining a task sequence of the cloud phone; for the current task in the task sequence, performing the following steps: using a visual language model to infer the current task and the current interface information of the cloud phone to generate a current operation instruction; after executing the current operation instruction on the cloud phone, loading the next interface information; in response to the current task being the last task in the task sequence, determining that the automated operation of the cloud phone is completed; in response to the current task not being the last task in the task sequence, using the next task in the task sequence as the new current task, using the next interface information as the new current interface information, and continuing to execute the operation steps.

[0008] In the second aspect, an embodiment of the present disclosure proposes a visual language model training method, including: obtaining a first training sample and a second training sample of a sample application, wherein the first training sample includes the previous task and current interface screenshots of the sample application, and the second training sample includes multiple rounds of conversations corresponding to the current task of the application; using the first training sample and the second training sample to perform interlaced training on the visual language model until the visual language model meets preset conditions.

[0009] In the third aspect, an embodiment of the present disclosure proposes a cloud phone automated operation device, including: an acquisition module, configured to acquire the task sequence of the cloud phone; an operation module, configured to perform the following steps for the current task in the task sequence: use a visual language model to infer the current task and the current interface information of the cloud phone to generate a current operation instruction; after executing the current operation instruction on the cloud phone, load the next interface information; in response to the current task being the last task in the task sequence, determine that the cloud phone automated operation is completed; an execution module, configured to, in response to the current task not being the last task in the task sequence, use the next task in the task sequence as the new current task, use the next interface information as the new current interface information, and continue to execute the operation steps.

[0010] In a fourth aspect, an embodiment of the present disclosure proposes a visual language model training device, comprising: an acquisition module, configured to acquire a first training sample and a second training sample of a sample application, wherein the first training sample includes a previous task and a current interface screenshot of the sample application, and the second training sample includes multiple rounds of conversations corresponding to the current task of the application; a training module, configured to use the first training sample and the second training sample to perform interlaced training on the visual language model until the visual language model meets preset conditions.

[0011] In a fifth aspect, an embodiment of the present disclosure proposes an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the first aspect or the second aspect.

[0012] In a sixth aspect, an embodiment of the present disclosure proposes a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to enable a computer to execute the method described in the first aspect or the second aspect.

[0013] In a seventh aspect, an embodiment of the present disclosure proposes a computer program product, including a computer program, which implements the method described in the first aspect or the second aspect when executed by a processor.

[0014] The key or important features of the embodiments of the present disclosure are not intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Other features, objects, and advantages of the present disclosure will become more apparent upon reading the detailed description of the non-limiting embodiments made with reference to the following drawings. The drawings are provided for a better understanding of the present disclosure and do not constitute a limitation of the present disclosure. Among them: Figure 1 This is a system architecture diagram of the cloud phone's automated operation method; Figure 2 is a flow chart of an embodiment of a cloud phone automated operation method according to the present disclosure; Figure 3 is a flowchart of an embodiment of a visual language model training method according to the present disclosure; Figure 4 It is a state iteration diagram of the cloud phone's automated operation method; Figure 5 It is a structural diagram of an embodiment of a cloud phone automatic operation device according to the present disclosure; Figure 6 is a structural diagram of an embodiment of a visual language model training device according to the present disclosure; Figure 7 3 is a block diagram of an electronic device used to implement the cloud phone automatic operation method of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0016] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0017] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in the present disclosure may be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0018] Figure 1 The system architecture diagram of the cloud phone automatic operation method is shown. Figure 1 As shown, the system architecture may include three layers and two loops, namely, a three-layer structure of an input layer 101, a decision layer 102, and an execution layer 103, as well as closed-loop control of a perception loop and an action loop.

[0019] The input layer 101 receives a task description. This is typically a natural language instruction that can describe a specific goal or a series of steps. This can be input by the user via text, or provided by the upper-layer application via an API (Application Programming Interface).

[0020] The decision layer 102 is the brain of the system, including the task parsing and planning module and the multimodal LLM (Large Language Model) controller.

[0021] The task analysis and planning module is responsible for breaking down high-level user instructions into executable task sequences and managing the process. In simple cases, instructions can be directly sent to the multimodal LLM controller for step-by-step execution. However, for complex tasks, particularly those involving multiple objectives, the task analysis and planning module can first generate a preliminary step plan based on rules or LLM. For example, for the task "Open the shopping app to buy a red dress," the task analysis and planning module might generate the steps: "Step 1: Open the shopping app; Step 2: Search for 'red dress'; Step 3: Filter / select the product and place the order." These steps are then handed off to the multimodal LLM controller for execution one by one. The task analysis and planning module also monitors the execution process and adjusts the subsequent plan as needed (for example, retrying or skipping a step if it fails).

[0022] The multimodal LLM controller is the core AI (Artificial Intelligence) decision-making unit, equipped with a fine-tuned multimodal large language model. It receives two inputs: the current natural language subtask description (from user instructions or the task parsing and planning module) and the current interface information from the execution layer 103 (including interface screenshots and possible UI structure data). The multimodal LLM controller combines the interface it "sees" with the instructions it understands to infer the next specific action (for example, "click coordinates (x, y)" or "enter 'hello' in the text box"). The LLM controller can be thought of as simulating the human cognitive process of viewing the screen, understanding it, and then taking action. Its output is a decision on the next action.

[0023] The execution layer 103 is used to implement the actual operation and interface collection of the cloud phone. It is the interface for the agent to interact with the environment, including the cloud phone instance and the device operation interface.

[0024] A cloud phone instance is the actual environment in which the application runs. It can be a virtual device hosted in a cloud data center, with screen output and input controls accessed or sent over the network. The system can launch multiple instances on demand, and each instance can be treated as a separate phone for AI operation.

[0025] The device operation interface is responsible for converting action commands from the multimodal LLM controller into actual input operations for the cloud phone. For example, if the LLM decides to "tap somewhere on the screen," the device operation interface will call an ADB (Android Debug Bridge) command or simulate a click event and send it to the cloud phone. If "swiping the screen" is required, a slide event is sent; if "entering text" is required, ADB's input text command or a simulated keyboard is used. Similarly, the device operation interface is responsible for collecting interface information from the cloud phone and feeding it back to the multimodal LLM controller, including interface screenshots and possible auxiliary information (such as the current UI control structure and element text).

[0026] The figure shows two closed loops: the perception loop (feedback from the cloud phone to the LLM) and the action loop (control from the LLM to the cloud phone), which correspond to the cycles of interface perception and action execution respectively.

[0027] The entire architecture forms a closed-loop control system: the LLM acquires information from environmental perception and continuously adjusts its decisions; its actions, in turn, change the state of the environment, and this new state is fed back to the LLM until the task is completed. This design is similar to the agent-environment loop in reinforcement learning, except that the decision-making core is a specially trained LLM, rather than a hand-coded policy. The architecture introduces an intelligent decision-making layer and visual feedback, giving the system adaptive capabilities: if a step does not meet expectations, the AI ​​can decide to perform other actions or retry based on the new interface content, rather than blindly following a fixed process.

[0028] Figure 2 A process 200 of an embodiment of a cloud phone automated operation method according to the present disclosure is shown. The cloud phone automated operation method includes the following steps: Step 201: Obtain the task sequence of the cloud phone.

[0029] In this embodiment, the execution subject of the cloud phone automation operation method can obtain the task sequence of the cloud phone.

[0030] In some embodiments, task description information is received and parsed to generate a task sequence. The task description is typically a natural language instruction, which can be a specific goal (e.g., "Please post a Weibo post with the content 'The weather is great today' on the Weibo app") or a description of a series of steps. Instructions can be entered by the user via text, or provided by a higher-level application via an API.

[0031] During the initialization phase, the task description information is parsed to generate a task sequence. If the task description information contains multiple subtasks (such as the "...then...then..." structure), it is broken down into a task sequence. If the goal in the task description information is abstract, the knowledge base or rule base can be queried to supplement the specific steps (such as "help me book movie tickets tonight", you need to first determine the theater application, select the showtime, etc., and then plan to book movie tickets). The result of the initialization phase is to determine the preliminary execution plan of the task and the target judgment conditions (how to judge whether the task is completed). At the same time, allocate or start an idle cloud phone instance from the resource pool to prepare the application environment to be operated. For example, you can restore to a clean initial interface by pre-installing relevant applications or using a snapshot.

[0032] Step 202: Use the visual language model to infer the current task and the current interface information of the cloud phone to generate the current operation instruction.

[0033] In this embodiment, for the current task in the task sequence, the visual language model can be used to infer the current task and the current interface information of the cloud phone to generate the current operation instruction.

[0034] The current interface information can be obtained from the cloud phone through the device operation interface. This information mainly includes interface screenshots and can also include structured information about UI elements. This information can be organized into a form suitable for visual language model input and provided to the visual language model. For example, the interface screenshot can be encoded as an image embedding and provided with a text description (such as "The current interface contains XX button, YY text") together with the visual language model.

[0035] In some embodiments, a preset prompt, current task, and current interface information are input into a visual language model, and a current operation instruction can be output.

[0036] The current task is combined with the current interface information and provided to the visual language model. For example, the prompt might be "User intent: Click the login button to log in. Interface information: [image]." Using a specific prompt template, the visual language model is guided to follow the user intent while considering the visual content. Because the visual language model is specifically trained, it can effectively associate interface information with the task and locate the corresponding location on the screen for the target mentioned in the task.

[0037] In some embodiments, the visual language model may include a visual sub-model and a language sub-model; by inputting the current interface information into the visual sub-model, the object information on the current interface can be output; by inputting the object information on the current interface and the current task into the language sub-model, the object information corresponding to the current task can be output; based on the object information corresponding to the current task, the current operation instruction can be generated.

[0038] After receiving the prompt word, the visual language model can first understand the interface through the visual sub-model. This step is similar to OCR and image feature extraction. The visual sub-model will identify key information on the interface, such as text, icons, and layout. Then, the language sub-model can find matching interface elements in the context according to the task requirements and generate action decisions. Here, the visual language model can be trained to generate output in the form of an action description, for example: <action type="tap" target="登录按钮" coordinates="[100,200]" / > , which means clicking the "Login" button at coordinates (100,200). Or <action type="input" target="用户名输入框" text="testuser" / > , indicating entering "testuser" into the username field. This output format can be learned through model fine-tuning, using a structure similar to JSON (JavaScript Object Notation) or XML (Extensible Markup Language) for machine parsing. While the actual output can also be a natural language description of the action, for reliability, a structured output can be agreed upon to reduce ambiguity.

[0039] Step 203: After executing the current operation instruction on the cloud phone, load the next interface information.

[0040] In this embodiment, the above-mentioned execution subject can load the next interface information after executing the current operation instruction on the cloud phone.

[0041] The device operation interface parses the current operation instructions output by the visual language model and executes the actual operation on the cloud phone instance through ADB or automated tools. For example, calling adb shell input tap 100 200 clicks (100, 200) or calling adb shell input text "testuser" inputs input. After executing, wait a short time for the interface to update.

[0042] In some embodiments, in response to the next interface information not meeting a preset condition, automatic correction may be performed based on the next interface information. For example, the error type may be determined based on the next interface information, and correction may be performed using a correction method corresponding to the error type.

[0043] If the interface after the operation output by the visual language model does not show the expected changes (such as remaining on the same page or an error prompt), fault tolerance can be performed. Specifically, the correction strategy is automatically attempted based on the new interface. For example, when a click does not produce an effect, the visual language model infers that "it may be necessary to close the pop-up window first", so it outputs an action to close the pop-up dialog box. This ability comes from the fact that the visual language model has added examples of abnormal situations during training, so that it learns some common error recovery modes (such as "If the 'password error' prompt appears, take a screenshot and end it"). Secondly, if the visual language model fails to handle it on its own, it can also take measures when it detects that there has been no progress for a long time, such as resetting the interface or trying alternative solutions. The multi-level error handling mechanism can ensure that the task is completed as much as possible in various situations or provide clear failure feedback, rather than hanging without response.

[0044] Step 204: Determine whether the current task is the last task in the task sequence.

[0045] In this embodiment, the execution subject can determine whether the current task is the last task in the task sequence. If it is the last task, execute step 205; if not, execute step 206.

[0046] Step 205: Determine whether the cloud phone automated operation is complete.

[0047] In this embodiment, if the current task is the last task in the task sequence, the execution entity may determine that the cloud phone automation operation is completed.

[0048] Step 206: Set the next task in the task sequence as the new current task, and set the next interface information as the new current interface information.

[0049] In this embodiment, if the current task is not the last task in the task sequence, the execution entity may use the next task in the task sequence as the new current task, use the next interface information as the new current interface information, and return to continue executing step 202.

[0050] After the action is executed, the next interface information can be obtained from the cloud phone again through the device operation interface. If the previous step was to enter text or click to jump to a page, a new interface will be loaded. At this point, it is necessary to check whether the task is completed or whether to proceed to the next task. This can be determined based on preset conditions. For example, detecting the "Login Successful" prompt indicates that the current task is completed. If the task is not completed, the loop returns to step 202 and the new current interface information is given to the visual language model to continue deciding the next step.

[0051] Steps 202-206 are the loop execution phase, which continues until the task completion condition is met or the maximum step limit is reached. Task completion can be determined based on the user's initial request. For example, if the user requests to "complete the order," the task is considered complete upon detection of the order confirmation page. If the user requests to "query and return certain information," the task is considered complete after the information is retrieved and returned via structured output, ending the task. Furthermore, operation logs and interface screenshots taken during execution can be compiled into a report and provided to the user or archived as test evidence to demonstrate task execution.

[0052] The disclosed embodiment provides a method for automated operation of a cloud phone, which adopts a one-step decision-making mode in the "perception-decision-action" cycle. The visual language model at each step only determines the action of the current single step. This is similar to the step-by-step operation when a person performs a task, and is combined with feedback to adjust the strategy. This step-by-step closed loop helps to correct errors in a timely manner and deal with uncertainty. For example, in human-computer interaction, each step may encounter a situation that does not meet expectations. By allowing the model to view the interface at each step, it can ensure that subsequent decisions are based on the latest status.

[0053] Figure 3 A process 300 of an embodiment of a visual language model training method according to the present disclosure is shown. The visual language model training method includes the following steps: Step 301: Obtain a first training sample and a second training sample of a sample application.

[0054] In this embodiment, the execution subject of the visual language model training method may obtain a first training sample and a second training sample of the sample application.

[0055] The first training sample may include screenshots of the previous task and current interface of the sample application, and the second training sample may include multiple rounds of dialogue corresponding to the current task of the application.

[0056] In some embodiments, the current interface screenshot is divided into multiple grids, and the multiple grids are clustered to remove grids corresponding to repeated background areas.

[0057] Typically, directly feeding the complete screenshot of the current interface into the visual language model generates a large number of visual tokens, which occupy the attention window. To reduce invalid visual tokens, the current interface screenshot can be preprocessed. First, the current interface screenshot is divided into a fixed-size grid. Then, repeated background areas are detected through clustering, and only the important and unique parts of the current interface screenshot are retained. This reduces a large number of visual tokens, allowing the visual language model to focus on key UI elements. Without significantly sacrificing understanding accuracy, the inference speed is improved to adapt to real-time interaction.

[0058] In some embodiments, an interface element is selected from the user interface tree of the sample application, and a task corresponding to the interface element is automatically generated through a script; after the task corresponding to the interface element is executed in the sample application, a screenshot of the current interface of the sample application is captured; and a first training sample is generated based on the task corresponding to the interface element and the current interface screenshot.

[0059] Building a small but high-quality GUI (Graphical User Interface) command dataset allows for further fine-tuning of visual language models based on pre-training. The sources of this GUI command dataset include: interface element location data, application navigation data, and question-answering and assertion data.

[0060] Interface element location data can include, for example, "Click the red login button" or "Find the 'Settings' icon on the screen." A large number of screenshots from various applications were collected, and several elements of interest and corresponding instructions were annotated manually or using rules. Some of this data referenced the open-source Screenshots to Description dataset, which was processed and organized to train the visual language model's OCR and location capabilities. In particular, Chinese application interfaces were manually annotated to ensure that the visual language model learned to recognize common Chinese UI text (Chinese characters on buttons, such as "OK" and "Cancel") and typical control styles.

[0061] App navigation data can include multi-step demonstrations such as "Open the Alipay app and enter the balance page." To obtain this type of data, we can automatically generate synthetic tasks using scripts. Using the real-world app's UI tree, we write a program that randomly selects interface elements to form instructions, such as "Click the XX menu and then open the YY function." This is then executed through a UI automation framework, with intermediate screenshots and actions recorded to generate data with process supervision signals. These synthetic tasks are numerous and cover a wide range of topics, helping visual language models learn general cross-interface operation strategies.

[0062] Given that automated cloud phone operations sometimes require information to be returned, a visual language model needs to be trained to read answers from the interface. For example, if we ask, "Please tell me the price of the product displayed on the current page," the visual language model can learn to recognize and answer the price text. Using the aforementioned location data, we generate additional question-and-answer format data, enabling the visual language model to not only output action commands but also, when needed, text answers. Furthermore, we provide the visual language model with descriptions of judgment conditions, enabling it to determine whether the interface state meets the conditions and output a Boolean result. These two types of data give the visual language model basic screen reading and judgment capabilities, meeting more complex testing requirements.

[0063] Step 302 : alternately train the visual language model using the first training sample and the second training sample until the visual language model meets a preset condition.

[0064] In this embodiment, the execution entity may alternately train the visual language model using the first training sample and the second training sample until the visual language model meets a preset condition, wherein the preset condition may be that the model effect reaches a preset effect, the number of iterations reaches a preset number, etc.

[0065] Here, we selected an open-source visual language model as the underlying multimodal model. This model is capable of understanding and describing image content, and supports both Chinese and English. Without specific fine-tuning, the visual language model already demonstrates excellent OCR and general knowledge question-answering capabilities. The model's open model weights allow for local deployment, making it suitable for customized training in private environments.

[0066] In order to enable the visual language model to support automated operations on cloud phones, the visual language model can be trained (fine-tuned) in a targeted manner. During fine-tuning, by designing training data in an interactive dialogue format, the visual language model can undergo multiple perception and action alternations in a single round of dialogue. When constructing training samples, instead of just providing the visual language model with a picture and asking it to answer where to click, a complete task dialogue is simulated, such as "User: Please log in on the App interface. AI: {Image 1 Description} I see the login interface. User: OK, please click the login button. AI:<action tap...> Clicked Login. {Image 2 Description} The login success page is displayed. ”

[0067] This interleaved format enables the visual language model to learn to process multiple rounds of perception and action streams within a single context. Specifically, two training modes are mixed: one is action-vision: at each step, the previous step's action and the current screenshot are interleaved as input, allowing the visual language model to only predict the loss of the action portion; the other is action-question-answer: some originally single-step questions are strung together into multiple rounds of dialogue to improve training efficiency. The combined use of these two methods significantly improves the visual language model's memory and consistency for long tasks. After interleaved training, the visual recognition model is less likely to experience contextual confusion when performing multi-step tasks and can correctly utilize information from previous steps.

[0068] In some embodiments, a parallel low-rank matrix pair is added to the visual language model; during backpropagation, the weights of the visual language model are frozen and the parallel low-rank matrix pair is updated.

[0069] Model fine-tuning efficiency can be improved through LoRA (Low-Rank Adaptation) and 8-bit quantization fine-tuning. LoRA is a technology for efficiently fine-tuning large pre-trained models. Its core goal is to significantly reduce the number of parameters and computing resources required for training while maintaining or approaching the performance of full parameter fine-tuning. INT8 quantization inference can accelerate model execution and reduce graphics memory usage. Quantization has minimal impact on performance, yet it enables a consumer-grade GPU (Graphics Processing Unit) or even a CPU (Central Processing Unit) to run the model, paving the way for parallel deployment of multiple instances in the cloud.

[0070] The disclosed embodiments provide a visual language model training method that focuses on the actual operational needs of the cloud phone environment, with training data and output formats close to real-world application scenarios. The model has small parameters but has undergone more targeted training, resulting in higher accuracy and faster speed on many tasks. More Chinese application data was added to the training, giving the model a deeper understanding of the details of domestic application interfaces. In addition, the output actions use structured instructions, making it easier for model decisions to interface with programs and reducing parsing ambiguity.

[0071] Figure 4 The following diagram shows the state iteration diagram of the cloud phone automation operation method. Figure 4 As shown, the state iteration includes the following steps: Step 401: Prepare the node to initialize the environment. During the initialization phase, the task description information is parsed into a task sequence. At the same time, an idle cloud phone instance is allocated or started from the resource pool to prepare the application environment for operation.

[0072] Step 402: The knowledge_check node searches the knowledge base. The knowledge base stores historical experience, and experience enhancement is performed by searching the knowledge base.

[0073] Step 403: Knowledge base hit determination.

[0074] Step 404: knowledge_enhance node experience enhancement.

[0075] Step 405: Model node AI reasoning and decision making. Reasoning is performed using the visual language model.

[0076] Step 406: The tool_valid node tool verifies the analysis.

[0077] Step 407, the should_tool_exec_continue tool verifies the result.

[0078] Step 408, the special tool ends.

[0079] Step 409: The tool node performs device operations.

[0080] Step 410: the knowledge_save node saves the experience.

[0081] Step 411, should_react_continue iterative control judgment.

[0082] To illustrate the workflow of this solution more intuitively, we will use a specific application scenario as an example to provide the entire process of completing a task using automated cloud phone operations.

[0083] An example task could be that the user wants to purchase an item using a certain e-commerce app on a cloud phone. The specific instruction description is: "Please open a certain e-commerce app, search for 'a certain model of mobile phone', add the item with the highest sales volume to the shopping cart, and take a screenshot of the shopping cart page."

[0084] This task involves cross-page and multi-step operations and requires making a selection (the item with the highest sales volume) based on interface information. The automated operation process of the cloud phone is as follows: The first step, task assignment: The user submits the above natural language instruction through the system interface. The task parsing module reads that this is a multi-step composite task and thus disassembles it into several subtasks: 1. Open a certain e-commerce app; 2. Search for 'a certain model of data'; 3. Select the item with the highest sales volume from the search results; 4. Add the item to the shopping cart; 5. Open the shopping cart page and take a screenshot.

[0085] At the same time, the module identifies that the final result of the task requires a screenshot output.

[0086] The second step, start the app: The system selects an idle cloud phone instance and executes "open a certain e-commerce app". First, directly start the Activity (an application component) of the app through ADB. If it cannot be directly located, the LLM can also be used to identify the desktop icon. Assume that the e-commerce app is successfully started and the home page interface of the e-commerce app is displayed on the cloud phone. The device operation interface obtains a screenshot of the home page and passes it to the multi-modal LLM controller.

[0087] The third step, search for items: The multi-modal LLM controller receives subtask 1 "search for 'a certain model of mobile phone'" at this time. Based on the screenshot of the home page, the model identifies a search box and other content on the interface. The model outputs an action <action type="tap" target="搜索框" coordinates="[x1,y1]" / > , and the system executes this click. After the click, the page switches to the search input state and the keyboard appears. The LLM detects the interface change and believes that text needs to be input, so the next step outputs <actiontype="input" target="search box" text="iPhone 13" / >. The system executes, enters the specified text in the search box and sends it. The e-commerce app starts searching and a search result list page appears.

[0088] Step 4: Product Selection: The multimodal LLM controller is presented with a screenshot of the search results list and is tasked with selecting the product with the highest sales volume. The model must read the sales volume information for each product from the results list and compare it. Typically, e-commerce app list items display sales volume or the number of reviews. The model has been trained on similar comparison tasks and therefore possesses a certain level of reasoning ability. The model might parse the sales figures for multiple result items and locate the item with the highest sales volume. It then outputs the corresponding click action: <action type="tap" target="第1项商品" coordinates="[x2,y2]" / > (Assume the highest-selling item is the first one in the list.) The system executes the click and goes to the product's details page.

[0089] Step 5: Add to Cart: After the details page is loaded, LLM receives the screenshot. The subtask is "Add to Cart". There is an "Add to Cart" button on the model recognition page. Output <action type="tap" target="加入购物车按钮" coordinates="[x3,y3]" / > After clicking the system, if you are not logged in, a login prompt may pop up. However, if you are logged in, a "Join Successful" prompt will pop up after clicking. LLM may see a pop-up prompt or a red dot prompt, which means the joining operation is complete.

[0090] In the sixth step, the model knows that the next task is to "open the shopping cart page and take a screenshot", and many e-commerce applications will have an entrance after adding the shopping cart. <action type="tap" target="购物车图标" coordinates="[x4,y4]" / > To enter the shopping cart. The page jumps to the shopping cart list.

[0091] Step 7: Screenshot Output: After the shopping cart page appears, the task planning module detects that the last step requires a screenshot. Instead of requiring the LLM to click, the task planning module directly calls an API to save the current screenshot image file. This also notifies the LLM of the task completion. The LLM might log "Item added to shopping cart and screenshot taken." The system returns the screenshot file and the execution log to the user as a result.

[0092] Throughout the process, the AI ​​operator successfully completed complex tasks across multiple pages. Several points are worth noting: Decision-making autonomy: In the fourth step, the AI ​​selects the "highest-selling" item, an action that requires understanding and comparing interface data. LLM makes this decision directly in one step through visual and language understanding, demonstrating its strong semantic understanding and reasoning capabilities.

[0093] Adaptive dynamics: If an unexpected situation occurs during a step, such as a login screen popping up when clicking "Add to Cart" in step 5, the model will recognize the login screen and either pause the task or continue the login process. If the task stops, the system can notify the user that they need to log in first. This fault-tolerant approach doesn't require a rigid shutdown; instead, it allows for continued progress or issues a reasonable termination message.

[0094] Multimodal Fusion: The example relies heavily on text and data in the UI. The model effectively integrates OCR results with visual localization. For example, even with very long product titles, the model can infer which product the sales figures correspond to based on the layout. Furthermore, even when buttons have icons but no text, the model can identify their meaning through visual features and associate them with the instructions.

[0095] Speed: In actual testing, all actions in the above scenario took approximately 10 seconds. Model inference took approximately five calls, each about one second, and ADB execution and page loading took about 5 seconds. Compared to manual operations, which can take 20-30 seconds, AI has achieved considerable efficiency. Using higher computing power could further accelerate the model, bringing the experience to near real-time.

[0096] Logs and reports: When the system performs the above tasks, it will record each model output action and the interface screenshot at that time. Ultimately, these can generate the following log snippets: 1 [Step 1] Click the search box (coordinates x1, y1) - Success 2 [Step2] Enter "certain model of mobile phone" into the search box - Success 3 [Step3] Click on the first item in the search result (coordinates x2, y2) - Success 4 [Step4] Click to add to cart (coordinates x3, y3) - Success 5 [Step5] Click the shopping cart icon (coordinates x4, y4) - Success 6 [Result] Task completed, screenshot saved to image_123.png The report also includes screenshots of key steps. This allows users or testers to easily verify whether the AI ​​correctly performs the expected operations. If errors are found, they can adjust the model or provide feedback accordingly.

[0097] Through this example, the cloud phone automation operation not only correctly completed the complex task, but also demonstrated obvious advantages over existing solutions: zero code, automatic understanding, high adaptability to changes, support for Chinese environment and good traceability.

[0098] The embodiments of the present disclosure provide a cloud phone automated operation solution based on multimodal LLM, the main technical effects of which are as follows: First, it significantly lowers the threshold for automation and improves development efficiency: users don't need programming skills or to learn complex test script syntax. They simply describe their requirements in natural language, and the corresponding operations are automatically completed. This expands automation beyond professional testers to a wider audience. For example, product managers and operations personnel can use this system to orchestrate mobile phone operation tasks. For enterprises, the speed of writing new test cases will be significantly increased, and the automation coverage of complex scenarios will be expanded. At the same time, since developers don't need to worry about device compatibility and underlying implementation, they can focus on the test logic itself, increasing development efficiency by at least several times.

[0099] Second, it's intelligent and robust, adapting to interface changes: Cloud phone automated operations exhibit human-like flexibility. When the app interface is updated, as long as the visual differences don't completely alter the semantics of the elements, the model can still identify the correct action (for example, a slight change in button position or color won't affect it, as the model focuses on text and icon features). Even if unexpected pop-ups or error messages appear, the model can assess the interface content and take appropriate action, such as attempting to close the pop-up or taking a screenshot to document the error. This adaptability and fault tolerance far surpass fixed scripts, significantly reducing maintenance costs.

[0100] Third, it boasts high versatility and cross-application operational capabilities: It's not limited to a single application or scenario. Because the model has learned general principles of GUI operations, it can execute tasks across different applications. For example, a cross-application workflow like "reading a message from WeChat and saving it in the Memo app" can be implemented with a single command. Furthermore, it supports any application, eliminating the need for specialized customization for each one. This broad applicability makes it a universal mobile automation platform, providing underlying support for various applications, such as testing, RPA, and assistance.

[0101] Fourth, excellent performance in the Chinese environment: Taking into account domestic scenarios, both model selection and training have been optimized for Chinese. Special enhancements have been made to Chinese in OCR and language understanding to improve the accuracy of text recognition in Chinese interfaces. The model can correctly understand Chinese UI terms such as "collect," "like," and "pay." In addition, the model itself has a grasp of Chinese semantics and is adept at understanding user instructions (Chinese descriptions) and generating Chinese output (such as answer content), ensuring that Chinese information can be processed smoothly from instructions to interfaces. For complex fonts, mixed traditional and simplified Chinese, English abbreviations, and other situations commonly seen in domestic applications, training can also be carried out through data coverage to improve robustness.

[0102] Fifth, it is resource-efficient and can be deployed on a large scale in the cloud: The model used has only approximately 2.5B parameters and has been optimized to run on consumer-grade GPUs and even CPUs. Deployment and operating costs are significantly reduced, enabling the parallel launch of numerous instances in the cloud to serve multiple users or multiple testing tasks. For example, a server equipped with sufficient GPUs can run dozens of cloud phone automation operation instances simultaneously, meeting the requirements of concurrent testing in CI / CD (Continuous Integration / Continuous Delivery) pipelines. On the terminal side, the miniaturization of the model also lays the foundation for future embedding in mobile devices and running in local private environments. Furthermore, the low resource usage directly benefits from fast response times. This real-time experience is critical for enhancing user confidence and expanding application scenarios, such as interactive assistants.

[0103] Sixth, it is open source and easy to integrate and expand: It is built based on open source models and open interfaces, and has no dependence on cloud services from specific vendors. The models and frameworks are the results of the open source community, and innovation is made on this basis, but openness is retained. This allows users to privately deploy the system to ensure data security, or further customize the model according to their own business. The system provides a clear API interface that can be easily integrated with existing testing frameworks, monitoring systems or business process management platforms. For example, testers can trigger a cloud phone operation task through a simple HTTP (Hypertext Transfer Protocol) call and obtain a result report, which greatly facilitates continuous integration. In addition, due to the modular architecture, it is easy to expand new functions, such as replacing more powerful models, adding plug-ins for special applications, etc., without affecting the overall operation.

[0104] Seventh, it improves testing quality and user experience: Cloud phone automation not only improves efficiency but also enhances testing quality. It can simulate real user behavior, including random or unusual operations, helping to identify issues with applications under unusual circumstances. Similarly, in RPA applications, AI's autonomous judgment can handle edge cases and reduce process bottlenecks. Furthermore, because the operation process is thoroughly documented and captured, it is easier to trace the source of problems and quickly locate the cause of errors. For end users, cloud phone automation can also serve as a personal mobile assistant, enhancing the user experience. It can understand vague high-level requirements and complete specific tasks, truly achieving "what you see is what you get" and "what you say is what you get." This represents a leap forward in human-computer interaction.

[0105] Further references Figure 5 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a cloud phone automatic operation device, which is similar to Figure 2Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0106] like Figure 5 As shown, the cloud phone automation operation device 500 of this embodiment may include: an acquisition module 501, an operation module 502, and an execution module 503. Among them, the acquisition module 501 is configured to obtain the task sequence of the cloud phone; the operation module 502 is configured to perform the following steps for the current task in the task sequence: using the visual language model to infer the current task and the current interface information of the cloud phone to generate the current operation instruction; after executing the current operation instruction on the cloud phone, loading the next interface information; in response to the current task being the last task in the task sequence, determining that the cloud phone automation operation is completed; the execution module 503 is configured to, in response to the current task not being the last task in the task sequence, set the next task in the task sequence as the new current task, set the next interface information as the new current interface information, and continue to execute the operation steps.

[0107] In this embodiment, in the cloud phone automatic operation device 500, the specific processing of the acquisition module 501, the operation module 502 and the execution module 503 and the technical effects thereof can be referred to respectively. Figure 2 The relevant descriptions of steps 201-203 in the corresponding embodiment are not repeated here.

[0108] In some optional implementations of this embodiment, the acquisition module 501 is further configured to: receive task description information; and parse the task description information to generate a task sequence.

[0109] In some optional implementations of this embodiment, the operation module 502 is further configured to: input a preset prompt word, a current task, and current interface information into the visual language model, and output a current operation instruction.

[0110] In some optional implementations of this embodiment, the visual language model includes a visual sub-model and a language sub-model; and the operation module 502 is further configured to: input the current interface information into the visual sub-model, and output the object information on the current interface; input the object information on the current interface and the current task into the language sub-model, and output the object information corresponding to the current task; and generate the current operation instruction based on the object information corresponding to the current task.

[0111] In some optional implementations of this embodiment, the cloud phone automated operation device 500 further includes: a correction module, configured to perform automatic correction based on the next interface information in response to the next interface information not meeting a preset condition.

[0112] In some optional implementations of this embodiment, the correction module is further configured to: determine the error type based on the next interface information, and perform correction using a correction method corresponding to the error type.

[0113] Further references Figure 6 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a visual language model training device. Figure 3 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0114] like Figure 6 As shown, the visual language model training device 600 of this embodiment may include: an acquisition module 601 and a training module 602. The acquisition module 601 is configured to acquire a first training sample and a second training sample of a sample application, wherein the first training sample includes screenshots of the previous task and the current interface of the sample application, and the second training sample includes multiple rounds of dialogue corresponding to the current task of the application; the training module 602 is configured to use the first training sample and the second training sample to interleave training the visual language model until the visual language model meets preset conditions.

[0115] In this embodiment, in the visual language model training device 600, the specific processing of the acquisition module 601 and the training module 602 and the technical effects thereof can be referred to in Figure 3 The relevant descriptions of steps 301-302 in the corresponding embodiment are not repeated here.

[0116] In some optional implementations of this embodiment, the visual language model training device 600 further includes: a preprocessing module configured to divide the current interface screenshot into multiple grids; cluster the multiple grids, and remove grids corresponding to repeated background areas.

[0117] In some optional implementations of this embodiment, the visual language model training device 600 also includes: a generation module, configured to select interface elements from the user interface tree of the sample application, and automatically generate tasks corresponding to the interface elements through a script; after executing the tasks corresponding to the interface elements in the sample application, taking a screenshot of the current interface of the sample application; and generating a first training sample based on the tasks corresponding to the interface elements and the current interface screenshot.

[0118] In some optional implementations of this embodiment, the training module 602 is further configured to: add parallel low-rank matrix pairs to the visual language model; and during the backpropagation process, freeze the weights of the visual language model and update the parallel low-rank matrix pairs.

[0119] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0120] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0121] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0122] like Figure 7 As shown, device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. RAM 703 may also store various programs and data required for the operation of device 700. Computing unit 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to bus 704.

[0123] Various components in device 700 are connected to I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, optical disk, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0124]

[01] The computing unit 701 may be a variety of general and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as the cloud phone automation operation method. For example, in some embodiments, the cloud phone automation operation method may be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the cloud phone automation operation method described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform the cloud phone automation operation method in any other appropriate manner (e.g., by means of firmware).

[0125] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0126] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0127] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0128] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0129] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0130] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0131] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions provided by this disclosure can be achieved. This is not limited herein.

[0132] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A cloud phone automated operation method, comprising: Get the task sequence of the cloud phone; For the current task in the task sequence, the following steps are performed: using a visual language model to infer the current task and the current interface information of the cloud phone to generate a current operation instruction; After executing the current operation instruction on the cloud phone, loading next interface information; in response to the current task being the last task in the task sequence, determining that the cloud phone automated operation is completed; In response to the current task not being the last task in the task sequence, the next task in the task sequence is used as a new current task, the next interface information is used as new current interface information, and the operation steps are continued.

2. The method according to claim 1, wherein The acquisition task sequence includes: Receive task description information; The task description information is parsed to generate the task sequence.

3. The method according to claim 1, wherein The using of the visual language model to infer the current task and the current interface information of the cloud phone to generate the current operation instruction includes: The preset prompt word, the current task and the current interface information are input into the visual language model, and the current operation instruction is output.

4. The method according to claim 1, wherein The visual language model includes a visual sub-model and a language sub-model; as well as The using of the visual language model to infer the current task and the current interface information of the cloud phone to generate the current operation instruction includes: Inputting the current interface information into the visual sub-model and outputting the object information on the current interface; Inputting the object information on the current interface and the current task into the language sub-model, and outputting the object information corresponding to the current task; The current operation instruction is generated based on the object information corresponding to the current task.

5. The method according to claim 1, wherein The method further comprises: In response to the next interface information not meeting a preset condition, automatic correction is performed based on the next interface information.

6. The method according to claim 5, wherein: The automatic correction based on the next interface information includes: The error type is determined based on the next interface information, and correction is performed using a correction method corresponding to the error type.

7. A visual language model training method, comprising: Obtaining a first training sample and a second training sample of a sample application, wherein the first training sample includes a screenshot of a previous task and a current interface of the sample application, and the second training sample includes multiple rounds of dialogue corresponding to the current task of the application; The visual language model is alternately trained using the first training sample and the second training sample until the visual language model meets a preset condition.

8. The method according to claim 7, wherein: The method further comprises: Dividing the current interface screenshot into multiple grids; Clustering is performed on the multiple grids, and grids corresponding to repeated background areas are removed.

9. The method according to claim 7, wherein: The method further comprises: Selecting an interface element from the user interface tree of the sample application, and automatically generating a task corresponding to the interface element through a script; After executing the task corresponding to the interface element in the sample application, capturing a screenshot of the current interface of the sample application; The first training sample is generated based on the task corresponding to the interface element and the current interface screenshot.

10. The method according to claim 7, wherein: The interleaving training of the visual language model using the first training sample and the second training sample includes: Adding parallel low-rank matrix pairs to the visual language model; During the back-propagation process, the weights of the visual language model are frozen and the parallel low-rank matrix pairs are updated.

11. A cloud phone automated operation device, comprising: The acquisition module is configured to acquire the task sequence of the cloud phone; An operation module is configured to perform the following steps for a current task in the task sequence: using a visual language model to infer the current task and current interface information of the cloud phone to generate a current operation instruction; After executing the current operation instruction on the cloud phone, loading next interface information; in response to the current task being the last task in the task sequence, determining that the cloud phone automated operation is completed; The execution module is configured to, in response to the current task not being the last task in the task sequence, take the next task in the task sequence as the new current task, take the next interface information as the new current interface information, and continue to execute the operation steps.

12. A visual language model training device, comprising: an acquisition module configured to acquire a first training sample and a second training sample of a sample application, wherein the first training sample includes a screenshot of a previous task and a current interface of the sample application, and the second training sample includes multiple rounds of dialogue corresponding to the current task of the application; The training module is configured to perform interlaced training on the visual language model using the first training sample and the second training sample until the visual language model meets a preset condition.

13. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6 or 7-10.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6 or 7 to 10.

15. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1-6 or 7-10.

Citation Information

Cited By

  • Mobile terminal application operation proxy method, system and device, medium and program product

    CN121387422A