Terminal operation method, storage medium, program product, electronic apparatus, and vehicle

By responding to the user's operation intentions at the terminal and generating operation control instructions using multimodal models, the problem of precise automation of multiple applications or dynamically changing target objects in smart cockpit technology is solved, and high reliability and accuracy automated operations are achieved.

CN120104022APending Publication Date: 2025-06-06BYD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510112522.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In smart cockpit technology, it is difficult for large-model-based on-vehicle intelligent voice systems to accurately automate multiple applications or dynamically changing target objects.

Method used

By responding to the user's operation intention at the terminal, the intention information and the image of the terminal user interface are obtained, and the operation control instructions are generated using a multimodal model to operate the terminal user interface.

Benefits of technology

It realizes precise operation of multiple or dynamically changing target objects, improves the reliability and accuracy of application automation operations, and expands the applicability of automated operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104022A_ABST
    Figure CN120104022A_ABST
Patent Text Reader

Abstract

The present application relates to a terminal operation method, a storage medium, a program product, an electronic device and a vehicle, and relates to the technical field of artificial intelligence, the method comprising: when a terminal responds to an operation intention of a user, acquiring intention information of the operation intention and a first interface image of a terminal user interface; and obtaining an operation control instruction according to the intention information and the first interface image, and operating the terminal user interface through the operation control instruction. According to the method, the operation control instruction is determined through the intention information and the interface image, the multiple application programs can be operated, meanwhile, global path planning with high accuracy and action information of each step can be obtained easily, and the reliability and accuracy of automatic operation of the application programs are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a terminal operation method, a storage medium, a program product, an electronic device and a vehicle. Background Art

[0002] In the development of smart cockpit technology, in-vehicle intelligent voice systems based on large models have become the mainstream direction. During the driving process, for a variety of applications (Apps) installed by users, the in-vehicle intelligent voice system can automatically operate the applications of the in-vehicle terminal based on the user's voice commands. At present, when performing automated operations on applications, when there are multiple target objects to be operated on a user interface, or when the target objects to be operated change dynamically, it is impossible to perform accurate operations on the target objects, which limits the coverage of the automated operations of applications. Summary of the invention

[0003] Embodiments of the present application provide a terminal operation method, a storage medium, a program product, an electronic device, and a vehicle to solve the above-mentioned problems.

[0004] In order to achieve the above object, according to a first aspect of the present application, a terminal operation method is provided, the method comprising:

[0005] When the terminal responds to the user's operation intention, acquiring the intention information of the operation intention and the first interface image of the terminal user interface;

[0006] An operation control instruction is obtained according to the intention information and the first interface image, and the terminal user interface is operated according to the operation control instruction.

[0007] Optionally, the operation control instruction includes application information and an action sequence, and obtaining the operation control instruction according to the intention information and the first interface image, and operating the terminal user interface through the operation control instruction includes:

[0008] Inputting the intention information and the first interface image into a multimodal model to obtain the application information and the action sequence;

[0009] The terminal user interface is operated according to the application information and the action sequence.

[0010] Optionally, the action sequence includes at least one action to be executed, and operating the terminal user interface according to the application information and the action sequence includes:

[0011] Acquire a second interface image, where the second interface image is the first interface image or an image of the terminal user interface after a preset action is performed;

[0012] Inputting the application information, the action sequence and the second interface image into the multimodal model to obtain action information of the action to be executed, wherein the action information of the action to be executed includes at least one of an action name and coordinate information;

[0013] The terminal user interface is operated according to the action information of the action to be executed and the application information.

[0014] Optionally, the application information includes at least one target application, and the to-be-executed action is an execution action of the target application corresponding to the to-be-executed action;

[0015] The operating the terminal user interface according to the action information of the action to be executed and the application information includes:

[0016] The target application corresponding to the action to be executed is operated according to the action information of the action to be executed.

[0017] Optionally, inputting the intention information and the first interface image into a multimodal model to obtain application information and an action sequence includes:

[0018] Based on the first interface image, the intent information and the auxiliary information, semantic understanding is performed through the multimodal model to obtain the application information and the action sequence.

[0019] Optionally, the auxiliary information includes at least one of offline document information and / or online document information of the target application, a reference action trajectory and a user action trajectory, wherein the reference action trajectory is an action sequence for the application information self-generated by the terminal system, and the user action trajectory is an action sequence for the application information customized by the user.

[0020] Optionally, the inputting the application information, the action sequence and the second interface image into the multimodal model to obtain the action information of the action to be performed includes:

[0021] In a case where the action to be performed is the first target action in the action sequence, inputting the application information, the action sequence and the second interface image into the multimodal model to obtain action information of the action to be performed;

[0022] In a case where the action to be performed is a second target action in the action sequence, inputting the application information, the action sequence, the second interface image, and the historical action record into the multimodal model to obtain action information of the action to be performed;

[0023] The second target action is the next action of the first target action, and the historical action record includes action information of an executed action before the second target action and an execution result of the executed action.

[0024] Optionally, operating a target application corresponding to the action to be executed by using the action information of the action to be executed includes:

[0025] Based on the action information, a preset tool is controlled to operate a target application corresponding to the action to be executed to obtain a task status, where the task status is used to indicate an execution status of the action to be executed.

[0026] Optionally, the method further comprises:

[0027] Control the terminal user interface to display the execution process of the action to be executed and / or the task status.

[0028] Optionally, the operating a target application corresponding to the to-be-performed action to obtain a task status includes:

[0029] When the second interface image does not satisfy the execution condition, the task state is the first task state, and the first task state is used to indicate to re-capture the terminal user interface.

[0030] Optionally, the method further comprises:

[0031] When the task state is the first task state, based on a preset area, re-capturing the terminal user interface to obtain an interface image of the preset area;

[0032] The second interface image is updated based on the interface image of the preset area, and the operation of inputting the application information, the action sequence and the second interface image into the multimodal model is re-executed.

[0033] Optionally, the to-be-executed action includes a first to-be-executed action and a second to-be-executed action, the second to-be-executed action is a next action of the first to-be-executed action, and the target application includes a first target application and a second target application;

[0034] The operating the target application corresponding to the to-be-performed action to obtain the task status includes:

[0035] When the first action to be executed is an action to operate on the first target application, and the second action to be executed is an action to operate on the second target application, the task state is a second task state, and the second task state is used to indicate switching from the first target application to the second target application.

[0036] Optionally, the operating a target application corresponding to the to-be-performed action to obtain a task status includes:

[0037] In the case where the action to be executed is a sensitive action, the task state is a third task state, and the third task state is used to indicate that the action to be executed is suspended;

[0038] Based on the third task state, obtaining first feedback input by a user;

[0039] Based on the first feedback, the preset tool is controlled to continue executing the action to be performed.

[0040] Optionally, the operating a target application corresponding to the to-be-performed action to obtain a task status includes:

[0041] After executing the action to be executed, evaluating the execution status of the action to be executed to obtain the execution result of the action to be executed;

[0042] Based on the execution result, the task status is determined.

[0043] Optionally, before determining the task status based on the execution result, the method further includes:

[0044] After executing the to-be-executed action, capturing an interface image after the execution;

[0045] Based on the interface image after execution, the execution status of the action to be executed is evaluated to obtain the execution result of the action to be executed.

[0046] Optionally, after operating the target application corresponding to the to-be-performed action to obtain the task status, the method further includes:

[0047] When the task status satisfies a preset condition, the operation of inputting the application information, the action sequence and the second interface image into the multimodal model is re-executed.

[0048] Optionally, the method further comprises:

[0049] Detecting and obtaining second feedback from the user, where the second feedback is used to indicate intention information fed back by the user during the execution process, and updating the intention information based on the second feedback from the user;

[0050] Based on the second feedback from the user, taking a screenshot of the current terminal user interface to update the first interface image;

[0051] Based on the intention information and the first interface image, the operation of obtaining the operation control instruction according to the intention information and the first interface image and operating the terminal user interface according to the operation control instruction is re-executed.

[0052] Optionally, the inputting the intention information and the first interface image into a multimodal model to obtain the application information and the action sequence includes:

[0053] Extracting features of the intent information based on the first network layer of the multimodal model to obtain a first feature sequence;

[0054] Extracting features of the first interface image based on the second network layer of the multimodal model to obtain a second feature sequence;

[0055] Based on the third network layer of the multimodal model, feature fusion is performed on the first feature sequence and the second feature sequence to obtain the application information and the action sequence.

[0056] According to the second aspect of the present application, an embodiment of the present application also provides a computer-readable storage medium on which a computer program is stored, and the computer-readable storage medium stores instructions, which, when executed by a computer, enable the computer to implement any one of the terminal operation methods provided in the embodiments of the present application.

[0057] According to the third aspect of the present application, an embodiment of the present application further provides a computer program product, which stores instructions, and when the instructions are executed by a computer, the computer implements any one of the terminal operation methods provided in the embodiments of the present application.

[0058] According to a fourth aspect of the present application, an embodiment of the present application further provides an electronic device, including:

[0059] a memory having a computer program stored thereon;

[0060] The processor is used to execute the computer program in the memory to implement any one of the terminal operation methods provided in the embodiments of the present application.

[0061] According to a fifth aspect of the present application, an embodiment of the present application also provides a vehicle, comprising the electronic device described above.

[0062] Some embodiments of the present specification include at least the following beneficial effects: by determining operation control instructions through intention information and real-time interface images, it is possible to operate multiple applications, and at the same time help to obtain highly accurate global path planning and action information for each step, thereby improving the reliability and accuracy of the application's automated operations. In this way, it is possible to achieve precise operations on multiple or dynamically changing target objects, thereby improving the applicability of the application's automated operations.

[0063] Other features and advantages of the present application will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application, and those skilled in the art can obtain other drawings based on these drawings without creative work.

[0065] In order to more completely understand the present application and its beneficial effects, the following description will be given in conjunction with the accompanying drawings, wherein the same figure numbers represent the same parts in the following description.

[0066] Figure 1 It is an application scenario diagram of the terminal operation method shown in some embodiments of this specification;

[0067] Figure 2 is an exemplary flow chart of a terminal operation method according to some embodiments of this specification;

[0068] Figure 3 is an exemplary flow chart of a multimodal model according to some embodiments of the present specification;

[0069] Figure 4 is an exemplary schematic diagram of iterative execution according to some embodiments of this specification;

[0070] Figure 5 is an exemplary structural diagram of a multimodal model according to some embodiments of this specification;

[0071] Figure 6 It is a schematic diagram of the structure of an electronic device according to some embodiments of this specification. DETAILED DESCRIPTION

[0072] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0073] In order to facilitate understanding of the implementation scheme provided in the embodiment of the present application, the relevant application background of the terminal operating system provided in the embodiment of the present application is first described.

[0074] At present, when automating the application of the vehicle terminal, the relevant technology mainly uses the semantic analysis module and the intention recognition module to perform natural language processing to obtain the user intention and keywords, and uses the mobile phone / tablet terminal automation testing tools such as uiautomator2 to write a series of rules to implement the jump, slide and other operations in the application. However, this method relies on fixed rules and lacks generalization. It often cannot effectively handle the rapidly updated applications. Moreover, the use of automation testing tools such as uiautomator2 requires the use of Android scene data to implement the operation of interface elements. In fact, the quality of Android scene data is uneven. In some applications, there will be errors such as scene data loss and overlap, and these errors can only be solved by the application developer, and they are often impossible to coordinate, resulting in the inability to implement automated operations in the above application scenarios.

[0075] In addition, even though some large multimodal models can output single-step actions on some applications, they cannot implement multi-step action planning, support a small number of applications, and do not support any Chinese applications. Moreover, these large multimodal models are usually trained based on supervised fine-tuning methods, so they cannot obtain enough data to achieve automated operations on different applications, resulting in deviations in the global path planning of these large multimodal models, and the inability to simultaneously guarantee the accuracy of global path planning and the accuracy of operations during training.

[0076] In view of this, some embodiments of the present specification provide a terminal operation method, which uses a multimodal model to achieve global path planning and switching between different applications; the multimodal model is used to cyclically perform operations on the selected application until the task requested by the user is completed, which helps to simultaneously obtain highly accurate global path planning and action information for each step, thereby improving the reliability and accuracy of the application's automated operation.

[0077] Figure 1 It is an application scenario diagram of the terminal operation method shown in some embodiments of this specification.

[0078] like Figure 1 As shown, the application scenario 100 of the terminal operation method may include a processor 110 , a terminal 120 , and a network 130 .

[0079] The processor 110 may process data and / or information obtained from other devices or system components. The processor may execute program instructions based on these data, information and / or processing results to perform one or more functions described in this specification. For example, the processor 110 may obtain an interface image of the terminal 120 based on the user's operation instructions, and determine an operation control instruction for the terminal, and control the application in the terminal according to the operation control instruction. For more information, please refer to the relevant description below.

[0080] In some embodiments, the processor 110 may be a single server or a server group. The server group may be centralized or distributed. In some embodiments, the processor 110 may be local or remote. In some embodiments, the processor 110 may be implemented on a cloud platform. As an example only, the cloud platform may include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an internal cloud, a multi-layer cloud, etc., or any combination thereof.

[0081] In some embodiments, the processor 110 may be integrated in the terminal 120. For example, the processor 110 may be a computing device (e.g., a vehicle-mounted computer) installed in a vehicle. The vehicle may include, but is not limited to, various types of vehicles such as fuel vehicles, electric vehicles, and trucks.

[0082] The terminal 120 can realize the interaction between the user and the processor 110. The user can refer to an operator or an administrator. In some embodiments, the terminal 120 can be used to display the execution process, execution results or other information of the action. In some embodiments, the terminal 120 can be installed on a vehicle. In some embodiments, the terminal 120 can receive operation control instructions from the processor 110 via a network and display them to the user through a screen. In some embodiments, the terminal 120 can include a mobile device 120-1, a tablet computer 120-2, a laptop computer 120-3, other devices with input and / or output functions, etc. or any combination thereof. The above examples are only used to illustrate the wide range of the terminal 120 device range and not to limit its range.

[0083] In some embodiments, the terminal 120 may include a display component (eg, a display screen), an interactive component (eg, a mouse, a keyboard, etc.), and the like.

[0084] The network 130 may include any suitable network capable of facilitating information and / or data exchange. In some embodiments, one or more components (e.g., the processor 110, the terminal 120, etc.) of the application scenario 100 may exchange information via the network 130. For example, the processor 110 may send a determined operation control instruction to the terminal 120 via the network 130.

[0085] In some embodiments, the application scenario 100 of the terminal operation method can be applicable to various fields, such as daily life assistant, office assistant, entertainment assistant, smart home control, transportation system, etc. For example, based on the terminal operation method, various functions of the vehicle can be controlled, such as adjusting the air conditioner, opening the window, controlling the in-vehicle entertainment system (such as playing music, adjusting the volume), etc.

[0086] It is worth noting that the application scenario 100 of the terminal operation method is provided for illustrative purposes only and is not intended to limit the scope of this specification. For those of ordinary skill in the art, various changes and modifications can be made according to the description of this specification. For example, the application scenario 100 may also include a database, an information source, etc. For another example, the application scenario 100 may be implemented on other devices to achieve similar or different functions. However, these changes and modifications will not deviate from the scope of this specification.

[0087] Figure 2 is an exemplary flow chart of a terminal operation method according to some embodiments of this specification. In some embodiments, process 200 can be executed by a processor based on an application scenario. Figure 2 As shown, process 200 includes the following steps.

[0088] Step S210: When the terminal responds to the user's operation intention, the intention information of the operation intention and the first interface image of the terminal user interface are acquired.

[0089] Action intent refers to the task that a user expects to accomplish when interacting with an application or device. For example, action intent.

[0090] For example, the user may input one or more operation intentions by voice input, text input, etc. For example, the user may input the operation intention by voice or other means, and the processor determines the intention information based on the operation intention. The operation intention may be: open WeChat and send a message to Zhang San, the content of which is "Hello, see you tomorrow", etc. Intent information refers to information related to the user's operation intention for the terminal. For example, the intent information may include a detailed description of the operation intention.

[0091] In some embodiments, the user's operation intention on the terminal can be obtained through the terminal to determine the user's intention information.

[0092] In some embodiments, the user's intention information may include the user's initial intention information, or the user's feedback intention information based on the execution result of at least one action.

[0093] An interface image refers to an image obtained by taking a screenshot of a terminal user interface. A first interface image refers to an image obtained by taking a screenshot of a terminal user interface when obtaining the user's intention information. For example, the first interface image includes various interface elements (such as buttons, text boxes, icons, etc.) in the terminal user interface when the user's initial intention information is obtained.

[0094] The terminal user interface is used to present data acquired and / or generated by the terminal or processor (e.g., the action execution process, the action execution result, the change of the target application, etc.). For example, the terminal user interface can display the target application. The terminal user interface can also be used to receive feedback information input by the user.

[0095] Feedback information refers to instruction information related to user input. In some embodiments, user input can be key, mouse, text, voice input, image input, touch screen, gesture instruction, EEG, eye movement or any other feasible way.

[0096] In some embodiments, the processor may be connected to the terminal for communication, and upon detecting user intent information or at specified intervals, may send an operation control instruction to the terminal to control the terminal to take a screenshot of the terminal user interface to obtain a first interface image.

[0097] In some embodiments, the terminal is covered by the field of view of the corresponding surveillance camera, and the surveillance camera can obtain the first interface image by collecting video or images under certain circumstances (e.g., based on the user's intention information trigger, etc.). After collecting the first interface image, the surveillance camera can upload the first interface image to the processor. The processor can directly use the video or image of the terminal collected by the surveillance camera as the first interface image.

[0098] Step S220, obtaining an operation control instruction according to the intention information and the first interface image, and operating the terminal user interface through the operation control instruction.

[0099] The operation control instruction is used to implement the user's operation intention on the terminal. For example, the operation control instruction may include a series of specific operation steps generated by completing the specific intention information. In some embodiments, the operation control instruction includes the action corresponding to each operation step and the location information of the action.

[0100] In some embodiments, the terminal may obtain the operation control instruction in a variety of ways based on the intention information and the first interface image. For example, the operation instruction may be determined by a machine learning model or the like.

[0101] In some embodiments, the terminal may perform corresponding actions on one or more application programs in the terminal user interface according to the operation control instruction.

[0102] In some embodiments, the operation control instruction may be determined based on a language model.

[0103] In some embodiments, the language model can be a large language model, etc. A large language model (Large Language Model, LLM) refers to a machine learning model formed by training based on deep learning technology, large-scale data and computing resources, mainly for natural language processing, but can also evolve to process other forms of data. Exemplary large language models may include, but are not limited to, a language representation model (Bidirectional Encoder Representation from Transformers, BERT), a generative pre-trained transformation model (Generative Pre-Trained Transformer, GPT), a natural language pre-training model based on a Transformer architecture (Extreme Language Model based on Transformer XL, XL-Net), a Chinese dialogue pre-training model (General Language Model for Chat-6Billion Parameters, ChatGLM-6B), etc.

[0104] In some embodiments of the present specification, by capturing intent information and real-time interface images, operational control instructions are obtained, which helps to determine global path planning and operations, improve the accuracy of automated operations, and enhance user experience.

[0105] It should be noted that the above description of the relevant process is only for example and explanation, and does not limit the scope of application of this specification. For those skilled in the art, various modifications and changes can be made to the process under the guidance of this specification. However, these modifications and changes are still within the scope of this specification.

[0106] Figure 3 is an exemplary flow chart of a multimodal model according to some embodiments of this specification. In some embodiments, process 300 can be executed by a processor based on an application scenario. Figure 3 As shown, process 300 includes the following steps.

[0107] In some embodiments, the operation control instruction includes application information and an action sequence, the operation control instruction is obtained according to the intention information and the first interface image, and the terminal user interface is operated through the operation control instruction, including:

[0108] Step S310: input the intention information and the first interface image into the multimodal model to obtain application information and an action sequence.

[0109] The multimodal model is a model used to determine operational control instructions.

[0110] In some embodiments, the input of the multimodal model may be intention information and a first interface image, and the output may be an operation control instruction.

[0111] In some embodiments, the processor may obtain a pre-trained model, and perform fine-tuning training based on the pre-trained model to obtain a multimodal model.

[0112] The pre-training stage refers to the stage of training the language model on large-scale data through unsupervised learning methods.

[0113] Unsupervised learning methods refer to the training methods in which language models learn the laws of language and language representation by reading a large amount of unlabeled text. For example, unsupervised learning methods include Masked Language Modeling (MLM) and Autoregressive Language Modeling (ALM). During the pre-training process, the language model can learn the contextual information of the language, such as the grammatical structure of sentences and the relationship between words. This information can help the language model better understand the language and perform better in subsequent tasks. At the same time, pre-training can also reduce the amount of labeled data required for specific tasks and reduce training costs.

[0114] In some embodiments, the pre-trained model can be a language model or a large language model pre-trained on the basis of a pre-training dataset. The pre-training dataset can be a general text corpus. The pre-training dataset can be composed of countless text sources, including books, articles, and websites. These data are carefully curated to ensure that human knowledge, language nuances, and cultural perspectives are fully reflected. The pre-training dataset is usually a large-scale dataset with rich features and samples.

[0115] The pre-trained model is suitable for a variety of scenarios. In actual applications, fine-tuning training can be used to obtain a language model or a large language model dedicated to a specific task.

[0116] In some embodiments, the processor can directly obtain an existing pre-trained pre-trained model through the network. For example, BERT, GPT or XL-Net can be used as a pre-trained model.

[0117] In some embodiments, the processor can obtain a pre-trained model through pre-training. For example, the pre-training data can be used to train GPT, etc., and the language expression ability of GPT can be further improved by using technologies such as supervised fine-tuning, feedback self-help, and human feedback reinforcement learning.

[0118] In some embodiments, the processor may construct a preset training corpus, perform fine-tuning training based on the initial multimodal model, and determine the multimodal model.

[0119] The preset training corpus refers to an information library including text data related to the automated operation of the application. In some embodiments, the preset training corpus may include sample operation instructions (i.e., training samples) and sample screenshots and their corresponding operation control instructions (i.e., training labels). In some embodiments, the processor may construct an automated operation training corpus for the application based on historical data and the operation control instructions corresponding to the historical data. In some embodiments, the operation control instructions corresponding to the sample data may be determined by manual annotation.

[0120] In some embodiments, when the processor constructs a preset training corpus, it can try to ensure the integrity of the corpus (such as whether the text content is complete, whether the sentences are complete, data index, etc.).

[0121] The initial multimodal model refers to a pre-trained model used to train a multimodal model. In some embodiments, the pre-trained model obtained in the above manner can be used as the initial multimodal model.

[0122] Fine-tuning training refers to the training phase in which the language model that has undergone the pre-training phase is fine-tuned based on a specific task through a supervised learning method.

[0123] Supervised learning method refers to a training method in which a language model that has undergone a pre-training phase is trained based on labeled data for a specific task.

[0124] In some embodiments, the processor may fine-tune the initial multimodal model through various fine-tuning strategies to determine the multimodal model. For example, the fine-tuning strategies include adaptive fine-tuning, multi-task learning, etc.

[0125] In some embodiments, the processor can be based on a plurality of first training samples with a first label, and can be trained by various methods to update the parameters of the initial multimodal model to obtain a multimodal model. For example, training can be performed based on a gradient descent method. As an example only, a plurality of sample medical text data with labels can be input into the initial multimodal model, a loss function can be constructed by the labels and the output results of the initial multimodal model, and the parameters of the initial multimodal model can be iteratively updated based on the loss function. When the loss function of the initial multimodal model meets the corresponding conditions, the model training is completed, and a trained multimodal model is obtained. Among them, the corresponding conditions can be that the loss function converges, the number of iterations reaches a threshold, etc.

[0126] Application information refers to information related to the target application. For example, application information may include the name and type of the target application. The target application refers to the specific application that needs to be operated, such as WeChat, alarm clock, navigation and other applications.

[0127] In some embodiments, the target application may include one or more, and the action sequence may include a set of multiple actions corresponding to one target application, or may include a set of multiple actions corresponding to different target applications.

[0128] An action sequence is a sequence of different actions arranged in chronological order.

[0129] Step S320: operating the terminal user interface according to the application information and the action sequence.

[0130] In some embodiments, the processor may perform responsive actions on the target application in the terminal user interface in a chronological order based on the application information and the action sequence.

[0131] In some embodiments of this specification,

[0132] In some embodiments, the action sequence includes at least one action to be performed,

[0133] Operate the terminal user interface based on application information and action sequences, including:

[0134] Acquire a second interface image, where the second interface image is the first interface image or an image of the terminal user interface after executing a preset action;

[0135] Inputting the application information, the action sequence and the second interface image into the multimodal model to obtain action information of the action to be executed, where the action information of the action to be executed includes at least one of an action name and coordinate information;

[0136] The terminal user interface is operated according to the action information and application information of the action to be executed.

[0137] The pending action refers to an action in the action sequence that currently needs to be executed.

[0138] The action information of the action to be executed includes at least one of a semantic description of the action to be executed and an action parameter of the action to be executed.

[0139] The semantic description of the action to be performed refers to a textual description of a specific operation or action. For example, the semantic description of the action to be performed may include the purpose, action, condition, etc. The action may include open, click, page down, page up, type, click and enter, return, stop, etc.

[0140] The action parameters of the action to be executed refer to a set of specific parameters corresponding to the execution of a specific action. For example, the action parameters may include action information, such as the action name, target element, coordinate information, element attributes, etc.

[0141] The target element refers to the interface element corresponding to a certain action. The target element can be a button, link, picture, etc.

[0142] The coordinate information refers to the specific location of the action, and the coordinate information can be expressed by screen coordinates (x, y).

[0143] An element ID is a unique identifier for an interface element, such as the ID of a button.

[0144] Element attributes refer to the properties of interface elements, such as class name, tag name, etc.

[0145] In some embodiments, the multimodal model can output at least one of a text description of the first interface image, application information, an action sequence, a semantic description of the action to be performed, and an initial task state based on the first interface image and the intent information.

[0146] The preset action is the previous action of the action to be executed.

[0147] When the action to be executed is the first action in the action sequence, the second interface image is the first interface image. When the action to be executed is other actions in the action sequence, the second interface image is the image of the terminal user interface after the preset action is executed. Other actions refer to actions after the first action in the action sequence.

[0148] The initial task state is the execution status before the to-be-executed action is executed. For example, the initial task state may include the execution status before a certain operation or action is executed, such as the execution status of the previous action before the to-be-executed action is executed.

[0149] The text description of the first interface image is a detailed description of the screen content. For example, the text description of the first interface image may include various elements, text content, buttons, input boxes, etc. on the interface.

[0150] In some embodiments, the multimodal model can determine the first prompt information based on the first interface image and the intention information; based on the first prompt information, output the application information and the action sequence; the multimodal model can determine the second prompt information based on the application information, the action sequence and the second interface image; based on the second prompt information, output the action information of the action to be performed.

[0151] The first prompt information refers to text input or instructions obtained by semantically understanding the first interface image and intention information.

[0152] The second prompt information refers to text input or instructions obtained by semantically understanding the application information, action sequence, and the second interface image.

[0153] In some embodiments of the present specification, global actions are planned and switched across applications through a multimodal model; anthropomorphic operations are cyclically performed on a selected target application through a multimodal model until the user's intention information is completed, thereby achieving both a highly accurate global plan and specific click parameters for each step, which helps to improve the accuracy of automated operations.

[0154] In some embodiments, the application information includes at least one target application, and the action to be performed is an execution action of the target application corresponding to the action to be performed;

[0155] The terminal user interface is operated according to the action information and application information of the action to be executed, including:

[0156] The target application corresponding to the action to be executed is operated through the action information of the action to be executed.

[0157] In some embodiments, based on the action information of the action to be performed, the corresponding target application performs a specific action by simulating user operations, for example: opening the WeChat application, entering text in the chat box, clicking the "Send" button, etc.

[0158] In some embodiments of the present specification, automating operations on target applications can improve efficiency, reduce manual operation errors, and help achieve user intent.

[0159] In some embodiments, the intention information and the first interface image are input into the multimodal model to obtain application information and an action sequence, including:

[0160] Based on the first interface image, intention information and auxiliary information, semantic understanding is performed through a multimodal model to obtain application information and action sequences.

[0161] Auxiliary information refers to document or paragraph information related to the target application.

[0162] In some embodiments, document fragments related to the target application can be retrieved from external knowledge bases (such as Wikipedia, web pages, search engines, etc.) as auxiliary information based on a retrieval model (such as DPR, Dense Passage Retrieval) in a Retrieval-Augmented Generation (RAG) framework.

[0163] In some embodiments, when the input of the multimodal model includes auxiliary information, the first training sample also includes sample auxiliary information.

[0164] In some embodiments of the present specification, auxiliary information is used to enable the multimodal model to output correct actions, which is not limited to supervised fine-tuning training data, greatly enhancing the generalization of the multimodal model, allowing the multimodal model to perform correct operations on trained models and user intentions, thereby reducing development and maintenance costs.

[0165] In some embodiments, the auxiliary information includes offline document information and / or online document information of the target application, at least one of a reference action trajectory and a user action trajectory. The reference action trajectory is an action sequence for application information generated by the terminal system, and the user action trajectory is an action sequence for application information customized by the user.

[0166] The offline document information of the target application refers to documents, manuals, help files, etc. related to the target application that are stored locally.

[0167] The online document information of the target application refers to documents, manuals, help files, etc. related to the target application stored on the Internet.

[0168] In some embodiments, the offline document information of the target application may include static resources such as offline usage instructions of the target application, target application operation instructions, etc. The online document information of the target application includes dynamic resources such as question and answer information related to the target application, official support pages, forum discussions, and update logs.

[0169] In some embodiments, offline document information of the target application may be obtained through a storage device, and online document information of the target application may be obtained through a web page.

[0170] The reference action trajectory is a standard step or recommended path predefined by the terminal to complete specific intent information. For example, if the target application is image editing software, the reference action trajectory includes action sequences such as opening a picture, adjusting brightness, and saving a picture.

[0171] A user action trajectory is a personalized path of steps or operations that users can create based on their preferences or workflow habits. For example, some users may prefer to adjust the contrast first and then the brightness. The user action sequence can be the same as or different from the reference action trajectory.

[0172] A path is a series of steps or actions recorded during the execution of a task.

[0173] A task usually refers to the intention information that the user wants to complete through the terminal. A task can be a single step, such as opening a file or sending an email; it can also involve multiple steps, such as creating and editing a document, performing data analysis, etc.

[0174] In some embodiments, reference motion trajectories and user motion trajectories of different applications can be obtained in a variety of ways, such as conducting experiments based on different applications, and determining the reference motion trajectory and user motion trajectory of a certain application through log recording, screen recording, user behavior analysis, etc. In some embodiments, the reference motion trajectories and user motion trajectories of different applications can be stored in a storage device, and the processor retrieves the corresponding reference motion trajectory and user motion trajectory related to a certain target application from the storage device based on the first interface image and the intent information.

[0175] In some embodiments, document information of the target application may be retrieved from an external knowledge base (eg, Wikipedia, web pages, search engines, etc.).

[0176] In some embodiments of the present specification, enhancing the learning ability of the multimodal model through sources such as offline application instruction documents, online search engine documents, self-exploration trajectories of the multimodal model, and user demonstration trajectories helps to enhance the generalization ability of the model.

[0177] In some embodiments, the application information, the action sequence, and the second interface image are input into the multimodal model to obtain action information of the action to be performed, including:

[0178] In the case where the action to be performed is the first target action in the action sequence, the application information, the action sequence and the second interface image are input into the multimodal model to obtain action information of the action to be performed;

[0179] In the case where the action to be performed is the second target action in the action sequence, the application information, the action sequence, the second interface image and the historical action record are input into the multimodal model to obtain action information of the action to be performed;

[0180] The second target action is the next action of the first target action, and the historical action record includes action information of the executed action before the second target action and the execution result of the executed action.

[0181] The first target action is an action in an action sequence. It is the first specific operation that the user or system plans to perform.

[0182] The first target action does not need to rely on the executed action. For example, when processing the first target action, the input of the multimodal model includes application information, the entire action sequence, and the current interface image (ie, the second interface image).

[0183] The second target action is an action after the first target action in the action sequence. The second target action can be any action after the first target action, or the second target action can be the next action after the first target action.

[0184] The execution of the second target action depends on the result of the first target action. The second target action can be the next operation based on the state or result after the previous action is completed. For example, when processing the first target action, the input of the multimodal model includes application information, action sequence, and current interface image, as well as historical action records, that is, the action information of the first target action and its execution result, which helps to more accurately understand and predict the next operation.

[0185] Historical action records refer to the records of various historical operations.

[0186] It should be noted that when executing the next action, the action information of the previous action and the execution result corresponding to the action are historical operations.

[0187] In some embodiments of the present specification, by combining historical action records, action information of the action to be executed that conforms to the actual situation can be output, so that the action information of the executed action is more accurate and fine-grained.

[0188] In some embodiments, the historical action record includes action information of the historical action and the execution result of the historical action.

[0189] The execution result refers to the final result or status after executing a certain operation or action. For example, the execution result may include execution completion information, execution failure information, and corresponding error information.

[0190] In some embodiments, the historical action record includes action information of a previous action and an execution result of the previous action.

[0191] In some embodiments of the present specification, through the execution results of historical actions, the multimodal model can avoid repeated operations and invalid operations and improve the accuracy of task execution; at the same time, the execution result records can help users optimize the behavior of the multimodal model and discover and solve potential problems; the action information of historical actions can help the multimodal model manage multiple tasks and ensure coordination and consistency between tasks.

[0192] In some embodiments, operating a target application corresponding to the action to be performed using the action information of the action to be performed includes:

[0193] Based on the action information, the preset tool is controlled to operate the target application corresponding to the action to be executed to obtain a task status, and the task status is used to indicate the execution status of the action to be executed.

[0194] The task status is the feedback to the user on the progress of the action to be executed and the next operation suggestion based on the execution result of the action to be executed and the user's request during the execution of the task. For example, the task status may include the task status of the completed action after executing a certain operation or action, before, during or after the execution of the action to be executed, etc. For example, the task status may include "continue", "complete" or "pause", etc.

[0195] "Continue" means that the action to be executed has been completed and the next action will be executed.

[0196] "Complete" means that the entire action sequence has been executed.

[0197] "Pause" means that the action to be executed is temporarily stopped, waiting for further user operation or waiting for specific conditions to be met before continuing.

[0198] A preset tool refers to a tool or module used to perform a specific operation or action. For example, a preset tool may be an execution module, which is used to actually interact with an application, such as simulating user clicks or keyboard input. The execution module may use toolkits such as Android simulation assistance to detect and operate UI (User Interface) controls, and actually perform the action to be performed (such as clicking text, icons, inputting text, sliding the screen, etc.) based on the action information of the action to be performed, and display the execution process on the terminal user interface of the terminal.

[0199] In some embodiments, the processor may determine and update the current task status based on the execution result of the pending action after executing the pending action. For example, when it is detected that the execution result of the pending action is completed, it is determined whether there are any unexecuted actions after the pending action, and the task status is determined to be "continued" or "completed".

[0200] In some embodiments of the present specification, the task status helps the multimodal model determine the progress of the action execution.

[0201] In some embodiments, the terminal user interface may be controlled to display the execution process and / or task status of the action to be executed.

[0202] In some embodiments of this specification, by displaying the execution process of the action to be executed and / or displaying the task status, the user can understand the operation progress of the multimodal model at any time, improving the transparency of action execution and user experience. The task status enables the user to better interact with the multimodal model and ensure the smooth progress of the action sequence.

[0203] In some embodiments, the preset tool may be controlled to play preset voice information based on the task status.

[0204] A preset voice message refers to a predefined voice message.

[0205] In some embodiments, in response to the execution result of the action to be executed being completed, the execution module can feedback the execution result to the multimodal model, and broadcast a preset voice message corresponding to the execution result through the terminal, such as "The corresponding page has been opened for you" and other content, to remind the user of the execution status of the action to be executed or to prompt the user for information that needs manual confirmation.

[0206] In some embodiments of the present specification, by presetting voice information, the user can understand the operation progress of the multimodal model at any time and intervene or adjust when necessary, thereby enhancing the user's sense of control.

[0207] In some embodiments, operating the target application corresponding to the action to be executed to obtain the task status includes:

[0208] When the second interface image does not meet the execution condition, the task state is the first task state, and the first task state is used to indicate re-capturing the terminal user interface.

[0209] The execution condition is a judgment condition for evaluating whether the second interface image can be executed. For example, the execution condition may include that there are unrecognized control items in the second interface image, the target control item is blocked, the definition of the screenshot is not enough, etc.

[0210] The target control item refers to the specific interface element where the action is to be performed, such as a button, input box, etc.

[0211] In some embodiments, the first task status may be displayed as a “screenshot” icon or prompt on the terminal user interface.

[0212] In some embodiments, a screenshot tool (such as an Android Debug Bridge (ADB) command, a screenshot interface, etc.) can be controlled to obtain a second interface image; control items on the complete image of the current screen, such as buttons, input boxes, etc., can be identified through various methods such as template matching, feature extraction, or deep learning; and a judgment is made as to whether the second interface image meets the execution conditions: if not, the first task status is displayed through the terminal user interface, for example, feedback information of "screenshot" is displayed through voice, text, images, etc.

[0213] In some embodiments of the present specification, by obtaining a screenshot of a smaller area, interference information can be reduced and the recognition accuracy of the control item can be improved.

[0214] In some embodiments, the method further comprises:

[0215] When the task state is the first task state, based on the preset area, re-capturing the terminal user interface to obtain an interface image of the preset area;

[0216] The second interface image is updated based on the interface image of the preset area, and the operation of inputting the application information, the action sequence and the second interface image into the multimodal model is re-executed.

[0217] In some embodiments, the area where the target control item is located can be determined as a preset area, and a screenshot tool (such as an ADB command, a screenshot interface, etc.) can be controlled to capture the preset area, obtain an interface image of the preset area, and update the second interface image based on the interface image of the preset area.

[0218] A preset area refers to a specified area in the terminal user interface.

[0219] In some embodiments, the updated second interface image, application information, and action sequence may be input into a multimodal model, and action information of the action to be performed may be output.

[0220] In some embodiments of the present specification, by executing actions based on screenshots of smaller areas, the risk of misoperation can be reduced, especially when there are multiple similar control items on the interface, thereby improving the accuracy of action execution and thus improving the model's work efficiency and user experience.

[0221] In some embodiments, the to-be-executed action includes a first to-be-executed action and a second to-be-executed action, the second to-be-executed action is the next action of the first to-be-executed action, and the target application includes a first target application and a second target application;

[0222] The target application corresponding to the action to be executed is operated to obtain the task status, including:

[0223] When the first action to be executed is an action to operate on the first target application, and the second action to be executed is an action to operate on the second target application, the task state is the second task state, and the second task state is used to indicate switching from the first target application to the second target application.

[0224] The second task state is used to indicate switching from the first target application in the application information to the second target application.

[0225] In some embodiments, the second task state can be displayed as a "switch" icon or prompt on the terminal user interface. In some embodiments, the preset tool can be controlled to execute the first action to be executed on the first target application. After the execution result of the first action to be executed is completed, the action information of the second action to be executed is output based on the multimodal model. Based on the action information of the second action to be executed, the preset tool is controlled to execute the second action to be executed on the second target application, and so on, until the actions in all action sequences are executed.

[0226] In some embodiments of the present specification, by indicating the application switching, the user can clearly understand the current operating status of the multimodal model, ensure the continuity of actions between different applications, and avoid interruptions and omissions.

[0227] In some embodiments, operating the target application corresponding to the action to be executed to obtain the task status includes:

[0228] When the action to be executed is a sensitive action, the task state is a third task state, and the third task state is used to indicate that the action to be executed is suspended;

[0229] Based on the third task state, obtaining a first feedback of the user input;

[0230] Based on the first feedback, the preset tool is controlled to continue executing the action to be executed.

[0231] Sensitive actions refer to operations involving user-related information, etc. For example, sensitive actions may include sending emails and other sensitive actions.

[0232] In some embodiments, the third task status may be displayed as a “pause” icon or prompt on the terminal user interface.

[0233] The first feedback refers to the user's confirmation or rejection of a sensitive action.

[0234] In some embodiments, if the action to be executed includes a sensitive action, before executing the sensitive action, the preset tool can be controlled to pause execution, and the terminal user interface can be controlled to display a third task status (such as "pause, confirm whether to send an email", etc.) to prompt the user to confirm through the terminal; after detecting and obtaining the confirmation operation, the preset tool is controlled to execute the action to be executed based on the action information of the action to be executed, and the task status is updated based on the execution result of the action to be executed to obtain an updated task status.

[0235] In some embodiments of the present specification, when an action requires user confirmation or provision of additional information, the multimodal model may pause the operation and wait for user input to ensure the accuracy of the action execution.

[0236] In some embodiments, operating the target application corresponding to the action to be executed to obtain the task status includes:

[0237] After executing the action to be executed, the execution status of the action to be executed is evaluated to obtain the execution result of the action to be executed;

[0238] Based on the execution results, the task status is determined.

[0239] Evaluation is the process of determining whether the execution of an action to be performed meets expectations.

[0240] The execution result is used to indicate the evaluation result of the execution status of the action after the execution of the action to be executed. For example, the execution result may include one or a combination of execution completion, execution failure, corresponding error information, etc.

[0241] In some embodiments, the execution result of the action to be executed can be obtained in a variety of ways. For example, the terminal user interface can be controlled to display the execution process of the action to be executed, and the execution result of the action to be executed can be determined by manual input after the action to be executed. For another example, after the action to be executed is executed, the log record of the terminal can be obtained to determine whether the action to be executed is completed and the execution result of the action to be executed can be determined.

[0242] In some embodiments, the task status can be updated based on the execution result of the action to be executed to obtain an updated task status. For example, in response to the successful execution of the action to be executed, if there is no unexecuted action after the execution of the action to be executed, the task status is updated, and the updated task status is "completed"; if there is still an unexecuted action after the execution of the action to be executed, the task status is updated, and the updated task status is "continued"; for another example, in response to the failure of the execution of the action to be executed, the task status is updated, and the updated task status is "failed", etc.

[0243] In some embodiments of the present specification, by evaluating the execution results of the task, the multimodal model can verify whether the action is performed as expected, thereby ensuring the correctness of the action execution.

[0244] In some embodiments, before determining the task status based on the execution result, the method further includes:

[0245] After executing the action to be executed, capture the interface image after execution;

[0246] Based on the interface image after execution, the execution status of the action to be executed is evaluated to obtain the execution result of the action to be executed.

[0247] The post-execution interface image refers to the image of the terminal user interface captured after a certain action is executed.

[0248] In some embodiments, after the action to be executed is executed, the execution module may be controlled to capture an interface image of the terminal user interface to obtain an interface image after execution.

[0249] In some embodiments, the execution of the action to be executed can be evaluated in a variety of ways based on the interface image after execution to obtain the execution result of the action to be executed. For example, objects in the interface image after execution can be identified by various image processing methods or machine learning methods. For example, the processor can extract multiple objects in the interface image after execution based on an image processing algorithm (such as a support vector machine, a neural network, etc.), and determine whether the multiple objects include a target object. If the target object is included, the execution result is determined to be completed; if the target object is not included, the execution result is determined to be execution failure.

[0250] The target object refers to the expected result after executing the action to be executed. For example, the target object may include keywords, key features, etc. corresponding to the action.

[0251] Different actions to be performed correspond to different target objects, and their corresponding relationships can be determined based on prior knowledge or historical data.

[0252] In some embodiments, when it is detected that the execution result is an execution failure, the corresponding error information when the execution fails (such as a crash, an application crash, etc.) can be determined based on the log records of the terminal.

[0253] In some embodiments, the error message may include slow page loading or application crash.

[0254] In some embodiments of the present specification, by capturing and evaluating interface screenshots, the multimodal model can promptly discover errors in action execution and improve the overall success rate of action execution.

[0255] Figure 4It is an exemplary schematic diagram of iterative execution according to some embodiments of the present specification.

[0256] In some embodiments, after operating the target application corresponding to the action to be executed to obtain the task status, the method further includes:

[0257] When the task status satisfies the preset conditions, the operation of inputting the application information, the action sequence and the second interface image into the multimodal model is re-executed.

[0258] In some embodiments, Figure 4 As shown, the operation of inputting the application information 410-1, the action sequence 410-2 and the second interface image 410-3 into the multimodal model 420 can be performed based on at least one round of iteration, and the at least one round of iteration includes: processing the application information 410-1, the action sequence 410-2, the second interface image 410-3 and the historical action record 410-4 based on the multimodal model, generating the action information 430 of the action to be executed in the current round of iteration, and controlling the preset tool to execute the action of the current round of iteration to obtain the execution result 440 of the current round; based on the execution result 440 of the current round of iteration, determining the task status 450 of the current round of iteration.

[0259] In some embodiments, the input of the multimodal model is related to the iteration round. In the first iteration round, the input of the multimodal model includes the intent information and the first interface image, and the first interface image is a screenshot of the terminal user interface corresponding to the acquisition of the user intent information; in subsequent iteration rounds, the input of the multimodal model includes application information, action sequence, historical action record and second interface image.

[0260] In some embodiments, in at least one non-last iteration of the iteration, the output of the multimodal model is the action information of the action to be performed in the current round. In the last iteration, the output of the multimodal model is the prompt information of completion.

[0261] In some embodiments, in the first round of iterations, the multimodal model can be used to process the intention information and the first interface image to generate application information and action sequences for the first round of iterations, and then the application information, action sequences, historical action records, and the second interface image are input into the multimodal model to obtain action information for the first round of iterations, and the preset tool is controlled to perform the actions for the first round of iterations to obtain the iteration results for the first round of iterations (including the action information for the first round of iterations and the corresponding task status), and whether the preset conditions are met can be determined based on the corresponding task status of the first round; in response to the corresponding task status of the current round being to continue, the next round of iterations is continued based on the application information, action sequence, historical action records, and the second interface image.

[0262] In some embodiments, in at least one non-last iteration of an iteration, the multimodal model processes the application information, action sequence, historical action record, and second interface image, generates action information of the current iteration, and controls the preset tool to execute the action of the current iteration to obtain the iteration result of the current iteration (including the action information of the current iteration and the corresponding task status); based on the task status corresponding to the current round, it can be judged whether the preset conditions are met; in response to the task status corresponding to the current round being continued, based on the execution result of the current round, the historical action record is updated, and the second interface image of the next iteration is intercepted, and the updated historical action record, application information, action sequence, and the second interface image of the next iteration are used as the input of the next iteration. In the last iteration, it can be judged whether the preset conditions are met based on the task status corresponding to the last round; in response to the task status corresponding to the last round being completed or failed to execute, a termination prompt signal is issued, and the iteration is terminated.

[0263] The preset condition is a judgment condition for evaluating whether the iteration is terminated. For example, the preset condition may include the corresponding task status being failed, the corresponding task status being completed, or the user's intention information being reacquired.

[0264] In some embodiments of the present specification, by traversing and executing each action in the action sequence, it can be ensured that the multimodal model accurately executes each action in the action sequence and effectively completes the user's intention.

[0265] In some embodiments, the method further comprises:

[0266] Detecting and obtaining a second feedback from the user, where the second feedback is used to indicate the intention information fed back by the user during the execution process, and updating the intention information based on the second feedback from the user;

[0267] Based on the second feedback from the user, taking a screenshot of the current terminal user interface to update the first interface image;

[0268] Based on the intention information and the first interface image, the operation of obtaining the operation control instruction according to the intention information and the first interface image and operating the terminal user interface through the operation control instruction is re-executed.

[0269] The current time is the moment or time period for capturing the terminal user interface. The current time can be any time point, depending on the time of executing the relevant instructions of each round of actions. For example, the terminal user interface can be re-captured at a specified time after executing each action to obtain the first interface image.

[0270] In some embodiments, the second feedback may include relevant content replied by the user after evaluation based on the execution result. For example, the second feedback may include intention information re-entered by the user, etc.

[0271] In some embodiments, the second feedback may occur in any round of iteration. The user may input the second feedback into the terminal by interacting with the terminal user interface. For example, the user may interact with the terminal user interface by text input, voice input, picture input, etc., and input the second feedback into the terminal. The terminal may communicate with the processor and send the second feedback. The processor may determine the feedback intention information based on the operation intention in response to reacquiring the user's operation intention.

[0272] The updated first interface image is the interface image corresponding to the second feedback moment. For example, the updated first interface image is an image of the terminal user interface obtained by taking a screenshot of the current terminal user interface at the same time as or at a specified interval when the second feedback is obtained.

[0273] In some embodiments, based on the updated intent information and the updated first interface image, the corresponding application information and action sequence can be obtained through the trained multimodal model in a similar manner as described above, and the corresponding application information, action sequence and second interface image can be input into the multimodal model to obtain the action information of the action to be executed as the operation control instruction. For more information, please refer to the relevant description above.

[0274] In some embodiments of the present specification, through user feedback, the multimodal model can re-output execution information to ensure the accuracy and flexibility of the task.

[0275] Figure 5 It is an exemplary structural diagram of a multimodal model according to some embodiments of the present specification.

[0276] In some embodiments, Figure 5 As shown, the first network layer 510 of the multimodal model 420 can be used to perform feature extraction on the intent information 511 to obtain a first feature sequence 530; the second network layer 520 of the multimodal model 420 can be used to perform feature extraction on the first interface image 512 to obtain a second feature sequence 540; the third network layer 550 of the multimodal model 420 can be used to perform feature fusion on the first feature sequence 530 and the second feature sequence 540 to obtain application information and action sequence 560.

[0277] In some embodiments, the input of the first network layer is intent information, and the output is a first feature sequence.

[0278] The first feature sequence is a feature vector or feature matrix obtained by extracting features from the intention information.

[0279] The first feature sequence contains the semantic information of the intent information, which can reflect the content and intent of the intent information.

[0280] In some embodiments, the first network layer may be BERT, LSTM (Long Short-Term Memory), etc.

[0281] In some embodiments, the input of the second network layer is the first interface image, and the output is the second feature sequence.

[0282] The second feature sequence is a feature vector or a feature matrix obtained by extracting features from the first interface image.

[0283] The second feature sequence includes visual information of the first interface image, and can reflect the state and content of the current screen.

[0284] In some embodiments, the second network layer may include a visual encoder and a projection layer. The visual encoder is used to preprocess the first interface image. For example, the preprocessing may include operations such as adjusting the image size, removing noise, and enhancing contrast to ensure the clarity and quality of the image data, and obtain a first interface image of a preset size. The visual encoder may be a Vision Transformer (ViT), etc. The projection layer is used to encode the first interface image of a preset size. For example, a ResNet50 model is used to perform a convolution operation on the image to extract high-level visual features. The projection layer outputs a feature vector of a fixed length as a second feature sequence. The second feature sequence may include different elements on the screen and their position information.

[0285] In some embodiments, the input of the third network layer includes the first feature sequence and the second feature sequence, and the output includes application information and an action sequence.

[0286] In some embodiments, the third network layer is a pre-trained large language model (LLM).

[0287] In some embodiments, the first network layer, the second network layer, and the third network layer can be obtained by training based on various feasible methods, for example, a gradient descent method, etc. For more information about the training, please refer to the above description.

[0288] In some embodiments of the present specification, through the first network layer, the second network layer, and the third network layer, the multimodal model can capture contextual information between different modalities and enhance the richness and accuracy of feature representation.

[0289] Figure 6 is a schematic diagram of the structure of an electronic device according to some embodiments of this specification. Figure 6As shown, the electronic device 600 may include: a processor 601, a memory 602. The electronic device 600 may also include one or more of a multimedia component 603, an input / output (I / O) component 604, and a communication component 605. In this embodiment, the electronic device 600 may be a device for implementing the terminal operation method provided in this embodiment.

[0290] The processor 601 is used to control the overall operation of the electronic device 600 to complete all or part of the steps in the terminal operation method described above. The memory 602 is used to store various types of data to support the operation of the electronic device 600, which may include instructions for any application or method used to operate on the electronic device 600, as well as application-related data, such as contact data, messages sent and received, pictures, audio, video, etc. The memory 602 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read only memory (PROM), read only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The multimedia component 603 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signal may be further stored in the memory 602 or sent through the communication component 605. The audio component also includes at least one speaker for outputting audio signals. The I / O component 604 provides an interface between the processor 601 and other interface modules, and the above-mentioned other interface modules may be keyboards, mice, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 605 is used for wired or wireless communication between the electronic device 600 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, Narrow Band Internet of Things (NB-IOT), Enhanced Narrow Band Internet of Things technology (Enhanced Machine-Type Communication, eMTC), or other 5G, etc., or a combination of one or more of them, is not limited here. Therefore, the corresponding communication component 605 may include: Wi-Fi module, Bluetooth module, NFC module, etc.

[0291] In an exemplary embodiment, the electronic device 600 can be implemented by one or more application specific integrated circuits (ASIC), digital signal processors (DSP), digital signal processing devices (DSPD), programmable logic devices (PLD), field programmable gate arrays (FPGA), controllers, microcontrollers, microprocessors or other electronic components to execute the above-mentioned terminal operation method.

[0292] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, and when the program instructions are executed by a processor, the steps of the terminal operation method described above are implemented. For example, the computer-readable storage medium may be the memory 602 including the program instructions described above, and the program instructions may be executed by the processor 601 of the electronic device 600 to implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present application;

[0293] Or, when the instructions are executed by a computer, they are executed to implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present application.

[0294] The present application also provides a vehicle, on which is disposed the electronic device provided by any of the above embodiments, and the electronic device is used to execute the terminal operation method provided by any of the above embodiments.

[0295] In one embodiment, the vehicle may be configured to be in a fully or partially automated driving mode. For example, the vehicle may control itself while in automated driving mode, and may determine the current state of the vehicle and its surroundings through human operation, determine the possible behavior of at least one other vehicle in the surroundings, and determine the confidence level corresponding to the possibility of the other vehicle performing the possible behavior, and control the vehicle based on the determined information. When the vehicle is in automated driving mode, the vehicle may be set to operate without human interaction.

[0296] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this application, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0297] The embodiments, implementation methods and related technical features of the present application can be combined and replaced with each other without conflict.

[0298] The above are only preferred embodiments of the present application and do not constitute any form of limitation to the present application. Although the descriptions of various embodiments in the embodiments of the present application have different focuses, for parts not described in detail in a certain embodiment, reference can be made to the relevant embodiments of other embodiments. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application are still within the scope of the technical solution of the present application.

Claims

1. A terminal operation method, characterized in that: The method comprises: When the terminal responds to the user's operation intention, acquiring the intention information of the operation intention and the first interface image of the terminal user interface; An operation control instruction is obtained according to the intention information and the first interface image, and the terminal user interface is operated according to the operation control instruction.

2. The method according to claim 1, characterized in that The operation control instruction includes application information and an action sequence, and obtaining the operation control instruction according to the intention information and the first interface image, and operating the terminal user interface through the operation control instruction includes: Inputting the intention information and the first interface image into a multimodal model to obtain the application information and the action sequence; The terminal user interface is operated according to the application information and the action sequence.

3. The method according to claim 2, characterized in that The action sequence includes at least one action to be executed, and the terminal user interface is operated according to the application information and the action sequence, including: Acquire a second interface image, where the second interface image is the first interface image or an image of the terminal user interface after a preset action is performed; Inputting the application information, the action sequence and the second interface image into the multimodal model to obtain action information of the action to be executed, wherein the action information of the action to be executed includes at least one of an action name and coordinate information; The terminal user interface is operated according to the action information of the action to be executed and the application information.

4. The method according to claim 3, characterized in that The application information includes at least one target application, and the to-be-executed action is an execution action for the target application corresponding to the to-be-executed action; The operating the terminal user interface according to the action information of the action to be executed and the application information includes: The target application corresponding to the action to be executed is operated according to the action information of the action to be executed.

5. The method according to claim 3, characterized in that: The step of inputting the intention information and the first interface image into a multimodal model to obtain application information and an action sequence includes: Based on the first interface image, the intent information and the auxiliary information, semantic understanding is performed through the multimodal model to obtain the application information and the action sequence.

6. The method according to claim 5, characterized in that The auxiliary information includes at least one of offline document information and / or online document information of the target application, a reference action trajectory and a user action trajectory, wherein the reference action trajectory is an action sequence for the application information generated by the terminal itself, and the user action trajectory is an action sequence for the application information customized by the user.

7. The method according to claim 3, characterized in that The step of inputting the application information, the action sequence and the second interface image into the multimodal model to obtain the action information of the action to be executed includes: In a case where the action to be performed is the first target action in the action sequence, inputting the application information, the action sequence and the second interface image into the multimodal model to obtain action information of the action to be performed; In a case where the action to be performed is a second target action in the action sequence, inputting the application information, the action sequence, the second interface image, and the historical action record into the multimodal model to obtain action information of the action to be performed; The second target action is the next action of the first target action, and the historical action record includes action information of an executed action before the second target action and an execution result of the executed action.

8. The method according to claim 4, characterized in that The operating the target application corresponding to the action to be executed by using the action information of the action to be executed includes: Based on the action information, a preset tool is controlled to operate a target application corresponding to the action to be executed to obtain a task status, where the task status is used to indicate an execution status of the action to be executed.

9. The method according to claim 8, characterized in that The method further comprises: Control the terminal user interface to display the execution process of the action to be executed and / or the task status.

10. The method according to claim 8, characterized in that The operating the target application corresponding to the to-be-performed action to obtain the task status includes: When the second interface image does not satisfy the execution condition, the task state is the first task state, and the first task state is used to indicate to re-capture the terminal user interface.

11. The method according to claim 10, characterized in that The method further comprises: When the task state is the first task state, based on a preset area, re-capturing the terminal user interface to obtain an interface image of the preset area; The second interface image is updated based on the interface image of the preset area, and the operation of inputting the application information, the action sequence and the second interface image into the multimodal model is re-executed.

12. The method according to claim 8, characterized in that The to-be-executed action includes a first to-be-executed action and a second to-be-executed action, the second to-be-executed action is the next action of the first to-be-executed action, and the target application includes a first target application and a second target application; The operating the target application corresponding to the to-be-performed action to obtain the task status includes: When the first action to be executed is an action to operate on the first target application, and the second action to be executed is an action to operate on the second target application, the task state is a second task state, and the second task state is used to indicate switching from the first target application to the second target application.

13. The method according to claim 8, characterized in that The operating the target application corresponding to the to-be-performed action to obtain the task status includes: In the case where the action to be executed is a sensitive action, the task state is a third task state, and the third task state is used to indicate that the action to be executed is suspended; Based on the third task state, obtaining first feedback input by a user; Based on the first feedback, the preset tool is controlled to continue executing the action to be performed.

14. The method according to claim 8, characterized in that The operating the target application corresponding to the to-be-performed action to obtain the task status includes: After executing the action to be executed, evaluating the execution status of the action to be executed to obtain the execution result of the action to be executed; Based on the execution result, the task status is determined.

15. The method according to claim 14, characterized in that Before determining the task status based on the execution result, the method further includes: After executing the to-be-executed action, capturing an interface image after the execution; Based on the interface image after execution, the execution status of the action to be executed is evaluated to obtain the execution result of the action to be executed.

16. The method according to claim 8, characterized in that After the target application corresponding to the to-be-performed action is operated to obtain the task status, the method further includes: When the task status satisfies a preset condition, the operation of inputting the application information, the action sequence and the second interface image into the multimodal model is re-executed.

17. The method according to any one of claims 2 to 16, characterized in that: The method further comprises: Detecting and obtaining second feedback from the user, where the second feedback is used to indicate intention information fed back by the user during the execution process, and updating the intention information based on the second feedback from the user; Based on the second feedback from the user, taking a screenshot of the current terminal user interface to update the first interface image; Based on the intention information and the first interface image, the operation of obtaining the operation control instruction according to the intention information and the first interface image and operating the terminal user interface according to the operation control instruction is re-executed.

18. The method according to claim 2, characterized in that The step of inputting the intention information and the first interface image into a multimodal model to obtain the application information and the action sequence includes: Extracting features of the intent information based on the first network layer of the multimodal model to obtain a first feature sequence; Extracting features of the first interface image based on the second network layer of the multimodal model to obtain a second feature sequence; Based on the third network layer of the multimodal model, feature fusion is performed on the first feature sequence and the second feature sequence to obtain the application information and the action sequence.

19. A computer-readable storage medium having a computer program stored thereon, characterized in that: The computer-readable storage medium stores instructions, which, when executed by a computer, enable the computer to implement the terminal operation method according to any one of claims 1 to 18.

20. A computer program product, characterized in that The computer program product stores instructions, which, when executed by a computer, enable the computer to implement the terminal operation method according to any one of claims 1 to 18.

21. An electronic device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the terminal operation method described in any one of claims 1 to 18.

22. A vehicle, characterized in that: An electronic device comprising the electronic device described in claim 21.