Multi-modal data processing method and device, equipment, storage medium and program product
By combining interface and feedback models, interface information is automatically acquired and processed to generate action information, solving the problem of users performing multiple operations on electronic devices and achieving more efficient task completion and model iteration.
Patent Information
- Application Number
- CN202510867080.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-11-18
AI Technical Summary
Users need to perform multiple operations through the graphical user interface when controlling electronic devices, which makes the work difficult and labor-intensive.
By using multimodal data processing methods, interface information is automatically acquired using interface and feedback models to generate action information and control device execution. The accuracy of the model is continuously improved by combining training data.
It reduces the number of interactions between users and devices, decreases the difficulty and workload of task completion, and improves the accuracy of task processing.
Smart Images

Figure CN120973271A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure belongs to the technical field of artificial intelligence, and in particular, relates to a multi-modal data processing method and device, electronic equipment, non-transitory computer-readable storage medium, and computer program product. BACKGROUND
[0002] With the development of science and technology, electronic equipment is increasingly widely used and has more and more functions, and has become one of the necessities in people's daily life. Electronic equipment generally integrates a display screen, so that a user can control the electronic equipment through a graphical user interface (GUI) displayed by the display screen. However, when controlling the electronic equipment to process some tasks, the user usually needs to control the electronic equipment through the GUI multiple times, resulting in a large work difficulty and workload of the user. SUMMARY
[0003] To overcome the problems in the related art, the present disclosure provides a multi-modal data processing method and device, electronic equipment, non-transitory computer-readable storage medium, and computer program product, which can reduce the work difficulty and workload of the user and improve the accuracy of task processing.
[0004] According to a first aspect of an embodiment of the present disclosure, a multi-modal data processing method is provided, and the method includes: obtaining first task information to be executed by a first device; in response to the first task information, obtaining first interface information of the first device; processing the first interface information through an interface model to obtain first action information; controlling the first device to execute the first action information; obtaining second interface information of the first device; processing the second interface information through a feedback model to determine first execution information of the first task information; determining, according to the first execution information of the first task information, first execution track information of the first task information as first training data, the first execution track information including the first interface information, the first action information, and the second interface information; and training the interface model using the first training data.
[0005] In some possible implementation manners, processing the second interface information through the feedback model to determine the first execution information of the first task information includes: if it is determined that a first task corresponding to the first task information is executed completely by processing the second interface information through the feedback model, first indication information indicating that the first task is executed completely is generated; and if it is determined that the first task is not executed completely by processing the second interface information through the feedback model, second indication information indicating that the first task is not executed completely is generated. The first execution information includes the first indication information and the second indication information.
[0006] In some possible implementation manners, processing the second interface information through the feedback model to determine the first execution information of the first task comprises: processing the first interface information, the second interface information and the first task information through the feedback model to determine the first execution information of the first task.
[0007] In some possible implementation manners, processing the first interface information, the second interface information and the first task information through the feedback model to determine the first execution information of the first task comprises: if it is determined through the feedback model that the first task information corresponds to a completed first task, generating first indication information indicating that the first task is completed; or if it is determined through the feedback model that the first task information corresponds to an uncompleted first task, generating second indication information indicating that the first task is uncompleted. The first execution information comprises the first indication information and the second indication information.
[0008] In some possible implementation manners, determining, according to the first execution information of the first task information, that the first execution trajectory information of the first task information is first training data comprises: processing the first execution information through a first screening model to determine the first execution trajectory information matching the interface model as the first training data. The first execution information comprises first indication information indicating that the first task is completed and first execution path information of the first task.
[0009] In some possible implementation manners, determining, according to the first execution information of the first task information, that the first execution trajectory information of the first task information is first training data comprises: processing the first execution information and the first execution trajectory information through a first screening model to determine the first execution trajectory information as the first training data. The first execution information comprises first indication information indicating that the first task is completed.
[0010] In some possible implementation manners, the method further comprises: training the interface model by using a training data set, wherein the training data set comprises training task information and training execution trajectory information thereof, and the training execution trajectory information comprises training interface information and labeled training action information of a training task corresponding to the training task information.
[0011] In some possible implementation manners, the first execution information and the first execution track information are processed by the first screening model, and the first execution track information is determined as the first training data, including: based on the first execution information being first indication information indicating that the first task execution is completed, the interface information in the first execution track information and the first task information are processed by the first screening model; it is determined that the interface information in the first execution track information is different from the training interface information in the training data set matched with the first task information, and the first execution track information is determined as the first training data; wherein the interface information in the first execution track information includes the first interface information and the second interface information.
[0012] In some possible implementation manners, according to the first execution information of the first task information, the first execution track information of the first task information is determined as the first training data, including: the first execution information and the first task information are processed by the second screening model, and the first execution track information matched with the interface model is determined as the first training data.
[0013] In some possible implementation manners, according to the first execution information of the first task information, the first execution track information of the first task information is determined as the first training data, including: the first execution information is processed by the first screening model, and a first index of the first execution track information is determined; the first execution information and the first task information are processed by the second screening model, and a second index of the first execution track information is determined; according to the first index and the second index, the first execution track information matched with the interface model is determined as the first training data.
[0014] In some possible implementation manners, the method further includes: obtaining second task information to be executed by a second device; in response to the second task information, obtaining third interface information of the second device; processing the third interface information by the interface model to obtain second action information; controlling the second device to execute the second action information; obtaining fourth interface information of the second device; processing the fourth interface information by the interface model to determine second execution information of the second task information.
[0015] According to a second aspect of the embodiments of the present disclosure, a multi-modal data processing apparatus is provided, including: an obtaining unit configured to obtain first task information to be executed by a first device; the obtaining unit is further configured to obtain first interface information of the first device in response to the first task information; a processing unit configured to process the first interface information by using an interface model to obtain first action information; the processing unit is further configured to control the first device to execute the first action information; the obtaining unit is further configured to obtain second interface information of the first device; the processing unit is further configured to process the second interface information by using a feedback model to determine first execution information of the first task information; the processing unit is further configured to determine, according to the first execution information of the first task information, first execution track information of the first task information as first training data, the first execution track information including the first interface information, the first action information and the second interface information; and the processing unit is further configured to train the interface model by using the first training data.
[0016] In some possible implementation manners, the processing unit is further configured to process the first interface information, the second interface information and the first task information by using the feedback model to determine the first execution information of the first task.
[0017] In some possible implementation manners, the processing unit is further configured to process the first execution information by using a first screening model to determine the first execution track information matching the interface model as the first training data. The first execution information includes first indication information used to indicate that the first task is executed completely and first execution path information of the first task.
[0018] In some possible implementation manners, the processing unit is further configured to process the first execution information and the first task information by using a second screening model to determine the first execution track information matching the interface model as the first training data.
[0019] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, including: a processor; a memory configured to store processor-executable instructions; and wherein the processor is configured to implement the steps of any of the multi-modal data processing methods.
[0020] According to a fourth aspect of the embodiments of the present disclosure, a non-transitory computer-readable storage medium is provided, when instructions in the storage medium are executed by a processor of a mobile terminal, the mobile terminal is enabled to perform any of the multi-modal data processing methods.
[0021] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements any of the multi-modal data processing methods.
[0022] The technical solutions provided by the embodiments of the present disclosure can have the following beneficial effects: on one hand, in response to the obtained first task information to be executed by the first device, the first interface information of the first device is automatically acquired, the acquired first interface information is processed through the interface model to obtain first action information, so that the first device can be controlled to automatically execute the first action information, without the need for the user to manually interact with the first device or reducing the number of times of manual interaction of the user with the first device, thereby reducing the difficulty and workload of the user in completing the task; on the other hand, after the first device is controlled to execute the first action information, the second interface information of the first device is continuously acquired, the second interface information is processed through the feedback model to determine the first execution information of the first task information, the first execution track information of the first task information is determined as the first training data according to the first execution information of the first task information, and then the interface model is trained using the first training data, so that the interface model can be continuously iterated, and the accuracy of the interface model can be improved.
[0023] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and are not limiting to the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0024] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure.
[0025] Figure 1 is a schematic diagram of a multi-modal data processing method according to an exemplary embodiment of the present disclosure.
[0026] Figure 2 is a flowchart of a multi-modal data processing method according to an exemplary embodiment of the present disclosure.
[0027] Figure 3 is a schematic diagram of a multi-modal data processing method according to another exemplary embodiment of the present disclosure.
[0028] Figure 4 is a reinforcement learning architecture diagram according to an exemplary embodiment of the present disclosure.
[0029] Figure 5 is another reinforcement learning architecture diagram according to an exemplary embodiment of the present disclosure.
[0030] Figure 6is a schematic diagram of a training environment according to an example embodiment of the present disclosure.
[0031] Figure 7 is yet another reinforcement learning architecture diagram according to an example embodiment of the present disclosure.
[0032] Figure 8 is still another reinforcement learning architecture diagram according to an example embodiment of the present disclosure.
[0033] Figure 9 is still another reinforcement learning architecture diagram according to an example embodiment of the present disclosure.
[0034] Figure 10 is a block diagram of a multi-modal data processing apparatus according to an example embodiment of the present disclosure.
[0035] Figure 11 is a schematic diagram of an application scenario of the method provided by an embodiment of the present disclosure.
[0036] Figure 12 is a schematic diagram of another application scenario of the method provided by an embodiment of the present disclosure.
[0037] Figure 13 is a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure.
[0038] Figure 14 is a functional block diagram schematic of a vehicle according to an example embodiment of the present disclosure. DETAILED DESCRIPTION
[0039] Some embodiments of the present disclosure will be described in detail herein with reference to the attached drawings, wherein the example embodiments represent only some of the embodiments consistent with the present disclosure. The following description is made with reference to the accompanying drawings in which like reference numerals refer to like elements in the several figures. The description of the methods, devices, and / or systems described herein can be modified in a variety of ways, and alternatives are also within the scope of the present disclosure. For example, the order in which the operations are described is merely an example and not limiting, as any of the described operations can be performed in a different order from those described herein or can be performed concurrently, except where the order of the operations is essential to the function of the application. Also, the description just given as to the methods describes only some embodiments. Functional and / or structural descriptions of various elements in the embodiments can be omitted in order to promote clarity and brevity.
[0040] The example embodiments described in some embodiments of the present disclosure do not represent all implementations consistent with the present disclosure. Instead, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0041] In the embodiments of the present disclosure, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an integral module or unit that contains the function of the module or unit.
[0042] First, some terms involved in the embodiments of the present disclosure are explained and described.
[0043] Artificial Intelligence (AI) is the use of digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.
[0044] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-training model technology, operation / interaction system, mechatronics, etc.
[0045] The pre-training model (PTM) is also called a large model, a basic model, or a cornerstone model. It refers to a deep neural network (DNN) with a large number of parameters. The PTM is trained on a large amount of unlabeled data. The PTM extracts common features from the data by using the function approximation capability of the large parameter DNN. After fine-tuning, parameter efficient fine-tuning (PEFT), prompt-tuning, and other techniques, the PTM is suitable for downstream tasks in various directions of artificial intelligence. Therefore, the pre-training model can achieve ideal results in few-shot or zero-shot scenarios. The PTM can be divided into language models, visual models, speech models, and multi-modal models according to the data modalities processed. The multi-modal model refers to a model that establishes feature representations of two or more data modalities. The pre-training model is an important tool for outputting artificial intelligence generated content (AIGC) and can also serve as a general interface connecting multiple specific task models.
[0046] The artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0047] With the research and progress of artificial intelligence technology, artificial intelligence technology is researched and applied in multiple fields, such as common smart home, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned vehicles, autonomous vehicles, drones, digital twins, virtual humans, robots, AIGC, conversational interaction, intelligent medical care, intelligent customer service, game AI, and the like. With the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0048] In the embodiments of the present disclosure, machine learning technology in artificial intelligence is used to implement task processing.
[0049] Machine learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, and other disciplines. It is a specialized research on how a computer simulates or implements human learning behavior to acquire new knowledge or skills, and reorganizes existing knowledge structure to continuously improve its performance. Machine learning is the core of artificial intelligence and the fundamental approach to enabling computers to have intelligence. It is applied in various fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and inductive learning. The pre-training model is a development result of deep learning, which integrates the above technologies.
[0050] The method provided by the embodiments of the present disclosure can be executed by any electronic device or computer device with information processing capability, such as a terminal and / or a server, and the present disclosure does not limit this.
[0051] Figure 1 is a schematic diagram of a multi-modal data processing method according to an exemplary embodiment of the present disclosure. In some embodiments, the method can be used to develop a UI Agent (User Interface Agent) that can process multi-modal data, such as processing UI screenshots, predicting actions to be performed, and achieving automatic execution of tasks to reduce the workload of users. Wherein N and M are both positive integers greater than or equal to 1. Figure 1
[0052] As shown in the method, the method includes 2 parts, namely model training and online inference, and specifically includes the following steps. Figure 1
[0053] In the training phase, step 1 is first executed, that is, screenshot (such as UI screenshot of an electronic device) collection, for example, the UI screenshot of an APP (application) installed on the electronic device can be obtained using the screenshot tool provided by the electronic device and saved to the disk. Then step 2 is executed, that is, manual annotation. For example, an annotation personnel can be hired to annotate the UI screenshot according to the task description (task information input by the user on the electronic device), generate actions and parameters, and save the UI screenshot and the annotated actions and parameters to generate a training set and a test set. Then step 3 is executed, that is, training the model, using the training data in the training set annotated manually to train and optimize the model included in the UI Agent, to generate a new model file or model parameters. Then step 4 is executed, that is, evaluating the model, using the test set to evaluate the new model parameters, checking the indicators and badcase (bad case) analysis, and determining whether to continue collecting screenshots. If it is necessary to continue collecting, return to step 1; if it is not necessary to continue collecting, the model file can be deployed online for online inference process. It is assumed that steps 1 to 4 are executed for N times, and then the UI Agent is deployed online.
[0054] The online inference stage includes the following steps: S11. Screenshot collection, which is the same as the screenshot collection in the model training. S12. Model inference, inputting the UI screenshot collected in S11, directly generating actions and parameters using the model parameters generated in the model training. S13. Action execution, according to the actions and parameters generated by the model inference, calling the system API (application program interface) of the electronic device (the electronic device used in the model training and the online inference stage can be the same or different) to execute the actions, and determining whether the task is completed according to the actions; if the task is not completed, the next screenshot collection is performed, for example, the second round of S11 to S13 is performed; if the task is completed, the task is completed. It is assumed here that the inference stage performs M rounds for the task to complete the task.
[0055] Figure 1 The scheme provided by the embodiments can be referred to as an SFT (Supervised Fine-Tuning) scheme. Although the UIAgent has broad prospects, it faces many challenges in actual application. First, the sequence decision problem has a delayed benefit, which means that the agent may not be able to immediately know the effectiveness of the current operation during the execution of the task, and the final benefit can be determined only after the task is completed. Second, the frequent updates of websites and applications cause the online observation results (such as the UI screenshots intercepted in the online inference stage) to be inconsistent with the offline data (such as the UI screenshots intercepted in the offline training stage), which brings difficulties to the learning and decision-making of the agent. In addition, various unpredictable interference terms, such as pop-up ads, login requests, and random order of search results, will affect the normal operation of the agent. In terms of technology, problems such as incomplete web page loading or temporary access restrictions to some websites also occur, which puts higher requirements on the performance and stability of the UIAgent. The above reasons result in a very low success rate of completing the task using the UIAgent in the scenarios of ad loading, network reason loading, and UI dynamic display. Because some APP ads are temporary or played in time periods, the offline screenshots collected in the model training stage may not necessarily encounter such UI screenshots. The network reason loading and UI dynamic display scenarios are similar, and some UI screenshots are occasional scenes.
[0056] Figure 1 When the scheme encounters such occasional scenes, there is no or insufficient training data, so the task completion rate is reduced in the online inference stage. Because the UI screenshots encountered in the online inference stage may be dynamically changed and inconsistent with the training data in the training stage. For example, the product display UI screenshot in the training data has no ad, while the product display UI screenshot intercepted in the online inference stage has an ad.
[0057] In this embodiment of the disclosure, dynamic UI display refers to a situation in the GUI interface where a certain page element (e.g., a box) is fixed, but the content contained within the box changes dynamically over time. For example, the content displayed when opening the page on the first day and the second day is different.
[0058] Figure 2 This is a flowchart illustrating a multimodal data processing method according to an exemplary embodiment of the present disclosure. In some embodiments, the method is applied to, for example... Figure 10 The multimodal data processing apparatus shown and the electronic device 100 equipped with the multimodal data processing apparatus (e.g., Figure 13 The electronic device in this embodiment includes the first device and the second device described below. The first device and the second device can be the same device or different devices. The electronic device used in this embodiment can include smartphones, tablets, wearable electronic devices, vehicles, etc., and is not limited thereto. Figure 2 As shown, the method provided in this disclosure embodiment may include the following steps.
[0059] In S210, the first task information to be executed by the first device is obtained.
[0060] In this embodiment of the disclosure, task information refers to a task instruction or task description issued to an electronic device. The first task information is any of these task instructions or task descriptions. For example, the first task information refers to a task instruction or task description issued to a first device during the model training phase, instructing the first device to complete a specific task. When the first task corresponding to the first task information is completed, the first execution trajectory information during the completion process of the first task can be collected as training data for model training (e.g., the first training data described below) to continue training the model and improve its accuracy.
[0061] In some embodiments, this disclosure can be applied to a program (Agent, such as UIAgent) installed in an electronic device. UIAgent can be understood as a program capable of thinking and interacting with the environment. Given a task (such as any one of the first task, second task, etc.), UIAgent will think about how to solve it, formulate a plan, take a series of actions (Action, such as first action information) on the environment (such as first interface information), receive feedback from the environment (Observation, such as second interface information), and repeat these steps until the task is completed.
[0062] In some embodiments, a first task that the first device needs to complete can be determined. In some embodiments, the first task that the first device needs to complete can refer to a task that needs to be completed through a graphical user interface on a display screen of the first device. For example, the first task that the first device needs to complete can include "send a message to Xiao Zhang through XX (the name of an application on the first device), the content of which is that there is a meeting at 5 o'clock today", "book a ticket to XX (the name of a place) for tomorrow morning", "call a car to go home at 6 o'clock this afternoon", etc., without limitation.
[0063] The first task information refers to any information related to the first task, such as one or more of the time, location, person, task content, etc. of the first task. In some embodiments, the first task information can be input into the first device by the user or the electronic device (such as a server and / or a terminal) in the form of interaction. For example, the first task information can be input into the first device by the user or the electronic device in the form of voice, the first task information can be input into the first device by the user or the electronic device in the form of text, the first task information can be input into the first device by the user or the electronic device in the form of touch operation, etc., without limitation. For example, information input by the user or the electronic device in the form of interaction can be obtained, and the information is determined as the first task information that needs to be completed; or information input by the user or the electronic device in the form of interaction can be obtained, and the first task information that needs to be completed is determined from the field of the information, etc., without limitation.
[0064] In some embodiments, the first task information can be generated when a preset condition is met. For example, if the current time meets the preset time, the first task that the first device needs to complete is automatically generated; and / or, if the current location where the first device is located meets the preset location, the first task that the first device needs to complete is automatically generated; if the current posture of the first device meets the preset posture, the first task that the first device needs to complete is automatically generated, etc., without limitation.
[0065] For example, the first task information refers to a voice instruction issued by the user or the electronic device through the vehicle-mounted voice assistant for the first time in a certain round of voice interaction, including a certain function control, information query or service request of the vehicle. For example, "turn on the air conditioner", "what's the weather like today?" and "I need the nearest gas station", etc.
[0066] In S220, in response to the first task information, first interface information of the first device is obtained.
[0067] Exemplarily, in the case of determining a first task to be completed by a first device, the first device obtains first interface information of the first device in response to the first task information, for example, screen information of a display screen of the first device. In some embodiments, the interface information includes a UI screenshot or UI screenshot related information of the electronic device. Exemplarily, the first interface information includes the first device in the UI screenshot or UI screenshot related information. In some embodiments, the UI Agent converts the UI screenshot into structured information as the interface information (including the first interface information, the second interface information, the third interface information, the fourth interface information, etc.) by using technologies such as screenshot XML, screenshot picture, OCR (Optical Character Recognition), abstract extraction, icon detection and capture, etc. In some embodiments, the interface information can include text, i.e., the UI screenshot can be described by text.
[0068] In some embodiments, the screen information of the display screen of the electronic device can be obtained through accessibility services. The accessibility services are a mechanism provided by the operating system of the electronic device, aiming to help users better use the electronic device. These services can access and manipulate interface elements of the device, such as text, buttons, pictures, etc. For example, through the accessibility services, at least the following operations can be completed: reading screen content: the accessibility services can access all the text and interface elements displayed on the display screen. This can be achieved by traversing the AccessibilityNodeInfo tree, which represents all the interactive elements on the current display screen. The accessibility services can simulate screen clicks, long presses, swipes and other gestures or actions to allow applications to interact with elements on the display screen automatically. The accessibility services can know which application or window is currently active to determine which application the user is interacting with. The accessibility services can also perform other operations, such as changing the device volume, interacting with notifications, simulating keyboard input, etc., to further expand the use and scope of the accessibility services.
[0069] In some embodiments, the screen information of the display screen of the electronic device can be obtained through APIs. Exemplarily, the display screen provides a programming interface that allows access and manipulation of interface elements through code. For example, some display screen frameworks or libraries provide APIs for reading and modifying interface content, so these APIs can be used to obtain the screen information of the display screen, such as the properties, status, content, etc. of the user interface on the display screen, without limitation.
[0070] In some embodiments, the screen information of the display screen of the electronic device can be obtained by a software tool. Some software tools can be used to analyze the interface content of the display screen, such as UI automation testing tools, screenshot tools, UI element recognition tools, etc. These tools can be used to capture the screen information of the display screen, such as the attributes, status, content, etc. of the user interface on the display screen, which are not limited herein.
[0071] In some embodiments, the screen information of the display screen of the electronic device can be obtained by system logs and debugging information. In some cases, the operating system or firmware of the display screen can provide system logs or debugging information, which can contain clues or data about the content of the screen information of the display screen. Therefore, by analyzing these system logs or debugging information, the screen information of the display screen can be obtained, such as the attributes, status, content, etc. of the user interface on the display screen, which are not limited herein.
[0072] In some embodiments, the screen information is multi-modal, containing text, image and structural hierarchical information, etc. In some embodiments, the hierarchical data of the view hierarchy of the screen contains too much information, such as detailed attributes of each UI element, which can exceed the input length limit of the interface model and interfere with the inference of the interface model.
[0073] In some embodiments, the view hierarchy of the user interface elements on the display screen of the first device can be obtained. In some embodiments, the view hierarchy of the user interface elements on the display screen can be obtained by accessibility services. In some embodiments, the root view of the current Activity can be obtained, and then the entire view tree can be traversed by recursion or iteration to obtain the view hierarchy of the user interface elements on the display screen. In some embodiments, the view hierarchy is converted into a text modal form to obtain the screen information. In this embodiment, when the view hierarchy of the user interface elements on the display screen is obtained, the view hierarchy can be converted into a text modal form to obtain the screen information.
[0074] In some embodiments, in the case of obtaining the view hierarchy of the user interface element on the display screen, the view hierarchy can be converted into the form of HyperText Markup Language (HTML) to obtain the screen information. The screen information in the form of HTML has at least the following advantages: first, token usage can be reduced: HTML is a relatively efficient way to represent the view tree structure, and can effectively reduce the use of tokens by means of invisible elements and other means to avoid exceeding the context limit of the interface model; second, easy to understand: when training the interface model, most of the training data is obtained by web crawling. A large part of it is in the form of HTML representation, so the interface model can well understand HTML and achieve good reasoning effect. The present embodiment can convert the view hierarchy of the UI into HTML syntax, so as to convert the screen information into text modal and simplify the screen UI information to improve the decision accuracy of the interface model.
[0075] In some embodiments, the first interface information includes screen information in text modal and image modal.
[0076] In some embodiments, the first interface information includes current display content information of the vehicle display component, such as the information content currently presented on the in-vehicle display screen or other display device. For example, navigation-related content such as current driving route, destination, estimated arrival time, and traffic conditions, or entertainment system-related content such as music playlist, track information, radio frequency, and the like, or communication-related content such as display of incoming call information, SMS content, contact list, and the like. The display content information is an important part of the vehicle intelligent control system, which can provide the driver with key driving information and entertainment services, and is also an important interface for system and user interaction. In the present embodiment, the in-vehicle display screen is taken as an example of the vehicle display component, i.e., the current display content information is a screenshot of the in-vehicle display screen.
[0077] In S230, the first interface information is processed by the interface model to obtain first action information.
[0078] The interface model in the embodiments of the present disclosure refers to a machine learning model or a deep learning model capable of processing input interface information to predict the action information performed for the interface information. For example, the interface model is included in the UI Agent. In some embodiments, the interface model includes a Vision-Language Model (VLM) capable of understanding and processing the association between images and text, so that the machine can understand visual content and natural language text at the same time. In some embodiments, the interface model includes a LLM model.
[0079] The UI Agent (User Interface Agent) in the embodiments of the present disclosure is an intelligent entity that actively or passively assists users in completing tasks through a user interface using artificial intelligence technology. The UI Agent uses large model technology (such as VLM / LLM) to realize the automatic operation of the intelligent agent on electronic devices such as mobile phones or computers, simulates human behavior to complete designated tasks, and covers various application scenarios such as WebGUI and Mobile GUI. The VLM undertakes multiple tasks of perception, planning, and decision-making in the UI Agent. It has UI task execution and reasoning capabilities, including global understanding capabilities (such as UI interface understanding) and local detail understanding capabilities (such as element positioning and reference capabilities) to meet various needs in UI operations. The embodiments of the present disclosure combine the interactivity of the user interface (UI) and the autonomy of the agent (Agent) to improve user experience and efficiency. In some embodiments, the UI Agent has autonomy and can actively perform tasks (such as predicting needs and automatically filling out forms) without explicit instructions from the user; it can respond to user operations or environmental changes in real time (such as dynamically adjusting the layout of the interface); and it can learn from user behavior and optimize interaction logic (such as personalized recommendations).
[0080] The UI Agent in the embodiments of the present disclosure can simulate human operations and automatically perform tasks. In some scenarios, it can assist in performing complex operation processes (such as "ordering + payment" on e-commerce platforms). In other embodiments, the UI Agent can automatically identify image content, recommend filter combinations, and preview the effects, and the user only needs to click to confirm to complete professional-level photo editing.
[0081] Through the UI Agent, the embodiments of the present disclosure can implement a "goal-driven" interaction in terms of interaction form, where the user only needs to give a goal or task, and the goal or task can be decomposed and gradually completed; support interaction with GUI, understand the current interaction interface of the user, and help the user complete the interaction of the user interface. Tasks related to the user interface can be decomposed into actions through the interface model, and then the electronic device is intelligently controlled to perform the decomposed actions to complete the required tasks, which can reduce the difficulty and workload of the user.
[0082] The action information in the embodiments of the present disclosure refers to the related actions and parameters of operating the user interface of the electronic device by controlling the electronic device. The action can be used to indicate what operation is performed on the user interface of the electronic device. The parameter can be used to indicate the coordinates, input messages, etc. of performing the action on the user interface of the electronic device. For example, the first action information is the action information predicted for the first interface information.
[0083] Exemplarily, the first action information refers to any information related to the first action. The first action refers to a specific operation determined by the interface model according to the first task information and the first interface information to control the first device to perform, including a click operation, a sliding operation, and a typing operation, and the like. The click operation refers to an operation of clicking an icon or a button, and the like, for activating a specific function or service. For example, when the user says "open music", the system can need to perform an operation of clicking an icon of a music application in a vehicle entertainment system. The sliding operation refers to an operation of sliding a screen or turning a page to view more information when a current page information list is folded or too long. For example, when a navigation system displays route information that is too long, the system can need to perform an operation of sliding a screen to display the complete route. The typing operation refers to an operation of clicking a search box and typing content, for searching for specific information or service in the system. For example, when the user says "search for nearby restaurants", the system can need to perform an operation of clicking a search box and inputting "nearby restaurants". The typing operation can also indicate inputting and sending text information in the vehicle system, such as sending a message or making a comment on social media. For example, when the user says "send a message to Zhang San", the system can need to perform an operation of opening a message application, selecting a contact Zhang San, and opening a chat window.
[0084] In the embodiments of the present disclosure, the action space can be defined first to determine the atomic operation instruction. For example, basic actions such as "click (x, y)", "slide", and "input text" are defined, and parameterization (such as coordinates and text content) is supported.
[0085] In some embodiments, the interface model can predict the action description or action in the first action information, and the parameter of the action. In some embodiments, the interface model includes a VLM model and a positioning model. The VLM model is used to generate the action description, and the positioning model (such as TinyClick) is used to calculate the accurate coordinates as the parameter of the action.
[0086] For example, when a task instruction such as "XX (the name of an APP) sends a message to Xiaoming: 'Have you eaten?' is given, the UI Agent can understand the task like a human being, and then perform a series of operations on a mobile phone or a computer, such as opening XX, finding a chat window of Xiaoming, inputting the message 'Have you eaten?' and sending. This process involves perception, understanding, and accurate operation of the UI interface, which is a Partially Observable Markov Decision Process (POMDP) problem. The agent cannot observe all state information, makes a decision according to the currently observable state (such as a UI screenshot and corresponding XML), and outputs an operation instruction such as "CLICK (100, 200)", where "CLICK" is the action name, and "(100, 200)" is the action parameter, that is, the coordinates of the click.
[0087] In S240, the first device is controlled to perform the first action information.
[0088] In the embodiments of the present disclosure, when the first action information corresponding to the first interface information is determined through the interface model, it is determined that the action needs to be taken in the current environment state, and then the first device can be controlled to perform the first action corresponding to the first action information.
[0089] For example, assuming that the first task information is a task published in natural language, such as "help me send XX (APP name) content to Xiao Zhang at 5 o'clock this afternoon". The UIAgent can obtain the screen information as the first interface information through the accessibility service, call the interface model to process the first interface information to determine the first action information, and help complete the click, text input, screen sliding, return, task completion and other actions (Action) through the accessibility service. According to the executed first action information, the first device can generate feedback (Observation), such as second interface information. Exemplarily, the second interface information can be the screen information of the first device after the first action information is executed, can be the interface screenshot of the first device, or can be the window hierarchy information obtained by the accessibility.
[0090] Exemplarily, the electronic device can include a display screen for displaying information input by a user, information provided to a user, and various graphical user interfaces of the electronic device, which can be composed of graphics, text, icons, numbers, videos, and any combination thereof. In one example, the display screen can be a liquid crystal display (LCD), or an organic light-emitting diode (OLED), which is not limited herein.
[0091] As an example, the action performed by the electronic device for the task can include the following forms:
[0092] 1. {"action_type": "click", "id": "<element_id>"} indicates clicking a certain control, and the id information of the control is required.
[0093] 2. {"action_type": "type", "id": "<element_id>", "text": "<input_text>"} indicates text input to an EditText control, and the control id information and the text information to be input are required.
[0094] 3. {"action_type": "scroll", "id": "<element_id>", "direction": "<UP|DOWN|LEFT|RIGHT>"} represents up, down, left or right screen sliding.
[0095] 4. {"action_type": "navigate_home"} represents pressing the home key to enter the desktop.
[0096] 5. {"action_type": "navigate_back"} represents pressing the back key to return.
[0097] 6. {"action_type": "status_complete"} status action, indicating that the task is completed, and the task can be terminated.
[0098] 7. {"action_type": "status_impossible"} status action, indicating that the task is impossible to complete, and the task needs to be terminated.
[0099] In some embodiments, in the case of obtaining the first task information and the first interface information, the first task information and the first interface information can be input into the interface model to predict the first action required to complete the first task. In the case of obtaining the first action output by the interface model, the first device can be controlled to execute the first action, so that the first device completes the first task that needs to be completed, thereby providing the GUI application with the ability of AI interaction in a unified manner without invasion, and providing the user with intelligent experience. As an example, if the first action is "send", the first device can be controlled to execute the "send" action; if the first action is "search", the first device can be controlled to execute the "search" action, which is not limited herein.
[0100] In some embodiments, in the case of obtaining the first action output by the interface model, the first action can be executed through the accessibility service, so that the first device completes the first task that needs to be completed.
[0101] In S250, the second interface information of the first device is obtained.
[0102] In some embodiments, the second interface information refers to the interface information converted from the first interface information after the first device executes the first action information. The specific content contained in the second interface information depends on the first interface information, the first action information and the first task information to be executed.
[0103] In some embodiments, after the first device is controlled to perform the first action, it can be determined whether the first device completes the first task. If it is determined that the first device completes the first task, no further operation is performed. If it is determined that the first device does not complete the first task, the steps of obtaining the screen information (e.g., the second interface information) of the display screen of the first device are repeated until the step of controlling the first device to perform the second action (corresponding to the second action information) until it is determined that the first device completes the first task, so that the user only needs to give a first task, and the first task can be decomposed and completed step by step.
[0104] In S260, the second interface information is processed by a feedback model to determine the first execution information of the first task information.
[0105] The feedback model in the embodiments of the present disclosure refers to a model in which the electronic device obtains or determines the information fed back by the environment after performing the action information. In some embodiments, the feedback model can be used to determine whether the task (e.g., the first task) is completed. In some embodiments, the information fed back by the environment includes the second interface information. In some embodiments, the feedback model includes a reward model. The reward model is a model used to evaluate and guide the training of the interface model. The reward model in the present solution is essentially to determine whether the task is completed, which is the feedback given by the environment, so it is named as the feedback model.
[0106] In some embodiments, the feedback model can use a pre-trained model, which does not need to be fine-tuned and trained, and can directly determine whether the task is completed, thereby reducing the cost. In other embodiments, a feedback model can be independently trained to determine whether the task is completed, thereby improving the accuracy of determining whether the task is completed.
[0107] In the embodiments of the present disclosure, the first execution information is used to indicate one or more of the completion degree, whether completed, whether correctly completed, and the like of the first task corresponding to the first task information.
[0108] In S270, according to the first execution information of the first task information, the first execution trajectory information of the first task information is determined as the first training data, and the first execution trajectory information includes the first interface information, the first action information, and the second interface information.
[0109] In the embodiments of the present disclosure, the execution trajectory information refers to the interface information collected during the execution of the task and the corresponding action information predicted by the interface model when the feedback model determines that the corresponding task has been executed by the electronic device. For example, the first execution trajectory information refers to the interface information (including the first interface information and the second interface information, and the specific interface information depends on the number of steps required to complete the first task, and is not limited to the first interface information and the second interface information exemplified herein) collected during the execution of the first task and the corresponding action information (including the first action information, and the specific action information depends on the number of steps required to complete the first task, and is related to the number of interface information) predicted by the interface model when the feedback model determines that the corresponding first task has been executed by the first device.
[0110] For example, if the first task to be completed by the first device is "help me send XX (APP name) content to Xiaozhang, which is a meeting at 5 pm this afternoon", the action information executed by the first device included in the first execution trajectory information can include:
[0111] {"step_idx": 1, "action_description": click[XX]};
[0112] {"step_idx": 2, "action_description": click[search]};
[0113] {"step_idx": 3, "action_description": type[Xiaozhang] in[search]};
[0114] {"step_idx": 4, "action_description": click[Xiaozhang]};
[0115] {"step_idx": 5, "action_description": type[meet at 5 pm this afternoon] in[null]};
[0116] {"step_idx": 6, "action_description": click[send]}.
[0117] In S280, the interface model is trained using the first training data.
[0118] In some embodiments, the method provided by the embodiments of the present disclosure can be applied to a training phase, i.e., continuously accumulating new training data to retrain the interface model in the UI Agent in the training phase. In other embodiments, the method provided by the embodiments of the present disclosure can be applied to an online inference phase, i.e., continuously accumulating new training data to retrain the interface model in the UI Agent in the online inference phase.
[0119] The method provided by the embodiments of the present disclosure can automatically obtain the first interface information of the first device in response to the obtained first task information to be executed by the first device, process the obtained first interface information through the interface model, and obtain the first action information, so as to realize the control of the first device to automatically execute the first action information, without the need for the user to manually interact with the first device or to reduce the number of times of manual interaction of the user with the first device, thereby reducing the difficulty and workload of the user in completing the task. On the other hand, after the control of the first device to execute the first action information, the second interface information of the first device is continuously obtained, the second interface information is processed through the feedback model to determine the first execution information of the first task information, the first execution track information of the first task information is determined as the first training data according to the first execution information of the first task information, and then the interface model is trained using the first training data, so as to realize the continuous iteration of the interface model and improve the accuracy of the interface model.
[0120] In the example embodiments, the processing of the second interface information through the feedback model to determine the first execution information of the first task information includes: if it is determined through the processing of the second interface information through the feedback model that the first task corresponding to the first task information is executed and completed, first indication information indicating that the first task is executed and completed is generated; and if it is determined through the processing of the second interface information through the feedback model that the first task is not executed and completed, second indication information indicating that the first task is not executed and completed is generated. The first execution information includes the first indication information and the second indication information.
[0121] In some embodiments, the feedback model can determine whether the first task is completed through the second interface information (e.g., the next UI screenshot after the first device executes the first action). For example, if the first task is an ordering task, the payment completion UI screenshot is input into the feedback model, and it is determined that the first task is completed. In other embodiments, the feedback model can determine whether the first task is completed through the second interface information and the first task information, so as to obtain the first execution information.
[0122] In some embodiments, when the feedback model is a reward model, the first execution information can be a reward output by the reward model. In some embodiments, when the first task is completed, a reward of 1 point is given, i.e., the first indication information is equal to 1 at this time; when the UI screenshot of any step before the completion of the first task is input into the reward model, a reward of 0 point is given, i.e., the second indication information is equal to 0. For example, for an order placement task, if only the corresponding goods are placed in the shopping cart, the reward is 0. However, the present disclosure is not limited thereto, as long as the first indication information and the second indication information can distinguish whether the first task is completed or not, and the specific value or identifier is not limited. In other embodiments, different values or identifiers of the first execution information can be used to represent the execution degree or execution progress information of the first task, for example, the reward output is incremented by 1 point for each action completed in the first task. In the following examples, the first indication information is taken as 1 and the second indication information is taken as 0, but the present disclosure is not limited thereto.
[0123] The method provided by the embodiments of the present disclosure can, on the one hand, in the training phase or online inference phase, add an additional feedback model to determine the first execution information of the first task information according to the second interface information, for example, to determine whether the first task is completed, and can accumulate new training data (including first training data of first execution track information) to retrain the interface model in the UI Agent according to the first execution information, thereby improving the accuracy of the UI Agent. On the other hand, by using the feedback model independent of the interface model to accumulate new training data, compared with directly using the interface model to accumulate new training data, more diversified training data can be obtained, for example, UI screenshots loaded with advertisements and their corresponding action information, thereby enriching the training samples for training the interface model, so that the interface model can more accurately predict dynamically changing UI screenshots, and improve the task completion rate and accuracy.
[0124] In an example embodiment, the second interface information is processed by the feedback model to determine the first execution information of the first task, including: the first interface information, the second interface information and the first task information are processed by the feedback model to determine the first execution information of the first task.
[0125] In some embodiments, the feedback model can determine not only whether the first task is completed, but also whether the first task is correctly completed according to the information of the historical UI screenshot (including the first interface information, and optionally, the first action information), the second interface information (for example, the next UI screenshot after executing the first action information), and the task instruction (for example, the first task information).
[0126] In the example embodiment, processing the first interface information, the second interface information and the first task information by the feedback model to determine the first execution information of the first task includes: if it is determined that the first task information corresponds to a completed first task by processing the first interface information, the second interface information and the first task information by the feedback model, first indication information indicating that the first task is completed is generated; if it is determined that the first task is not completed by processing the first interface information, the second interface information and the first task information by the feedback model, second indication information indicating that the first task is not completed is generated. The first execution information includes the first indication information and the second indication information.
[0127] In some embodiments, the input of the reward model is the next UI screenshot of the first task, the page information description of the UI screenshots of all previous steps (including the first interface information), and the task description / task instruction of the first task (i.e., the first task information). That is, the reward model not only determines whether the first task is completed, but also determines whether the first task is correctly completed. When the reward model determines that the UI screenshot of a step indicates that the first task is not completed, it extracts the page information description in the UI screenshot and stores it as the input of the reward model at the next step. Then, the reward model determines whether the executed first task is consistent with the first task in the task description based on the next UI screenshot and the page information description of the UI screenshots of all previous steps. For example, the task instruction is to order a latte, but a milk tea is ordered. Although the payment page is finally reached, if only the payment page is determined to be completed, the task is actually incorrectly completed.
[0128] In some embodiments, the page information description in the UI screenshot extracted by the reward model can be abstract information or text information extracted by the reward model from the analysis and processing of the first page information. In some embodiments, the page information description of the UI screenshots of all previous steps can be summarized as the history information of the previous step (here, the previous step is relative to the second interface information) as output.
[0129] In some embodiments, the feedback model outputs the first indication information when the first task is correctly completed. The feedback model outputs the second indication information when the first task is not completed or incorrectly completed. That is, the correct completion of the first task is considered as the completion of the first task, and if the first task is incorrectly completed, it is also considered as not completed. That is, the feedback model provided in the embodiments of the present disclosure can not only identify whether the task is completed, but also identify whether the task is correctly completed. The execution track information of the correctly completed task is used as the training data of the interface model, so that the task completion rate and the correctness of the interface model can be further improved.
[0130] In an example embodiment, the method provided by the embodiments of the present disclosure further includes: obtaining second task information to be executed by a second device; in response to the second task information, obtaining third interface information of the second device; processing the third interface information through the interface model to obtain second action information; controlling the second device to execute the second action information; obtaining fourth interface information of the second device; processing the fourth interface information through the interface model to determine second execution information of the second task information.
[0131] For example, the second device can be any electronic device in a real user environment after the interface model or the UI Agent is trained and deployed online. In the online deployment stage, the feedback model or the reward model can no longer be needed. In the online inference stage, the UI Agent or the interface model itself can determine whether the task is completed.
[0132] In the embodiments of the present disclosure, the second execution information is used to indicate whether the second task corresponding to the second task information is completed. When the second execution information indicates that the second task is not completed, the fourth interface information is continuously processed through the interface model to predict third action information, and then the fifth interface information is obtained, and the fifth interface information is processed through the interface model to determine whether the second task is completed. In this way, the process is repeated until the second task is completed.
[0133] The method provided by the embodiments of the present disclosure first calls a pre-trained model as the interface model in the UI Agent. Then, the pre-trained model is processed by SFT, that is, the pre-trained model is fine-tuned by using a static training data set, so that the pre-trained model has the ability to preliminarily identify whether a task is completed. For example, in the training data set, whether each UI screenshot indicates that the task is completed can be manually annotated. However, this identification can be inaccurate because the static training data set is not real-time. Then, the fine-tuned UI Agent is further trained by using the first training data obtained by the feedback model and the screening model. After the training, the UI Agent or the interface model is deployed online for online inference by users. In the online deployment, the reward model can no longer be needed, that is, whether the task is completed is determined by the ability of the UI Agent or the interface model, so as to determine whether the next UI screenshot needs to be intercepted to continue to predict the action and the parameters thereof.
[0134] Figure 3 FIG. 1 is a schematic diagram of a multi-modal data processing method according to another example embodiment of the present disclosure. Figure 3 The method provided by the embodiments is also divided into two processes of model training and online inference, which is similar to the method provided by the embodiments of the present disclosure. Figure 1 The main difference between the embodiments is in the process of model training.
[0135] As shown in FIG. 1, the method provided by the embodiments of the present disclosure includes the following steps. Figure 3As shown, during the training phase of the model (e.g., the interface model), it is assumed that n rounds from S311 to S31n are executed for a task (e.g., the first task) to generate execution trajectory information (e.g., the first execution trajectory information) for this task. The training data (e.g., the first training data) for training the interface model is determined by the reward model and automatic filtering, where n is a positive integer greater than or equal to 1.
[0136] In S311, screenshot acquisition begins. For example, a screenshot tool can be used to capture UI screenshots of the app and save them to disk. Then, the UI screenshots are input into the interface model for model inference. In this round, model parameters generated during the fine-tuning phase can be used to directly generate actions and parameters from the UI screenshots. Next comes action execution. For example, based on the actions and parameters generated by model inference, the system API is called to execute the actions. Then, a reward model is used to determine if the task is complete. For example, after action execution, the next UI screenshot is captured, and this screenshot, along with historical information from the previous step (including the page information description of the previous step's UI screenshot extracted by the interface model), is sent to the reward model to determine if the task is complete. If the task is complete, all UI screenshots, actions, and their parameters from all steps are stored as training data; otherwise, similar to online inference, the loop continues to the next step, returning to the screenshot acquisition step. An automatic filtering step is also included, for example. Once a sufficient amount of training data has been collected through the reward model (e.g., a predetermined number of training data, which can be set according to actual needs, such as 1,000 or 10,000; or, training data collection for a predetermined period of time, such as one week or one month), the training data is automatically filtered by the filtering model. The model is then trained using the filtered training data.
[0137] For example, determining when to stop model training can be done by checking if the newly added training data is different from the training data introduced during the fine-tuning phase; and / or by providing an evaluation set to test the accuracy of the UIAgent's inference. Furthermore, it's also possible to set how many iterations to stop training, or when the loss function converges, etc.
[0138] Online reasoning stage and Figure 1 The online inference phase in the embodiment is similar. Here, it is assumed that m rounds were executed for a task (e.g., the second task), namely S321, S322 up to S32m, where m is a positive integer greater than or equal to 1, which will not be elaborated here.
[0139] The method provided in this disclosure can solve the problem of low task completion rate of UIAgent in scenarios such as advertising, loading due to network issues, and dynamic UI displays. This is because some app advertisements are temporary and played in time slots, and offline screenshots may not necessarily encounter such UI screenshots. Loading due to network issues and dynamic UI displays are similar to intermittent scenarios. When the SFT technology solution encounters such intermittent scenarios, the lack of training data will reduce the task completion rate.
[0140] The method provided in this disclosure achieves autonomous decision-making in dynamic environments through reinforcement learning and judges task completion through a reward model, thereby accumulating training data for online training, which can be directly used in UI Agent interaction scenarios. For example, the method provided in this disclosure is an online reinforcement learning UI Agent training method. The technical field of online reinforcement learning UI Agent belongs to autonomous decision-making and intelligent agent technology in the field of artificial intelligence, focusing on the combination of human-computer interaction and intelligent agent technology. In the SFT scheme, model training and online inference are completely separated, and the model relies entirely on historical data for training, which can lead to the inability to reflect changes in the APP's UI when these changes occur. The online reinforcement learning UI Agent deploys the model (interface model) online while simultaneously deploying an online reinforcement learning training environment to learn the APP's UI changes in near real-time and periodically updates the online inference model parameters to maintain the task completion rate of the UI Agent model (including the interface model within it).
[0141] Figure 4 This is a diagram illustrating a reinforcement learning architecture according to an exemplary embodiment of this disclosure. For example... Figure 4 As shown, reinforcement learning is a framework for solving control (or decision-making) tasks. It learns by trial and error from the environment (env 410) and receiving rewards (positive or negative, represented by Rt in the diagram), then treats these rewards as feedback. The intelligent agent 420 responsible for decision-making and trial and error is called the Agent. It can be compared to a machine learning or deep learning model in supervised learning and is a learnable function, such as the interface model mentioned above.
[0142] refer to Figure 4, the agent 420 generates an action and a parameter (denoted as At) at the t-th moment according to the state St (for example, first interface information) output by the environment 410 at the t-th moment, and the corresponding action At can be performed by the environment 410, and the next state or the next step state (here, the state St+1 at the t+1-th moment, for example, the second interface information) is obtained. The agent 420 is an action and parameter reasoning prediction model, and is responsible for the decision-making and trial-and-error of the agent, that is, the model parameter at the online reasoning time. The state is a set of variables (observation) used to express the state of the environment. In the embodiments of the present disclosure, the UI screenshot is the state, and the state is different according to the situation of each APP or each task to be executed. The action is an action that can be selected at the time of decision-making. The reward is the feedback signal received by the environment 410 after the action of the agent 420. p(At / St) indicates the probability that the agent 420 outputs the action At at the state St. p(St+1, Rt / St, At) indicates the probability that the environment 410 outputs the state St+1 and the reward Rt at the state St and the action At. When the reward indicates done, it indicates that the interaction for the current task has been completed, for example, the reward is divided into 1.
[0143] In the embodiments of the present disclosure, the agent 420 can observe the environment 410 around it, and the agent can calculate the observed environment through a policy network to determine the action to be performed at each decision-making period. For example, when the agent is a soccer robot, the playing field where the soccer robot is located is the environment where each agent is located, and winning the game is the task set for the soccer robot. The processing performed by the soccer robot at each decision-making period can include performing different specified actions, etc. The agent can be an entity that makes independent decisions in a specified environment. The entity can exist in the form of hardware, for example, can be a drone, a robot, etc. It can also exist in the form of software, such as a software program, for example, an AI game character in an electronic game, etc. The environment where the corresponding agent is located can be a real environment. It can also be a virtual environment, for example, information about other AI game characters within the attack range of an AI game character, etc.
[0144] When the agent is a hardware device, the agent can include a processor, a memory, etc. inside. The memory is used to store data related to the processing of the agent, for example, program code related to the processing of the agent. The processor can load and process the data stored in the memory, and then implement the method provided in the present application, etc. When the agent is a hardware device, the agent can be a drone, a robot, a vehicle, etc.
[0145] When the agent is a software program, the computer device running the agent can include a processor, a memory, etc. The memory is used to store data related to the processing of the agent, which can be program code related to the agent, for example. The processor can load and process the data stored in the memory, thereby implementing the method provided in the present application, etc. The agent can run on a computer device or an electronic device, and the computer device or the electronic device can start a corresponding processing process for the agent for running the corresponding agent. When the agent is a software program, the agent can be an AI game character, etc.
[0146] When the agent is a hardware device, each step in the method flow can be executed by the agent, and when the agent is a software program, each step in the method flow can be executed by a computer device or an electronic device running the agent, i.e., through a processing process started by the computer device for the agent.
[0147] The method provided in the embodiments of the present disclosure can construct an intelligent interface with real-time interaction and autonomous optimization capability by combining online reinforcement learning and UI Agent, thereby improving user experience. In the future, with the further integration of large models and reinforcement learning, UI Agent will play a role in more complex scenarios.
[0148] Figure 5 is another reinforcement learning architecture diagram according to an example embodiment of the present disclosure. As shown in Figure 5 The environment module is an important module of online reinforcement learning. Figure 5 In the embodiments, the environment 410 can include the following parts: action space definition and parameters, simulator 412 and real machine 413, action execution, reward model 414.
[0149] For example, the action space definition and parameters include the definition and parameter types of all action spaces of the UI Agent, which are consistent with the SFT scheme and consistent with the client execution. The parameter types are, for example, type (input text), coordinates, etc.
[0150] Exemplarily, the following action spaces and parameters are defined: Idle means doing nothing; Dualpoint means double-clicking; Type means inputting text; Goback means going back; Gohome means going to the homepage; Enter means inputting the enter key; TaskComplete means task completion; TaskImpossible means task impossibility, such as execution error, ending the task; Sensitive means sensitive data, not allowed to take screenshots, and the task is interrupted at this time; OpenApp means opening an App; Wait means waiting; Speak means voice playing, which means that the user input task description is incomplete, and the voice is played to ask the user to supplement the corresponding information; Request means obtaining the user-supplemented information; and ToolUse means calling a third-party tool to execute a task or action.
[0151] Exemplarily, the online reinforcement learning scheme needs to truly complete action execution in the simulator 412 or the real machine 413, and the App to be optimized is installed on a mobile phone, a vehicle machine, or a simulator. Exemplarily, offline production can be performed through the virtual device 411 to form the simulators 412. Exemplarily, the real machine and the simulator are simultaneously deployed in the test environment. The real machine is a real electronic device such as a mobile phone, but a real mobile phone can only open one App, which will result in a high cost, and therefore, a simulator can be used to open a virtual mobile phone or a simulation device or a software mobile phone on a server or an electronic device, so that multiple APPs for testing can be opened on one physical machine. In this way, on the one hand, the cost can be reduced, and on the other hand, for some small mobile phones and small APPs, some mobile phones cannot open these APPs, but the simulator can open these APPs.
[0152] Exemplarily, after the action and parameter prediction are completed, an action is executed by calling an adb tool and the next state (UI screenshot) is obtained.
[0153] Exemplarily, the next state (UI screenshot) and the historical information of the previous step are fed into the reward model 414 to determine whether the current task is completed, and the corresponding reward is output.
[0154] Exemplarily, the environment 410 further includes a cross-platform device connection service for mapping the action At predicted by the agent 420 to an operation instruction that can be recognized and executed by the real machine and the simulator. Because the action predicted by the agent 420 is one or more of the defined action spaces and parameters, the machine cannot directly execute it, and a mapping process is needed. That is, it is converted into an instruction in the App.
[0155] The embodiment of the present disclosure simulates various possible operation conditions of the APP to be optimized in a test environment through a simulator and a real machine, also has advertisements, and may encounter network loading and UI dynamic display, thereby obtaining new UI screenshots different from the static training data set, stores these UI screenshots, and when the preset number, for example, 1000 or 10,000, is reached, re-trains the UI Agent model, so as to realize the quasi-real-time update of the UI Agent model. For example, update once in a few days or weeks. After training by such training data, the loading of advertisements, network loading, and UI dynamic display can be more accurately recognized. Then, the UI Agent is deployed to the real user environment. The embodiment of the present disclosure does not take the UI Agent itself as a reward model, but gives a reward through an independent reward model to judge whether the task is completed, and the training data collected in this way is more likely to be different from the original training data set in the fine-tuning stage, and the discrimination is greater.
[0156] In the test environment shown in the above Figure 5 In the test environment shown in the above
[0157] Online reinforcement learning has cross-domain universality. The trial-and-error mechanism of reinforcement learning and the autonomous decision-making ability of the agent make it suitable for any scene that needs dynamic interaction and continuous optimization, such as embodied intelligence, Figure 6 An example of building an online reinforcement learning environment in the embodiment of the present disclosure is shown.
[0158] Because a large amount of training data is needed, multiple real machines / simulators must be run in parallel to collect training data. Running multiple real machines / simulators in parallel has great challenges because when one simulator encounters an unknown error, thread synchronization is inefficient and fault propagation is frequent. In order to solve this challenge, as shown in FIG. 6, the embodiment of the present disclosure uses a message queue to synchronize the communication between the simulator and the real machine. Figure 6As shown, the embodiments of the present disclosure set up a server-client system, in which all simulator processes run in independent server processes. For example, real machine / simulator 1 process runs in working server 1, real machine / simulator 2 process runs in working server 2, and real machine / simulator k process runs in working server k, where k is a positive integer greater than or equal to 1. Each simulator process communicates with the master training process in the master server through a different UIAutomotor server (i.e., the corresponding working server). The master training process sends high-level instructions (such as reset and step) to the UIAutomotor server, and the UIAutomotor server parses the high-level instructions into low-level UI commands (such as typing characters and tapping coordinates), which are executed by the simulator process. When an exception is triggered in the simulator, the UIAutomotor checks whether it can be recovered (for example, the UI command is executed in the simulator for too long), and if it cannot be recovered, the simulator process is reset. When an exception is triggered in the UIAutomotor server, the master training process stops and resets the UIAutomotor server to ensure the correctness of the data. The overall design architecture is as shown in Figure 6 As shown, this design can be easily scaled up to a multi-machine setup.
[0159] Figure 6 In an embodiment, a host or master server configured with a GPU accelerator has a local copy of the current policy πtand distributes the policy to all worker machines or servers equipped with one GPU and multiple CPUs. Then, each worker machine or server will collect trajectories for different tasks using πt. For example, the master server assigns task 1, task 2, to task k to working server 1, working server 2, and working server k, respectively. Task 1 trajectory is collected through real machine / simulator 1, task 2 trajectory is collected through real machine / simulator 2, and task k trajectory is collected through real machine / simulator k. After all the collection processes are synchronized, the host or master server collects all the trajectories together and updates the policy to πt+1. This process is iterated until the policy converges.
[0160] In an example embodiment, according to the first execution information of the first task information, determining the first execution trajectory information of the first task information as first training data comprises: processing the first execution information through a first screening model to determine the first execution trajectory information matching the interface model as the first training data; wherein the first execution information comprises first indication information for indicating completion of the first task execution and first execution path information of the first task.
[0161] In the embodiments of the present disclosure, the screening model refers to a machine learning or deep learning model used to further screen the training data obtained through the feedback model, and determine the machine learning or deep learning model used to continue training the interface model. The screening model can also be referred to as a critic model, that is, an evaluation or assessment model, which, in the embodiments of the present disclosure, refers to a model for automatically screening the training data. The screening model includes a first screening model and / or a second screening model. For example, the first screening model refines the reward (for example, the first indication information) output by the reward model, so as to automatically screen the accumulated training data, and select the training data suitable for the current capability of the U I Agent to continue training the U I Agent.
[0162] For example, the capability of the U I Agent can be different at different times or stages of training. For example, assuming that the training is set to iterate for 9 rounds, the capability of the U I Agent at the first to third rounds can be determined as being in the initial stage, the capability of the U I Agent at the fourth to sixth rounds can be determined as being in the middle stage, and the capability of the U I Agent at the seventh to ninth rounds can be determined as being in the later stage.
[0163] In some embodiments, the first execution information includes first indication information indicating completion of the first task execution and first execution path information of the first task. The first execution path information is used to indicate the number of steps performed and / or the execution path taken when the first task execution is completed. For example, when the reward output by the reward model is equal to 1, the reward is refined based on the reward according to the first execution path information.
[0164] In some embodiments, the first execution path information includes the total number of steps performed after the first task is completed, that is, the execution steps of the first task. In the first stage (for example, the initial stage) of training of the U I agent or the interface model, the model capability is low, and it can not be able to process the training data with too many execution steps. For example, in the first stage of training of the U I agent or the interface model, the reward is decreased according to the increase of the execution steps of each execution trajectory information (including the first execution trajectory information). For example, initially, the reward = 1, for the execution trajectory information completed in one step, the reward remains 1, for the execution trajectory information completed in two steps, the reward = 1 x 0.9, and for the execution trajectory information completed in three steps, the reward = 1 x 0.9 2and so on. Exemplarily, the reward of the first execution path information updated by the first screening model is referred to as the first index. Exemplarily, in the first stage of training the UI agent or the interface model, the execution trajectory information with a larger first index, i.e., a relatively smaller number of execution steps, is selected as the training data for training the UI agent or the interface model. Exemplarily, in the second stage (e.g., the middle stage) of training the UI agent or the interface model, the execution trajectory information with a smaller first index than the first stage, i.e., a relatively larger number of execution steps than the first stage, is selected as the training data for training the UI agent or the interface model. Exemplarily, in the third stage (e.g., the later stage) of training the UI agent or the interface model, the execution trajectory information with a smaller first index than the second stage, i.e., a relatively larger number of execution steps than the second stage, is selected as the training data for training the UI agent or the interface model. Here, the training of the UI agent or the interface model is exemplarily divided into three stages, but the present disclosure is not limited thereto, and the training stages can be divided according to actual scenarios and needs. That is, the present disclosure adopts a training method with gradually increasing difficulty, gradually improving the inference ability of the UI agent or the interface model, so that the training data selected in each stage can match the model ability of the UI agent or the interface model in the corresponding stage. For example, the initial stage can correctly infer the inside of a single UI screenshot, the middle stage can correctly infer the relationship between two UI screenshots, and the later stage can realize end-to-end task inference ability. It can be understood that the data here is only used for exemplification, and as long as the first screening model outputs a smaller first index as the number of execution steps increases.
[0165] In some embodiments, the first execution path information includes the path executed after the first task is completed. In some cases, the first task is completed, but the execution path can be different even if the same number of execution steps is used. In other cases, the first task is completed, and both the number of execution steps and the execution path can be different. For example, for the task of ordering a single latte, although it is completed in three steps, the first execution path is to enter the latte page from the home page, then put the latte into the shopping cart, and then enter the payment page; the second execution path is to directly enter the user's common shopping list page, then put the latte into the shopping cart, and then enter the payment page. Different execution paths reflect the proficiency or difficulty of the task of ordering a latte. In some embodiments, the first screening model can adjust the reward output by the reward model according to the different execution paths of the task, and output different first indexes. The corresponding execution trajectory information is selected according to the first index to train the UI agent or the interface model.
[0166] In an example embodiment, according to the first execution information of the first task information, determining the first execution track information of the first task information as the first training data comprises: processing the first execution information and the first execution track information by a first screening model to determine the first execution track information as the first training data. The first execution information includes first indication information indicating completion of the first task execution.
[0167] In some embodiments, the first screening model comprehensively processes the first execution information and the first execution track information to obtain more diversified training data to train the UI agent or the interface model.
[0168] In an example embodiment, the method provided by the embodiments of the present disclosure further comprises: training the interface model using a training data set, the training data set including training task information and training execution track information thereof, the training execution track information including training interface information of a training task corresponding to the training task information and labeled training action information thereof.
[0169] In some embodiments, before training the interface model using the first training data obtained by the feedback model, a pre-training model is obtained, and the pre-training model is fine-tuned using a labeled training data set to obtain an interface model with preliminary reasoning ability of action information. For example, the training data set can be obtained by manually labeling training action information corresponding to training interface information of training task information.
[0170] In some embodiments, the training data set further labels whether each training interface information indicates that the corresponding training task information has been executed. In this way, after the fine-tuning stage, the interface model also has the ability to judge whether the task is completed.
[0171] In an example embodiment, processing the first execution information and the first execution track information by the first screening model to determine the first execution track information as the first training data comprises: based on the first execution information being first indication information indicating completion of the first task execution, processing interface information in the first execution track information and the first task information by the first screening model; determining that the interface information in the first execution track information is different from the training interface information in the training data set matching the first task information, and determining the first execution track information as the first training data. The interface information in the first execution track information includes the first interface information and the second interface information.
[0172] In some embodiments, the first screening model, upon receiving the first indication information indicating that the first task execution is completed, finds, through the first task information, training task information in the training data set that matches the first task information, for example, both are tasks of ordering a cup of coffee through an APP, and then compares the training interface information in the training task information with the interface information of the first task information to indicate, by the first index, whether new interface information (e.g., new UI screenshot) different from the training interface information appears in the interface information of the first task information. For example, the first index when new interface information appears is greater than the first index when no new interface information appears. For example, the screening first index indicates the first execution track information in which new interface information appears as the first training data. Thus, new training data different from the training data set in the fine-tuning stage can be obtained, which enriches the diversity of the training data, for example, UI screenshots of different advertisements played at different times and the predicted action information of clicking the "X" in the upper right corner of the advertisement to close the advertisement and then continuing to normally perform other actions are obtained. UI screenshots loaded under different network conditions or UI screenshots of dynamic UI display can also be obtained, thereby improving the inference ability and task completion rate of the trained interface model.
[0173] Figure 7 is another reinforcement learning architecture diagram according to an exemplary embodiment of the present disclosure. As shown in Figure 7 The agent 420 includes an interface model 421 and a first screening model 422. The agent 420 is an online inference part, and in the online reinforcement learning scheme, the first screening model 422 is additionally added as a critic model for screening training data.
[0174] For example, the interface model 421 is an Actor model, which is an online inference model, i.e., the model parameters generated by training can directly predict actions and parameters. The first screening model 422 is a critic model, which is a value model for screening training data. For example, the first screening model 422 screens the matching training data according to the current model capability to iterate the model. The training data that does not match the current model capability is discarded as discarded data and is not used to iterate the model.
[0175] Figure 7The embodiment includes a critic model, which is a step-level model. The critic model is an evaluator that smoothes and subdivides the reward. The step-level critic model scores the reward discount output by the reward model, i.e., the reward outputs only 0 or 1 points and cannot distinguish how many steps it takes to complete the task. The step-level critic model deeply processes the reward. It filters the training samples or training data according to the current ability of the UI Agent model. For example, when the UI Agent model is in the early stage, it cannot handle complex tasks, so it filters tasks with fewer steps (e.g., one step or two steps) as training samples for the UI Agent model. When the UI Agent model is in the middle stage, it filters training samples that can be completed in general steps as training samples for the UI Agent model. When the UI Agent model is in the late stage, it gives tasks that can be completed with more steps as training samples for the UI Agent model. Because the reward itself indicates whether the task is completed. Therefore, the input of the critic model includes the reward and how many steps it takes to complete the task.
[0176] In some embodiments, the first filtering model filters good-quality training data, i.e., training data with fewer steps, for the iterative model. The scoring mechanism of the critic model can be that the more steps it takes to complete a task instruction, the lower the score. That is, the more skilled, the faster and shorter path to complete a task. But it is not absolute, because some tasks must have 2 steps, i.e., it is not said that the shorter path is absolutely better.
[0177] In some embodiments, the step-level critic model is a small model, and its training sample is the UI screenshot and its label indicating whether the task is completed, e.g., the current UI screenshot and the next UI screenshot after the action is performed. The other models mentioned in the embodiments of the present disclosure are a trajectory as a training sample, i.e., all UI screenshots and action information required to complete a task as a training sample.
[0178] In some embodiments, the input of the step-level critic model is the UI screenshot, and whether the UI screenshot is a new UI screenshot relative to the training data set of the fine-tuning stage is determined according to the UI screenshot. If it is, the trajectory corresponding to the UI screenshot is taken as a training sample, otherwise, it is not taken as a training sample, so that new samples that are not the same can be continuously added to train the UI Agent. Or when the reward is equal to 1, each UI screenshot in the task is input into the critic model in turn to determine whether a new UI screenshot appears.
[0179] In an exemplary embodiment, determining the first execution trajectory information of the first task information as the first training data based on the first execution information of the first task information includes: processing the first execution information and the first task information through a second filtering model to determine the first execution trajectory information that matches the interface model as the first training data.
[0180] For example, the second filtering model is a machine learning or deep learning model used to filter the first training data based on task information or the complexity of the task description. For example, the second filtering model is a task-level critic model, that is, it filters training data that matches the capabilities of the interface model based on the complexity of the task description.
[0181] In this embodiment of the disclosure, when the second filtering model receives a first execution information indicating that the first task has been completed, such as when it receives the first instruction information, it outputs a second indicator based on the complexity of the first task information. For example, the higher the complexity of the first task information, the larger the second indicator; the lower the complexity of the second task information, the smaller the second indicator. For example, the complexity of the first task information can be determined based on the length of the natural language contained in the first task information, the number of steps required to complete the task, or the difficulty of the task operation. For example, assuming the first task information is "Order me a coffee" or "Order me a coffee, no ice, 30% sugar, delivered in 20 minutes," the second filtering model can determine that the former task description is simpler than the latter, and the corresponding second indicator is smaller. For example, in the first stage of interface model training, execution trajectory information with a relatively small second indicator can be selected as training data. In the second stage of interface model training, execution trajectory information with a larger second indicator than in the first stage can be selected as training data. In the third stage of interface model training, execution trajectory information with a larger second indicator than in the second stage can be selected as training data.
[0182] In an exemplary embodiment, determining the first execution trajectory information of the first task information as the first training data based on the first execution information of the first task information includes: processing the first execution information through a first filtering model to determine a first indicator of the first execution trajectory information; processing the first execution information and the first task information through a second filtering model to determine a second indicator of the first execution trajectory information; and determining the first execution trajectory information that matches the interface model as the first training data based on the first indicator and the second indicator.
[0183] Figure 8 This is another reinforcement learning architecture diagram illustrated according to an exemplary embodiment of the present disclosure. For example... Figure 8As shown, the intelligent agent 420 includes an interface model 421 and a second filtering model 423. The second filtering model 423 is a task-level critic model, whose input is a task description or task instruction. Some task descriptions are simple, such as ordering a latte, while others are complex, such as ordering a latte without sugar, at room temperature, etc. Based on the current capabilities of the UI Agent, training samples with simple task descriptions are matched in the initial stage, and training samples with complex task descriptions are matched later. This is because the UI Agent cannot handle overly complex tasks in the early stages.
[0184] In some embodiments, the capability of the UI Agent can be determined based on the task completion rate, or based on the training execution time.
[0185] Figure 9 This is another reinforcement learning architecture diagram illustrated according to an exemplary embodiment of the present disclosure. For example... Figure 9 As shown, the first screening model 422 and the second screening model 423 can be combined to screen training data. For example, the first indicator output by the first screening model 422 and the second indicator output by the second screening model 423 can be weighted and summed to obtain a fusion indicator. Training data that matches the capabilities of the current UIAgent can then be selected based on the fusion indicator.
[0186] The online reinforcement learning AI agent technology integrates multimodal large-model inference with reinforcement learning dynamic decision-making mechanisms, forming the following technical framework: On the one hand, a perception and decision-making closed loop, based on a visual language model (VLM), allows the agent to perceive the environmental state by parsing the UI and generate action commands such as clicks and inputs, along with their corresponding parameters. For example, the current UI screenshot and the user's current task input are input into the VLM, which outputs the action commands and their corresponding parameters based on the current UI screenshot, the current task input, and the current policy. On the other hand, reinforcement learning drives optimization, employing a partially observable Markov decision process (POMDP) model. Through a trial-and-error mechanism, the policy is dynamically adjusted to address challenges such as task reward delays and environmental interference. For example, the POMDP model uses the aforementioned critic model, primarily referring to its borrowing of this idea. In some embodiments, the UI agent itself determines whether a task is completed. However, because the training samples used by the UI agent during the fine-tuning phase are relatively static, it cannot accurately identify whether a dynamic UI screenshot indicates task completion. For instance, a task might actually be completed in two steps, but due to misidentification, it might be considered to have been completed in the 20th step, leading to task reward delays. The embodiments disclosed herein introduce an additional reward model, independent of UIAgent. Due to the additional constraints, it can increase the number of training samples that are different from those in the fine-tuning stage, thereby more accurately identifying whether the task has been completed, that is, accurately identifying whether the task has been completed in advance.
[0187] The method provided in this disclosure can improve the completion rate of automated operations, such as sending messages and ordering coffee, by using an online reinforcement learning UI agent, thereby enhancing the accessibility and interactive experience. Table 1 provides experimental comparison data.
[0188] Table 1
[0189]
[0190] Figure 10 This is a block diagram illustrating a multimodal data processing apparatus according to an exemplary embodiment of the present disclosure. Figure 10As shown, the multi-modal data processing apparatus 1000 provided by the embodiments of the present disclosure includes an obtaining unit 1010 and a processing unit 1020. The obtaining unit 1010 is configured to obtain first task information to be executed by a first device. The obtaining unit 1010 is further configured to acquire first interface information of the first device in response to the first task information. The processing unit 1020 is configured to process the first interface information by using an interface model to obtain first action information. The processing unit 1020 is further configured to control the first device to execute the first action information. The obtaining unit 1010 is further configured to obtain second interface information of the first device. The processing unit 1020 is further configured to process the second interface information by using a feedback model to determine first execution information of the first task information. The processing unit 1020 is further configured to determine, according to the first execution information of the first task information, first execution track information of the first task information as first training data, the first execution track information including the first interface information, the first action information and the second interface information. The processing unit 1020 is further configured to train the interface model by using the first training data.
[0191] In exemplary embodiments, the processing unit 1020 is further configured to: if it is determined that a first task corresponding to the first task information is executed completely by processing the second interface information by using the feedback model, generate first indication information indicating that the first task is executed completely; and if it is determined that the first task is not executed completely by processing the second interface information by using the feedback model, generate second indication information indicating that the first task is not executed completely. The first execution information includes the first indication information and the second indication information.
[0192] In exemplary embodiments, the processing unit 1020 is further configured to process the first interface information, the second interface information and the first task information by using the feedback model to determine first execution information of the first task.
[0193] In exemplary embodiments, the processing unit 1020 is further configured to: if it is determined that a first task corresponding to the first task information is executed completely by processing the first interface information, the second interface information and the first task information by using the feedback model, generate first indication information indicating that the first task is executed completely; and if it is determined that the first task is not executed completely by processing the first interface information, the second interface information and the first task information by using the feedback model, generate second indication information indicating that the first task is not executed completely. The first execution information includes the first indication information and the second indication information.
[0194] In an example embodiment, the processing unit 1020 is further configured to process the first execution information by a first screening model, and determine the first execution trajectory information matching the interface model as the first training data. The first execution information includes first indication information indicating completion of the first task execution and first execution path information of the first task.
[0195] In an example embodiment, the processing unit 1020 is further configured to process the first execution information and the first execution trajectory information by a first screening model, and determine the first execution trajectory information as the first training data. The first execution information includes first indication information indicating completion of the first task execution.
[0196] In an example embodiment, the processing unit 1020 is further configured to train the interface model by using a training data set, the training data set including training task information and training execution trajectory information thereof, the training execution trajectory information including training interface information and labeled training action information of a training task corresponding to the training task information.
[0197] In an example embodiment, the processing unit 1020 is further configured to process interface information in the first execution trajectory information and the first task information by the first screening model based on the first execution information being the first indication information indicating completion of the first task execution, and determine that the interface information in the first execution trajectory information is different from the training interface information in the training data set matching the first task information, and determine the first execution trajectory information as the first training data. The interface information in the first execution trajectory information includes the first interface information and the second interface information.
[0198] In an example embodiment, the processing unit 1020 is further configured to process the first execution information and the first task information by a second screening model, and determine the first execution trajectory information matching the interface model as the first training data.
[0199] In an example embodiment, the processing unit 1020 is further configured to process the first execution information by a first screening model, and determine a first index of the first execution trajectory information; process the first execution information and the first task information by a second screening model, and determine a second index of the first execution trajectory information; and determine the first execution trajectory information matching the interface model as the first training data according to the first index and the second index.
[0200] In an example embodiment, the obtaining unit 1010 is further configured to obtain second task information to be executed by a second device. The obtaining unit 1010 is further configured to obtain third interface information of the second device in response to the second task information. The processing unit 1020 is further configured to process the third interface information by using the interface model to obtain second action information. The processing unit 1020 is further configured to control the second device to execute the second action information. The obtaining unit 1010 is further configured to obtain fourth interface information of the second device. The processing unit 1020 is further configured to process the fourth interface information by using the interface model to determine second execution information of the second task information.
[0201] Figure 10 Other contents of the embodiments can refer to other embodiments and will not be described herein.
[0202] Figure 11 is a network interaction architecture diagram of the method provided by the embodiments of the present disclosure. As shown in Figure 11 , the application scenario includes a user 1101, an electronic device 1102, and a server 1103.
[0203] As an example, the training of the model provided by the embodiments of the present disclosure can be performed by the server 1103, and the server 1103 can save the trained model as an interface model or send the trained interface model to the electronic device 1102.
[0204] As an example, the method provided by the embodiments of the present disclosure can be implemented in the electronic device 1102 (for example, a mobile phone, a computer, a smart voice interaction device, a smart home appliance, a vehicle terminal, an aircraft, etc.). As shown in Figure 11 , the electronic device 1102 can perform task automatic execution based on the task information issued by the user 1101 and without relying on the server 1103, by using the processor and the memory of the electronic device 1102.
[0205] As another example, the method provided by the embodiments of the present disclosure can be implemented in the cloud. As shown in Figure 11 , after the electronic device 1102 receives the task information (for example, second task information) issued by the user 1101, the electronic device 1102 sends the task information to the server 1103, the server 1103 processes the task information, and after obtaining the action information matched with the task information, the server 1103 returns the action information to the electronic device 1102, so that the electronic device 1102 executes the action corresponding to the action information.
[0206] It can be understood that the electronic device mentioned in the embodiments of the present disclosure includes but is not limited to a terminal or a server. In other words, the electronic device can be a server or a terminal, or a system composed of a server and a terminal. Among them, the terminal can be an electronic device, including but not limited to a mobile phone, a tablet computer, a desktop computer, a notebook computer, a palm computer, a vehicle-mounted device, an augmented reality / virtual reality (AR / VR) device, a head-mounted display, a smart television, a wearable device, a smart speaker, a digital camera, a camera, and other mobile internet devices (MID) with network access capability, or a terminal in a train, a ship, a flight, and the like.
[0207] Among them, the server mentioned above can be a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, vehicle-road cooperation, content delivery networks (CDN), and basic cloud computing services such as big data and artificial intelligence platforms.
[0208] Optionally, the data involved in the embodiments of the present disclosure can be stored in a computer device, or can be stored based on cloud storage technology, which is not limited here.
[0209] Figure 12 is a schematic diagram of an application scenario of the method provided by the embodiments of the present disclosure. Figure 12 is a schematic diagram of the method provided by the embodiments of the present disclosure applied to a vehicle-mounted scenario. As shown in Figure 12 The vehicle-mounted terminal 1202 can send task information to the electronic device 1201, the electronic device 1201 calls the interface model obtained by training to obtain action information, and returns the action information to the vehicle-mounted terminal 1202 for execution. The electronic device 1201 can be a server where an application program is located, or can belong to the vehicle-mounted terminal 1202 (i.e., the background of the vehicle-mounted terminal 1202), and the like, which is not limited here.
[0210] Among them, the vehicle-mounted terminal 1202 can be arranged in the vehicle 1203, which is not limited here. The sound collecting component can be a microphone arranged on the steering wheel of the vehicle 1203, which is used to collect voice information of the user, convert the voice information into text form, and obtain task information.
[0211] The application can be displayed in the vehicle terminal 1202. The application in the embodiment of the present disclosure can be any application program. The application programs used in different scenarios can be different, such as remote video conference, education, message, travel, audiobook, and advertisement, etc. Various applications that can be applied on the vehicle. Among them, the travel scenario can be further divided into different sub-scenarios such as commuting, traveling, and congestion; the social scenario can be further divided into different sub-scenarios such as gathering and interaction; and the entertainment scenario can be further divided into different sub-scenarios such as listening to songs, novels, and games.
[0212] Further, the embodiment of the present disclosure also provides an electronic device, comprising: a processor; a memory for storing processor executable instructions; wherein the processor is configured to implement the steps of the multi-modal data processing method of any of the above embodiments.
[0213] Referring to Figure 13 , Figure 13 is a structural schematic diagram of an electronic device provided by the embodiment of the present disclosure. As Figure 13 shown, the electronic device 1300 in the embodiment of the present disclosure can include one or more processors 1301, a memory 1302, and an input / output interface 1303. The processor 1301, the memory 1302, and the input / output interface 1303 are connected through a bus 1304. The memory 1302 is used to store a computer program, the computer program including program instructions, the input / output interface 1303 is used to receive data and output data, such as used for data interaction between a host computer and a computer device, or used for data interaction between various virtual machines in the host computer; the processor 1301 is used to execute the program instructions stored in the memory 1302.
[0214] Among them, the processor 1301 can perform the following operations: obtaining first task information to be executed by a first device; in response to the first task information, obtaining first interface information of the first device; processing the first interface information through an interface model to obtain first action information; controlling the first device to execute the first action information; obtaining second interface information of the first device; processing the second interface information through a feedback model to determine first execution information of the first task information; according to the first execution information of the first task information, determining the first execution track information of the first task information as first training data, the first execution track information including the first interface information, the first action information, and the second interface information; training the interface model using the first training data.
[0215] The memory 1302 can include read-only memory and random access memory, and provide instructions and data for the processor 1301 and the input / output interface 1303. A part of the memory 1302 can also include non-volatile random access memory.
[0216] Figure 14 This is a functional block diagram illustrating a vehicle according to an exemplary embodiment of the present disclosure. For example, vehicle 1400 can be a hybrid vehicle, a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicle. Vehicle 1400 can be an autonomous vehicle, a semi-autonomous vehicle, or a non-autonomous vehicle. (Refer to...) Figure 14 The vehicle 1400 may include various subsystems, such as an infotainment system 1410, a perception system 1420, a decision control system 1430, a drive system 1440, and a computing platform 1450. The vehicle 1400 may also include more or fewer subsystems, and each subsystem may include multiple components. Furthermore, each subsystem and component of the vehicle 1400 can be interconnected via wired or wireless means. In some embodiments, the infotainment system 1410 may include a communication system, an entertainment system, and a navigation system. The perception system 1420 may include several sensors for sensing information about the environment surrounding the vehicle 1400. For example, the perception system 1420 may include a global positioning system (which may be GPS, BeiDou, or other positioning systems), an inertial measurement unit (IMU), lidar, millimeter-wave radar, ultrasonic radar, and a camera device. The decision control system 1430 may include a computing system, a vehicle controller, a steering system, a throttle, and a braking system. The drive system 1440 may include components that provide power to the vehicle 1400. In one embodiment, the drive system 1440 may include an engine, an energy source, a transmission system, and wheels. The engine may be one or a combination of internal combustion engines, electric motors, and compressed air engines. The engine is capable of converting energy provided by the energy source into mechanical energy. Some or all of the functions of the vehicle 1400 are controlled by a computing platform 1450. The computing platform 1450 may include at least one processor 1451 and a memory 1452, the processor 1451 being capable of executing instructions 1453 stored in the memory 1452.
[0217] The processor 1451 can be any conventional processor, such as a commercially available CPU. The processor can also include a Graphics Processing Unit (GPU), a Field Programmable Gate Array (FPGA), a System on Chip (SOC), an Application Specific Integrated Circuit (ASIC), or a combination thereof. The memory 1452 can be realized by any type of volatile or nonvolatile storage devices or a combination thereof. In addition to the instructions 1453, the memory 1452 can also store data, such as road maps, route information, the position, direction, speed, and the like of the vehicle. The data stored in the memory 1452 can be used by the computing platform 1450. In the embodiments of the present disclosure, the processor 1451 can execute the instructions 1453 to complete all or part of the steps of the methods described above.
[0218] The embodiments of the present disclosure provide an electronic device, comprising: a processor, an input output interface, a memory, obtaining a computer program in the memory by the processor, and executing each step of the method shown in any of the above embodiments.
[0219] The embodiments of the present disclosure also provide a computer readable storage medium storing a computer program, which is adapted to be loaded by the processor and execute the text processing method provided by each step of any of the above embodiments. For details, refer to the implementation manner provided by each step of any of the above embodiments, which will not be described here. In addition, the beneficial effects of using the same method will not be described here. For technical details of the computer readable storage medium embodiments involved in the present disclosure, refer to the description of the method embodiments of the present disclosure. As an example, the computer program can be deployed to execute on one computer device, or on multiple computer devices located in one place, or on multiple computer devices distributed in multiple places and interconnected through a communication network.
[0220] The computer readable storage medium can be an internal storage unit of the device or the electronic device provided by any of the above embodiments, such as the hard disk or memory of the computer device. The computer readable storage medium can also be an external storage device of the computer device. Further, the computer readable storage medium can include both the internal storage unit and the external storage device of the computer device. The computer readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer readable storage medium can also be used to temporarily store data that has been output or will be output.
[0221] The embodiment of the present disclosure further provides a computer program product, comprising a computer program, wherein the computer program is executed by a processor to implement the multi-modal data processing method according to any one of the above embodiments.
[0222] The embodiment of the present disclosure further provides a computer program product or a computer program, comprising computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to enable the computer device to perform the method provided in any one of the various optional manners described above.
[0223] As to the apparatus in the above embodiments, the specific manners in which the various modules / units perform operations have been described in detail in the embodiments of the method, and will not be described in detail here. It should be noted that the acquisition, storage, use, processing, etc. of data in the technical solutions of the present disclosure all comply with relevant provisions of national laws and regulations.
[0224] The terms "first", "second", etc. in the description, claims and drawings of the embodiments of the present disclosure are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. It should be understood that the present disclosure is not limited to the precise structures already described and shown in the drawings, and various modifications and changes can be made without departing from the scope thereof.
Claims
1. A multimodal data processing method, characterized in that, include: Obtain information on the first task to be executed by the first device; In response to the first task information, obtain the first interface information of the first device; The first interface information is processed by the interface model to obtain the first action information; Control the first device to execute the first action information; Obtain the second interface information of the first device; The second interface information is processed by a feedback model to determine the first execution information of the first task information; Based on the first execution information of the first task information, the first execution trajectory information of the first task information is determined to be used as the first training data. The first execution trajectory information includes the first interface information, the first action information and the second interface information. The interface model is trained using the first training data.
2. The method according to claim 1, characterized in that, The second interface information is processed through a feedback model to determine the first execution information of the first task information, including: If the feedback model processes the second interface information and determines that the first task corresponding to the first task information has been completed, then a first indication message is generated to indicate that the first task has been completed. If the feedback model processes the second interface information and determines that the first task has not been completed, then a second indication message is generated to indicate that the first task has not been completed. The first execution information includes the first instruction information and the second instruction information.
3. The method according to claim 1, characterized in that, The second interface information is processed through a feedback model to determine the first execution information of the first task, including: The feedback model processes the first interface information, the second interface information, and the first task information to determine the first execution information of the first task.
4. The method according to claim 3, characterized in that, The feedback model processes the first interface information, the second interface information, and the first task information to determine the first execution information of the first task, including: If the feedback model processes the first interface information, the second interface information, and the first task information, and determines that the first task corresponding to the first task information has been completed, then a first indication information is generated to indicate that the first task has been completed. If the feedback model processes the first interface information, the second interface information, and the first task information, and determines that the first task has not been completed, then a second indication message is generated to indicate that the first task has not been completed. The first execution information includes the first instruction information and the second instruction information.
5. The method according to claim 1, characterized in that, Based on the first execution information of the first task information, the first execution trajectory information of the first task information is determined to be used as the first training data, including: The first execution information is processed by the first screening model to determine the first execution trajectory information that matches the interface model as the first training data; The first execution information includes first indication information for indicating the completion of the first task and first execution path information for the first task.
6. The method according to claim 1, characterized in that, Based on the first execution information of the first task information, the first execution trajectory information of the first task information is determined to be used as the first training data, including: The first execution information and the first execution trajectory information are processed by the first screening model to determine the first execution trajectory information as the first training data. The first execution information includes first indication information for indicating that the first task has been completed.
7. The method according to claim 6, characterized in that, Also includes: The interface model is trained using a training dataset, which includes training task information and training execution trajectory information. The training execution trajectory information includes training interface information and labeled training action information for the training task corresponding to the training task information.
8. The method according to claim 7, characterized in that, The first execution information and the first execution trajectory information are processed by a first screening model to determine the first execution trajectory information as the first training data, including: Based on the first execution information being a first indication information indicating the completion of the first task, the interface information and the first task information in the first execution trajectory information are processed by the first filtering model; If the interface information in the first execution trajectory information is determined to be different from the training interface information in the training dataset that matches the first task information, the first execution trajectory information is used as the first training data. The interface information in the first execution trajectory information includes the first interface information and the second interface information.
9. The method according to claim 1, characterized in that, Based on the first execution information of the first task information, the first execution trajectory information of the first task information is determined to be used as the first training data, including: The first execution information and the first task information are processed by the second screening model to determine the first execution trajectory information that matches the interface model as the first training data.
10. The method according to claim 1, characterized in that, Based on the first execution information of the first task information, the first execution trajectory information of the first task information is determined to be used as the first training data, including: The first execution information is processed by the first screening model to determine the first indicator of the first execution trajectory information; The first execution information and the first task information are processed by the second screening model to determine the second indicator of the first execution trajectory information; Based on the first indicator and the second indicator, the first execution trajectory information that matches the interface model is determined as the first training data.
11. The method according to claim 1, characterized in that, Also includes: Obtain information on the second task to be executed by the second device; In response to the second task information, obtain the third interface information of the second device; The second action information is obtained by processing the third interface information through the interface model. Control the second device to execute the second action information; Obtain the fourth interface information of the second device; The fourth interface information is processed through the interface model to determine the second execution information of the second task information.
12. A multimodal data processing device, characterized in that, include: The acquisition unit is used to acquire information about the first task to be executed by the first device. The obtaining unit is also configured to respond to the first task information and obtain the first interface information of the first device; The processing unit is used to process the first interface information through the interface model to obtain the first action information; The processing unit is also configured to control the first device to execute the first action information; The obtaining unit is further configured to obtain the second interface information of the first device; The processing unit is further configured to process the second interface information through a feedback model to determine the first execution information of the first task information; The processing unit is further configured to determine, based on the first execution information of the first task information, to use the first execution trajectory information of the first task information as the first training data, wherein the first execution trajectory information includes the first interface information, the first action information and the second interface information; The processing unit is also used to train the interface model using the first training data.
13. The apparatus according to claim 12, characterized in that, The processing unit is further configured to process the first interface information, the second interface information, and the first task information through the feedback model to determine the first execution information of the first task.
14. The apparatus according to claim 12, characterized in that, The processing unit is further configured to process the first execution information through a first screening model, and determine the first execution trajectory information that matches the interface model as the first training data; The first execution information includes first indication information for indicating the completion of the first task and first execution path information for the first task.
15. The apparatus according to claim 12, characterized in that, The processing unit is further configured to process the first execution information and the first task information through a second filtering model, and determine the first execution trajectory information that matches the interface model as the first training data.
16. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured as follows: The steps of implementing the multimodal data processing method according to any one of claims 1 to 11.
17. A non-transitory computer-readable storage medium, wherein when instructions in the storage medium are executed by a processor of a mobile terminal, the mobile terminal is enabled to perform the multimodal data processing method of any one of claims 1 to 11.
18. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the multimodal data processing method as described in any one of claims 1 to 11.