Task execution method and device, electronic equipment and medium
By combining the graphical user interface with artificial intelligence, using the prediction model to generate interactive information, the task planning model to split tasks, and the exception handling model to identify and handle exceptions, the problems of task execution interference and information loss in the existing technology are solved, and more efficient and safe task execution is achieved.
Patent Information
- Application Number
- CN202510655609.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-09-12
AI Technical Summary
In the field of combining graphical user interfaces with artificial intelligence, existing technologies have difficulty effectively solving problems that interfere with task execution, such as pop-up ads or messages interfering with normal task execution, missing information or requiring multiple rounds of interaction during system operation, and operations involving user safety requiring user confirmation.
Using an AI-based prediction model, the system generates at least one prediction, including interaction content, interaction attributes, and interaction instructions, to assist the terminal in completing the current task. Furthermore, the task planning model breaks down user instructions into multiple subtasks, and an exception handling model identifies and handles exceptions to ensure smooth task execution.
It improves the completion rate and efficiency of task execution, reduces task interruptions due to interference or information loss, and enhances the intelligence and security of user interaction.
Smart Images

Figure CN120631483A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, and in particular to a task execution method and device, electronic equipment, and medium. Background Art
[0002] In the field of combining graphical user interfaces (GUIs) with artificial intelligence (AI), back-end services are required to automatically identify UI operations to be performed based on the GUI and send them to the client for execution. However, in actual scenarios, there may be interference with task execution or information loss, and operational issues involving user safety remain to be resolved. Summary of the Invention
[0003] The present disclosure provides a task execution method and device, electronic equipment, and medium.
[0004] The first aspect embodiment of the present disclosure proposes a task execution method, including: based on current image information and a current task, generating at least one prediction information through a prediction model, the current image information being image information corresponding to a current display interface obtained during the terminal's execution of the current task, and the at least one prediction information including at least one of interactive content, interactive attributes, and interactive instructions; based on the instruction information corresponding to the first prediction information in the at least one prediction information and the next image information, using the prediction model to determine the execution status of the current task until the current task is in a completed state.
[0005] In some embodiments of the present disclosure, the method also includes: generating at least one task based on user instructions through a task planning model, the task planning module is obtained by training the initial task planning model using a task planning data set, at least one task includes task information and a task sequence number, and at least one task is used for the terminal to perform the corresponding task according to the task sequence number; determining current image information, the current image information is obtained when the display interface of the terminal meets preset conditions during the execution of the current task, and the current image information includes at least one missing information or prompt information.
[0006] In some embodiments of the present disclosure, the preset conditions include at least one of the following: there is at least one missing information or prompt information that makes it impossible to continue the current task; there is a target operation, and the target operation is an operation that requires the user to confirm through the terminal.
[0007] In some embodiments of the present disclosure, based on the current image information and the current task, at least one prediction information is generated through a prediction model, including at least one of the following: through the prediction model, second prediction information is generated according to the first missing information in the at least one missing information and the task information corresponding to the current task, and the interaction content corresponding to the second prediction information includes the first missing information, and / or the interaction instruction corresponding to the second prediction information is selection or confirmation; through the prediction model, third prediction information is generated according to the target operation in the current image information and the task information corresponding to the current task, the interaction content corresponding to the third prediction information includes the target operation, and the interaction instruction corresponding to the third prediction information is confirmation or rejection.
[0008] In some embodiments of the present disclosure, the method further includes: determining the execution status of the current task using a prediction model based on at least one prompt information in the current image information; when the execution status of the current task is unable to continue, identifying the abnormality of the current image information through an exception handling model, determining the abnormality type and the abnormality handling and / or abnormal attributes corresponding to the abnormality type, so that the terminal can perform the abnormal handling, or the terminal can perform the abnormal handling based on the abnormal attributes.
[0009] In some embodiments of the present disclosure, the method further includes: obtaining a first training data set, the first training data set including multiple image information and at least one of interactive content, interactive attributes, and interactive instructions corresponding to each image information; and using the first training data set to train an initial prediction model to obtain a prediction model.
[0010] In some embodiments of the present disclosure, the method further includes: obtaining a second training data set, the second training data set including multiple image information and an abnormality type corresponding to each image information, and abnormality processing and / or abnormal attributes corresponding to each abnormality type; using the second training data set to train an initial abnormality processing model to obtain an abnormality processing model.
[0011] In the above embodiment, by splitting the user instructions into tasks, the terminal determines the current image information with missing information or important operations during the execution of the corresponding task, so as to generate corresponding interactive actions, realize multiple rounds of human-computer interaction, and improve the completion rate and efficiency of task execution.
[0012] The second embodiment of the present disclosure provides a task execution device, which is configured to execute the task execution method provided in the first embodiment of the present disclosure.
[0013] An embodiment of the third aspect of the present disclosure proposes an electronic device, comprising: a processor and a memory for storing a computer program that can be run on the processor, wherein when the processor is used to run the computer program, it executes the method described in any one of the embodiments of the first aspect of the present disclosure.
[0014] The fourth aspect of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to enable a computer to execute the method described in any one of the embodiments of the first aspect of the present disclosure.
[0015] The fifth aspect of the present disclosure provides a program product, including computer instructions, which are used to enable a computer to execute the method described in any one of the embodiments of the first aspect of the present disclosure.
[0016] In summary, the task execution method and device, electronic device and medium proposed in the present disclosure include: based on the current image information and the current task, generating at least one prediction information through a prediction model, the current image information is the image information corresponding to the current display interface obtained during the terminal's execution of the current task, and the at least one prediction information includes at least one of interaction content, interaction attributes, and interaction instructions; based on the instruction information corresponding to the first prediction information in the at least one prediction information and the next image information, using the prediction model to determine the execution status of the current task until the current task is completed. The prediction model is obtained by training based on AI or artificial intelligence to generate multiple rounds of interaction based on the current screen screenshot of the terminal and the currently executed task, thereby improving the completion rate and efficiency of task execution.
[0017] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.
[0019] Figure 1 This is an interactive diagram of a task execution method proposed in an embodiment of the present disclosure;
[0020] Figure 2 A flowchart of a task execution method proposed in an embodiment of the present disclosure;
[0021] Figure 3 A schematic diagram of the exception handling process proposed in the embodiment of the present disclosure;
[0022] Figure 4 A flowchart of the prediction model proposed in the embodiment of the present disclosure;
[0023] Figure 5 A flowchart of the exception handling model proposed in the embodiment of the present disclosure;
[0024] Figure 6A system architecture diagram for a task execution method;
[0025] Figure 7 A schematic diagram of the structure of a task execution device proposed in an embodiment of the present disclosure;
[0026] Figure 8 This is a schematic diagram of the structure of an electronic device proposed in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] The following describes in detail embodiments of the present disclosure, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present disclosure, and should not be construed as limiting the present disclosure.
[0028] A GUI Agent (Graphical User Interface Agent) is an AI-based automated program that can complete tasks by understanding and manipulating graphical user interfaces (GUIs). Current solutions for GUI Agents mostly focus on improving the model's discriminative capabilities from an algorithmic perspective. However, there are few solutions to address the issues of abnormal input and real-world interactions in real-world application scenarios. For example, unusual pages like pop-up ads or messages can interfere with normal task execution, leading to errors. Alternatively, system operations may encounter situations where key information is missing or multiple options are available, requiring multiple rounds of interaction with the user to supplement key information or select the next step. Furthermore, critical operations involving user security, such as payment, require user confirmation before execution to avoid the serious consequences of incorrect operations.
[0029] In order to solve the above-mentioned problems of abnormal input and real interaction, the present disclosure proposes a task execution method, which uses artificial intelligence / AI models to identify abnormal pages and issue corresponding processing operations, as well as identify user instructions and understand intentions, and then generate multiple rounds of dialogue to achieve the purpose of assisting task completion.
[0030] The task execution method provided in this application will be described in detail below with reference to the accompanying drawings.
[0031] The task execution method proposed in the present disclosure can be applied to the fields of smart cars and smart cockpits for human-computer dialogue, intelligent recognition, and intelligent interaction on vehicle computers. Furthermore, it can be applied to the application fields of artificial intelligence / AI models in vehicles, such as intelligent interaction scenarios between terminals (clients) and the cloud, which are not limited by the present disclosure.
[0032] The task execution method proposed in this disclosure is performed on the basis of user authorization in operations related to user privacy or security, such as obtaining user voice, obtaining current image information, taking screenshots, password-free payment, obtaining training data sets, etc., and strictly abides by relevant laws and regulations such as privacy and security.
[0033] In some embodiments, the task execution method proposed in the present disclosure can be applied to a task execution system, which includes a cloud and a terminal. The cloud and the terminal can communicate with each other. For example, the terminal can upload user instructions or image information of the terminal's current display interface to the cloud, and the cloud can send tasks, prediction information, exception handling and / or exception attributes to the terminal.
[0034] In some embodiments, the task execution method proposed in the present disclosure can be applied to a terminal, which is deployed with a prediction model and an exception handling model, so that the terminal can perform task splitting, generate prediction information, identify exceptions, etc. based on user instructions or image information of the current display interface.
[0035] In some embodiments, the terminal is a device with a display interface, such as a vehicle, a mobile phone, or a computer.
[0036] Figure 1 This is an interactive diagram of a task execution method proposed in an embodiment of the present disclosure. The method is executed by the cloud and the terminal, wherein the prediction model, task planning model, and exception handling model are deployed in the cloud, and the terminal is used to execute tasks, feedback image information, send user instructions, etc. Figure 1 As shown, the method includes the following steps:
[0037] Step 101: The terminal sends a user instruction to the cloud.
[0038] In some embodiments, the user instruction may be sending text information input by the user through the display interface of the terminal to the cloud, or it may be sending audio information input by the user through the audio device of the terminal to the cloud, and the cloud transcribing the audio information into text information to obtain the user instruction; or it may be sending text information and audio information to the cloud, and after the cloud transcribes the audio information into text information, the two parts of text information are integrated to obtain the user instruction. The method of obtaining user instructions is not limited by this disclosure.
[0039] For example, for a specific request made by a user, such as "Please help me order a cup of coffee A using the XXX app", the client records and uploads the audio of the request, and the cloud transcribes the voice into text "Please help me order a cup of coffee A using the XXX app".
[0040] Step 102: The cloud generates at least one task.
[0041] In some embodiments, the current image information is sent by the terminal to the cloud, specifically including: receiving a user instruction, generating at least one task through a task planning model, where the task planning model is obtained by training an initial task planning model using a task planning dataset, and the at least one task includes task information and a task sequence number;
[0042] In some embodiments, the task planning model is obtained by training an initial task planning model using a task planning data set, wherein the initial task planning model can be a traditional machine learning model, such as a linear model, a tree model, a Bayesian model, etc., or a lightweight neural network model, or a large model with strong inference, which is not limited in this disclosure. In the present disclosure, the initial task planning model used by the task planning model is a large model with strong inference, so that the splitting of at least one task obtained can be more reasonable and a higher completion rate can be achieved.
[0043] In some embodiments, a task planning model is used to generate at least one task according to a user instruction, which may be to split the user instruction into multiple ordered subtasks, each subtask corresponding to an operation.
[0044] For example, the task planning module is responsible for global planning, breaking down the task into multiple subtasks to be executed sequentially. This decomposition ensures that the task is progressing in the right direction at a macro level, while also making the subtasks easier to complete, resulting in a higher completion rate. The task planning module is implemented based on a large model with strong reasoning capabilities.
[0045] For example, the cloud-based task planning module breaks down the user's instructions into multiple subtasks: (1) open a certain APP; (2) search for "A Coffee" coffee drink; (3) select the drink configuration; (4) add to the shopping cart; (5) complete the payment.
[0046] Step 103: The cloud sends at least one task to the terminal.
[0047] In some embodiments, the cloud sends at least one task to the terminal, so that the terminal executes the corresponding task according to the task sequence number. The terminal may execute the tasks in sequence according to the order of the multiple subtasks obtained by splitting.
[0048] For example, the cloud task planning module sends the split multiple subtasks to the client, and the client executes the corresponding subtasks in sequence.
[0049] Step 104: The terminal executes at least one task.
[0050] In some embodiments, the terminal executes corresponding tasks according to the task sequence number.
[0051] Step 105: The terminal sends the current image information to the cloud.
[0052] In some embodiments, the current image information is obtained when the display interface of the terminal during the execution of the current task meets the preset conditions, and the current image information includes at least one missing information or prompt information.
[0053] In some embodiments, the preset conditions include at least one of the following: there is at least one missing information or prompt information that makes it impossible to continue the current task; there is a target operation, and the target operation is an operation that requires the user to confirm through the terminal.
[0054] In some embodiments, when the terminal is executing at least one task in sequence, that is, when the terminal is executing the current task, when the display interface meets the preset conditions, the current image information corresponding to the current display interface is sent to the cloud, and the cloud generates at least one prediction information through a prediction model based on the current image information and the task information corresponding to the current task.
[0055] In some embodiments, the display interface meets the preset conditions when there is at least one missing information or prompt information in the display interface that makes it impossible to continue the current task, or there is a target operation in the display interface, and the target operation is an operation that requires user confirmation.
[0056] In some embodiments, the missing information may be necessary information for executing the current task, for example, multiple options that are consistent with the execution of the current task appear in the display interface, or options that are inconsistent with the execution of the current task appear, resulting in the lack of necessary confirmation information for executing the current task. The user is required to fill in the relevant confirmation information before continuing to execute the current task.
[0057] In some embodiments, the target operation may be an operation involving user privacy or security, such as a payment operation, a message sending operation, a red envelope sending operation, etc.
[0058] In some embodiments, the prompt information may be a pop-up window of the abnormal page, such as a system message window or an advertising window; or, the prompt information may be a window corresponding to the target operation, such as a payment operation window or a message sending window or a red envelope sending window; or, the prompt information may be a prompt that the current task is completed.
[0059] In some embodiments, when a page with missing information or target operation appears on the display interface, the terminal needs to send the current image information corresponding to the current display interface to the cloud. The terminal can obtain the current image information by taking a screenshot, recording the screen, etc.
[0060] For example, the client uploads a screenshot, and the cloud understands the screenshot and decides the next action to perform.
[0061] For example, when the client performs the subtask of selecting a beverage configuration, it does not clearly indicate the required cup type (large / medium / small), whether to add ice, and other information, which are crucial for task completion. When there are multiple target drinks (A1 coffee, A2 coffee) in the search results, the user needs to confirm to avoid placing the wrong order.
[0062] For example, when the client is executing a subtask, some important operations, such as payment, sending messages, sending red envelopes, etc., need to be confirmed by the user. To this end, a user confirmation link is added to the interaction process, and the client sends the current screenshot to the cloud, so that the cloud generates corresponding interaction information and sends it to the client to obtain user information or rejection information.
[0063] Step 106: The cloud generates at least one prediction information.
[0064] In some embodiments, the cloud generates at least one prediction information through a prediction model based on current image information and current task.
[0065] In some embodiments, the current image information is image information corresponding to the current display interface obtained during the terminal's execution of the current task, and the at least one prediction information includes at least one of interaction content, interaction attributes, and interaction instructions.
[0066] In some embodiments, while executing a current task, the terminal may, under certain conditions, take a screenshot or photo of the current display interface to obtain current image information, and then send the current image information to the cloud, which uses a prediction model to identify and analyze the current image information and generate at least one prediction information. The specific condition may be that the current display interface requires interaction.
[0067] In some embodiments, at least one prediction information is used to generate content or instructions required for interaction based on the current image information when there is a need for interaction in the current image information.
[0068] In some embodiments, at least one prediction information is used to assist the terminal in completing a current task.
[0069] In some embodiments, based on the current image information and the current task, at least one prediction information is generated through a prediction model, including at least one of the following: through the prediction model, second prediction information is generated according to the first missing information in the at least one missing information and the task information corresponding to the current task, and the interaction content corresponding to the second prediction information includes the first missing information, and / or the interaction instruction corresponding to the second prediction information is selection or confirmation; through the prediction model, third prediction information is generated according to the target operation in the current image information and the task information corresponding to the current task, the interaction content corresponding to the third prediction information includes the target operation, and the interaction instruction corresponding to the third prediction information is confirmation or rejection.
[0070] In some embodiments, the second prediction information and the third prediction information are one or more of the at least one prediction information, the second prediction information is generated when the first missing information is present in the current image information, and the third prediction information is generated when the target operation is present in the current image information. The first missing information is any one of the at least one missing information, and the target operation corresponds to the prompt information in the current image information.
[0071] In some embodiments, the prediction model generates second prediction information based on the first missing information and task information corresponding to the current task. The second prediction information is used to inquire the user about the first missing information to obtain the user's selection information or confirmation information.
[0072] For example, when a client searches for coffee A in a certain APP, since the search results include A1 coffee, A2 coffee, three cup types (large / medium / small), and options such as whether to add ice, the client uploads the screenshot to the cloud. The cloud generates a script based on the screenshot and the currently executed subtask to ask the user which type of coffee to choose, which cup type to choose, and whether to add ice, etc.
[0073] In some embodiments, the prediction model generates third prediction information based on the target operation and task information corresponding to the current task. The third prediction information is used to obtain the user's confirmation or rejection to complete the target operation or interrupt the target operation.
[0074] For example, when the client executes the subtask "payment", since it involves user confirmation, the client uploads a screenshot to the cloud. The cloud generates a script based on the screenshot and the currently executed subtask to obtain the user's confirmation or rejection.
[0075] Step 107: The cloud sends at least one prediction information to the terminal.
[0076] In some embodiments, the cloud sends the first prediction information of the at least one prediction information to the terminal, which may be sending the at least one prediction information to the terminal in sequence, so that the terminal executes the corresponding interactive content and / or interactive instructions in sequence.
[0077] In some embodiments, the first prediction information may be any one of the at least one prediction information.
[0078] In some embodiments, the cloud sends the first prediction information to the terminal, which may be displayed through a display interface of the terminal and / or output through an audio device, which is not limited in this disclosure.
[0079] In some embodiments, the cloud sends second prediction information to the terminal, the interaction content corresponding to the second prediction information includes the first missing information, and / or the interaction instruction corresponding to the second prediction information is selection or confirmation.
[0080] In some embodiments, after receiving the second prediction information, the terminal will output the interactive content and / or interactive instructions corresponding to the second prediction information through the display interface and / or audio device, so as to receive confirmation information or selection information input by the user for the interactive content and / or interactive instructions through the display interface and / or audio device.
[0081] In some embodiments, the cloud sends the second prediction information to the terminal, and the second prediction information is used to inquire the user about the first missing information to obtain the user's selection information or confirmation information.
[0082] For example, when a client searches for coffee A in a certain APP, since the search results include A1 coffee, A2 coffee, three cup types (large / medium / small), and options such as whether to add ice, the cloud will ask the user through voice or screen display which type of coffee to choose, which cup type to choose, and whether to add ice, etc.
[0083] In some embodiments, the cloud sends third prediction information to the terminal, the interaction content corresponding to the third prediction information includes a target operation, and the interaction instruction corresponding to the third prediction information is confirmation or rejection.
[0084] In some embodiments, after receiving the third prediction information, the terminal will output the interaction instruction corresponding to the third prediction information through the display interface and / or audio device, so as to receive the confirmation information or rejection information input by the user for the interaction instruction through the display interface and / or audio device.
[0085] In some embodiments, the cloud sends the third prediction information to the terminal, and the third prediction information is used to obtain the user's confirmation or rejection to complete the target operation or interrupt the target operation.
[0086] For example, when the client executes the subtask "payment", since it involves user confirmation, the cloud obtains the user's confirmation or rejection through the client's voice output or screen display, such as outputting voice or pop-up window.
[0087] In step 108 , the cloud determines the execution status of the current task.
[0088] In some embodiments, the cloud uses a prediction model to determine the execution status of the current task based on the instruction information and next image information corresponding to the first prediction information sent by the terminal until the current task is completed.
[0089] In some embodiments, the prediction model is also used to determine the execution status of the current task, and when the instruction information corresponding to the first prediction information and the next image information sent by the terminal are received, it is determined whether the current task is completed.
[0090] In some embodiments, after receiving the user's confirmation information, selection information, rejection information or other instructions, the terminal sends it to the cloud, so that the cloud generates the next prediction information or ends the current task based on the user's instructions.
[0091] In some embodiments, the terminal sends the first instruction information and first image information corresponding to the second prediction information to the cloud. Based on the first instruction information and the first image information, the cloud generates the next prediction information using a prediction model, sends the next prediction information to the terminal for execution, and determines whether the current task has been completed. The next prediction information combines the first instruction information with the first image information to obtain a corresponding operation that the terminal needs to perform. The prediction model can determine whether the current task has been completed based on the image information uploaded by the terminal after executing the prediction information.
[0092] For example, the user inputs audio through the client's audio device and selects A1 coffee, medium cup, and no ice, etc. The client uploads the audio to the cloud and uploads the current screenshot to the cloud. The cloud generates the next GUI operation based on the audio uploaded by the client and the current screenshot, that is, generates the GUI operation of selecting A1 coffee, medium cup, and no ice, until the subtask of selecting the beverage configuration currently being executed is completed.
[0093] In some embodiments, the terminal sends second instruction information and second image information corresponding to the third prediction information to the cloud. The cloud generates next prediction information based on the second instruction information and the second image information, sends the next prediction information to the terminal for execution, and determines whether the current task is completed. The next prediction information is the operation that the terminal needs to perform, obtained by combining the second instruction information and the second image information.
[0094] For example, the user inputs audio through the client's audio device to confirm payment, and the client uploads the audio to the cloud and the current screenshot to the cloud, so that the cloud can generate the next GUI operation based on the user's payment confirmation information and the current screenshot, that is, generate an operation to confirm payment. If the user has authorized the client to pay without a password, the client can directly complete the payment based on the operation to confirm payment; if the user has not authorized the client to pay without a password, the client can continue to take screenshots based on the operation to confirm payment and upload the current screenshot to the cloud. The cloud's prediction model can generate the next prediction information based on the screenshot to prompt the user to enter the payment password until the payment subtask is completed.
[0095] In some embodiments, the generated at least one task executes the above process in a loop until the prediction model determines that the execution status of each task is completed, and the user instruction is executed.
[0096] For example, for a user instruction of "operate a certain APP to order A coffee takeout", the user instruction is completed after the client completes the subtasks of opening the certain APP, searching for "A coffee" coffee drink, selecting the drink configuration, adding it to the shopping cart, and completing the payment.
[0097] In some embodiments, the cloud uses a prediction model to determine the execution status of the current task based on at least one prompt information in the current image information, where the execution status can be any of a completed state, an uncompleted state, and an inability to continue execution.
[0098] In some embodiments, when the prompt information indicates that the task has been completed, or the prompt information shows that the task execution result is consistent with the preset result, the prediction model determines that the task execution status is a completed status.
[0099] In some embodiments, when the prompt information is different from the task information of the current task, the prediction model outputs an unfinished state and can generate corresponding prediction information based on at least one missing information or target operation in the current image information.
[0100] In some embodiments, when the prompt information is a system pop-up window, a message pop-up window, or other information, the prediction model determines that the execution status of the current task is unable to continue.
[0101] In the above embodiment, the cloud analyzes user instructions and understands intent through a task planning model to split the user instructions into at least one task, and sends at least one task to the terminal for execution. When a display interface that meets preset conditions appears during the client execution process, the current image information sent by the receiving terminal is used to generate at least one prediction information using the prediction model to assist in the execution of the current task, thereby achieving the purpose of improving task completion rate and efficiency. The prediction model can also be used to determine the task execution status. In particular, the task planning module is trained based on a strong inference large model, which improves the rationality of task splitting and further improves task completion rate and efficiency.
[0102] Step 109 : The cloud determines exception handling and / or exception attributes.
[0103] In some embodiments, when the execution status of the current task is unable to continue, the cloud uses an exception handling model to identify the exception of the current image information, and determines the exception type and the exception handling and / or exception attributes corresponding to the exception type.
[0104] In some embodiments, based on the current image information, the prediction model can determine the execution status of the current task. When the current task cannot continue to be executed, it indicates that the prompt information is abnormal, and the abnormal prompt information is identified and processed by the exception handling model.
[0105] In some embodiments, the exception handling model is obtained by training an initial exception handling model using training data. The cloud can train the initial exception handling model to obtain the exception handling model, or another system or entity can train the initial exception handling model and provide it to the cloud for use, and this disclosure does not limit this.
[0106] In some embodiments, the prediction model determines the execution status of the current task and the results obtained include: completed status, uncompleted status, and unable to continue execution.
[0107] In some embodiments, the cloud uses an exception handling model to identify the exception of the current image information, and determines the exception type and the exception handling and / or exception attributes corresponding to the exception type. This can be to use the exception handling model to identify the first prompt information in the current image information, obtain the first exception type corresponding to the first prompt information, and output the corresponding first exception handling and / or first exception attribute according to the first exception type.
[0108] In some embodiments, exception types include: ordinary system pop-ups, system pop-ups that can be closed automatically, system pop-ups that cannot be closed automatically, message pop-ups, etc. The number of exception types identified by the exception handling model can be increased according to the training data set during the training process, and this disclosure is not limited to this.
[0109] In some embodiments, exception handling includes: closing, waiting, pausing execution, etc., and exception attributes include: preset duration, coordinates, user manual closing required, etc. There is a one-to-one correspondence between exception handling, exception attributes and exception types.
[0110] In some embodiments, the exception handling corresponding to the ordinary system pop-up window is waiting, and the exception attribute is the preset duration. That is, when an ordinary system pop-up window appears in the current image information, the cloud-based exception handling model outputs waiting + preset duration.
[0111] For example, a common system pop-up window may be a countdown advertisement. For a countdown advertisement, you can wait for a period of time; or it may be a splash screen page, such as the splash screen page that appears when you open a certain APP. Since the splash screen page is displayed for a specific length of time, you can wait for a period of time for the splash screen page to disappear automatically.
[0112] In some embodiments, the exception handling corresponding to the system pop-up window that can be closed automatically is close, and the exception attribute is the coordinates of the pop-up window closing position, that is, when a system pop-up window that can be closed automatically appears in the current image information, the exception handling model on the cloud side outputs close + coordinates; the exception handling corresponding to the system pop-up window that cannot be closed automatically is pause execution, and the exception attribute is a prompt that requires the user to manually close or manually input, that is, when a system pop-up window that cannot be closed automatically appears in the current image information, the exception handling model on the cloud side outputs pause execution + a prompt that requires the user to manually close or manually input.
[0113] For example, a system pop-up window that can be automatically closed is a system notification. For a system notification, it can be automatically closed, and the output is the coordinate position of the close button corresponding to the system notification pop-up window and the close operation.
[0114] For example, a system pop-up window that cannot be closed automatically is a system update pop-up window. The pop-up window requires the user to choose whether to update now or be reminded again after a period of time, and / or requires the user to manually enter a password to log in. The system output is to suspend execution, and a prompt requiring the user to manually close it or manually enter a password.
[0115] In some embodiments, the exception handling corresponding to the message pop-up window is closing or pausing execution, and the exception attribute is coordinates, that is, when a message pop-up window appears in the current image information, such as a new message in a social software, the exception handling model outputs pausing execution + the coordinates of the new message reply box, or outputs closing + the closing position coordinates of the new message pop-up window.
[0116] In some embodiments, the types of exceptions that can be identified by the exception handling model can be increased by increasing the data in the training data set during the training process, which is not limited by the present disclosure.
[0117] For example, the action prediction module not only completes the normal interaction path of the app, but also correctly handles abnormal pages such as randomly appearing advertisements, splash screens, message pop-ups, page pop-ups, and system notifications. The system trains a separate model for abnormal pages, identifies different types of abnormal pages, and performs predefined actions, such as a waiting period for countdown ads.
[0118] Step 110: The cloud sends exception handling and / or exception attributes to the terminal.
[0119] In some embodiments, the cloud sends the exception handling and / or exception attributes to the terminal, so that the terminal performs the exception handling / performs the exception handling based on the exception attributes.
[0120] In some embodiments, the exception handling corresponding to the ordinary system pop-up window is waiting, the exception attribute is duration, and the cloud outputs waiting + preset duration, then the client will execute the operation of waiting for the preset duration.
[0121] For example, a common system pop-up window may be a countdown advertisement. For the countdown advertisement, the client waits for a period of time; or it may be a splash screen page, such as a splash screen page that appears when a certain APP is opened. The client waits for a period of time for the splash screen page to disappear automatically.
[0122] In some embodiments, the exception handling corresponding to the system pop-up window that can be closed automatically is close, and the exception attribute is the coordinates of the pop-up window closing position. The exception handling model on the cloud side outputs close + coordinates, and the client executes to click the close button at the coordinate position; the exception handling corresponding to the system pop-up window that cannot be closed automatically is pause execution, and the exception attribute is a prompt that requires the user to manually close or manually input. The exception handling model on the cloud side outputs pause execution + a prompt that requires the user to manually close or manually input, and the client executes to pause the current task and outputs a prompt that requires the user to manually close or manually input through the audio device, so that the user can manually close the system pop-up window or manually input the corresponding information.
[0123] For example, a system pop-up window that can be automatically closed is a system notification. For a system notification, it can be automatically closed. The output is the coordinate position of the close button corresponding to the system notification pop-up window and the close operation. The client performs the automatic close operation according to the coordinate position.
[0124] For example, a system pop-up window that cannot be closed automatically is a system update pop-up window. The pop-up window requires the user to choose whether to update now or be reminded again after a period of time, and / or requires the user to manually enter a password to log in. The system output is to suspend execution, and a prompt requiring the user to manually close it or manually enter a password. The client manually closes it or manually enters the password according to the prompt.
[0125] In some embodiments, the exception handling corresponding to the message pop-up window is closing or pausing execution, and the exception attribute is the coordinates, that is, when a message pop-up window appears in the current image information, such as a new message in a social software, the exception handling model outputs pausing execution + the coordinates of the new message reply box, or outputs closing + the coordinates of the close button of the new message pop-up window. The client can suspend execution and move the input key to the coordinate position of the new message reply box, or click the close button of the new message pop-up window.
[0126] In the above embodiment, an exception handling model is used to identify the abnormal prompt information appearing in the current image information, thereby outputting the corresponding exception handling and / or abnormal attributes, so that the terminal can process it according to the exception handling and / or abnormal attributes, thereby avoiding the abnormal prompt information interfering with or hindering the execution of the current task, thereby improving the task completion rate and efficiency.
[0127] In the above embodiment, step 109 and step 110 are optional steps, that is, when the execution status of the current task is unable to continue, the cloud can use the exception handling model to handle the abnormal situation accordingly to improve the task completion rate and efficiency, thereby achieving the purpose of assisting task execution.
[0128] In the above embodiment, user instructions can be split through the task planning model, and multiple tasks can be executed in sequence by the terminal. During the execution process, if there is missing information or prompt information, the relevant interaction process is generated through the prediction model, and the task execution status is determined through the prediction model to assist task execution, thereby improving the task completion rate and efficiency, and further improving the system security and privacy protection.
[0129] Figure 2 This is a flow chart of a task execution method proposed in an embodiment of the present disclosure. The method is executed by a terminal, wherein the prediction model, task planning model, and exception handling model are deployed on the terminal. Figure 1 As shown, the method includes the following steps:
[0130] Step 201: Generate at least one prediction information through a prediction model based on current image information and current task.
[0131] In some embodiments, the current image information is image information corresponding to the current display interface obtained during the execution of the current task, and the at least one prediction information includes at least one of interaction content, interaction attributes, and interaction instructions.
[0132] In some embodiments, while executing a current task, the terminal may, under certain conditions, take a screenshot or photo of the current display interface to obtain current image information, and then use a prediction model to identify and analyze the current image information to generate at least one piece of prediction information. The specific condition may be that the current display interface requires human-computer interaction.
[0133] In some embodiments, at least one prediction information is used to generate content or instructions required for interaction based on the current image information when human-computer interaction is required in the current image information.
[0134] In some embodiments, at least one prediction information is used to assist the terminal in completing a current task.
[0135] In some embodiments, the current image information is determined by the terminal, specifically including: generating at least one task through a task planning model according to user instructions, the task planning model is obtained by training an initial task planning model using a task planning data set, and at least one task includes task information and a task number; the terminal executes the corresponding task according to the task number; when the display interface during the execution of the current task meets the preset conditions, the current image information is determined, and the current image information includes at least one missing information or prompt information.
[0136] In some embodiments, the user instruction may be text information input by the user through the display interface of the terminal, or it may be audio information input by the user through the audio device of the terminal, and the terminal transcribes the audio information into text information to obtain the user instruction; or it may be transcribing the audio information into text information and then integrating the two parts of text information to obtain the user instruction. The method of obtaining the user instruction is not limited by this disclosure.
[0137] For example, in response to a specific request issued by a user, such as "Help me order a cup of coffee A using the XXX APP", the client records and transcribes the voice into text "Help me order a cup of coffee A using the XXX APP".
[0138] In some embodiments, the task planning model may be provided to the terminal by other devices, or may be obtained after the terminal performs model training.
[0139] In some embodiments, the task planning model is obtained by training an initial task planning model using a task planning data set, wherein the initial task planning model can be a traditional machine learning model, such as a linear model, a tree model, a Bayesian model, etc., or a lightweight neural network model, or a large model with strong inference, which is not limited in this disclosure. In the present disclosure, the initial task planning model used by the task planning model is a large model with strong inference, so that the splitting of at least one task obtained can be more reasonable and a higher completion rate can be achieved.
[0140] In some embodiments, a task planning model is used to generate at least one task according to a user instruction, which may be to split the user instruction into multiple ordered subtasks, each subtask corresponding to an operation.
[0141] For example, the task planning module is responsible for global planning, breaking down the task into multiple subtasks to be executed sequentially. This decomposition ensures that the task is progressing in the right direction at a macro level, while also making the subtasks easier to complete, resulting in a higher completion rate. The task planning module is implemented based on a large model with strong reasoning capabilities.
[0142] For example, the task planning module of the terminal breaks down the user's instructions into multiple subtasks: (1) open a certain APP; (2) search for "A Coffee" coffee drink; (3) select the drink configuration; (4) add to the shopping cart; (5) complete the payment.
[0143] In some embodiments, the terminal executes the corresponding task according to the task sequence number, which may be performed by the terminal in sequence according to the order of the multiple subtasks obtained by splitting.
[0144] In some embodiments, the preset conditions include at least one of the following: there is at least one missing information or prompt information that makes it impossible to continue the current task; there is a target operation, and the target operation is an operation that requires the user to confirm through the terminal.
[0145] In some embodiments, when the terminal is executing at least one task in sequence, that is, when the terminal is executing the current task, when the display interface meets the preset conditions, that is, the current image information corresponding to the current display interface is determined, at least one prediction information is generated through the prediction model.
[0146] In some embodiments, the display interface meets the preset conditions when there is at least one missing information or prompt information in the display interface that makes it impossible to continue the current task, or there is a target operation in the display interface, which is an operation that requires user confirmation through the terminal.
[0147] In some embodiments, the missing information may be necessary information for executing the current task, for example, multiple options that are consistent with the execution of the current task appear in the display interface, or options that are inconsistent with the execution of the current task appear, resulting in the lack of necessary confirmation information for executing the current task. The user is required to fill in the relevant confirmation information before continuing to execute the current task.
[0148] In some embodiments, the target operation may be an operation involving user privacy or security, such as a payment operation, a message sending operation, a red envelope sending operation, etc.
[0149] In some embodiments, the prompt information may be a pop-up window of the abnormal page, such as a system message window or an advertising window; or, the prompt information may be a window corresponding to the target operation, such as a payment operation window or a message sending window or a red envelope sending window; or, the prompt information may be a prompt that the current task is completed.
[0150] In some embodiments, the terminal may determine the current image information by performing operations such as screenshot and screen recording.
[0151] For example, when the client performs the subtask of selecting a beverage configuration, it does not clearly indicate the required cup type (large / medium / small), whether to add ice, and other information, which are crucial for task completion. When there are multiple target drinks (A1 coffee, A2 coffee) in the search results, the user needs to confirm to avoid placing the wrong order.
[0152] For example, when the client is executing a subtask, some important operations, such as payment, sending messages, sending red envelopes, etc., need to be confirmed by the user. To this end, a user confirmation link is added to the interaction process, and the client obtains the current screenshot.
[0153] In some embodiments, based on the current image information and the current task, at least one prediction information is generated through a prediction model, including at least one of the following: through the prediction model, second prediction information is generated according to the first missing information in the at least one missing information and the task information corresponding to the current task, and the interaction content corresponding to the second prediction information includes the first missing information, and / or the interaction instruction corresponding to the second prediction information is selection or confirmation; through the prediction model, third prediction information is generated according to the target operation in the current image information and the task information corresponding to the current task, the interaction content corresponding to the third prediction information includes the target operation, and the interaction instruction corresponding to the third prediction information is confirmation or rejection.
[0154] In some embodiments, the second prediction information and the third prediction information are one or more of the at least one prediction information, the second prediction information is generated when the first missing information is present in the current image information, and the third prediction information is generated when the target operation is present in the current image information. The first missing information is any one of the at least one missing information, and the target operation corresponds to the prompt information in the current image information.
[0155] In some embodiments, the prediction model generates second prediction information based on the first missing information and task information corresponding to the current task. The second prediction information is used to inquire the user about the first missing information to obtain the user's selection information or confirmation information.
[0156] For example, when a client searches for coffee A in a certain APP, since the search results include A1 coffee, A2 coffee, three cup types (large / medium / small), and options such as whether to add ice. The client obtains a screenshot and generates a script based on the screenshot and the currently executed subtask to ask the user which type of coffee to choose, which cup type to choose, and whether to add ice, etc.
[0157] In some embodiments, the prediction model generates third prediction information based on the target operation and task information corresponding to the current task. The third prediction information is used to obtain the user's confirmation or rejection to complete the target operation or interrupt the target operation.
[0158] For example, when the client executes the subtask "payment", since it involves the need for user confirmation, the client obtains a screenshot and generates a script based on the screenshot and the currently executed subtask to obtain the user's confirmation or rejection.
[0159] In the above-described embodiment, the terminal analyzes user instructions and understands their intent through a task planning model, splitting the user instructions into at least one task and sending at least one task to the terminal for execution. When a display interface that meets preset conditions appears during client execution, the terminal determines the current image information and uses the prediction model to generate at least one prediction information to assist in the execution of the current task, thereby improving task completion rate and efficiency. In particular, the task planning module is trained based on a large, strong inference model, which improves the rationality of task splitting and further enhances task completion rate and efficiency.
[0160] Step 202 : Based on the instruction information and next image information corresponding to the first prediction information in the at least one prediction information, use the prediction model to determine the execution status of the current task until the current task is in a completed state.
[0161] In some embodiments, the terminal executes corresponding interactive content and / or interactive instructions in sequence based on at least one prediction information.
[0162] In some embodiments, the first prediction information may be any one of the at least one prediction information.
[0163] In some embodiments, the first prediction information may be displayed through a display interface of the terminal and / or output through an audio device, which is not limited in this disclosure.
[0164] In some embodiments, the second prediction information is used to inquire the user about the first missing information to obtain the user's selection information or confirmation information.
[0165] For example, when a client searches for coffee A in a certain APP, since the search results include A1 coffee, A2 coffee, three cup types (large / medium / small), and options such as whether to add ice, the user is asked through voice or screen display which type of coffee to choose, which cup type to choose, and whether to add ice, etc.
[0166] In some embodiments, the third prediction information is used to obtain the user's confirmation or rejection to complete the target operation or interrupt the target operation.
[0167] For example, when the client executes the subtask "payment", since it involves the need for user confirmation, the client obtains the user's confirmation or rejection by outputting voice or displaying on the screen, such as outputting voice or a pop-up window.
[0168] In some embodiments, the prediction model is also used to determine the execution status of the current task, and to determine whether the current task has been completed based on the instruction information corresponding to the first prediction information and the next image information.
[0169] In some embodiments, the terminal may generate next prediction information using a prediction model based on the first instruction information and the first image information corresponding to the second prediction information, and determine whether the current task has been completed. The next prediction information combines the first instruction information with the first image information to obtain a corresponding operation that the terminal needs to perform. The prediction model may determine whether the current task has been completed based on the image information determined by the terminal after executing the prediction information.
[0170] For example, the user inputs audio through the client's audio device and selects A1 coffee, medium cup, and no ice, etc. The client generates the next GUI operation based on the audio and the current screenshot, that is, generates the GUI operation of selecting A1 coffee, medium cup, and no ice, until the currently executed subtask of selecting the beverage configuration is completed.
[0171] In some embodiments, the terminal generates next prediction information based on the second instruction information and the second image information corresponding to the third prediction information, and determines whether the current task has been completed. The next prediction information is an operation that the terminal needs to perform obtained by combining the second instruction information and the second image information.
[0172] For example, the user inputs audio through the client's audio device to confirm payment. The client generates the next GUI operation based on the user's audio information confirming payment and the current screenshot, that is, generates an operation to confirm payment. If the user has authorized the client to pay without a password, the client can directly complete the payment based on the operation to confirm payment; if the user has not authorized the client to pay without a password, the client can continue to take screenshots based on the operation to confirm payment and the prediction model can generate the next prediction information based on the screenshot to prompt the user to enter the payment password until the payment subtask is completed.
[0173] In some embodiments, the generated at least one task executes the above process in a loop until the prediction model determines that the execution status of each task is completed, and the user instruction is executed.
[0174] For example, for a user instruction of "operate a certain APP to order A coffee takeout", the user instruction is completed after the client completes the subtasks of opening the certain APP, searching for "A coffee" coffee drink, selecting the drink configuration, adding it to the shopping cart, and completing the payment.
[0175] In the above embodiment, the terminal can split user instructions through the task planning model and execute multiple tasks in sequence. During the execution process, the prediction model is used to generate related interaction processes and judge the execution status of the current task to improve the task completion rate and efficiency, and further improve the system security and privacy protection.
[0176] Figure 3 This is a flowchart of the exception handling process proposed in the embodiment of the present disclosure. Figure 3 based on Figure 2 The embodiment shown, as Figure 3 As shown, the method further includes the following steps:
[0177] Step 301: Determine the execution status of the current task using a prediction model based on at least one prompt information in the current image information.
[0178] In some embodiments, the prediction model can determine the execution status of the current task based on the current image information.
[0179] In some embodiments, the execution status includes: a completed status, an uncompleted status, and a status in which execution cannot continue.
[0180] In some embodiments, based on at least one prompt information in the current image information, when the prompt information shows completion or is the same as the task information of the current task, the prediction model outputs a completion status.
[0181] In some embodiments, based on at least one prompt information in the current image information, when the prompt information is different from the task information of the current task, the prediction model outputs an unfinished state and can generate corresponding prediction information based on at least one missing information or target operation in the current image information.
[0182] In some embodiments, based on at least one prompt information in the current image information, when the prompt information contains abnormal or irrelevant information compared to the task information of the current task, the prediction model outputs a state in which execution cannot continue.
[0183] Step 302: When the execution status of the current task is that the task cannot be continued, the current image information is subjected to abnormality recognition by using an abnormality handling model to determine the abnormality type and the abnormality handling and / or abnormality attributes corresponding to the abnormality type.
[0184] In some embodiments, exception handling and / or exception attributes are determined for the terminal to perform exception handling, or the terminal performs exception handling based on the exception attributes.
[0185] In some embodiments, based on the current image information, the prediction model can determine the execution status of the current task. When the current task cannot continue to be executed, it indicates that the prompt information is abnormal, and the abnormal prompt information is identified and processed by the exception handling model.
[0186] In some embodiments, the exception handling model is obtained by training an initial exception handling model using training data. The terminal can train the initial exception handling model to obtain the exception handling model, or another system or entity can train the initial exception handling model and provide it to the terminal for use, which is not limited by this disclosure.
[0187] In some embodiments, the prediction model determines the execution status of the current task and the results obtained include: completed status, uncompleted status, and unable to continue execution.
[0188] In some embodiments, the terminal uses an exception handling model to identify the exception of the current image information, and determines the exception type and the exception handling and / or exception attributes corresponding to the exception type. The terminal can use the exception handling model to identify the first prompt information in the current image information, obtain the first exception type corresponding to the first prompt information, and output the corresponding first exception handling and / or first exception attribute according to the first exception type.
[0189] In some embodiments, exception types include: ordinary system pop-ups, system pop-ups that can be closed automatically, system pop-ups that cannot be closed automatically, message pop-ups, etc. The number of exception types identified by the exception handling model can be increased according to the training data set during the training process, and this disclosure is not limited to this.
[0190] In some embodiments, exception handling includes: closing, waiting, pausing execution, etc., and exception attributes include: preset duration, coordinates, user manual closing required, etc. There is a one-to-one correspondence between exception handling, exception attributes and exception types.
[0191] In some embodiments, the exception handling corresponding to the ordinary system pop-up window is waiting, and the exception attribute is the preset duration. That is, when an ordinary system pop-up window appears in the current image information, the cloud-based exception handling model outputs waiting + preset duration.
[0192] For example, a common system pop-up window may be a countdown advertisement. For a countdown advertisement, you can wait for a period of time; or it may be a splash screen page, such as the splash screen page that appears when you open a certain APP. Since the splash screen page is displayed for a specific length of time, you can wait for a period of time for the splash screen page to disappear automatically.
[0193] In some embodiments, the exception handling corresponding to the system pop-up window that can be closed automatically is close, and the exception attribute is the coordinates of the pop-up window closing position, that is, when a system pop-up window that can be closed automatically appears in the current image information, the exception handling model outputs close + coordinates; the exception handling corresponding to the system pop-up window that cannot be closed automatically is pause execution, and the exception attribute is a prompt that requires the user to manually close or manually input, that is, when a system pop-up window that cannot be closed automatically appears in the current image information, the exception handling model outputs pause execution + a prompt that requires the user to manually close or manually input.
[0194] For example, a system pop-up window that can be automatically closed is a system notification. For a system notification, it can be automatically closed, and the output is the coordinate position of the close button corresponding to the system notification pop-up window and the close operation.
[0195] For example, a system pop-up window that cannot be closed automatically is a system update pop-up window. The pop-up window requires the user to choose whether to update now or be reminded again after a period of time, and / or requires the user to manually enter a password to log in. The system output is to suspend execution, and a prompt requiring the user to manually close it or manually enter a password.
[0196] In some embodiments, the exception handling corresponding to the message pop-up window is closing or pausing execution, and the exception attribute is coordinates, that is, when a message pop-up window appears in the current image information, such as a new message in a social software, the exception handling model outputs pausing execution + the coordinates of the new message reply box, or outputs closing + the closing position coordinates of the new message pop-up window.
[0197] In some embodiments, the types of exceptions that can be identified by the exception handling model can be increased by increasing the data in the training data set during the training process, which is not limited by the present disclosure.
[0198] For example, the action prediction module not only completes the normal interaction path of the app, but also correctly handles abnormal pages such as randomly appearing advertisements, splash screens, message pop-ups, page pop-ups, and system notifications. The system trains a separate model for abnormal pages, identifies different types of abnormal pages, and performs predefined actions, such as a waiting period for countdown ads.
[0199] In some embodiments, the exception handling corresponding to the ordinary system pop-up window is waiting, the exception attribute is duration, and the exception handling model outputs waiting + preset duration, then the operation of waiting for the preset duration will be executed.
[0200] For example, a common system pop-up window may be a countdown advertisement. For the countdown advertisement, the client waits for a period of time; or it may be a splash screen page, such as a splash screen page that appears when a certain APP is opened. The client waits for a period of time for the splash screen page to disappear automatically.
[0201] In some embodiments, the exception handling corresponding to the system pop-up window that can be closed automatically is close, and the exception attribute is the coordinates of the pop-up window closing position. The exception handling model outputs close + coordinates, and the client executes to click the close button at the coordinate position; the exception handling corresponding to the system pop-up window that cannot be closed automatically is pause execution, and the exception attribute is a prompt that the user needs to manually close or manually input. The exception handling model outputs pause execution + a prompt that the user needs to manually close or manually input, and the client executes to pause the current task and outputs a prompt that the user needs to manually close or manually input through the audio device, so that the user can manually close the system pop-up window or manually input the corresponding information.
[0202] For example, a system pop-up window that can be automatically closed is a system notification. For a system notification, it can be automatically closed. The output is the coordinate position of the close button corresponding to the system notification pop-up window and the close operation. The client performs the automatic close operation according to the coordinate position.
[0203] For example, a system pop-up window that cannot be closed automatically is a system update pop-up window. The pop-up window requires the user to choose whether to update now or be reminded again after a period of time, and / or requires the user to manually enter a password to log in. The system output is to suspend execution, and a prompt requiring the user to manually close it or manually enter a password. The client manually closes it or manually enters the password according to the prompt.
[0204] In some embodiments, the exception handling corresponding to the message pop-up window is closing or pausing execution, and the exception attribute is the coordinates, that is, when a message pop-up window appears in the current image information, such as a new message in a social software, the exception handling model outputs pausing execution + the coordinates of the new message reply box, or outputs closing + the coordinates of the close button of the new message pop-up window. The client can suspend execution and move the input key to the coordinate position of the new message reply box, or click the close button of the new message pop-up window.
[0205] In the above embodiment, the terminal uses an exception handling model to identify the abnormal prompt information appearing in the current image information, and outputs the corresponding exception handling and / or abnormal attributes to process according to the exception handling and / or abnormal attributes, so as to avoid the abnormal prompt information interfering with or hindering the execution of the current task, thereby improving the task completion rate and efficiency.
[0206] Figure 4 This is a flow chart of the prediction model proposed in the embodiment of the present disclosure. Figure 4 based on Figure 1 or Figure 2 The embodiment shown, as Figure 4 As shown, the following steps are included:
[0207] Step 401: Obtain a first training data set.
[0208] In some embodiments, the first training data set includes a plurality of image information and at least one of interaction content, interaction attributes, and interaction instructions corresponding to each image information.
[0209] In some embodiments, the image information may include at least one piece of missing information, and the interaction content and / or interaction attributes and / or interaction instructions correspond one-to-one to the at least one piece of missing information.
[0210] In some embodiments, the image information may be the absence of missing information or the presence of prompt information, and the interaction attribute may be the execution status of the current task, such as completion status, unfinished status, status that cannot continue execution, etc.
[0211] In some embodiments, the first training data set may be obtained by integrating historical interaction data, or may be obtained from other public databases, or may be provided by a specific application party, which is not limited by the present disclosure.
[0212] In some embodiments, the first training data set can be modified according to different scenarios or actual needs, such as adding training data for the current scenario, or deleting training data that is not suitable for the current scenario to avoid occupying computing power of the training process.
[0213] Step 402: Use the first training data set to train the initial prediction model to obtain a prediction model.
[0214] In some embodiments, the initial prediction model can be a traditional machine learning model, such as a linear model, a tree model, a Bayesian model, etc., or a lightweight neural network model, or a large model with strong inference, which is not limited by the present disclosure.
[0215] In some embodiments, training the initial prediction model using the first training dataset may include data preprocessing, model training, model optimization, model evaluation, and other processes to obtain a prediction model. The present disclosure does not limit the specific model training process, and the training process may be adjusted and optimized accordingly based on different initial prediction models.
[0216] In some embodiments, the prediction model can be used to generate multiple rounds of prediction information and to determine the execution status of the current task.
[0217] In the above embodiment, the initial prediction model can be trained based on the acquired first training data set to obtain a prediction model applied to the task execution method proposed in the present disclosure. The prediction model can generate multiple rounds of interactions based on missing information during the task execution process and can judge the execution status of the current task to improve the task completion rate and efficiency.
[0218] Figure 5 This is a flow chart of the exception handling model proposed in the embodiment of the present disclosure. Figure 5 based on Figure 1 、 Figure 2 、 Figure 3 The embodiment shown, as Figure 5 As shown, the following steps are included:
[0219] Step 501: Obtain a second training data set.
[0220] In some embodiments, the second training dataset includes a plurality of image information, an abnormality type corresponding to each image information, and an abnormality treatment and / or abnormality attribute corresponding to each abnormality type.
[0221] In some embodiments, the image information includes prompt information, which may be a pop-up window of an abnormal page, such as a system message window or an advertisement window.
[0222] In some embodiments, the prompt information corresponds one-to-one with the exception type, exception handling and / or exception attributes.
[0223] In some embodiments, the second training data set may be obtained by integrating historical abnormal data, or may be obtained from other public databases, or may be provided by a specific application party, which is not limited by the present disclosure.
[0224] In some embodiments, the second training data set can be modified according to different scenarios or actual needs, such as adding training data for the current scenario, or deleting training data that is not suitable for the current scenario to avoid occupying computing power of the training process.
[0225] Step 502: Use the second training data set to train the initial exception handling model to obtain the exception handling model.
[0226] In some embodiments, the initial exception handling model can be a traditional machine learning model, such as a linear model, a tree model, a Bayesian model, etc., or a lightweight neural network model, or a large model with strong inference, which is not limited by the present disclosure.
[0227] In some embodiments, training the initial exception handling model using the second training dataset may include data preprocessing, model training, model optimization, model evaluation, and other processes to obtain the exception handling model. The present disclosure does not limit the specific model training process; the training process may be adjusted and optimized accordingly based on different initial exception handling models.
[0228] In some embodiments, the exception handling model is used to identify prompt information appearing in the current image information, determine the corresponding exception type, and output the exception type and the exception handling and / or exception attributes corresponding to the exception type.
[0229] In the above embodiment, by obtaining a second training data set and using the second training data set to train an exception handling model, the exception handling model is applied to the task execution method proposed in the present disclosure to realize the identification of abnormal pages during the task execution process, so as to avoid abnormal pages interfering with task execution or causing execution errors, thereby improving the task completion rate and efficiency.
[0230] In summary, the task execution method proposed in the present disclosure can process user instructions by splitting tasks, generating interactive processes, and identifying abnormal pages based on task planning models, prediction models, and exception handling models, thereby avoiding execution interruptions or exceptions and improving task completion rate and efficiency.
[0231] The following is a specific implementation of a task execution method provided by the present disclosure:
[0232] Figure 6 This is the system framework diagram of the solution, such as Figure 6As shown, the solution continues to use the voice human-computer interaction framework, including modules such as speech recognition, intent understanding and speech synthesis, and adds task planning, visual modal perception, action prediction and action execution.
[0233] The system framework can be roughly divided into two parts: client app (application) and cloud service (service).
[0234] The client app's responsibilities include: recording user audio, capturing screen images, uploading audio streams and images to the cloud, and executing commands issued by the cloud.
[0235] The responsibilities of the cloud service include: speech recognition, intent understanding, task planning, and action prediction.
[0236] Module Description:
[0237] 1. Mission Planning:
[0238] The task planning module is responsible for global planning, breaking down tasks into multiple subtasks to be executed sequentially. This decomposition ensures that the task progresses in the right direction at a macro level, while also making subtasks easier to complete, resulting in a higher completion rate. The task planning module is implemented based on a large, strongly inferential model. Its input is user instructions, and its output is a sequence of subtasks.
[0239] 2. Action Prediction
[0240] The input for action prediction is the subtask description, the current screenshot, and the historical actions. The output is the next action to be performed. The prediction is based on a supervised fine-tuned vision language model (VLM).
[0241] 2.1 Action Prediction-Abnormal Page Processing
[0242] In addition to completing normal app interaction paths, the action prediction module also correctly handles anomalous pages such as randomly appearing ads, splash screens, message pop-ups, page pop-ups, and system notifications. The system trains a separate model for anomalous pages, identifying different types of anomalous pages and executing predefined actions. For example, for countdown ads, a waiting period is executed.
[0243] 2.2 Action Prediction-Multi-round Interaction
[0244] How to continue the task when information is missing during the interaction. For example, in the example above, "Please use the XXX app to order a cup of coffee A," it is not clear what type of cup (large, medium, or small) is required, whether it should be iced, etc., which are crucial for completing the task. If the search results include multiple target drinks (such as "A1 coffee" and "A2 coffee"), the user needs to confirm to avoid placing the wrong order.
[0245] In this solution, we've added a multi-round interaction design for these situations. Based on the user's instructions and the beverage information currently displayed on the screen, the predictive model generates a response to the user, asking them to clarify or confirm the information.
[0246] 3. Action execution:
[0247] During the execution phase, some important operations, such as payment, sending messages, and sending red envelopes, require user confirmation. To address this, a user confirmation step is added to the interaction. The prediction model identifies important operations, outputs actions requiring user confirmation, and sends them to the client.
[0248] Detailed process introduction:
[0249] Specifically, it is a specific instruction issued to the user, such as "help me order a cup of coffee A using the XXX APP."
[0250] 1. Voice recognition: The client records and uploads the audio of the command, and the cloud transcribes the voice into text: "Help me order a cup of coffee A using the XXX app."
[0251] 2. Intent understanding: The cloud-based intent understanding module recognizes that the user’s intention is to “use a certain app to order coffee takeout.”
[0252] 3. Task Planning: The cloud-based task planning module breaks down the user's instructions into multiple basic subtasks:
[0253] (1) Open a certain app;
[0254] (2) Search for coffee drinks “A coffee”;
[0255] (3) Select beverage configuration;
[0256] (4) Add to cart;
[0257] (5)Complete payment.
[0258] 4. Action Prediction and Execution: The client uploads a screenshot; the cloud interprets the image, determines the next action, and sends instructions to the client. The client performs the corresponding GUI operation based on the instructions. If the task is not completed, the client continues to upload screenshots. This process repeats until each subtask is completed. During execution, the client interacts with the user to confirm key information.
[0259] The following is a detailed description of the execution process
[0260] (1) Open a certain app;
[0261] (2) Search for coffee drinks “A coffee”;
[0262] (3) Select the drink configuration. Ask the user to clarify any key configurations that are unclear, such as cup size (large or medium), temperature (with ice or not), etc.
[0263] (4) Add to cart;
[0264] (5) Complete payment: If the app supports password-free payment, user confirmation is required before payment.
[0265] In summary, the beneficial effects of the above scheme are as follows:
[0266] (1) Using a strong inference model to train the model of the task planning module and splitting complex tasks into multiple subtasks can improve the task completion rate;
[0267] (2) We have processed abnormal pages, which can effectively avoid the interference caused by advertisements, page pop-ups, and message pop-ups in actual usage scenarios;
[0268] (3) Add multiple rounds of interactions to confirm necessary information.
[0269] Figure 7 FIG. 7 is a structural diagram of a task execution device 700 according to an embodiment of the present disclosure. Figure 7 As shown, the device includes: a prediction module 701 and a determination module 702.
[0270] The prediction module 701 is used to generate at least one prediction information through a prediction model based on the current image information and the current task. The current image information is the image information corresponding to the current display interface obtained during the execution of the current task. The at least one prediction information includes at least one of the interactive content, interactive attributes, and interactive instructions.
[0271] The determination module 702 is configured to determine the execution status of the current task using a prediction model based on the instruction information and next image information corresponding to the first prediction information in the at least one prediction information until the current task is in a completed state.
[0272] In some embodiments, the system also includes a planning module, which is used to generate at least one task based on user instructions through a task planning model. The task planning module is obtained by training an initial task planning model using a task planning data set. At least one task includes task information and a task sequence number. At least one task is used for the terminal to perform the corresponding task according to the task sequence number; determine the current image information. The current image information is obtained when the display interface of the terminal meets the preset conditions during the execution of the current task. The current image information includes at least one missing information or prompt information.
[0273] In some embodiments, the preset conditions include at least one of the following: there is at least one missing information or prompt information that makes it impossible to continue the current task; there is a target operation, and the target operation is an operation that requires the user to confirm through the terminal.
[0274] In some embodiments, the prediction module is used to generate second prediction information through a prediction model based on first missing information in at least one missing information and task information corresponding to the current task, where the interaction content corresponding to the second prediction information includes the first missing information, and / or the interaction instruction corresponding to the second prediction information is selection or confirmation; and generate third prediction information through a prediction model based on the target operation in the current image information and task information corresponding to the current task, where the interaction content corresponding to the third prediction information includes the target operation, and the interaction instruction corresponding to the third prediction information is confirmation or rejection.
[0275] In some embodiments, the prediction module is also used to determine the execution status of the current task using a prediction model based on at least one prompt information in the current image information; when the execution status of the current task is unable to continue, the current image information is identified as abnormal through the exception handling model, and the exception type and the exception handling and / or exception attributes corresponding to the exception type are determined, so that the terminal can perform exception handling, or the terminal performs exception handling based on the exception attributes.
[0276] In some embodiments, the prediction module is also used to obtain a first training data set, which includes multiple image information and at least one of the interactive content, interactive attributes, and interactive instructions corresponding to each image information; the initial prediction model is trained using the first training data set to obtain a prediction model.
[0277] In some embodiments, the prediction module is further used to obtain a second training data set, the second training data set including multiple image information and the exception type corresponding to each image information, and the exception handling and / or exception attributes corresponding to each exception type; the initial exception handling model is trained using the second training data set to obtain the exception handling model. In summary, the task execution method and device proposed in the present disclosure, by utilizing artificial intelligence / AI models, identify exception pages or task interruption pages and determine corresponding processing operations, as well as identify user instructions and understand intent, thereby generating multiple rounds of dialogue to achieve the purpose of assisting task completion, thereby improving task completion rate and efficiency.
[0278] Figure 8 is a structural diagram of an electronic device 800 for implementing the above-mentioned task execution method according to an exemplary embodiment.
[0279] Reference Figure 8 , the electronic device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .
[0280] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 802 may include one or more modules to facilitate interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate interaction between the multimedia component 808 and the processing component 802.
[0281] The memory 804 is configured to store various types of data to support operations on the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0282] The power supply component 806 provides power to the various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 800.
[0283] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.
[0284] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.
[0285] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.
[0286] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the electronic device 800. For example, the sensor assembly 814 can detect the open / closed state of the electronic device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also detect changes in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and temperature changes of the electronic device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0287] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, 4G LTE, 5G NR (NewRadio) or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0288] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.
[0289] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the instructions can be executed by the processor 820 of the electronic device 800 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0290] The embodiments of the present disclosure further provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the task execution method described in the above embodiments of the present disclosure.
[0291] The embodiments of the present disclosure further provide a computer program product, including a computer program, which executes the task execution method described in the above embodiments of the present disclosure when a processor executes the computer program.
[0292] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices, systems, and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.
[0293] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "illustrative embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with an embodiment or example is included in at least one embodiment or example of the present disclosure. In this specification, the illustrative use of the above terms does not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0294] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present disclosure includes additional implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present disclosure belong.
[0295] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processing module, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection having one or more wires (control method), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or otherwise processing it in a suitable manner if necessary, and then storing it in a computer memory.
[0296] It should be understood that the various parts of the embodiments of the present disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0297] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0298] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disk, etc.
[0299] Although the embodiments of the present disclosure have been shown and described above, it is understood that the above embodiments are exemplary and are not to be construed as limitations on the present disclosure. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present disclosure.
Claims
1. A task execution method, characterized in that: The method comprises: Based on the current image information and the current task, generating at least one piece of prediction information through a prediction model, wherein the current image information is image information corresponding to the current display interface obtained during the execution of the current task, and the at least one piece of prediction information includes at least one of interaction content, interaction attributes, and interaction instructions; Based on the instruction information and next image information corresponding to the first prediction information in the at least one prediction information, the prediction model is used to determine the execution status of the current task until the current task is in a completed state.
2. The method according to claim 1, characterized in that The method further comprises: Based on a user instruction, generating at least one task through a task planning model, wherein the task planning model is obtained by training an initial task planning model using a task planning dataset, the at least one task including task information and a task sequence number, and the at least one task is used by the terminal to perform a corresponding task according to the task sequence number; The current image information is determined, where the current image information is obtained when a display interface of the terminal during execution of the current task meets a preset condition, and the current image information includes at least one missing information or prompt information.
3. The method according to claim 2, characterized in that The preset conditions include at least one of the following: The existence of the at least one missing information or the prompt information makes it impossible to continue to perform the current task; There is a target operation, which is an operation that requires confirmation by the user through the terminal.
4. The method according to claim 3, characterized in that The generating of at least one prediction information by a prediction model based on the current image information and the current task includes at least one of the following: generating, by the prediction model, second prediction information based on first missing information in the at least one missing information and task information corresponding to the current task, wherein the interaction content corresponding to the second prediction information includes the first missing information, and / or the interaction instruction corresponding to the second prediction information is selection or confirmation; Through the prediction model, third prediction information is generated according to the target operation in the current image information and the task information corresponding to the current task. The interaction content corresponding to the third prediction information includes the target operation, and the interaction instruction corresponding to the third prediction information is confirmation or rejection.
5. The method according to claim 3, characterized in that The method further comprises: determining, based on at least one prompt information in the current image information, an execution state of the current task using the prediction model; When the execution status of the current task is that it cannot continue to be executed, the current image information is identified as abnormal through the exception handling model, and the exception type and the exception handling and / or exception attributes corresponding to the exception type are determined, so that the terminal can perform the exception handling, or the terminal can perform the exception handling based on the exception attributes.
6. The method according to claim 1, characterized in that The method further comprises: Acquire a first training data set, where the first training data set includes a plurality of image information and at least one of interaction content, interaction attributes, and interaction instructions corresponding to each image information; The initial prediction model is trained using the first training data set to obtain the prediction model.
7. The method according to claim 5, characterized in that The method further comprises: Acquire a second training data set, the second training data set including a plurality of image information, an abnormality type corresponding to each image information, and an abnormality treatment and / or abnormality attribute corresponding to each abnormality type; The initial exception handling model is trained using the second training data set to obtain the exception handling model.
8. A task execution device, characterized in that: The task execution device is configured to execute the method according to any one of claims 1 to 7.
9. An electronic device, characterized in that: include: A processor and a memory for storing a computer program that can be run on the processor, wherein the processor performs the method according to any one of claims 1 to 7 when running the computer program.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.
11. A program product, characterized in that The method comprises computer instructions for causing a computer to execute the method according to any one of claims 1 to 7.