Event execution method and device, electronic device, and storage medium

CN122816508APending Publication Date: 2026-09-25BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510353809.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]然而,相关技术在通过人工智能实现自动化操作时,常出现下一步操作预测不准确的问题,自动化效果较差

Benefits of technology

[0024]在本公开的技术方案中,在迭代执行目标事件的过程中,是基于上一步操作的操作描述和操作结果中的至少一者、上一步操作的触发位置、以及用户指令进行下一步操作的操作预测,进而得到下一步操作的触发位置。其中,在基于预测的触发位置触发当前所处界面的基础上,还会将该下一步操作更新为下一次迭代时的上一步操作,并记录更新后的上一步操作的触发位置,以及生成该上一步操作的操作描述和操作结果中的至少一者,进而用于下一次迭代操作的操作预测。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122816508A_ABST
    Figure CN122816508A_ABST
Patent Text Reader

Abstract

The present disclosure relates to an event execution method and device, electronic equipment and storage medium. The method comprises: in the case of receiving a user instruction for indicating execution of a target event, iteratively performing the following operations until the target event is executed: performing operation prediction based on at least one of operation description and operation result of the last operation, trigger position of the last operation and the user instruction to obtain trigger position of the next operation; triggering the current interface based on the trigger position of the next operation, and updating the next operation as the last operation for the next iteration; and recording the trigger position of the updated last operation, and generating at least one of operation description and operation result of the updated last operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence, and in particular to an event execution method and apparatus, electronic device, and storage medium. Background Technology

[0002] With the continuous development of artificial intelligence, AI technology is being widely applied in various fields. For example, AI is often used to automate the operation of equipment.

[0003] However, when related technologies use artificial intelligence to automate operations, they often encounter problems such as inaccurate prediction of the next operation, resulting in poor automation effects. Summary of the Invention

[0004] This disclosure provides a training method and apparatus, electronic device, and storage medium for a human-computer interaction model, which can improve the accuracy of operation prediction and achieve good operation automation.

[0005] According to a first aspect of this disclosure, an event execution method is provided, which, upon receiving a user instruction to instruct the execution of a target event, iteratively performs the following operations until the target event is completed:

[0006] Based on at least one of the operation description and operation result of the previous operation, the trigger position of the previous operation, and the user instruction, the operation prediction is performed to obtain the trigger position of the next operation.

[0007] The current interface is triggered based on the trigger position of the next operation, and the next operation is updated to the previous operation in the next iteration; and the trigger position of the updated previous operation is recorded, and at least one of the operation description and operation result of the updated previous operation is generated.

[0008] According to a second aspect of this disclosure, a method for training a human-computer interaction model is provided, comprising:

[0009] Obtain the target dataset of sample events; the sample events contain multiple operations, and the target dataset contains the standard inputs and standard outputs of each operation; wherein, the standard input of any operation includes at least one of the operation description and operation result of the previous operation, the user instruction used to trigger the sample event, and the trigger position of the previous operation; the standard output of any operation includes the trigger position of the operation.

[0010] The standard inputs for each operation are provided to the model to be trained, so that the model to be trained outputs the trigger position of the corresponding operation;

[0011] The trigger positions of the standard input and output of each operation are compared with the trigger positions contained in the respective standard output. The model to be trained is iteratively corrected according to the comparison results until the output obtained based on the standard input of the operation contained in the sample event meets the preset iteration conditions when the comparison results with the standard output of the corresponding operation are satisfied.

[0012] According to a third aspect of this disclosure, an event execution apparatus is provided, comprising: an iteration unit, which, upon receiving a user instruction to execute a target event, iteratively performs the following operations until the target event is completed:

[0013] Based on at least one of the operation description and operation result of the previous operation, the trigger position of the previous operation, and the user instruction, the operation prediction is performed to obtain the trigger position of the next operation.

[0014] The current interface is triggered based on the trigger position of the next operation, and the next operation is updated to the previous operation in the next iteration; and the trigger position of the updated previous operation is recorded, and at least one of the operation description and operation result of the updated previous operation is generated.

[0015] According to a fourth aspect of this disclosure, a training apparatus for a human-computer interaction model is provided, comprising:

[0016] The acquisition unit acquires a target dataset of sample events; the sample events include multiple operations, and the target dataset includes standard inputs and standard outputs for each operation; wherein, the standard input of any operation includes at least one of the operation description and operation result of the previous operation, a user instruction for triggering the sample event, and the trigger position of the previous operation; the standard output of any operation includes the trigger position of the operation.

[0017] The prediction unit provides the standard inputs of each operation to the model to be trained, so that the model to be trained outputs the trigger position of the corresponding operation;

[0018] The correction unit compares the trigger positions of the standard input and output of each operation with the trigger positions contained in the respective standard output, and iteratively corrects the model to be trained according to the comparison results until the output obtained based on the standard input of the operation contained in the sample event meets the preset iteration conditions when the comparison result with the standard output of the corresponding operation is satisfied.

[0019] According to a fifth aspect of this disclosure, an electronic device is provided, comprising:

[0020] processor;

[0021] Memory used to store processor-executable instructions;

[0022] The processor implements the method as described in the first or second aspect by running the executable instructions.

[0023] According to a sixth aspect of this disclosure, a computer-readable storage medium is provided that stores computer instructions thereon, which, when executed by a processor, implement the steps of the method as described in the first or second aspect.

[0024] In the technical solution disclosed herein, during the iterative execution of the target event, the operation prediction for the next operation is based on at least one of the operation description and operation result of the previous operation, the trigger position of the previous operation, and the user instruction, thereby obtaining the trigger position of the next operation. Specifically, in addition to triggering the current interface based on the predicted trigger position, the next operation is updated to the previous operation for the next iteration, and the updated trigger position of the previous operation is recorded, along with at least one of the operation description and operation result of that previous operation, which is then used for operation prediction in the next iteration.

[0025] It should be understood that during the automated execution of an event, any information about the previous operation helps in predicting the next operation. This disclosure, based on the trigger position and user instruction of the previous operation, further introduces at least one of the operation description and operation result of the previous operation. This allows the disclosure to more accurately predict the trigger position of the next operation during any iteration, avoiding the problem of inaccurate predictions of the next operation caused by related technologies relying solely on the trigger position and user instruction of the previous operation. Attached Figure Description

[0026] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0027] Figure 1 This is a flowchart illustrating an event execution method according to an exemplary embodiment of this disclosure;

[0028] Figure 2A This is a flowchart illustrating a training method for a human-computer interaction model according to an exemplary embodiment of this disclosure;

[0029] Figure 2B This is a flowchart illustrating an event execution method based on a human-computer interaction model, as shown in an exemplary embodiment of this disclosure;

[0030] Figure 3 This is a flowchart illustrating another event execution method based on a human-computer interaction model, as shown in an exemplary embodiment of this disclosure;

[0031] Figure 4 This is a schematic diagram illustrating the input and output of a model according to an exemplary embodiment of this disclosure;

[0032] Figure 5 This is a flowchart illustrating an automated operation based on a human-computer interaction model, as shown in an exemplary embodiment of this disclosure;

[0033] Figure 6 This is a block diagram illustrating an event execution apparatus according to an exemplary embodiment of the present disclosure;

[0034] Figure 7 This is a block diagram illustrating a training device for a human-computer interaction model, as shown in an exemplary embodiment of this disclosure;

[0035] Figure 8 This is a schematic diagram of the structure of an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation

[0036] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0037] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0038] It should be understood that although the terms first, second, third, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0039] With the continuous development of artificial intelligence, it has been widely applied in various fields. Among them, in order to improve the efficiency of users operating equipment, artificial intelligence has been used to automate the operation of equipment.

[0040] When automating the operation of equipment, related technologies typically use user commands and the trigger position of the previous operation as the basis for predicting the next operation. This allows for the prediction of the trigger position of the next operation. Based on this, the current page can be triggered according to the predicted trigger position. In this way, operation prediction can be repeatedly executed, and page operations can be triggered according to the prediction results until the automated operation is completed.

[0041] However, when users actually use the relevant technologies to automate operations, they often make mistakes, resulting in poor automation effects.

[0042] To address this, this disclosure proposes an event execution method to avoid the common problem of misoperation when automating event execution using related technologies.

[0043] Figure 1 This is a flowchart illustrating an event execution method as an exemplary embodiment of the present disclosure. Figure 1 As shown, the method may include the following steps:

[0044] Step 102: Upon receiving a user instruction to execute a target event, iteratively perform the following operations until the target event is completed:

[0045] Step 1021: Based on at least one of the operation description and operation result of the previous operation, the trigger position of the previous operation, and the user instruction, perform operation prediction to obtain the trigger position of the next operation.

[0046] As can be seen from the above introduction, the information used in the related technologies to predict the trigger position of the next operation only includes the user instruction and the trigger position of the previous operation. Not only is the amount of information used for prediction small, but the correlation between the information used for prediction and the prediction result is not very close. As a result, the related technologies often have the problem of inaccurate prediction when making operation predictions, making it difficult to achieve good automation.

[0047] In view of this, considering the limited amount of information available for operation prediction in related technologies and the weak correlation between information and output, this disclosure further introduces at least one of the operation description and operation result of the previous operation as the basis for operation prediction. The operation description elucidates the operation content, which can not only be directly used to infer the next operation but also corroborate user instructions to indirectly deduce the next operation. The operation result of the previous operation directly influences the next operation, obviously significantly improving the accuracy of predicting the next operation.

[0048] It is not difficult to see that the two pieces of information introduced in this disclosure as the basis for operational prediction are not randomly selected, but deliberately selected based on the continuous nature of automated operations. Both are pieces of information that are strongly related to the next operation, which can greatly improve the accuracy of operational prediction.

[0049] In this disclosure, an event refers to the object to be completed by the aforementioned automated operation, which consists of one or more operations executed in succession.

[0050] In this disclosure, the basis for predicting the next operation may include, in addition to the user instruction used to trigger the sample event and the trigger position of the previous operation, at least one of the operation description and operation result of the previous operation. In other words, the basis for predicting the next operation can be "user instruction + trigger position of the previous operation + operation description of the previous operation", "user instruction + trigger position of the previous operation + operation result of the previous operation", or "user instruction + trigger position of the previous operation + operation description of the previous operation + operation result of the previous operation".

[0051] Step 1022: Trigger the current interface based on the trigger position of the next operation, and update the next operation to the previous operation in the next iteration; and record the trigger position of the updated previous operation, and generate at least one of the operation description and operation result of the updated previous operation.

[0052] In this disclosure, considering the continuity between the steps of the automated operation, the next operation may be related not only to the previous operation but also to all the operations that have been executed. Therefore, the basis for predicting the next operation may include not only various information from the previous operation but also various information from each of the executed operations.

[0053] In this scenario, information obtained from executed operations, excluding user instructions, can be collectively referred to as historical information. In other words, the prediction basis for the next operation can include "user instructions and historical information," meaning the prediction of the next operation is based on user instructions and historical information. Historical information can include at least one of the operation description and operation result of each executed operation, as well as the trigger position of each executed operation. It's clear that in this case, the prediction basis for the next operation encompasses both information from the previous operation and information from other executed operations. Under this premise, after completing the operation execution based on the operation prediction result and updating the next operation of this iteration to the previous operation of the next iteration, the recorded historical information can be updated based on at least one of the updated operation description and operation result of the previous operation, as well as the recorded trigger position of the updated previous operation.

[0054] Furthermore, in this case, the historical information may include: the trigger position of each executed operation, the operation description of each executed operation, and the operation result of each executed operation. Based on this, the operation prediction based on user instructions and historical information can be: predicting the operation based on user instructions and historical information to obtain the trigger position and operation description of the next operation; correspondingly, the operation of updating historical information can be: updating the historical information based on the updated operation description and operation result of the previous operation, as well as the recorded trigger position.

[0055] Furthermore, the basis for operational prediction, in addition to user instructions and historical information, may also include historical summaries summarizing each executed operation. It should be noted that the historical summaries in this disclosure refer to information used to provide a general description of executed operations. For example, they may include the current stage of execution, the tasks completed by the executed operation, and the progress of events based on the executed operation. This disclosure does not impose any limitations on this.

[0056] In this scenario, operation prediction based on user instructions and historical information can be achieved by: predicting operations based on user instructions, historical information, and a historical summary summarizing each executed operation, thereby obtaining the trigger position of the next operation, the operation description of the next operation, and the updated historical summary based on the next operation. Furthermore, after updating the next operation of the current iteration to the previous operation for the next iteration, the updated historical summary can be recorded for operation prediction in the next iteration.

[0057] In this disclosure, the above operation prediction can be achieved based on a large language model. For example, this disclosure can pre-train a human-computer interaction model for automated operation. Based on this, the above operation prediction can be: inputting at least one of the operation description and operation result of the previous operation, the trigger position of the previous operation, and the user command into the pre-trained human-computer interaction model, so that the trigger position of the next operation can be output by the human-computer interaction model.

[0058] It should be noted that in cases where the target event in this disclosure involves multiple steps and requires multiple iterations, since the first operation has no preceding operation, the prediction of the first operation can be performed without information from the preceding operation, simply based on the user's instructions. Alternatively, the parameter value used as information for the preceding operation can be 0. Of course, to improve the accuracy of the first operation, besides directly setting the information from the preceding operation to none, the current interface can be considered the result of the preceding operation. For example, when presenting the result of the preceding operation as an image, a screenshot of the current interface can be used as the result of the preceding operation of the first operation. Alternatively, the description of the current interface can be considered the description of the preceding operation; for example, the interface information of the current interface can be recorded in text form and used as the description of the preceding operation of the first operation.

[0059] In addition, the target event in this disclosure may also contain only a single operation. In this case, it is equivalent to only one iteration operation being performed during the entire event execution process. Since it contains only a single operation, there are no previous and next steps. Therefore, similar to the first operation when the event contains multiple operations, the information about the previous operation used as the basis for prediction may be nonexistent, or the parameter value used as information about the previous operation may be 0. Of course, the current interface can also be regarded as the operation result of the previous operation. For example, when presenting the operation result of the previous operation in the form of an image, a screenshot of the current interface can be used as the operation result of the previous operation of this unique operation. Alternatively, the interface description of the current interface can be regarded as the operation description of the previous operation. For example, the interface information of the current interface can be recorded in text form and then used as the operation description of the previous operation of this unique operation.

[0060] Of course, the above examples are merely illustrative. When a specific operation does not include the previous operation, the specific information to be obtained as information for the previous operation and used for operation prediction of that specific operation can be determined by those skilled in the art according to actual needs, and this disclosure does not impose any restrictions on this.

[0061] It should also be stated that the event execution method described in this disclosure can be applied to any type of carrier. For example, the execution subject of this disclosure can be any type of electronic device, such as a smartphone, tablet computer, or other mobile terminal, or a smart TV, PC (personal computer), or other fixed terminal; as another example, the execution subject of this disclosure can be a server or server cluster; and as yet another example, the execution subject of this disclosure can be an event execution engine deployed in the cloud. It should be understood that as long as there are sufficient processing resources for event execution, any device can serve as the execution subject in this disclosure. The specific carrier to which this disclosure is applied can be determined by those skilled in the art based on actual needs, and this disclosure does not impose any restrictions on this.

[0062] As can be seen from the above technical solution, in the process of iteratively executing the target event, this disclosure predicts the next operation based on at least one of the operation description and operation result of the previous operation, the trigger position of the previous operation, and the user instruction, thereby obtaining the trigger position of the next operation. Specifically, based on triggering the current interface according to the predicted trigger position, the next operation is updated to the previous operation for the next iteration, and the updated trigger position of the previous operation is recorded, along with at least one of the operation description and operation result of that previous operation, which is then used for operation prediction in the next iteration.

[0063] It should be understood that during the automated execution of an event, any information about the previous operation helps in predicting the next operation. This disclosure, based on the trigger position and user instruction of the previous operation, further introduces at least one of the operation description and operation result of the previous operation. This allows the disclosure to more accurately predict the trigger position of the next operation during any iteration, avoiding the problem of inaccurate predictions of the next operation caused by related technologies relying solely on the trigger position and user instruction of the previous operation.

[0064] The following section uses "Operation prediction through a large language model" as an example to introduce the technical solution of this disclosure. Since the prediction basis used in this disclosure for operation prediction includes at least one of the operation description and operation result of the previous operation, using a large language model for operation prediction is equivalent to inputting at least one of the operation description and operation result of the previous operation.

[0065] First, the model training process disclosed herein will be introduced.

[0066] Figure 2A This is a flowchart illustrating a training method for a human-computer interaction model, as shown in an exemplary embodiment of this disclosure. Figure 2AAs shown, the method may include the following steps:

[0067] Step 202A: Obtain the target dataset of the sample events; the sample events include multiple operations, and the target dataset includes the standard inputs and standard outputs of each operation; wherein, the standard input of any operation includes at least one of the operation description and operation result of the previous operation, the user instruction used to trigger the sample event, and the trigger position of the previous operation; the standard output of any operation includes the trigger position of the operation.

[0068] When using models for operation prediction, related technologies rely solely on user instructions and the trigger position of the previous operation to predict the next operation's trigger location. Correspondingly, during model training, the sample data also only includes user instructions and the trigger positions of various operations. This results in limited information for prediction and a weak correlation between the information used and the prediction results. Consequently, models trained using these technologies often exhibit inaccurate predictions, hindering effective automation.

[0069] As mentioned above, considering the limited amount of input information and weak correlation between information and output in related technologies, this disclosure further introduces at least one of the operation description and operation result of the previous operation as model input for operation prediction. The operation description elucidates the operation content, which can be directly used to infer the next operation and can also be corroborated with user instructions to indirectly deduce the next operation. The operation result of the previous operation directly influences the next operation, significantly improving the prediction accuracy of the next operation.

[0070] As stated above, the two pieces of information introduced in this disclosure as model inputs are not randomly selected, but deliberately chosen based on the continuous nature of automated operations. Both are pieces of information that are strongly correlated with the model output, which can significantly improve the accuracy of the model output.

[0071] To facilitate understanding, before going into detail about the training method for human-computer interaction models, we will first introduce the concepts involved in this method.

[0072] In this disclosure, an event refers to the object to be completed by the aforementioned automated operation, consisting of at least one continuously executed operation. A sample event refers to the event corresponding to the sample data used for model training. The target dataset refers to the set of sample data obtained by executing the sample events, specifically the standard input and standard output of each operation contained in the sample event. Standard input refers to the sample data used to predict the trigger position of the next operation during model training; correspondingly, standard output refers to the standard sample data used to characterize the model output during model training, i.e., the label of the model output.

[0073] Step 204A: Provide the standard inputs of each operation to the model to be trained so that the model to be trained can output the trigger position of the corresponding operation.

[0074] In this disclosure, the standard input for any operation includes at least one of the operation description and operation result of the previous operation, as well as the user instruction for triggering the sample event and the trigger position of the previous operation; while the standard output for any operation includes the trigger position of the next operation.

[0075] In this disclosure, the standard inputs and standard outputs of each operation can be obtained by executing sample events. It should be noted that when actually executing sample events, either an automated operation trained in related technologies can be used, selecting an automated operation without errors, and recording the necessary information as standard inputs and standard outputs during each operation; or the sample events can be executed manually to ensure the accuracy of the entire sample event execution, and then recording the necessary information as standard inputs and standard outputs during each operation. Of course, the examples here are merely illustrative. How to obtain the various information used as standard inputs and standard outputs can be determined by those skilled in the art according to actual needs, as long as the accuracy of the obtained information and the absence of errors in the event are guaranteed. This disclosure does not impose any limitations in this regard.

[0076] In this disclosure, as described above, the standard input for any operation, in addition to the user instruction for triggering the sample event and the trigger position of the previous operation, may further include at least one of the operation description and operation result of the previous operation. In other words, the standard input for any operation can be "user instruction + trigger position of the previous operation + operation description of the previous operation", "user instruction + trigger position of the previous operation + operation result of the previous operation", or "user instruction + trigger position of the previous operation + operation description of the previous operation + operation result of the previous operation".

[0077] In this disclosure, the continuity between the steps of an automated operation is considered, such that the next operation may be related not only to the previous operation but also to all the operations that have been performed. Therefore, the standard input of any operation may contain various information of each of the previous operations in addition to the various information of each of the operations that have been performed.

[0078] In this case, information obtained from executed operations, excluding user instructions, can be collectively referred to as historical information. In other words, the standard input for any operation can include "user instructions and historical information." Historical information can include at least one of the operation description and operation result of each executed operation, and the triggering position of each executed operation. In other words, in this case, the standard information for any operation encompasses both the information used as standard input in the previous operation and the information used as standard input in other executed operations.

[0079] Furthermore, in this case, in addition to user instructions and historical information, a historical summary summarizing each executed operation may also be included. It should be noted that the historical summary in this disclosure refers to information used to provide a general description of the executed operation, such as the current stage of execution, the task completed by the executed operation, and the progress of events based on the executed operation. This disclosure does not impose any limitations on this.

[0080] Of course, the above examples are merely illustrative. The types of information to be used as standard input, and whether the standard input for any operation "contains only information from the previous operation" or "contains information from each of the operations that have been performed," can be determined by those skilled in the art based on actual needs. This disclosure does not impose any restrictions on this.

[0081] Corresponding to the standard input, the information types contained in the standard output of this disclosure can also be determined by those skilled in the art according to actual needs. For example, in addition to the triggering location of any operation, it may also include at least one of the following: an operation description of any operation, a historical summary updated based on any operation, historical information updated based on the corresponding operation, etc., and this disclosure does not impose any limitations on this.

[0082] Step 206A: The trigger positions of the standard input and output of each operation are compared with the trigger positions contained in the respective standard outputs, so as to iteratively correct the model to be trained according to the comparison results until the output obtained based on the standard input of the operation contained in the sample event meets the preset iteration conditions when the comparison result with the standard output of the corresponding operation is satisfied.

[0083] For ease of understanding, the following is a brief explanation based on two examples:

[0084] In one embodiment, the standard input for any operation may include: user instructions and historical information. The historical information includes: the trigger location of each executed operation, the operation description of each executed operation, and the operation result of each executed operation. The standard output for any operation may include: the trigger location of the operation and the operation description of the operation.

[0085] Based on this, during model training, the standard inputs for each operation can be provided to the model to be trained, so that the model can output the trigger position and operation description of the corresponding operation. During the comparison, the trigger position and operation description output based on the standard input of each operation are compared with the trigger position and operation description contained in the respective standard output, and then the model to be trained is iteratively corrected according to the comparison results.

[0086] It is easy to see that the types of information used as input to the model and the types of information used as output to the model are not completely consistent in this embodiment. The input includes the operation result, while the output does not include the operation result.

[0087] Furthermore, in this embodiment, the standard input for any operation can be further supplemented with the aforementioned historical summary, and correspondingly, the standard output for any operation can be supplemented with the historical summary updated based on that operation. In other words, the standard input for any operation can include: user instructions, historical information, and historical summary; while the standard output for any operation can include: the triggering location of that operation, the operation description of that operation, and the historical summary updated based on that operation.

[0088] Based on this, the standard inputs of each operation can be provided to the model to be trained, so that the model to be trained can output the trigger position of the corresponding operation, the operation description of the corresponding operation, and the historical summary updated based on the corresponding operation. The trigger position, operation description, and historical summary output based on the standard input of each operation are compared with the trigger position, operation description, and historical summary contained in their respective standard outputs, so as to iteratively correct the model to be trained based on the comparison results.

[0089] In this embodiment, since the standard input of each operation contains the operation result of the corresponding operation, the operation result can also be used as one of the comparison objects for iterative correction of the model.

[0090] For example, based on the trigger positions of each operation output by the model to be trained, the corresponding positions on the terminal screen can be triggered to obtain the operation results of the corresponding operations. The operation results obtained based on the trigger positions of each operation output by the model to be trained can be compared with the operation results contained in the standard input of the next operation of the corresponding operation. Based on this, the comparison results of the operation results can also be used as one of the bases for model iterative correction. For example, if the above output only contains the trigger position and operation description, the model to be trained can be iteratively corrected based on the comparison results of the trigger position, operation description, and operation results. For another example, if the above output contains the trigger position, operation description, and historical summary, the model to be trained can be iteratively corrected based on the comparison results of the trigger position, operation description, historical summary, and operation results.

[0091] In another embodiment, apart from user instructions, the types of information input to the model and the types of information output to the model can be completely consistent. For example, the information type of the standard input for any operation can be consistent with the previous embodiment, and may include: user instructions and historical information. The historical information includes: the triggering position of each executed operation, the operation description of each executed operation, and the operation result of each executed operation. The standard output of any operation may include: the triggering position of that operation, the operation description of that operation, and the historical information updated based on that operation.

[0092] Based on this, the standard inputs of each operation can be provided to the model to be trained, so that the model to be trained can output the trigger position of the corresponding operation and the updated historical information based on the corresponding operation. The trigger position and updated historical information output based on the standard input of each operation are compared with the trigger position and updated historical information contained in the respective standard output, and the model to be trained is iteratively corrected based on the comparison results.

[0093] Furthermore, in this embodiment, the standard input for any operation may also include a historical summary of the previous operation, while the standard output for any operation may include an updated historical summary based on that operation. Correspondingly, the comparison object used for iterative correction of the model to be trained can include this historical summary. The specific execution process is the same as other comparison objects, as described above, and will not be repeated here.

[0094] Of course, the above two embodiments are illustrative. The information contained in the input and output of the model, and how to compare the information to iteratively correct the model, can be determined by those skilled in the art according to actual needs. This disclosure does not impose any restrictions on this.

[0095] In this disclosure, a loss function can be pre-constructed for the model to be iteratively corrected. For example, if the standard output of any operation includes the trigger position and operation description of that operation, the above comparison and iterative correction operations can be performed as follows: the trigger position and operation description based on the standard input and output of each operation, and the trigger position and operation description contained in the respective standard output, are input into the loss function corresponding to the model to be trained, so as to iteratively correct the model to be trained based on the feedback of the loss function.

[0096] It is worth noting that the loss function can be set based on the model's output. For example, since the model output in this example includes two items: the trigger position and the operation description, the loss function corresponding to the model to be trained can be designed to include two parts: a position prediction sub-function and a description generation sub-function. Based on this, loss weights can be configured for the position prediction sub-function and the description generation sub-function respectively to optimize the model trained based on the loss function. For instance, since the actual execution of each operation is based on the predicted trigger position of the next operation, the trigger position output is the most important output of the model, while the operation description assists in predicting the trigger position of the next operation. Therefore, when setting the loss weights of the loss function, the loss weight of the position prediction sub-function can be set higher than the loss weight of the description generation sub-function.

[0097] Of course, the examples here are merely illustrative. How to construct the loss function and how to configure the loss weights for each part of the loss function when constructing the loss function based on the model output can be determined by those skilled in the art according to actual needs. For example, the loss function can also be constructed based on the model input. For instance, a sub-function can be configured for each type of information contained in the model input to represent the part of the model used to analyze and interpret the corresponding information. For another example, regarding the configuration of loss weights, when the model output contains three parts: "trigger position, operation description, and historical summary", the loss function can include a position prediction sub-function, a description generation sub-function, and a historical summary sub-function. The weight ratio of the loss weights of the three can be set as "loss weight of position prediction sub-function > loss weight of description generation sub-function, loss weight of position prediction sub-function > loss weight of historical summary sub-function". In most cases, the setting of the loss weights of each part in the loss function is positively correlated with the degree of influence of the corresponding part on the prediction accuracy of the trigger position of the next operation. This disclosure does not impose any restrictions on this.

[0098] In this disclosure, an iterative condition for determining whether the model has been trained can be preset. When the comparison result meets or satisfies the iterative condition, the model is considered to have been trained and the currently trained model is taken as the trained human interaction model.

[0099] In this disclosure, iteration conditions can be set from different dimensions.

[0100] In one scenario, the iteration mechanism of the model can be configured. For example, the number of iterations can be preset, and once the actual number of iterations reaches the preset number of iterations, the iteration condition is considered to be met. Alternatively, the upper limit of the training time for a single model can be preset, and once the training time for iteratively correcting the same model reaches the upper limit, the iteration condition is considered to be met.

[0101] In another scenario, the iteration condition can be set based on the comparison results. For example, the iteration condition can be determined to be satisfied when the comparison results for each operation are completely consistent; or the iteration condition can be determined to be satisfied when the comparison results for each operation show that the similarity is higher than the preset similarity; or the iteration condition can be determined to be satisfied when the comparison results show that the number of operations with similarity higher than the preset similarity reaches the preset proportion of all operations.

[0102] In another case, the iteration conditions can be tied to the convergence of the loss function; that is, once the loss function converges, the iteration conditions are considered to be satisfied.

[0103] Of course, the above examples are all illustrative. The specific way to set the iteration conditions can be determined by those skilled in the art according to actual needs, and this disclosure does not impose any restrictions on this.

[0104] It should be stated that the model training method described in this disclosure can be applied to any type of carrier. For example, the execution subject of this disclosure can be any type of electronic device, such as a smartphone, tablet computer, or other mobile terminal, or a smart TV, PC (personal computer), or other fixed terminal; as another example, the execution subject of this disclosure can be a server or server cluster; and as yet another example, the execution subject of this disclosure can be a model training engine deployed in the cloud. It should be understood that any carrier with sufficient processing resources for model training can serve as the execution subject in this disclosure. The specific carrier to which this disclosure is applied can be determined by those skilled in the art based on actual needs, and this disclosure does not impose any restrictions on this.

[0105] It should also be stated that the operation results mentioned above can be represented in any form. For example, the operation result can be recorded as an image. In this case, after each triggering of the current interface according to the triggering position output by the model, a screenshot of the switched interface can be taken as the operation result of this operation. Alternatively, the operation result can be recorded as text. In this case, after each triggering of the current interface according to the triggering position output by the model, text content describing the screen changes caused by the triggering operation can be generated as the operation result of this operation. Of course, the examples here are merely illustrative. The specific form in which the operation results in this disclosure are represented can be determined by those skilled in the art according to actual needs, and this disclosure does not impose any restrictions on this.

[0106] In addition, it is important to emphasize that since the first operation of a sample event usually does not have a previous operation, the model input corresponding to the first operation during model training can be only the user command. Alternatively, the input of at least one of the operation description or result of the previous operation, as well as the triggering position of the previous operation, can be absent, which is reflected as 0 in the parameters.

[0107] As can be seen from the above technical solution, in the model training process, this disclosure further introduces at least one of the operation description and operation result of the previous operation, in addition to the user command and the trigger position of the previous operation. Based on this, the current large language models have the initial ability to understand content, and content information such as operation description and operation result can be used as model input to help the model predict the trigger position of the next operation.

[0108] It should be understood that the technical solution of this disclosure, on the one hand, utilizes the characteristics of current large language models as described above; on the other hand, the use of operation descriptions and operation results as model inputs is not blind selection, but rather takes into account the continuous nature of specific events in automated operations, using information such as operation descriptions and operation results that are strongly correlated with previous and subsequent operations as inputs. Clearly, using at least one of the operation descriptions and operation results as model inputs can significantly improve the accuracy of the model's prediction of the next operation, avoiding the erroneous operation problems caused by inaccurate model predictions of the next operation in related technologies.

[0109] After training the aforementioned human-computer interaction model, it can be used to perform automated operations, or in other words, execute various events. Below, we will use the execution of a target event as an example to introduce the process of executing events based on the trained human-computer interaction model.

[0110] It should be emphasized that this example only introduces the technical solution of this disclosure from the perspective of the execution end. Most of the operation methods, such as what the model input includes, what the model output includes, and how to set the model loss function, have been explained in the training end described above. Please refer to the above introduction and they will not be repeated here.

[0111] Figure 2B This is a flowchart illustrating an event execution method based on a human-computer interaction model, as an exemplary embodiment of this disclosure. Figure 2B As shown, the method may include the following steps:

[0112] Step 202B: Upon receiving a user instruction to execute a target event, iteratively perform the following operations until the target event is completed:

[0113] Step 2021B involves taking at least one of the operation description and operation result of the previous operation, the trigger position of the previous operation, and the user instruction input into a pre-trained human-computer interaction model, so that the human-computer interaction model can output the trigger position of the next operation.

[0114] In this disclosure, upon receiving a user instruction to perform a target event, a pre-trained human-computer interaction model can predict the trigger position for each step of the operation, so that the device can perform the trigger operation at the corresponding position on the current interface based on the prediction result. In other words, the model prediction can be iteratively executed, and the trigger operation can be performed based on the prediction result.

[0115] It is important to emphasize that, as mentioned earlier, since the first operation does not have a preceding operation, when predicting the trigger position of the first operation, the parameters representing at least one of the operation description and result of the preceding operation, as well as the trigger position of the preceding operation, can all be 0. The input information used to predict subsequent operations can be obtained during the execution of the preceding operation.

[0116] Step 2022B: Trigger the current interface based on the trigger position of the next operation, and update the next operation to the previous operation in the next iteration; and record the trigger position of the updated previous operation, and generate at least one of the operation description and operation result of the updated previous operation.

[0117] In this disclosure, after the trigger position of any operation is predicted by the model, the current interface can be triggered based on that trigger position to execute that operation. After execution, the operation is no longer an unexecuted next step, but rather an executed previous step. Therefore, the original next step (i.e., any of the aforementioned operations) needs to be updated to the previous step for the next iteration. Specifically, for model prediction in the next iteration, the trigger position of the updated previous step, as well as at least one of the operation description and operation result of the updated previous step, must be recorded.

[0118] As mentioned above, the model input can further include historical information and historical summaries, and the model output can further include operation descriptions and historical summaries. The logic of the execution end remains consistent regardless of the model input and output; the only difference is the addition of relevant information during model input and output, and the addition or subtraction of recorded information after the operation is completed. Specific operations can be found in the training end's description and the execution end's description of model prediction and operation execution, which will not be elaborated upon here.

[0119] To facilitate understanding, let's further illustrate the steps involved in execution using the example where "model input includes: user instructions, a screenshot of the current interface, historical information, and historical summaries; model output includes: the trigger point for the next operation, a description of the next operation, and an updated historical summary." In this case, the steps executed by the execution end can be as follows: Figure 3 As shown.

[0120] Figure 3 This is a flowchart illustrating another event execution method based on a human-computer interaction model, as shown in an exemplary embodiment of this disclosure. Figure 3 As shown, the method may include the following steps:

[0121] Step 302: Upon receiving a user instruction to execute a target event, iteratively perform the following operations until the target event is completed:

[0122] Step 3021: Input the screenshot of the current interface, historical information, and historical summary into the pre-trained human-computer interaction model, so that the human-computer interaction model can output the trigger position of the next operation, the operation description of the next operation, and the updated historical summary; the historical information includes: the trigger position of each executed operation, the operation description, and the screenshot of the current interface updated based on the corresponding operation.

[0123] Step 3022: Trigger the current interface based on the trigger position of the next operation, so that the interface obtained by switching based on the trigger operation is the updated current interface, and take a screenshot of the updated current interface; and update the historical information based on the trigger position of the next operation, the operation description, and the screenshot of the updated current interface, so that the next operation is regarded as one of the executed operations for the next iteration operation.

[0124] In this embodiment, since the model input includes multiple parts such as a screenshot of the current interface (hereinafter referred to as the current screenshot for ease of description), historical information, and historical summary, the large language model to be trained can be a multimodal large language model, such as the Qwen2-VL (7B) model, the MiniCPM-V 2.6 (8B) model, etc., which have lightweight characteristics. Of course, the examples here are only illustrative, and any type of large language model can be used as the model to be trained in this embodiment. The specific large language model used can be determined by those skilled in the art according to actual needs, and this disclosure does not limit it.

[0125] Before training the model, the structure of the model input can be pre-built. For example, this structure can be as follows:

[0126] Picture 1:

[0127] / media / ml / ssd1 / data / aitw / aitw_images / general / 2.png

[0128] Please generate the next move according to the ui screenshot,instruction,previous actions and previous action results.Instruction:Search for the best pizza restaurants on Maps.

[0129] Previous actions:

[0130] Step0:{"picture": / media / ml / ssd1 / data / aitw / aitw_images / general / 0.png,"action_type":6,"action_description":"press the home button."}.

[0131] Step 1: {"picture": / media / ml / ssd1 / data / aitw / aitw_images / general / 1.png,"action_type":0,"action_description":"scroll up."}.

[0132] Previous action results:

[0133] "By doing so, the app drawer was accessed where various apps aredisplayed. This allows the user to find and open the Maps application to search for pizza restaurants."

[0134] In this input structure, "Picture1" is the current screenshot; "Please generate…" is the user instruction, indicating that the event to be automated is "search for the best pizza place on the map"; "Previous actions" is the historical information, which currently contains two executed operations, namely "Step0" and "Step1", both of which contain the current screenshot "Picture", action type "action_type", and action description "action_description" for the corresponding operation.

[0135] The action type "action_type" can record multiple fields. For example, the recorded fields can include: a "point" field, which is the trigger position of the corresponding executed operation, specifically coordinates on the screen; and a "type" field, where different parameters express different meanings. For example, a type field parameter of 0 indicates swiping down; a type field parameter of 1 indicates swiping up; a type field parameter of 8 indicates swiping left; a type field parameter of 9 indicates swiping right; a type field parameter of 4 indicates clicking; a type field parameter of 3 indicates inputting text; a type field parameter of 5 indicates pressing the back button; a type field parameter of 6 indicates pressing the Home button; a type field parameter of 7 indicates pressing the Enter key; a type field parameter of 10 indicates task completion; and a type field parameter of 11 indicates task failure. In the current example, Step0 has a type field of 6, which indicates that the operation type is pressing the Home button. The action description is the operation description of the above executed operation. "Previous action results" is the above historical summary.

[0136] Correspondingly, the structure of the model output can be as follows:

[0137] "action_type":4,"click_point":(0.49,0.35),"action_description":"clickon the Maps app located at the upper-middle part of the screen.","Previousaction results":"By doing so, the Maps application has been opened, allowing for a search for local pizza restaurants.The reason is that the Maps app provides location-based search results, which can help find the best pizzarestaurants in the area."

[0138] In this output structure, "action_type" represents the action type, with the "type" field set to "4", indicating a click operation. The "click_point" field represents the click coordinates, which are (0.49, 0.35), indicating the trigger position for the next action that has not yet been executed, i.e., the trigger position for any of the aforementioned actions. "action_description" is the operation description for the next action that has not yet been executed, i.e., the operation description for any of the aforementioned actions. "Previous action results" is the updated historical summary.

[0139] Based on this, the model's input and output can be as follows: Figure 4 As shown. The above input and output content, and Figure 4 As we can see, the model is currently predicting the trigger position of the third operation of the event. The executed operations already include Step0 and Step1. Therefore, the historical operation information includes screenshots and operation information of Step0 and Step1.

[0140] As mentioned above, after constructing the model's input and output, sample data can be acquired first, and then the model can be trained using the methods described above. During this process, a loss function can be constructed for iterative optimization of the model.

[0141] This embodiment can use an open-source dataset, which can be labeled by technical personnel themselves to serve as sample data for model training. For example, the loss function used in this embodiment can be cross-entropy loss:

[0142]

[0143] In this cross-entropy loss, C is the vocabulary after the sample data is converted into numerical form; y perd The token, or label, is the result of the model's prediction; while y tar The token is labeled, which is the standard output mentioned above.

[0144] Furthermore, as described above, this embodiment can also assign different loss weights to different parts of the loss function based on different modules of the model. In this case, the loss function is:

[0145] TotalLoss=w1*CE act +w2*CE act_desc +w3*CE perv_act

[0146] In this loss function, TotalLoss represents the overall loss function; CE act The loss in the model output represents the location of the operation trigger; CE act_descIt is the loss in the operation description section, CE prev_act The loss is the historical operation summary part. The weights assigned to these three parts must ensure that w1>w2 and w1>w3, so as to strengthen the model's learning of the specific operation part.

[0147] Once the human-computer interaction model is trained, automated operations can be performed based on it.

[0148] Figure 5 This is a flowchart illustrating an automated operation method based on a human-computer interaction model, which is an exemplary embodiment of this disclosure.

[0149] like Figure 5 As shown, after receiving the user's command, the system enters a loop within the dashed box to execute the operations. The first operation and the remaining operations are described below:

[0150] I. First Operation

[0151] The first operation instructed by the user, since there is no previous operation, therefore... Figure 5 The storage space shown, used to store historical operation information, did not originally record any information related to historical steps.

[0152] After receiving a user command, the device will obtain a screenshot of the current page. On the one hand, the current page screenshot and the user command will be preprocessed so that the preprocessed user command and the current page screenshot can be input into the multimodal large model for the next operation prediction. On the other hand, the current page screenshot will be stored in the storage space of historical operation information.

[0153] For example, let's take the scenario of a user ordering takeout via their smartphone. Assume the smartphone is currently on its home screen, and the user's command is command X, which reads, "Order meal B from restaurant A." Upon receiving this command, the smartphone can capture a screenshot of the current screen, i.e., a "home screen screenshot." Firstly, the "home screen screenshot" and "command X" are preprocessed and input into a multimodal large-scale model. Secondly, the "home screen screenshot" is recorded in the historical operation information storage space.

[0154] It should be noted that, for ease of expression, in the following text, "preprocessing the 'desktop screenshot' and 'instruction X'" will be simplified to "desktop screenshot" and "instruction X" as the model input. The other information used as model input is similar and will not be elaborated further.

[0155] In this example, after inputting the "desktop screenshot" and "command X" into the multimodal large model, the output of the multimodal large model can be the trigger position of the first operation, the operation description of the first operation, and the historical summary updated based on the first operation. Furthermore, the maintained historical operation information can be updated based on the trigger position of the first operation, the operation description of the first operation, and the historical summary updated based on the first operation.

[0156] In other words, if we name the first operation Step0, the trigger point of the first operation Point0, the operation description of the first operation Description0, and the historical summary updated after the first operation Results0, then the input and output of the human-computer interaction model, as well as the historical operation information before and after operation prediction, during the model prediction phase before executing Step0, can be shown in Table 1 below:

[0157]

[0158]

[0159] Table 1

[0160] After the multimodal large model outputs the prediction result of Step 0, the desktop can be triggered based on Point 0. In this example, this is usually the location of the "food delivery app" icon, thus entering the food delivery app's software page. At this point, it can be determined whether the "order food delivery" event corresponding to instruction X has been completed. If not, the operation loop for the next step is entered.

[0161] In this embodiment, determining whether the "ordering takeout event" is complete can be done in several ways. For example, the last operation corresponding to the event can be predetermined, and it can be determined whether the currently executed operation is the last operation. If so, the event is determined to be complete; otherwise, the event is determined to be incomplete. Another example is that the operation result of the event can be predetermined, and it can be determined whether the operation result of the currently executed operation matches the predetermined operation result. If they match, the event is determined to be complete; otherwise, the event is determined to be incomplete. Of course, these examples are merely illustrative. How to determine whether an event is complete, and thus whether to enter the next operation cycle, can be determined by those skilled in the art according to actual needs, and this disclosure does not impose any limitations on this.

[0162] II. Other Operations

[0163] After the first operation is completed, the execution process of each subsequent operation is the same as that of the first operation, with the only difference being the input. For example, compared to the first operation, each subsequent operation has a previous operation, so the input also includes "historical information" and "historical summary".

[0164] Continuing with the example above, the second operation can be called Step 1. Step 1 also includes two parts: model prediction and operation execution. In the model prediction process of Step 1, it is necessary to first take a screenshot of the operation page after Step 0, that is, to take a screenshot of the aforementioned food delivery app's page. On the one hand, this screenshot needs to be recorded in the storage space of historical operation information; on the other hand, this screenshot can be used as the model input for Step 1. Therefore, the model input for prediction is: the app screenshot, instruction X, historical operation information updated based on Step 0, and historical summary Results0 updated based on Step 0. The model output is: the trigger point Point1 of Step 1, the operation description Description1 of Step 1, and the historical summary Results1 updated based on Step 1.

[0165] In other words, during the model prediction phase before Step 1, the inputs and outputs of the human-computer interaction model, as well as the historical operation information before and after the operation prediction, can be shown in Table 1 below:

[0166]

[0167] Table 2

[0168] After the multimodal large model outputs the prediction results of Step 1, the software page can be triggered based on Point 1. In this example, this is usually the icon of "Shop A", which leads to Shop A's store page. At this point, it can be determined again whether the "order takeout event" corresponding to instruction X has been completed. If not, the operation loop for the next step is entered.

[0169] Continuing with the example above, the third operation can be called Step 2, which also includes two parts: model prediction and operation execution. In the model prediction process of Step 2, it is necessary to first take a screenshot of the operation page after Step 1, that is, to take a screenshot of the store page of the aforementioned store A. On the one hand, this screenshot needs to be recorded in the storage space of historical operation information; on the other hand, this screenshot can be used as the model input for Step 2. Therefore, the model input for prediction is: store page screenshot, instruction X, historical operation information updated based on Step 1, and historical summary Results1 updated based on Step 1. The model output is: the trigger point Point2 for Step 2, the operation description Description2 for Step 2, and the historical summary Results2 updated based on Step 2.

[0170] In other words, during the model prediction phase before Step 2, the inputs and outputs of the human-computer interaction model, as well as the historical operation information before and after the operation prediction, can be shown in Table 1 below:

[0171]

[0172]

[0173] Table 3

[0174] After the multimodal large model outputs the prediction results of Step 2, the store page can be triggered based on Point 2. In this example, this is usually the location of the "B Package" icon, which leads to the checkout page for Package B. At this point, it can be determined again whether the "order takeout event" corresponding to instruction X has been completed. If not, the operation loop for the next step can be entered.

[0175] Continuing with the example above, the fourth operation can be called Step 3, which also includes two parts: model prediction and operation execution. In the model prediction process of Step 3, it is necessary to first take a screenshot of the operation page after Step 2, that is, to take a screenshot of the settlement page of Package B mentioned above. On the one hand, this screenshot of the settlement page needs to be recorded in the storage space of historical operation information; on the other hand, this screenshot of the settlement page can be used as the model input for Step 3. At this time, the model input for prediction is: the screenshot of the settlement page, instruction X, historical operation information updated based on Step 2, and historical summary Results2 updated based on Step 2. The model output is: the trigger point Point3 of Step 3, the operation description Description3 of Step 3, and the historical summary Results3 updated based on Step 3.

[0176] In other words, during the model prediction phase before Step 3, the inputs and outputs of the human-computer interaction model, as well as the historical operation information before and after the operation prediction, can be shown in Table 4 below:

[0177]

[0178]

[0179] Table 4

[0180] After the multimodal large model outputs the prediction results of Step 3, the checkout page can be triggered based on Point 3. In this example, this is usually the icon position of the checkout control, thus completing the checkout operation for Shop A's Package B. At this point, it can be determined that the "order takeout event" of "Instruction X" has been completed and will not enter the next operation loop.

[0181] As can be seen from the above embodiments, the technical solution of this disclosure can automatically execute user instructions such as "ordering takeout," simplifying user operations. Specifically, when predicting the next operation using a large language model, the model input includes not only the trigger position of the previous operation but also historical operation information and a summary of executed operations. This significantly improves the accuracy of the trigger position of the next operation output by the model, avoiding the common problem of erroneous operations in the automated execution process of related technologies.

[0182] Figure 6 This is a block diagram illustrating an event execution apparatus according to an exemplary embodiment of this disclosure. (Refer to...) Figure 6 The device includes an iteration unit 601.

[0183] When the iteration unit 601 receives a user instruction to execute a target event, it iteratively performs the following operations until the target event is completed:

[0184] Based on at least one of the operation description and operation result of the previous operation, the trigger position of the previous operation, and the user instruction, the operation prediction is performed to obtain the trigger position of the next operation.

[0185] The current interface is triggered based on the trigger position of the next operation, and the next operation is updated to the previous operation in the next iteration; and the trigger position of the updated previous operation is recorded, and at least one of the operation description and operation result of the updated previous operation is generated.

[0186] Optional,

[0187] The operation prediction based on at least one of the operation description and operation result of the previous operation, the triggering position of the previous operation, and the user instruction includes: operation prediction based on the user instruction and historical information; the historical information includes: at least one of the operation description and operation result of each executed operation, and the triggering position of each executed operation;

[0188] The iteration unit 601 is also used to update the historical information based on at least one of the updated operation description and operation result of the previous operation and the recorded trigger position.

[0189] Optionally, the historical information includes: the triggering location of each executed operation, the operation description of each executed operation, and the operation result of each executed operation;

[0190] The operation prediction based on the user instructions and historical information includes: performing operation prediction based on the user instructions and historical information to obtain the trigger position of the next operation and the operation description of the next operation;

[0191] The step of updating the historical information based on at least one of the updated operation description and operation result of the previous operation and the recorded trigger position includes: updating the historical information based on the updated operation description and operation result of the previous operation and the recorded trigger position.

[0192] Optional,

[0193] The step of predicting the operation based on the user instruction and the historical information to obtain the trigger position of the next operation and the operation description of the next operation includes: predicting the operation based on the user instruction, the historical information and the historical summary used to summarize each executed operation to obtain the trigger position of the next operation, the operation description of the next operation, and the historical summary updated based on the next operation.

[0194] The iteration unit 601 is also used to record the updated history summary after updating the next step operation to the previous step operation when the next iteration is executed.

[0195] Optionally, the step of predicting the trigger position of the next operation based on at least one of the operation description and operation result of the previous operation, the trigger position of the previous operation, and the user instruction includes:

[0196] The human-computer interaction model is pre-trained and uses at least one of the operation description and operation result of the previous operation, the trigger position of the previous operation, and the user command input to output the trigger position of the next operation.

[0197] Figure 7 This is a block diagram illustrating a training apparatus for a human-computer interaction model, as shown in an exemplary embodiment of this disclosure. (Refer to...) Figure 7 The device includes an acquisition unit 701, a prediction unit 702, and a correction unit 703.

[0198] Acquisition unit 701 acquires the target dataset of the sample event; the sample event includes multiple operations, and the target dataset includes the standard input and standard output of each operation; wherein, the standard input of any operation includes at least one of the operation description and operation result of the previous operation, the user instruction used to trigger the sample event, and the trigger position of the previous operation; the standard output of any operation includes the trigger position of the operation.

[0199] The prediction unit 702 provides the standard inputs of each operation to the model to be trained, so that the model to be trained outputs the trigger position of the corresponding operation;

[0200] The correction unit 703 compares the trigger positions of the standard input and output of each operation with the trigger positions contained in the respective standard output, and performs iterative correction on the model to be trained according to the comparison results until the output obtained based on the standard input of the operation contained in the sample event meets the preset iteration conditions when the comparison results with the standard output of the corresponding operation are satisfied.

[0201] Optional,

[0202] The standard inputs for any of the operations include: the user instructions and historical information;

[0203] The historical information includes: at least one of the operation description of each executed operation and the operation result of each executed operation, and the triggering position of each executed operation.

[0204] Optional,

[0205] The historical information includes: the trigger location of each executed operation, the operation description of each executed operation, and the operation result of each executed operation;

[0206] The standard output of any operation also includes: an operation description of any operation;

[0207] The prediction unit 702 is further used to: provide the standard input of each operation to the model to be trained, so that the model to be trained outputs the trigger position of the corresponding operation and the operation description of the corresponding operation;

[0208] The correction unit 703 is further used to compare the trigger position and operation description of the standard input and output based on each operation with the trigger position and operation description contained in the respective standard output.

[0209] Optional,

[0210] The standard input for any operation also includes: a historical summary of all executed operations;

[0211] The standard output of any operation also includes: an updated historical summary based on any operation;

[0212] The prediction unit 702 is further used to: provide the standard input of each operation to the model to be trained, so that the model to be trained outputs the trigger position of the corresponding operation, the operation description of the corresponding operation, and the historical summary updated based on the corresponding operation;

[0213] The correction unit 703 is further used to compare the trigger position, operation description, and history summary of the standard input and output based on each operation with the trigger position, operation description, and history summary contained in the respective standard output.

[0214] Optionally, the correction unit 703 is also used for:

[0215] Based on the trigger positions of each operation output by the model to be trained, the corresponding positions on the terminal screen are triggered to obtain the operation results of the corresponding operations; and the operation results obtained based on the trigger positions of each operation output by the model to be trained are compared with the operation results contained in the standard input of the next operation of the corresponding operation.

[0216] Based on the comparison results of the trigger location, the operation description, and the operation result, the model to be trained is iteratively corrected.

[0217] Optionally, the standard output of any operation may also include: historical information updated based on the operation.

[0218] The prediction unit 702 is further used to: provide the standard input of each operation to the model to be trained, so that the model to be trained outputs the trigger position of the corresponding operation and the historical information updated based on the corresponding operation;

[0219] The correction unit 703 is further used to compare the trigger position and updated historical information of the standard input and output based on each operation with the trigger position and updated historical information contained in the respective standard output.

[0220] Optionally, the standard output of any of the operations further includes: an operation description of any of the operations; the correction unit 703 is further used for:

[0221] The trigger positions and operation descriptions based on the standard inputs and outputs of each operation, along with the trigger positions and operation descriptions contained in their respective standard outputs, are input into the loss function corresponding to the model to be trained, so as to iteratively correct the model to be trained based on the feedback of the loss function.

[0222] The loss function includes a location prediction subfunction and a description generation subfunction; the loss weight of the location prediction subfunction is higher than the loss weight of the description generation subfunction.

[0223] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0224] Accordingly, this disclosure also provides a training apparatus for a human-computer interaction model, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement a training method for a human-computer interaction model as described in any of the above embodiments, for example, the method may include: acquiring a target dataset of sample events; the sample events include multiple operations, and the target dataset includes standard inputs and standard outputs for each operation; wherein the standard input of any operation includes at least one of an operation description and an operation result of the previous operation, a user instruction for triggering the sample event, and a trigger position of the previous operation; the standard output of any operation includes: the trigger position of the any operation; providing the standard inputs of each operation to the model to be trained, so that the model to be trained outputs the trigger position of the corresponding operation; comparing the trigger position based on the standard input and output of each operation with the trigger position contained in the respective standard output, so as to iteratively correct the model to be trained according to the comparison result, until the output obtained based on the standard input of the operation contained in the sample event meets the preset iteration condition in the comparison result with the standard output of the corresponding operation.

[0225] Accordingly, this disclosure also provides an electronic device, which includes a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors. The programs contain instructions for implementing a training method for a human-computer interaction model as described in any of the above embodiments. For example, the method may include: acquiring a target dataset of sample events; the sample events include multiple operations, and the target dataset includes standard inputs and standard outputs for each operation; wherein the standard input of any operation includes at least one of an operation description and an operation result of the previous operation, a user instruction for triggering the sample event, and a trigger position of the previous operation; the standard output of any operation includes the trigger position of the operation; providing the standard inputs of each operation to a model to be trained, so that the model to be trained outputs the trigger position of the corresponding operation; comparing the trigger positions based on the standard inputs and outputs of each operation with the trigger positions contained in the respective standard outputs, and iteratively correcting the model to be trained according to the comparison results, until the output obtained based on the standard inputs of the operations contained in the sample events meets the preset iteration conditions when compared with the standard outputs of the corresponding operations.

[0226] Figure 8 This is a block diagram illustrating an apparatus 800 for implementing a training method for a human-computer interaction model, according to an exemplary embodiment. For example, apparatus 800 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.

[0227] Reference Figure 8 The device 800 may include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.

[0228] Processing component 802 typically controls the overall operation of device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 802 may include one or more processors 820 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.

[0229] Memory 804 is configured to store various types of data to support the operation of device 800. Examples of such data include instructions for any application or method operating on device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0230] Power supply component 806 provides power to various components of device 800. Power supply component 806 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to device 800.

[0231] Multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0232] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.

[0233] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0234] Sensor assembly 814 includes one or more sensors for providing status assessments of various aspects of device 800. For example, sensor assembly 814 may detect the on / off state of device 800, the relative positioning of components such as the display and keypad of device 800, changes in the position of device 800 or a component of device 800, the presence or absence of user contact with device 800, the orientation or acceleration / deceleration of device 800, and temperature changes of device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.

[0235] Communication component 816 is configured to facilitate wired or wireless communication between device 800 and other devices. Device 800 can access wireless networks based on communication standards, such as WiFi, 2G or 3G, 4G LTE, 5G NR (New Radio), or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0236] In an exemplary embodiment, the apparatus 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0237] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, which can be executed by a processor 820 of the device 800 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0238] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0239] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

[0240] The above description is merely a preferred embodiment of this disclosure and is not intended to limit this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. An event execution method, characterized in that, Upon receiving a user instruction to execute a target event, iteratively perform the following operations until the target event is completed: Based on at least one of the operation description and operation result of the previous operation, the trigger position of the previous operation, and the user instruction, the operation prediction is performed to obtain the trigger position of the next operation. The current interface is triggered based on the trigger position of the next operation, and the next operation is updated to the previous operation in the next iteration. In addition, record the trigger position of the updated previous operation, and generate at least one of the operation description and operation result of the updated previous operation.

2. The method according to claim 1, characterized in that, The operation prediction based on at least one of the operation description and operation result of the previous operation, the triggering position of the previous operation, and the user instruction includes: operation prediction based on the user instruction and historical information; the historical information includes: at least one of the operation description and operation result of each executed operation, and the triggering position of each executed operation; It also includes updating the historical information based on at least one of the updated operation description and operation result of the previous operation, and the recorded trigger position.

3. The method according to claim 2, characterized in that, The historical information includes: the trigger location of each executed operation, the operation description of each executed operation, and the operation result of each executed operation; The operation prediction based on the user instructions and historical information includes: performing operation prediction based on the user instructions and historical information to obtain the trigger position of the next operation and the operation description of the next operation; The step of updating the historical information based on at least one of the updated operation description and operation result of the previous operation and the recorded trigger position includes: updating the historical information based on the updated operation description and operation result of the previous operation and the recorded trigger position.

4. The method according to claim 3, characterized in that, The step of predicting the operation based on the user instruction and the historical information to obtain the trigger position of the next operation and the operation description of the next operation includes: predicting the operation based on the user instruction, the historical information and the historical summary used to summarize each executed operation to obtain the trigger position of the next operation, the operation description of the next operation, and the historical summary updated based on the next operation. It also includes: after updating the next operation to the previous operation in the next iteration, recording the updated history summary.

5. The method according to claim 1, wherein predicting the trigger position of the next operation based on at least one of the operation description and operation result of the previous operation, the trigger position of the previous operation, and the user instruction comprises: The human-computer interaction model is pre-trained and uses at least one of the operation description and operation result of the previous operation, the trigger position of the previous operation, and the user command input to output the trigger position of the next operation.

6. A training method for a human-computer interaction model, characterized in that, include: Obtain the target dataset of sample events; The sample event comprises multiple operations, and the target dataset comprises standard inputs and standard outputs for each operation; wherein, the standard input for any operation includes at least one of the operation description and operation result of the previous operation, the user instruction for triggering the sample event, and the trigger position of the previous operation; the standard output for any operation includes the trigger position of the operation. The standard inputs for each operation are provided to the model to be trained, so that the model to be trained outputs the trigger position of the corresponding operation; The trigger positions of the standard input and output of each operation are compared with the trigger positions contained in the respective standard output. The model to be trained is iteratively corrected according to the comparison results until the output obtained based on the standard input of the operation contained in the sample event meets the preset iteration conditions when the comparison results with the standard output of the corresponding operation are satisfied.

7. The method according to claim 6, characterized in that, The standard inputs for any of the operations include: the user instructions and historical information; The historical information includes: at least one of the operation description of each executed operation and the operation result of each executed operation, and the triggering position of each executed operation.

8. The method according to claim 7, characterized in that, The historical information includes: the trigger location of each executed operation, the operation description of each executed operation, and the operation result of each executed operation; The standard output of any operation also includes: an operation description of any operation; The step of providing the standard input of each operation to the model to be trained so that the model to be trained can output the trigger position of the corresponding operation includes: providing the standard input of each operation to the model to be trained so that the model to be trained can output the trigger position of the corresponding operation and the operation description of the corresponding operation; The step of comparing the trigger positions of the standard inputs and outputs based on each operation with the trigger positions contained in the respective standard outputs includes: comparing the trigger positions and operation descriptions of the standard inputs and outputs based on each operation with the trigger positions and operation descriptions contained in the respective standard outputs.

9. The method according to claim 8, characterized in that, The standard input for any operation also includes: a historical summary of all executed operations; The standard output of any operation also includes: an updated historical summary based on any operation; The step of providing the standard input of each operation to the model to be trained so that the model to be trained can output the trigger position of the corresponding operation and the operation description of the corresponding operation, includes: providing the standard input of each operation to the model to be trained so that the model to be trained can output the trigger position of the corresponding operation, the operation description of the corresponding operation, and the historical summary updated based on the corresponding operation. The step of comparing the trigger position and operation description of the standard input and output of each operation with the trigger position and operation description contained in the respective standard output includes: comparing the trigger position, operation description, and historical summary of the standard input and output of each operation with the trigger position, operation description, and historical summary contained in the respective standard output.

10. The method according to claim 8, characterized in that, It also includes: triggering the corresponding position on the terminal screen based on the trigger position of each operation output by the model to be trained, so as to obtain the operation result of the corresponding operation; and comparing the operation result obtained based on the trigger position of each operation output by the model to be trained with the operation result contained in the standard input of the next operation of the corresponding operation. The step of iteratively correcting the model to be trained based on the comparison results includes: iteratively correcting the model to be trained based on the comparison results of the trigger position, the operation description, and the operation result.

11. The method according to claim 7, characterized in that, The standard output of any operation also includes: historical information updated based on the operation. The step of providing the standard input of each operation to the model to be trained so that the model to be trained can output the trigger position of the corresponding operation includes: providing the standard input of each operation to the model to be trained so that the model to be trained can output the trigger position of the corresponding operation and the historical information updated based on the corresponding operation; The step of comparing the trigger positions of the standard inputs and outputs based on each operation with the trigger positions contained in the respective standard outputs includes: comparing the trigger positions of the standard inputs and outputs based on each operation with the updated historical information contained in the respective standard outputs.

12. The method according to claim 6, characterized in that, The standard output of any operation further includes: an operation description of any operation; the step of comparing the trigger positions based on the standard input and output of each operation with the trigger positions contained in their respective standard outputs, and iteratively correcting the model to be trained based on the comparison results, includes: The trigger positions and operation descriptions based on the standard inputs and outputs of each operation, along with the trigger positions and operation descriptions contained in their respective standard outputs, are input into the loss function corresponding to the model to be trained, so as to iteratively correct the model to be trained based on the feedback of the loss function. The loss function includes a location prediction subfunction and a description generation subfunction; the loss weight of the location prediction subfunction is higher than the loss weight of the description generation subfunction.

13. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor implements the method as described in any one of claims 1-12 by executing the executable instructions.

14. A computer-readable storage medium storing computer instructions thereon, characterized in that, When executed by the processor, this instruction implements the steps of the method as described in any one of claims 1-12.