Action decision based on self-tune mechanism
The self-tune mechanism for action decision in cloud computing resource management addresses inefficiencies by generating and optimizing actions in natural language and code, reducing human intervention and enhancing decision-making efficiency.
Patent Information
- Application Number
- PCT/US2024/055288
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-08
- Filing Date
- 2024-11-10
- Publication Date
- 2025-06-12
AI Technical Summary
Current methods for managing cloud computing resources require extensive human intervention and are inefficient, as they often necessitate defining different rules or designing machine learning algorithms for various scenarios or applications.
A self-tune mechanism for action decision that generates actions in natural language for a target application based on previous states and rewards, verifies the reasonableness of these actions using predefined rules, and adjusts hyperparameters or prompts to optimize action generation.
This approach reduces the need for human intervention, improves the quality and efficiency of action decision-making, and enables real-time optimization of actions and their corresponding code for cloud computing resource management.
Smart Images

Figure US2024055288_12062025_PF_FP_ABST
Abstract
Description
ACTION DECISION BASED ON SELF-TUNE MECHANISMBACKGROUND
[0001] Cloud computing is a technology widely used by modem businesses and organizations. Cloud computing is a network-based computing model that provides users with computing resources, storage resources, network resources, etc., through the Internet. Users can autonomously select, configure and use the required resources through the network. Cloud computing is based on virtualization technology. By creating virtual machines that can run various applications on physical machines, users can conveniently use cloud computing services.SUMMARY
[0002] This Summary is provided to introduce a selection of concepts that are further described below in the Detailed Description. It is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0003] Embodiments of the present disclosure propose a method, apparatus and computer- readable medium for action decision based on self-tune mechanism. A set of previous states associated with a target application and a set of previous rewards corresponding to the set of previous states may be obtained. An action in natural language for the target application may be generated based on the set of previous states and the set of previous rewards. It may be verified with a set of predefined rules whether the action is reasonable. Action code in computer language corresponding to the action may be generated in response to verifying that the action is reasonable. The target application may be caused to execute the action code.
[0004] It should be noted that the above one or more aspects comprise the features hereinafter fully described and particularly pointed out in the claims. The following description and the drawings set forth in detail certain illustrative features of the one or more aspects. These features are only indicative of the various ways in which the principles of various aspects may be employed, and this disclosure is intended to include all such aspects and their equivalents.BRIEF DESCRIPTION OF THE DRAWINGS
[0005] The disclosed aspects will hereinafter be described in connection with the appended drawings that are provided to illustrate and not to limit the disclosed aspects.
[0006] FIG. 1 illustrates an exemplary process for action decision based on self-tune mechanism according to an embodiment of the present disclosure.
[0007] FIG. 2 illustrates an exemplary process for determining a permissible action and an observation indicator according to an embodiment of the present disclosure.
[0008] FIG. 3 illustrates an exemplary process for generating an action in natural language fora target application according to an embodiment of the present disclosure.
[0009] FIG. 4 illustrates an exemplary process for calculating a reward according to an embodiment of the present disclosure.
[0010] FIG. 5 is a flowchart of an exemplary method for action decision for a target application according to an embodiment of the present disclosure.
[0011] FIG. 6 illustrates an exemplary apparatus for action decision for a target application according to an embodiment of the present disclosure.
[0012] FIG. 7 illustrates another exemplar,' apparatus for action decision for a target application according to an embodiment of the present disclosure.DETAILED DESCRIPTION
[0013] The present disclosure will now be discussed with reference to several example implementations. It is to be understood that these implementations are discussed only for enabling those skilled in the art to better understand and thus implement the embodiments of the present disclosure, rather than suggesting any limitations on the scope of the present disclosure.
[0014] The platform for providing cloud computing sendees may be referred to as a cloud computing platform. It is desirable to rationally manage resources on the cloud computing platform to better utilize limited resource capacity. Resources on the cloud computing platform may include, e.g., computing resource, storage resource, network resource, etc. These resources may be collectively referred to as cloud computing resources. The cloud computing resources may be managed through various applications, e.g., a Workload Management (WLM) application for managing the allocation of resources among different workloads, a predictive maintenance application for predicting potential machine failures and recommending maintenance measures, a network traffic management application for analyzing network traffic patterns to ensure optimal network performance, etc. These applications may perform corresponding actions to manage the cloud computing resources from different aspects. Since the system environment or user requirements are changing dynamically, the actions performed by each application may also need to be updated in real time. A decision about an action to be performed by an application may be made through a set of predefined rules or manually -designed machine learning algorithms. These methods usually require defining different rules or designing different machine learning algorithms for different scenarios or applications, which consumes a lot of human resources and is inefficient.
[0015] Embodiments of the present disclosure propose action decision based on self-tune mechanism. A set of permissible actions that a target application can perform and at least one observation indicator that should be watched during the operation of the target application may be predetermined. Herein, a target application refers to an application for which an action decision ismade. After determining the permissible action and the observation indicator, a decision about an action to be performed by the target application may be made. An action in natural language for the target application may be generated based on a set of previous states associated with the target application and a set of previous rewards corresponding to the set of previous states through a language model. Herein, a language model refers to a deep learning model that can understand the meaning of natural language, generate natural language texts, or perform other natural language tasks. It should be appreciated that language models include a multi-modal model that can perform processing tasks for natural language as well as other modalities. After the action in natural language is generated, it may be verified whether the action is reasonable through a heuristic method, e.g., with a set of predefined rules. The set of predefined rules may include: a rule for evaluating whether the action is included in a predetermined set of permissible actions, and / or a rule for evaluating whether the action is meaningless or untrustworthy. If it is verified that the action is unreasonable, a hyperparameter of the language model may be adjusted and / or a prompt provided to the language model may be modified, and an action may be regenerated. If it is verified that the action is reasonable, the action may be converted into action code in computer language. The target application may be caused to execute the action code, and a corresponding state may be collected. This state may be used together with the set of previous states and the set of previous rewards that obtained before generating the action, to calculate a reward corresponding to the action. Subsequently, a subsequent action in natural language for the target application may be generated based on some or all of the previously obtained states and rewards.
[0016] The operations following the operations for determining the permissible action and the observation indicator in the above process may be repeatedly performed in an online manner, so that the state generated after the target application performs each action may be continuously monitored, and the action and corresponding action code for the target application may be updated or optimized in real time through combining with previously collected states and previously calculated rewards. The above process reduces the need for human intervention, and aims to achieve action decision through self-tuning. This improves the quality and efficiency of action decision.
[0017] In addition, the predetermined set of permissible actions that the target application can perform may be used to verify whether the generated action is reasonable. The predetermined at least one observation indicator that should be watched during the operation of the target application may be used to determine the state corresponding to which indicator should be collected after the action code is executed. Permissible actions and observation indicators associated with different target applications may be different. By predetermining the permissible action and the observation indicator associated with the target application, the quality andefficiency of action decision for the target application may be further improved. This also makes the action decision process according to the embodiments of the present disclosure widely applicable to various applications requiring action decision, thereby implementing a unified action decision scheme.
[0018] In addition, in the above process, after the action in natural language for the target application is generated, the reasonableness of the actions is verified. In the case where it is verified that the action is reasonable, the action in natural language is converted into the action code in computer language, and the action code is then applied to the target application. This approach can avoid an unreasonable action from being performed and improve the robustness of action decision. In addition, the way of decomposing the action generation into first generating an action in natural language that is easy for humans to understand and then generating action code in computer language helps improve the accuracy of action generation and enables the use of a heuristic method to verify the reasonableness of the action.
[0019] The foregoing discussion and the following discussion may involve examples of action decision for an application used for cloud computing resource management. However, it should be appreciated that the embodiments of the present disclosure are not limited thereto, but action decision for applications in other fields may be performed in a similar manner. For example, for an energy’ management application used to optimize energy consumption and ensure stable energy supply, the solution of the embodiments of the present disclosure may also be used to make decision about an action to be performed by this application.
[0020] Various embodiments of the present disclosure will hereinafter be described in detail in connection with the appended drawings.
[0021] FIG. 1 illustrates an exemplary process 100 for action decision based on self-tune mechanism according to an embodiment of the present disclosure. The process 100 may be used to make decisions about an action to be performed by a target application 102. The target application 102 may be various applications that require action decision, e.g., a workload management application for managing the allocation of resources among different workloads, a predictive maintenance application for predicting potential machine failures and recommending maintenance measures, a network traffic management application for analyzing network traffic patterns to ensure optimal network performance, an energy management application for optimizing energy consumption and ensuring stable energy supply, etc. The process 100 may be performed online.
[0022] A task design specification 104 and / or an indicator description 106 associated with the target application 102 may be obtained. The task design specification 104 may be text content or a file that describe a task performed by the target application 102. The indicator description 106may be text content or a file that explains indicators related to the target application 102. The task design specification 104 and the indicator description 106 may be two separate files or a same file. The task design specification 104 and / or the indicator description 106 may be preprocessed through a preprocessing component 110, to generate a task description 112 and / or a set of candidate indicators 114 for the target application 102. The preprocessing component 110 is. e.g., a machine learning model capable of performing a text summarization task. Taking the target application 102 being a workload management application as an example, the task description 112 for the target application 102 generated through the preprocessing component 110 may be “The workload management application is responsible for managing the allocation of resources among different workloads. A corresponding priority is assigned to a workload based on the workload’s delay sensitivity. For example, sending an email should be delay sensitive, while backing up data might be delay insensitive. The workload management application provides a workload state signal to the client to increase or decrease the resources allocated to the workload at that client...'’ Similarly, taking the target application 102 being a workload management application as an example, the set of candidate indicators 1 14 for the target application 102 generated through the preprocessing component 110 may be: network smoothness, disk queue depth, network source address, network destination address, disk read latency, disk write latency, etc. These indicators may also be referred to as performance counters.
[0023] Subsequently, a set of permissible actions 122 that the target application 102 can perform and / or at least one observation indicator 124 that should be watched during the operation of the target application 102 may be determined based on the task description 112 and / or the set of candidate indicators 114 through a parameter determining component 120. The permissible action 122 and the observation indicator 124 may be determined through a language model. The language model may be a language model capable of generating an action in natural language and a recommended indicator, e.g., a Generative Pre-trained Transformer-4 (GPT-4) model, etc. An exemplary process for determining the permissible action 122 and the observation indicator 124 will be described below in conjunction with FIG. 2.
[0024] After the permissible action 122 and / or the observation indicator 124 is determined, a decision about an action to be performed by the target application 102 may be made through an action decision agent 130.
[0025] The action decision agent 130 may comprise an action planning component 140 for generating an action in natural language for the target application 102. For example, a set of previous states associated with the target application 102 and a set of previous rewards corresponding to the set of previous states may be obtained. The set of previous states may include states collected before the action planning component 140 performs the action generationoperation. The set of previous rewards may include rewards calculated before the action planning component 140 performs the action generation operation. The reward may be used to evaluate the quality of an action generated by the action planning component 140. The reward may be previously calculated by a reward evaluating component 190. Initially, e.g., when the action decision agent 130 or the action planning component 140 is just started, a reward may be randomly produced. The reward may be in numerical value. Since the state is collected from the environment outside the action decision agent 130, the state may also be referred to as external feedback. In contrast, since the reward is calculated by the reward evaluating component 190 inside the action decision agent 130, the reward may also be referred to as internal feedback. An action in natural language for the target application 102 may be generated based on the set of previous states and the set of previous rewards. The action may be generated through a language model. The language model may be a language model capable of generating an action in natural language, e.g., a GPT- 4 model, etc. An exemplary process for generating the action in natural language for the target application 102 will be described later in conjunction with FIG. 3.
[0026] The action decision agent 130 may comprise an action verifying component 150 for verifying the reasonableness of the action generated by the action planning component 140. The action verifying component 150 may comprise a short-term memory 152 for storing actions newly generated by the action planning component 140, etc. It may be verified whether the action is reasonable through a heuristic method, e g., with a set of predefined rules. The set of predefined rules may include a rule for evaluating whether the action is included in a predetermined set of permissible actions 122. The predetermined set of permissible actions 122 defines the action space that the target application can perform. If the action generated by the action planning component 140 is not included in the set of permissible actions 122, the action may be verified as unreasonable. Alternatively or additionally, the set of predefined rules may include a rule for evaluating whether the action is meaningless or untrustworthy. Such a rule is intended to determine whether the action has the hallucination issue. A generative language model sometimes produce a meaningless or untrustworthy output, such an issue referred to as the hallucination issue. It may be determined whether the action is meaningless or untrustworthy through evaluating whether it is inconsistent, contains incorrect information, etc. If the action generated by the action planning component 140 is determined to be meaningless or untrustworthy, the action may be verified as unreasonable. The rules described above may be used individually or in combination with each other. It should be appreciated that the rules described above for verifying whether the action is reasonable are merely exemplary. Depending on actual application requirements, other rules may also be used to verify the reasonableness of the action.
[0027] If it is verified that the action generated by the action planning component 140 isunreasonable, the action planning process may be optimized, and an action may be regenerated. In an implementation, a hyperparameter of the language model involved in the action planning component 140 may be adjusted. For example, the temperature parameter of the language model may be increased so that the output it produces is more rigorous. In another implementation, a prompt provided to the language model may be modified. For example, the length of a responses from the language model may be limited or the language model may be required to respond based on known facts. The implementations described above may be implemented individually or in combination with each other. It should be appreciated that the approaches described above for optimizing the action planning process are merely exemplary. Depending on actual application requirements, other approaches may be employed to optimize the action planning process to enhance the reasonableness of the generated action in natural language. Subsequently, an action in natural language for the target application 102 may be regenerated based on the set of previous states and the set of previous rewards through the language model. Preferably, if a reasonable action still cannot be obtained within a predetermined time or after the action planning component 140 performs a predetermined number of action generation operations, a default action may be invoked.
[0028] If it is verified that the action generated by the action planning component 140 is reasonable, the action may be used as an action 142, and the subsequent process may be performed. For example, action code 162 in computer language corresponding to the action 142 may be generated through a code generation component 160 in the action decision agent 130. The action code 162 in computer language may be generated through a language model. The language model may be a language model capable of generating program code, e.g., a GPT-4 model, a Code Llama model, etc. Before this, a prompt to be provided to the language model may be created first. Preferably, a prompt template associated with generating the action code may be designed in advance. The prompt template may include a part for loading an action. When creating a prompt, the action 142 generated by the action planning component 140 may be loaded into the corresponding part in the prompt template. In addition, the prompt template may include a response instruction for instructing the language model on how- to respond. As an example, a response instruction could be "I will provide an action plan. Please convert it into corresponding program code. Example code: import numpy as np; set_disk_latency (0.1)... ’' It should be appreciated that the response instruction described above is merely exemplary. Depending on actual application requirements, the response instruction in the prompt template associated with generating action code may has other forms and may include more or fewer content.
[0029] The target application 102 may be caused to execute the action code 162. Subsequently, a state 172 produced through the execution of the action code 162 by the target application 102may be collected through a monitoring component 170. The state 172 may correspond to the previously determined observation indicator 124. Operation and maintenance data produced after the target application 102 executing the action code 162 may be collected through the monitoring component 170. The operation and maintenance data may include log, monitoring information, application information, etc. Data corresponding to the observation indicator 124 may be extracted from the operation and maintenance data as the state 172. The operation and maintenance data may include data collected from different machines or at different time intervals. Preferably, when extracting the state 172, these data may be aggregated based on actual application requirements. In addition, the operation and maintenance data may include some meaningless noise data. Preferably, these noisy data may be filtered out from the operation and maintenance data when extracting the state 172.
[0030] After the state 172 is collected, it may be stored in a long-term memory 180 associated with the target application 102 for subsequent action decisions, e.g., for use as a historical state when generating a subsequent action. The long-term memoiy 180 may have stored a previously obtained set of historical states 182 and a set of historical rewards 184.
[0031] Next, a reward 192 corresponding to the action 142 may be calculated through a reward evaluating component 190 in the action decision agent 130. The reward 192 may be used to evaluate the uality of the action 142. The reward 192 may be in numerical value. The reward 192 corresponding to the action 142 may be calculated based on one or more of the action 142, the state 172, the set of historical states 182, and the set of historical rewards 184. ft should be appreciated that the set of historical states 182 may correspond to a set of previous states obtained when the action 142 is generated, and the set of historical rewards 184 may correspond to a set of previous rewards obtained when the action 142 is generated. The reward may be calculated through a language model. The language model may be a language model capable of scoring an input, e.g., the GPT-4 model, the Vicuna model, etc. An exemplary process for calculating the reward will be described later in conjunction with FIG. 4. The reward 192 may be stored in the long-term memory 180 for subsequent action decisions, e.g., for use as a historical reward when generating a subsequent action.
[0032] Subsequently, a subsequent action in natural language for the target application 102 may be generated based on one or more of the newly collected state 172. the newly calculated reward 192, and the set of historical states 182 and the set of historical rewards 184 stored in the long-term memory 180.
[0033] The operations following the operations for determining the permissible action 122 and the observation indicator 124 in the process 100 may be repeatedly performed in an online manner, so that the state generated after the target application 102 performs each action may becontinuously monitored, and the action and corresponding action code for the target application 102 may be updated or optimized in real time through combining with previously collected states and previously calculated rewards. The above process reduces the need for human intervention, and aims to achieve action decision through self-tuning. This improves the quality and efficiency of action decision.
[0034] In process 100, the set of permissible actions 122 that the target application 102 can perform and / or the at least one observation indicator 124 that should be watched during the operation of the target application 102 is predetermined through the preprocessing component 110 and the parameter determining component 120 on the right. The permissible action 122 may be used to verify whether the action generated by the action planning component 140 in the action decision agent 130 is reasonable. The observation indicator 124 may be used to determine the state corresponding to which indicator should be collected after the action code 162 is executed. Permissible actions and observation indicators associated with different target applications may be different. By predetermining the permissible action and the observation indicator associated with the target application, the qualify and efficiency of action decision for the target application may be further improved. This also makes the process 100 widely applicable to various applications requiring action decision, thereby implementing a unified action decision scheme.
[0035] In addition, in the process 100. after the action planning component 140 generates the action in natural language for the target application 102, the action verifying component 150 verifies whether the action is reasonable. In the case where it is verified that the action is reasonable, the action in natural language is converted into the action code in computer language, and the action code is then applied to the target application 102. This approach can avoid an unreasonable action from being performed and improve the robustness of action decision. In addition, the way of decomposing the action generation into first generating an action in natural language that is easy for humans to understand and then generating action code in computer language helps improve the accuracy of action generation and enables the use of a heuristic method to verify the reasonableness of the action.
[0036] It should be appreciated that the process for action decision based on self-tune mechanism described above in conjunction with FIG. 1 is merely exemplary . Depending on actual application requirements, the steps in the process for action decision may be replaced or modified in any manner, and the process may comprise more or fewer steps. For example, in the process 100, the step for determining the permissible action 122 and / or the observation indicator 124 is described. This step is not mandatory. In the case where the step for determining the permissible action 122 and / or the observation indicator 124 is not performed, subsequent steps may not consider the permissible action 122 and / or the observation indicator 124. In addition, the specificorder or hierarchy of the steps in the process 100 is merely exemplary, and the process for action decision may be performed in an order different from the described order.
[0037] FIG. 2 illustrates an exemplary' process 200 for determining a permissible action and an observation indicator according to an embodiment of the present disclosure. The process 200 may correspond to the operation at the parameter determining component 120 in FIG. 1. In the process 200, a permissible action 222 and an observation indicator 224 may be generated through a language model 220. The language model 220 may be a language model capable of generating an action in natural language and a recommendation indicator, e g., a GPT-4 model, etc. Before this, a prompt 212 to be provided to the language model 220 may be created first through a prompt creator 210.
[0038] A task description 202 and candidate indicators 204 of a target application may be obtained. The task description 202 and the candidate indicators 204 may correspond to the task description 112 and the candidate indicator 114 in FIG. 1. respectively.
[0039] The prompt creator 210 may create the prompt 212 based on the task description 202 and / or the candidate indicators 204 of the target application. Preferably, a prompt template 206 for the prompt creator 210 may be designed in advance. The prompt template 206 may include multiple parts for loading a task description and a candidate indicator, respectively. When creating the prompt 212, the task description 202 and the candidate indicators 204 of the target application may be loaded into the corresponding part in the prompt template 206, respectively. In addition, the prompt template 206 may include a response instruction for instructing the language model 220 on how to respond. Taking the target application being a workload management application as an example, the response instruction may be "I will provide a task description and a list of performance counters. I want you to generate appropriate actions and select appropriate performance counters to achieve the optimization goal..." It should be appreciated that the response instruction described above is merely exemplary. Depending on actual application requirements, the response instruction in the prompt template 206 may have other forms and may include more or fewer content.
[0040] The prompt 212 may be provided to the language model 220. The language model 220 may utilize its semantic understanding capability7, logical reasoning capability7, language expression capability, big data support capability, etc., to generate a set of permissible actions 222. Taking the target application being a workload management application as an example, a permissible action may be increasing disk read latency, reducing disk write latency, increasing input / output per second (IOPS), reducing IOPS, etc. Alternatively or additionally, the language model 220 may select at least one observation indicator 224 from the set of candidate indicators 204 based on the prompt 212.
[0041] Preferably, the language model 220 may generate an indicator code 226 in computer language corresponding to the observation indicator 224. The generated code may be stored, and invoked and executed when, e.g., the target application executes an action code or collects a state of the target application. In the case where the language model 220 generates the indicator code 226, the response instruction in the prompt template 206 may include an instruction for generating the indicator code, e g., “Convert the selected performance counters into program code with the following function: perf_counter. get_value(counter_name: str) . ”.
[0042] The permissible action 222, the observation indicator 224, and / or the indicator code 226 may be generated simultaneously through a single interaction with the language model 220. Alternatively, the permissible action 222, the observation indicator 224, and / or the indicator code 226 may be generated at different times through multiple interactions with the language model 220.
[0043] It should be appreciated that the process for determining the permissible action and the observation indicator described above in connection with FIG. 2 is merely exemplary. Depending on actual application requirements, the steps in the process for determining the permissible action and the observation indicator may be replaced or modified in any manner, and the process may comprise more or fewer steps. For example, in the process 200, the prompt 212 is created based on the task description 202. the candidate indicators 204, and the prompt template 206, but the embodiments of the present disclosure are not limited thereto. In some embodiments, only one or two of the task description 202, the candidate indicators 204, and the prompt template 206 may be considered. In the absence of the prompt template 206, the prompt 212 may be created through combining one or more of the task description 202, the candidate indicators 204, and the response instruction. Additionally, in the process 200, the language model 220 generates the permissible action 222, the observation indicator 224, and / or the indicator code 226, but the embodiments of the present disclosure are not limited thereto. In some embodiments, the language model 220 may generate only one or two of the permissible action 222, the observation indicator 224. and the indicator code 226.
[0044] FIG. 3 illustrates an exemplary process 300 for generating an action in natural language for a target application according to an embodiment of the present disclosure. The process 300 may correspond to the operation at the action planning component 140 in FIG. 1. In the process 300, an action 332 in natural language may be generated through a language model 330. The language model 330 may be a language model capable of generating an action in natural language, e.g., the GPT-4 model, etc. Before this, a prompt 322 to be provided to the language model 330 may be created first through a prompt creator 320.
[0045] A task description 302 for the target application may be obtained. The task description302 may correspond to the task description 112 in FIG. 1. Considering the task description of the target application when generating an action for the target application can make the generated action more accurate.
[0046] A set of previous states associated with the target application and a set of previous rewards corresponding to the set of previous states may be obtained. The set of previous states may include a newly collected recent state 304 and at least one historical state 310 stored in a long-term memory. The recent state 304 may be collected after the target application executes action code corresponding to a newly generated action. The long-term memon' is, e.g., the longterm memon 180 in FIG. 1. The set of previous rewards may include a newly calculated recent reward 306 and at least one historical reward 308 stored in a long-term memory. The recent reward 306 is a reward corresponding to the recent state 304. The historical reward 308 is a reward corresponding to the historical state 310. Initially, the recent state 304 and / or the recent reward 306 may be a randomly generated state and / or reward. Moreover, the historical state 310 and the historical reward 308 may not exist.
[0047] The prompt creator 320 may create the prompt 322 based on one or more of the task description 302, the set of previous states, and the set of previous rewards associated with the target application. Preferably, a prompt template 312 for the prompt creator 320 may be designed in advance. The prompt template 312 may include multiple parts for loading a task description, a previous state, and a previous reward, respectively. When creating the prompt 322, the task description 302, the set of previous states, and the set of previous rewards associated with the target application may be loaded into the corresponding part in the prompt template 312 respectively. In addition, the prompt template 312 may include a response instruction for instructing the language model 330 on how to respond. Taking the target application be a workload management application as an example, the response instruction may be "I will provide a task description, a set of previous states, and a set of previous rewards corresponding to the set of previous states. Please analyze the relationship between the previous states and their corresponding previous rewards first, and use the following format to provide a subsequent action that is able to increase the reward: 1. Increase disk latency by 1.0; 2. Decrease disk latency by 3.0. Please provide only an action list, and the action list is in Markdown format..." It should be appreciated that the response instruction described above is merely exemplary. Depending on actual application requirements, the response instruction in the prompt template 312 may have other forms and may include more or fewer content.
[0048] The prompt 322 may be provided to the language model 330. The language model 330 may utilize its semantic understanding capability, logical reasoning capability, language expression capability, big data support capability, etc., to analyze the relationship between the setof previous states and the set of previous rewards included in the prompt 322, and further generate an action 332 in natural language. As an example, the generated action 332 in natural language may be "1. Reduce disk latency by 1.0; 2. Increase IOPS by 3.0."
[0049] It should be appreciated that process for generating the action in natural language for the target application described above in conjunction with FIG. 3 is merely exemplary. Depending on actual application requirements, the steps in the process for generating the action in natural language may be replaced or modified in any manner, and the process may comprise more or fewer steps. For example, in the process 300, the prompt 322 is created based on the task description 302, the previous states, the previous rewards, and the prompt template 312, but the embodiments of the present disclosure are not limited thereto. In some embodiments, only one or two of the task description 302, the previous states, the previous rewards, and the prompt template 312 may be considered. In the absence of the prompt template 312, the prompt 322 may be created through combining one or more of the task description 302, the previous states, the previous rew ards, and the response instruction.
[0050] FIG. 4 illustrates an exemplary process 400 for calculating a reward according to an embodiment of the present disclosure. The process 400 may correspond to the operation at the reward evaluating component 190 in FIG. 1. In the process 400, a reward 432 corresponding to a recent action 404 for a target application may be calculated through a language model 430. The language model 430 may be a language model capable of scoring an input, e g., the GPT-4 model, the Vicuna model, etc. Before this, a prompt 422 to be provided to the language model 430 may be created first through a prompt creator 420. Referring back to FIG. 1, a recent action 404 may correspond to the action 142 in FIG. 1.
[0051] A task description 402 for the target application may be obtained. The task description 402 may correspond to the task description 112 in FIG. 1. Considering the task description of the target application when calculating the rew ard associated with the target application can make the calculated reward more accurate.
[0052] A recent state 406 corresponding to the recent action 404 may be obtained. Referring back to FIG. 1, the recent state 406 may correspond to the state 172 in FIG. 1, which may be obtained after the target application 102 executes the action code 162 corresponding to the action 142. A set of historical states 408 associated with the target application may be obtained. The set of historical states 408 may be stored in a long-term memory. The long-term memory is. e.g., the long-term memory 180 in FIG. 1. The historical states 408 may correspond to the historical state 182 in FIG. 1. The recent state 406 and the historical states 408 may be combined into previous states. Previous rewards may be obtained. The previous rewards may include a set of historical rewards 410 corresponding to the set of historical states 408. The set of historical rewards 410may be stored in a long-term memory. The long-term memory is, e.g., the long-term memory 180 in FIG. 1. The historical rewards 410 may correspond to the historical rewards 184 in FIG. 1. Initially, the historical states 408 and the historical rewards 410 may not exist.
[0053] The prompt creator 420 may create the prompt 422 based on one or more of the task description 402 of the target application, the recent action 404. the previous states including the recent state 406 and the historical states 408, and the previous rewards including the historical rewards 410. Preferably, a prompt template 412 for the prompt creator 420 may be designed in advance. The prompt template 412 may include multiple parts for loading a task description, a recent action, previous states, and previous rewards, respectively. When creating the prompt 422, the task description 406 of the target application, the recent action 406, the previous states including the recent state 406 and the historical states 408, and the previous rewards including the historical rewards 410 may be loaded into the corresponding part in the prompt template 412, respectively. In addition, the prompt template 412 may include a response instruction for instructing the language model 430 on how to respond. Taking the target application being a workload management application as an example, the response instruction may be "I will provide a task description, a recent action, previous states and previous rewards. Please provide a reward corresponding to the action. Please remember: 1. only return the reward: 2. the value of the reward should be between 1 and 10 ..." It should be appreciated that the response instruction described above is merely exemplary. Depending on actual application requirements, the response instruction in the prompt template 412 may have other forms and may include more or fewer content.
[0054] The prompt 422 may be provided to the language model 430. The language model 430 may utilize its semantic understanding capability, logical reasoning capability, language expression capability, big data support capability, etc., to analyze the relationships between the set of previous states and the set of previous rewards included in the prompt 422, and further generate a reward 432 for the recent action 404. Optionally, the language model 430 may provide an interpretation for the reward 432 to facilitate improvements for subsequent actions.
[0055] It should be appreciated that the process for calculating the reward described above in conjunction with FIG. 4 is merely exemplary. Depending on actual application requirements, the steps in the process for calculating the reward may be replaced or modified in any manner, and the process may comprise more or fewer steps. For example, in the process 400. the prompt 422 is created based on the task description 402, the recent action 404, the previous states, the previous rewards, and the prompt template 412, but the embodiments of the present disclosure are not limited thereto. In some embodiments, only one or two of the task description 402, the recent action 404, the previous states, the previous rewards, and the prompt template 412 may beconsidered. In the absence of the prompt template 412, the prompt 422 may be created through combining one or more of the task description 402, the recent action 404, the previous states, the previous rewards, and the response instruction.
[0056] FIG. 5 is a flowchart of an exemplary method 500 for action decision for a target application according to an embodiment of the present disclosure.
[0057] At 10, a set of previous states associated with a target application and a set of previous rewards corresponding to the set of previous states may be obtained.
[0058] At 520, an action in natural language for the target application may be generated based on the set of previous states and the set of previous rewards.
[0059] At 530, it may be verified with a set of predefined rules whether the action is reasonable.
[0060] At 540, in response to verifying that the action is reasonable, action code in computer language corresponding to the action may be generated.
[0061] At 550, the target application may be caused to execute the action code.
[0062] In an implementation, the set of predefined rules may include: a rule for evaluating whether the action is included in a predetermined set of permissible actions, and / or a rule for evaluating whether the action is meaningless or untrustworthy.
[0063] The set of permissible actions may be determined through: obtaining a task design specification and / or an indicator description associated with the target application; preprocessing the task design specification and / or the indicator description, to generate a task description and / or a set of candidate indicators for the target application; and generating the set of permissible actions based on the task description and / or the set of candidate indicators.
[0064] In an implementation, the action may be generated through a language model. The method 500 may further comprise: in response to verifying that the action is unreasonable, adjusting a hyperparameter of the language model and / or modifying a prompt provided to the language model; and regenerating an action in natural language for the target application based on the set of previous states and the set of previous rewards through the language model.
[0065] In an implementation, the method 500 may further comprise: collecting a state produced through the execution of the action code by the target application; calculating a reward corresponding to the action based on the action, the state, the set of previous states, and the set of previous rewards; and generating a subsequent action in natural language for the target application based on the state, the reward, the set of previous states, and the set of previous states.
[0066] The collecting a state produced through the execution of the action code by the target application may comprise: determining an observation indicator that should be watched during the operation of the target application; collecting operation and maintenance data produced after the execution of the action code by the target application; and extracting data corresponding to theobservation indicator from the operation and maintenance data as the state.
[0067] The determining an observation indicator that should be watched during the operation of the target application may comprise: obtaining a task design specification and / or an indicator description associated with the target application; preprocessing the task design specification and / or the indicator description, to generate a task description and / or a set of candidate indicators for the target application; and selecting the observation indicator from the set of candidate indicators based on the task description.
[0068] The method 500 may further comprise: storing the state and / or the rew ard in a longterm memory associated with the target application.
[0069] It should be appreciated that the method 500 may further comprise any other step / process for action decision for a target application according to the embodiments of the present disclosure as mentioned above.
[0070] FIG. 6 illustrates an exemplary apparatus 600 for action decision for a target application according to an embodiment of the present disclosure.
[0071] The apparatus 600 may comprise: a state and reward obtaining module 610, for obtaining a set of previous states associated with a target application and a set of previous rew ards corresponding to the set of previous states; an action generating module 620, for generating an action in natural language for the target application based on the set of previous states and the set of previous rewards; an action verifying module 630, for verifying whether the action is reasonable with a set of predefined rules; an action code generating module 640, for in response to verifying that the action is reasonable, generating action code in computer language corresponding to the action; and an action code executing module 650, for causing the target application to execute the action code. Moreover, the apparatus 600 may further comprise any other modules configured for action decision for a target application according to the embodiments of the present disclosure as mentioned above.
[0072] FIG. 7 illustrates another exemplary apparatus 700 for action decision for a target application according to an embodiment of the present disclosure.
[0073] The apparatus 700 may comprise a processor 710; and a memory 720 storing computer-executable instructions. The computer executable instructions, when executed, may cause the processor 710 to: obtain a set of previous states associated with a target application and a set of previous rewards corresponding to the set of previous states; generate an action in natural language for the target application based on the set of previous states and the set of previous rewards; verify- whether the action is reasonable with a set of predefined rules; in response to verifying that the action is reasonable, generate action code in computer language corresponding to the action; and cause the target application to execute the action code.
[0074] In an implementation, the set of predefined rules may include: a rule for evaluating whether the action is included in a predetermined set of permissible actions, and / or a rule for evaluating whether the action is meaningless or untrustworthy.
[0075] The set of permissible actions may be determined through: obtaining a task design specification and / or an indicator description associated with the target application; preprocessing the task design specification and / or the indicator description, to generate a task description and / or a set of candidate indicators for the target application; and generating the set of permissible actions based on the task description and / or the set of candidate indicators.
[0076] In an implementation, the action may be generated through a language model. The computer executable instructions, when executed, may further cause the processor 710 to: in response to verifying that the action is unreasonable, adjust a hyperparameter of the language model and / or modify a prompt provided to the language model; and regenerate an action in natural language for the target application based on the set of previous states and the set of previous rew ards through the language model.
[0077] In an implementation, the computer executable instructions, when executed, may further cause the processor 710 to: collect a state produced through the execution of the action code by the target application; calculate a reward corresponding to the action based on the action, the state, the set of previous states, and the set of previous rewards: and generate a subsequent action in natural language for the target application based on the state, the reward, the set of previous states, and the set of previous states.
[0078] The collecting a state produced through the execution of the action code by the target application may comprise: determining an observation indicator that should be watched during the operation of the target application; collecting operation and maintenance data produced after the execution of the action code by the target application; and extracting data corresponding to the observation indicator from the operation and maintenance data as the state.
[0079] The determining an observation indicator that should be w atched during the operation of the target application may comprise: obtaining a task design specification and / or an indicator description associated with the target application; preprocessing the task design specification and / or the indicator description, to generate a task description and / or a set of candidate indicators for the target application; and selecting the observation indicator from the set of candidate indicators based on the task description.
[0080] The computer executable instructions, when executed, may further cause the processor 710 to: store the state and / or the reward in a long-term memory associated with the target application.
[0081] It should be appreciated that the processor 710 may further perform any otherstep / process of the method for action decision for a target application according to the embodiments of the present disclosure as mentioned above.
[0082] The embodiments of the present disclosure propose a computer program product for action decision for a target application, comprising a computer program that is executed by a processor for: obtaining a set of previous states associated with a target application and a set of previous rewards corresponding to the set of previous states; generating an action in natural language for the target application based on the set of previous states and the set of previous rewards; verifying whether the action is reasonable with a set of predefined rules; in response to verifying that the action is reasonable, generating action code in computer language corresponding to the action; and causing the target application to execute the action code. Furthermore, the computer program may be further executed for implementing any other steps / processes of the method for action decision for a target application according to the embodiments of the present disclosure as mentioned above.
[0083] The embodiments of the present disclosure may be embodied in a computer-readable medium. The computer-readable medium may comprise instructions that, when executed, cause a processor to: obtain a set of previous states associated with a target application and a set of previous rewards corresponding to the set of previous states; generate an action in natural language for the target application based on the set of previous states and the set of previous rewards; verify whether the action is reasonable with a set of predefined rules; in response to verifying that the action is reasonable, generate action code in computer language corresponding to the action; and cause the target application to execute the action code. Furthermore, the instructions, when executed, may further cause the processor to perform any other steps / processes of the method for action decision for a target application according to the embodiments of the present disclosure as mentioned above.
[0084] It should be appreciated that all the operations in the methods described above are merely exemplary, and the present disclosure is not limited to any operations in the methods or sequence orders of these operations, and should cover all other equivalents under the same or similar concepts. In addition, the articles '‘a” and ‘’an” as used in this specification and the appended claims should generally be construed to mean “one” or “one or more” unless specified otherwise or clear from the context to be directed to a singular form.
[0085] It should also be appreciated that all the modules in the apparatuses described above may be implemented in various approaches. These modules may be implemented as hardware, software, or a combination thereof. Moreover, any of these modules may be further functionally divided into sub-modules or combined together.
[0086] Processors have been described in connection with various apparatuses and methods.These processors may be implemented using electronic hardware, computer software, or any combination thereof. Whether such processors are implemented as hardware or software will depend upon the particular application and overall design constraints imposed on the system. By way of example, a processor, any portion of a processor, or any combination of processors presented in the present disclosure may be implemented with a microprocessor, microcontroller, digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic device (PLD), a state machine, gated logic, discrete hardware circuits, and other suitable processing components configured for performing the various functions described throughout the present disclosure. The functionality of a processor, any portion of a processor, or any combination of processors presented in the present disclosure may be implemented with software being executed by a microprocessor, microcontroller, DSP, or other suitable platform.
[0087] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, threads of execution, procedures, functions, etc. The software may reside on a computer-readable medium. A computer-readable medium may include, by way of example, memory such as a magnetic storage device (e.g., hard disk, floppy disk, magnetic strip), an optical disk, a smart card, a flash memory device, random access memory (RAM), read only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), a register, or a removable disk. Although memory is shown separate from the processors in the various aspects presented throughout the present disclosure, the memory may be internal to the processors, e g., cache or register.
[0088] The previous description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein. All structural and functional equivalents to the elements of the various aspects described throughout the present disclosure that are known or later come to be known to those of ordinary skilled in the art are expressly incorporated herein and intended to be encompassed by the claims.
Claims
CLAIMS1 . A method for action decision for a target application, comprising: obtaining a set of previous states associated with a target application and a set of previous rewards corresponding to the set of previous states; generating an action in natural language for the target application based on the set of previous states and the set of previous rewards; verifying whether the action is reasonable with a set of predefined rules; in response to verifying that the action is reasonable, generating action code in computer language corresponding to the action; and causing the target application to execute the action code.
2. The method of claim 1, wherein the set of predefined rules includes: a rule for evaluating whether the action is included in a predetermined set of permissible actions, and / or a rule for evaluating whether the action is meaningless or untrustworthy.
3. The method of claim 2, wherein the set of permissible actions is determined through: obtaining a task design specification and / or an indicator description associated with the target application; preprocessing the task design specification and / or the indicator description, to generate a task description and / or a set of candidate indicators for the target application; and generating the set of permissible actions based on the task description and / or the set of candidate indicators.
4. The method of claim 1, wherein the action is generated through a language model, and the method further comprises: in response to verifying that the action is unreasonable, adjusting a hyperparameter of the language model and / or modifying a prompt provided to the language model; and regenerating an action in natural language for the target application based on the set of previous states and the set of previous rewards through the language model.
5. The method of claim 1, further comprising: collecting a state produced through the execution of the action code by the target application; calculating a reward corresponding to the action based on the action, the state, the set of previous states, and the set of previous rewards; and generating a subsequent action in natural language for the target application based on the state, the reward, the set of previous states, and the set of previous states.
6. The method of claim 5, wherein the collecting a state produced through the execution ofthe action code by the target application comprises: determining an observation indicator that should be watched during the operation of the target application; collecting operation and maintenance data produced after the execution of the action code by the target application; and extracting data corresponding to the observation indicator from the operation and maintenance data as the state.
7. The method of claim 6, wherein the determining an observation indicator that should be watched during the operation of the target application comprises: obtaining a task design specification and / or an indicator description associated with the target application; preprocessing the task design specification and / or the indicator description, to generate a task description and / or a set of candidate indicators for the target application: and selecting the observation indicator from the set of candidate indicators based on the task description.
8. The method of claim 5, further comprising: storing the state and / or the reward in a long-term memory associated with the target application.
9. An apparatus for action decision for a target application, comprising: a processor; and a memory storing computer-executable instructions that, when executed, cause the processor to: obtain a set of previous states associated with a target application and a set of previous rewards corresponding to the set of previous states, generate an action in natural language for the target application based on the set of previous states and the set of previous rewards, verify whether the action is reasonable with a set of predefined rules, in response to verifying that the action is reasonable, generate action code in computer language corresponding to the action, and cause the target application to execute the action code.
10. The apparatus of claim 9, wherein the set of predefined rules includes: a rule for evaluating whether the action is included in a predetermined set of permissible actions, and / or a rule for evaluating whether the action is meaningless or untrustworthy.
11. The apparatus of claim 10. wherein the set of permissible actions is determinedthrough: obtaining a task design specification and / or an indicator description associated with the target application; preprocessing the task design specification and / or the indicator description, to generate a task description and / or a set of candidate indicators for the target application: and generating the set of permissible actions based on the task description and / or the set of candidate indicators.
12. The apparatus of claim 9, wherein the action is generated through a language model, and the computer-executable instructions, when executed, further cause the processor to: in response to verifying that the action is unreasonable, adjust a hyperparameter of the language model and / or modifying a prompt provided to the language model; and regenerate an action in natural language for the target application based on the set of previous states and the set of previous rewards through the language model.
13. The apparatus of claim 9, wherein the computer-executable instructions, when executed, further cause the processor to: collect a state produced through the execution of the action code by the target application; calculate a reward corresponding to the action based on the action, the state, the set of previous states, and the set of previous rewards; and generate a subsequent action in natural language for the target application based on the state, the reward, the set of previous states, and the set of previous states.
14. The apparatus of claim 13, wherein the collecting a state produced through the execution of the action code by the target application comprises: determining an observation indicator that should be watched during the operation of the target application; collecting operation and maintenance data produced after the execution of the action code by the target application: and extracting data corresponding to the observation indicator from the operation and maintenance data as the state.
15. A computer-readable medium for action decision for a target application, comprising instructions that, when executed, cause a processor to: obtain a set of previous states associated with a target application and a set of previous rewards corresponding to the set of previous states; generate an action in natural language for the target application based on the set of previous states and the set of previous rewards; verify whether the action is reasonable with a set of predefined rules;in response to verifying that the action is reasonable, generate action code in computer language corresponding to the action; and cause the target application to execute the action code.