Action decision based on self-adjusting mechanism
By adopting action decision-making methods based on self-regulation mechanism in cloud computing resource management, allowable actions and observation indicators are predetermined, and actions are generated and verified using language models, the problems of large human resources consumption and low efficiency in the existing technology are solved, and efficient and automated action decisions are achieved.
Patent Information
- Application Number
- CN202311684190.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-08
- Publication Date
- 2025-06-10
AI Technical Summary
When managing cloud computing resources, the prior art needs to define different rules for different scenarios or applications or design different machine learning algorithms, resulting in high consumption and low efficiency of human resources.
The action decision-making method based on the self-adjustment mechanism is adopted to determine the allowable actions and observation indicators of the target application in advance, and the actions in natural language form are generated using the language model, and the action rationality is verified through predefined rules, and finally the action code in computer language form is generated.
It reduces the need for artificial intervention, improves the quality and efficiency of action decisions, realizes self-regulation of action decisions, and is suitable for various applications requiring action decisions.
Smart Images

Figure CN120123071A_ABST
Abstract
Description
Background Art
[0001] Cloud computing is a technology widely used by modern enterprises and organizations. Cloud computing is a network-based computing model that provides computing resources, storage resources, network resources, etc. to users via the Internet. Users can autonomously select, configure, and use the required resources through the network. Cloud computing is based on virtualization technology. By creating virtual machines that can run various applications on physical machines, users can conveniently use cloud computing services. Summary of the Invention
[0002] The present invention content is provided to introduce a set of concepts that will be further described in the following detailed implementation manners. The present invention content is not intended to identify the key features or essential features of the protected subject matter, nor is it intended to limit the scope of the protected subject matter.
[0003] Embodiments of the present disclosure propose methods, devices, and computer-readable media for action decision-making based on a self-regulating mechanism. A set of previous states associated with a target application and a set of previous rewards corresponding to the set of previous states can be obtained. Based on the set of previous states and the set of previous rewards, an action in natural language form for the target application can be generated. A set of predefined rules can be used to verify whether the action is reasonable. In response to verifying that the action is reasonable, an action code in computer language form corresponding to the action can be generated. The target application can be made to execute the action code.
[0004] It should be noted that one or more of the above aspects include the features described in detail below and specifically pointed out in the claims. The following description and drawings elaborate on certain illustrative features of the one or more aspects. These features merely indicate various ways in which the principles of the various aspects can be employed, and the present disclosure is intended to include all these aspects and their equivalent transformations. Brief Description of the Drawings
[0005] The following will describe the disclosed aspects in conjunction with the drawings, which are provided for illustration and not for limitation of the disclosed aspects.
[0006] Figure 1 An exemplary process for action decision-making based on a self-regulating mechanism according to an embodiment of the present disclosure is shown.
[0007] Figure 2 An exemplary process for determining allowable actions and observation metrics according to an embodiment of the present disclosure is shown.
[0008] Figure 3 An exemplary process for generating an action in natural language form for a target application according to an embodiment of the present disclosure is shown.
[0009] Figure 4 illustrates an exemplary process for calculating rewards according to an embodiment of the present disclosure.
[0010] Figure 5 is a flowchart of an exemplary method for action decision-making for a target application according to an embodiment of the present disclosure.
[0011] Figure 6 illustrates an exemplary apparatus for action decision-making for a target application according to an embodiment of the present disclosure.
[0012] Figure 7 illustrates another exemplary apparatus for action decision-making for a target application according to an embodiment of the present disclosure. Detailed implementation manners
[0013] The present disclosure will now be discussed with reference to several exemplary embodiments. It should be understood that the discussion of these embodiments is only for enabling those skilled in the art to better understand and thus implement the embodiments of the present disclosure, rather than teaching any limitation to the scope of the present disclosure.
[0014] A platform for providing cloud computing services may be referred to as a cloud computing platform. It is desirable to manage the resources on the cloud computing platform reasonably to better utilize the limited resource capacity. The resources on the cloud computing platform may include, for example, computing resources, storage resources, network resources, etc. These resources may be collectively referred to as cloud computing resources. The cloud computing resources may be managed through various applications, such as a Workload Management (WLM) application for managing the allocation of resources among different workloads, a predictive maintenance application for predicting potential failures of machines and suggesting maintenance measures, a network traffic management application for analyzing network traffic patterns to ensure optimal network performance, etc. These applications may perform corresponding actions to manage the cloud computing resources from different aspects. Since the system environment or user requirements are changing dynamically, the actions performed by each application also need to be updated in real time. The actions to be performed by the application may be decided through a set of predefined rules or manually designed machine learning algorithms. These methods usually require defining different rules or designing different machine learning algorithms for different scenarios or applications, thus consuming a large amount of human resources and being inefficient.
[0015] Embodiments of the present disclosure propose action decision-making based on a self-regulating mechanism. A set of allowable actions that the target application can execute and at least one observation index to be concerned about during the operation of the target application can be determined in advance. Herein, the target application refers to the application for which action decision-making is carried out. After determining the allowable actions and observation indexes, a decision can be made on the actions to be executed by the target application. A natural language form of action for the target application can be generated through a language model, based on a set of previous states associated with the target application and a set of previous rewards corresponding to the set of previous states. Herein, the language model refers to a deep learning model that can understand the meaning of natural language, generate natural language text, or perform other natural language processing tasks. It should be understood that the language model includes a multimodal model that can perform processing tasks of multiple modalities including natural language. After generating the action in natural language form, a heuristic method, such as using a set of predefined rules, can be used to verify whether the action is reasonable. The set of predefined rules can include rules for evaluating whether the action is included in the pre-determined set of allowable actions and / or rules for evaluating whether the action is meaningless or untrustworthy. If it is verified that the action is unreasonable, the hyperparameters of the language model for generating the action can be adjusted and / or the prompt words provided to the language model can be modified, and the action can be regenerated. If it is verified that the action is reasonable, the action can be converted into an action code in computer language form. The target application can be made to execute the action code, and the corresponding state can be collected. This state can be used together with a set of previous states and a set of previous rewards obtained before generating the action to calculate the reward corresponding to the action. Subsequently, a natural language form of subsequent action for the target application can be generated based on some or all of the previously obtained states and rewards.
[0016] The operations after the operation for determining the allowable actions and observation indexes in the above process can be repeatedly executed in an online manner, so that the state generated after the target application executes each action can be continuously monitored, and the actions and corresponding action codes for the target application can be updated or optimized in real time by combining the previously collected states and the previously calculated rewards. The above process reduces the need for human intervention and aims to achieve action decision-making in a self-regulating manner. This improves the quality and efficiency of action decision-making.
[0017] In addition, a set of allowable actions that a pre-determined target application can perform can be used to verify whether the generated actions are reasonable. At least one observation metric that should be concerned during the pre-determined operation of the target application can be used to determine which metrics' corresponding states should be collected after the action code is executed. The allowable actions and observation metrics associated with different target applications may be different. By pre-determining the allowable actions and observation metrics associated with the target application, the quality and efficiency of action decision-making for the target application can be further improved. This also enables the action decision-making process according to the embodiments of the present disclosure to be widely applicable to various applications that require action decision-making, thus realizing a unified action decision-making scheme.
[0018] In addition, during the above process, after generating an action in natural language form for the target application, the reasonableness of the action is verified. Only when it is verified that the action is reasonable, the action in natural language form is converted into an action code in computer language form, and the action code is then applied to the target application. This way can avoid the execution of unreasonable actions and improve the robustness of action decision-making. In addition, decomposing action generation into first generating an action in natural language form that is easy for humans to understand and then generating an action code in computer language form helps to improve the accuracy of action generation and enables the use of heuristic methods to verify the reasonableness of the action.
[0019] The foregoing discussion and the following discussion may involve examples of action decision-making for applications used in cloud computing resource management. However, it should be understood that the embodiments of the present disclosure are not limited thereto, but can perform action decision-making for applications in other fields in a similar manner. For example, for an energy management application used to optimize energy consumption and ensure stable energy supply, the solution of the embodiments of the present disclosure can also be used to make decisions on the actions to be performed.
[0020] The following will describe the various embodiments of the present disclosure in detail with reference to the accompanying drawings.
[0021] Figure 1 An exemplary process 100 for action decision-making based on a self-regulating mechanism according to an embodiment of the present disclosure is shown. The process 100 can be used to make decisions on the actions to be performed by a target application 102. The target application 102 can be various applications that require action decision-making, such as a workload management application for managing the allocation of resources among different workloads, a predictive maintenance application for predicting potential failures of machines and recommending maintenance measures, a network traffic management application for analyzing network traffic patterns to ensure optimal network performance, an energy management application for optimizing energy consumption and ensuring stable energy supply, etc. The process 100 can be executed online.
[0022] Task design specifications 104 and / or metric descriptions 106 associated with the target application 102 can be obtained. The task design specifications 104 can be text content or a document used to describe the tasks performed by the target application 102. The metric description 106 can be text content or a document used to explain the metrics related to the target application 102. The task design specifications 104 and the metric description 106 can be two separate documents or the same document. The task design specifications 104 and / or the metric description 106 can be preprocessed by the preprocessing component 110 to generate a task description 112 for the target application 102 and / or a set of candidate metrics 114. The preprocessing component 110 is, for example, a machine learning model capable of performing a text summarization task. Taking the target application 102 as a workload management application as an example, the task description 112 for the target application 102 generated by the preprocessing component 110 can be "The workload management application is responsible for managing the allocation of resources among different workloads. Corresponding priorities are assigned to the workloads according to the latency sensitivity of the workloads. For example, sending an email should be latency-sensitive, while backing up data may be latency-insensitive. The workload management application provides a workload status signal to the client to increase or decrease the resources allocated to the workloads at the client...". Also taking the target application 102 as a workload management application as an example, a set of candidate metrics 114 for the target application 102 generated by the preprocessing component 110 can be: network smoothness, disk queue depth, network source address, network destination address, disk read latency, disk write latency, etc. These metrics can also be referred to as performance counters.
[0023] Subsequently, the parameter determination component 120 can determine a set of allowable actions 122 that the target application 102 can perform and / or at least one observation metric 124 that should be concerned during the running of the target application 102 based on the task description 112 and / or the set of candidate metrics 114. The allowable actions 122 and the observation metrics 124 can be determined by a language model. The language model can be a language model capable of generating actions in natural language form and recommended metrics, such as the Generative Pre-trained Transformer-4 (GPT-4) model, etc. This will be described later in conjunction with Figure 2 to illustrate an exemplary process for determining the allowable actions 122 and the observation metrics 124.
[0024] After determining the allowable actions 122 and / or the observation metrics 124, the action decision agent 130 can make a decision on the actions to be performed by the target application 102.
[0025] The action decision-making agent 130 may include an action planning component 140 for generating actions in the form of natural language for the target application 102. For example, a set of previous states associated with the target application 102 and a set of previous rewards corresponding to the set of previous states may be obtained. The set of previous states may include the states collected before the action planning component 140 performs the action generation operation. The set of previous rewards may include the rewards calculated before the action planning component 140 performs the action generation operation. The rewards may be used to evaluate the quality of the actions generated by the action planning component 140. The rewards may be previously calculated by the reward evaluation component 190. Initially, for example, when the action decision-making agent 130 or the action planning component 140 is just started, the rewards may be randomly generated. The rewards may be in numerical form. Since the states are collected from the environment external to the action decision-making agent 130, the states may also be referred to as external feedback. In contrast, since the rewards are calculated by the reward evaluation component 190 inside the action decision-making agent 130, the rewards may also be referred to as internal feedback. An action in the form of natural language for the target application 102 may be generated based on the set of previous states and the set of previous rewards. The action may be generated by a language model. The language model may be a language model capable of generating actions in the form of natural language, such as the GPT-4 model, etc. The exemplary process for generating an action in the form of natural language for the target application 102 will be described later in conjunction with Figure 3 to illustrate an exemplary process for generating an action in the form of natural language for the target application 102.
[0026] The action decision-making agent 130 may include an action verification component 150 for verifying the rationality of the actions generated by the action planning component 140. The action verification component 150 may include a short-term memory 152 for storing the actions newly generated by the action planning component 140 and the like. The rationality of the actions can be verified by heuristic methods, such as using a set of predefined rules. The set of predefined rules may include rules for evaluating whether the action is included in a predefined set of allowable actions 122. The predefined set of allowable actions 122 defines the action space that the target application can execute. If the action generated by the action planning component 140 is not included in the set of allowable actions 122, the action can be verified as irrational. Alternatively or additionally, the set of predefined rules may include rules for evaluating whether the action is meaningless or untrustworthy. Such rules are designed to determine whether there is a hallucination problem with the action. Generative language models sometimes produce meaningless or untrustworthy outputs, and such problems are called hallucination problems. It can be determined whether the action is meaningless or untrustworthy by evaluating whether the action is contradictory or contains incorrect information. If it is determined that the action generated by the action planning component 140 is meaningless or untrustworthy, the action can be verified as irrational. The above rules can be used alone or in combination with each other. It should be understood that the above rules for verifying the rationality of actions are only exemplary. According to actual application requirements, other rules can also be used to verify the rationality of actions.
[0027] If it is verified that the action generated by the action planning component 140 is irrational, the action planning process can be optimized and the action can be regenerated. In one implementation, the hyperparameters of the language model involved in the action planning component 140 can be adjusted. For example, the temperature parameter of the language model can be increased to make its output more rigorous. In another implementation, the prompt words provided to the language model can be modified. For example, the response length of the language model can be restricted or the language model can be required to reply based on known facts. The above implementations can be executed alone or in combination with each other. It should be understood that the above ways of optimizing the action planning process are only exemplary. According to actual application requirements, other ways can also be adopted to optimize the action planning process to enhance the rationality of the actions in natural language form generated. Subsequently, based on a set of previous states and a set of previous rewards, the action in natural language form for the target application 102 can be regenerated through the language model. Preferably, if a reasonable action still cannot be obtained within a predetermined time or after the action planning component 140 has performed a predetermined number of action generation operations, a default action can be invoked.
[0028] If the action generated by the action planning component 140 is verified to be reasonable, the action can be taken as action 142 and the subsequent process can be carried out. For example, the code generation component 160 in the action decision-making agent 130 can generate the action code 162 in the form of a computer language corresponding to the action 142. The action code 162 in the form of a computer language can be generated by a language model. The language model can be a language model capable of generating program code, such as the GPT-4 model, the CodeLlama model, etc. Before that, the prompt words to be provided to the language model can be created first. Preferably, a prompt word template associated with generating the action code can be designed in advance. The prompt word template can include a part for loading the action. When creating the prompt words, the action 142 generated by the action planning component 140 can be loaded into the corresponding part of the prompt word template. In addition, the prompt word template can include response instructions for guiding how the language model responds. As an example, the response instructions can be "I will provide an action plan. Please convert it into the corresponding program code. Exemplary code: import numpy as np; set_disk_latency(0.1)……". It should be understood that the above response instructions are only exemplary. According to the actual application requirements, the response instructions in the prompt word template associated with generating the action code can have other forms and can include more or less content.
[0029] The target application 102 can be made to execute the action code 162. Subsequently, the status 172 generated by the target application 102 executing the action code 162 can be collected through the monitoring component 170. The status 172 can correspond to the previously determined observation metric 124. The operation and maintenance data generated after the target application 102 executes the action code 162 can be collected through the monitoring component 170. The operation and maintenance data can include logs, monitoring information, application information, etc. The data corresponding to the observation metric 124 can be extracted from the operation and maintenance data as the status 172. The operation and maintenance data may include data from different machines or collected at different time intervals. Preferably, when extracting the status 172, these data can be aggregated according to the actual application requirements. In addition, the operation and maintenance data may include some meaningless noise data. Preferably, when extracting the status 172, these noise data can be filtered out from the operation and maintenance data.
[0030] After the status 172 is collected, it can be stored in the long-term memory 180 associated with the target application 102 for subsequent action decision-making, such as being used as the historical status when generating subsequent actions. The long-term memory 180 may already have stored a set of previous historical statuses 182 and a set of historical rewards 184.
[0031] Next, the reward evaluation component 190 in the action decision-making agent 130 can calculate the reward 192 corresponding to the action 142. The reward 192 can be used to evaluate the quality of the action 142. The reward 192 can be in numerical form. The reward 192 corresponding to the action 142 can be calculated based on one or more of the action 142, the state 172, a set of historical states 182, and a set of historical rewards 184. It should be understood that the set of historical states 182 can correspond to a set of previous states obtained when generating the action 142, and the set of historical rewards 184 can correspond to a set of previous rewards obtained when generating the action 142. The reward can be calculated by a language model. The language model can be a language model capable of scoring the input, such as the GPT-4 model, the Vicuna model, etc. An exemplary process for calculating the reward will be described later in conjunction with Figure 4 to illustrate the exemplary process for calculating the reward. The reward 192 can be stored in the long-term memory 180 for subsequent action decision-making, such as being used as a historical reward when generating subsequent actions.
[0032] Subsequently, based on one or more of the newly acquired state 172, the newly calculated reward 192, and a set of historical states 182 and a set of historical rewards 184 stored in the long-term memory 180, a subsequent action in natural language form for the target application 102 can be generated.
[0033] The operations in the process 100 that are after the operations for determining the allowable actions 122 and the observation metrics 124 can be repeatedly executed online, so that the state generated after the target application 102 executes each action can be continuously monitored, and the actions and corresponding action codes for the target application 102 can be updated or optimized in real time by combining the previously acquired states and the previously calculated rewards. The above process reduces the need for human intervention and aims to achieve action decision-making in a self-regulating manner. This improves the quality and efficiency of action decision-making.
[0034] In the process 100, a set of allowable actions 122 that the target application 102 can execute and / or at least one observation metric 124 that should be concerned during the operation of the target application 102 are pre-determined by the preprocessing component 110 and the parameter determination component 120 on the right. The allowable actions 122 can be used to verify whether the actions generated by the action planning component 140 in the action decision-making agent 130 are reasonable. The observation metric 124 can be used to determine which metrics' corresponding states should be collected after the action code 162 is executed. The allowable actions and observation metrics associated with different target applications may be different. By pre-determining the allowable actions and observation metrics associated with the target application, the quality and efficiency of action decision-making for the target application can be further improved. This also enables the process 100 to be widely applicable to various applications requiring action decision-making, thus realizing a unified action decision-making scheme.
[0035] In addition, in Process 100, after the action planning component 140 generates an action in natural language form for the target application 102, the action verification component 150 verifies whether the action is reasonable. Only when it is verified that the action is reasonable, the action in natural language form is converted into an action code in computer language form, and the action code is then applied to the target application 102. This way can avoid the execution of unreasonable actions and improve the robustness of action decision-making. In addition, decomposing action generation into first generating an action in natural language form that is easy for humans to understand and then generating an action code in computer language form helps to improve the accuracy of action generation and enables the use of heuristic methods to verify the reasonableness of actions.
[0036] It should be understood that the process for action decision-making based on the self-regulating mechanism described above in conjunction with Figure 1 is only exemplary. According to actual application requirements, the steps in the process for action decision-making can be replaced or modified in any way, and the process can include more or fewer steps. For example, in Process 100, steps for determining the allowable action 122 and / or the observation metric 124 are described. This step is not necessary. In the case where the steps for determining the allowable action 122 and / or the observation metric 124 are not executed, the subsequent steps can ignore the allowable action 122 and / or the observation metric 124. In addition, the specific order or hierarchy of the steps in Process 100 is only exemplary, and the process for action decision-making can be executed in an order different from the described order.
[0037] Figure 2 An exemplary process 200 for determining the allowable action and the observation metric according to an embodiment of the present disclosure is shown. Process 200 can correspond to Figure 1 the operation at the parameter determination component 120. In Process 200, the allowable action 222 and the observation metric 224 can be generated by a language model 220. The language model 220 can be a language model capable of generating actions in natural language form and recommended metrics, such as the GPT-4 model, etc. Prior to this, a prompt 212 to be provided to the language model 220 can be created by a prompt creator 210.
[0038] The task description 202 and the candidate metrics 204 of the target application can be obtained. The task description 202 and the candidate metrics 204 can respectively correspond to Figure 1 the task description 112 and the candidate metrics 114 in
[0039] The prompt creator 210 can create a prompt 212 based on the task description 202 and / or the candidate metrics 204 of the target application. Preferably, a prompt template 206 for the prompt creator 210 can be designed in advance. The prompt template 206 can include multiple parts for loading the task description and candidate metrics respectively. When creating the prompt 212, the task description 202 and candidate metrics 204 of the target application can be loaded into the corresponding parts in the prompt template 206 respectively. Additionally, the prompt template 206 can include response instructions for guiding how the language model 220 should respond. Taking the target application being a workload management application as an example, the response instructions can be "I will provide a task description and a list of performance counters. I hope you generate appropriate actions and select appropriate performance counters to achieve the optimization goal...". It should be understood that the above response instructions are only exemplary. According to the actual application requirements, the response instructions in the prompt template 206 can have other forms and can include more or less content.
[0040] The prompt 212 can be provided to the language model 220. The language model 220 can utilize its semantic understanding ability, logical reasoning ability, language expression ability, big data support ability, etc. to generate a set of allowable actions 222. Taking the target application being a workload management application as an example, the allowable actions can be increasing disk read latency, reducing disk write latency, increasing input / output per second (IOPS), decreasing IOPS, etc. Alternatively or additionally, the language model 220 can select at least one observation metric 224 from a set of candidate metrics 204 based on the prompt 212.
[0041] Preferably, the language model 220 can generate a metric code 226 in the form of a computer language corresponding to the observation metric 224. The generated code can be stored and called and executed when, for example, the target application executes the action code or collects the state of the target application. In the case where the language model 220 generates the metric code 226, the response instructions in the prompt template 206 can include instructions for generating the metric code, such as "Convert the selected performance counter into a program code with the following function: perf_counter.get_value(counter_name:str)…".
[0042] The allowable actions 222, observation metrics 224, and / or metric codes 226 can be generated simultaneously through one interaction with the language model 220. Alternatively, the allowable actions 222, observation metrics 224, and / or metric codes 226 can be generated at different times through multiple interactions with the language model 220.
[0043] It should be understood that, as described above in connection withFigure 2 The described process for determining allowable actions and observation metrics is merely exemplary. According to actual application requirements, the steps in the process for determining allowable actions and observation metrics can be replaced or modified in any way, and the process can include more or fewer steps. For example, in process 200, the prompt 212 is created based on the task description 202, candidate metrics 204, and prompt template 206, but the embodiments of the present disclosure are not limited thereto. In some embodiments, only one or two of the task description 202, candidate metrics 204, and prompt template 206 may be considered. In the absence of the prompt template 206, the prompt 212 can be created by combining one or more of the task description 202, candidate metrics 204, and response instructions. Additionally, in process 200, the language model 220 generates the allowable actions 222, observation metrics 224, and / or metric codes 226, but the embodiments of the present disclosure are not limited thereto. In some embodiments, the language model 220 can generate only one or two of the allowable actions 222, observation metrics 224, and metric codes 226.
[0044] Figure 3 An exemplary process 300 for generating actions in natural language form for a target application according to an embodiment of the present disclosure is shown. Process 300 may correspond to Figure 1 the operations at the action planning component 140 in. In process 300, an action 332 in natural language form can be generated by a language model 330. The language model 330 can be a language model capable of generating actions in natural language form, such as the GPT-4 model, etc. Prior to this, a prompt 322 to be provided to the language model 330 can be created by a prompt creator 320.
[0045] The task description 302 of the target application can be obtained. The task description 302 may correspond to Figure 1 the task description 112 in. Considering the task description of the target application when generating actions for the target application can make the generated actions more accurate.
[0046] A set of previous states associated with the target application and a set of previous rewards corresponding to the set of previous states can be obtained. The set of previous states can include a newly acquired recent state 304 and at least one historical state 310 stored in the long-term memory. The recent state 304 can be acquired after the target application executes the action code corresponding to the newly generated action. The long-term memory is, for example, Figure 1The long-term memory 180. A set of previous rewards may include a newly computed new reward 306 and at least one historical reward 308 stored in the long-term memory. The new reward 306 is a reward corresponding to the new state 304. The historical reward 308 is a state corresponding to the historical state 310. Initially, the new state 304 and / or the new reward 306 may be randomly generated states and / or rewards. Also, there may be no historical state 310 and historical reward 308.
[0047] The prompt creator 320 can create a prompt 322 based on one or more of the task description 302 associated with the target application, a set of previous states, and a set of previous rewards. Preferably, a prompt template 312 for the prompt creator 320 can be pre-designed. The prompt template 312 may include multiple parts for loading the task description, previous states, and previous rewards respectively. When creating the prompt 322, the task description 302 associated with the target application, a set of previous states, and a set of previous rewards can be loaded into the corresponding parts in the prompt template 312 respectively. Additionally, the prompt template 312 may include response instructions for guiding how the language model 330 should respond. Taking the target application being a workload management application as an example, the response instructions may be "I will provide a task description, a set of previous states, and a set of previous rewards corresponding to the set of previous states. Please analyze the relationship between the previous states and their corresponding previous rewards, and use the following format to provide subsequent actions that can improve the reward: 1. Increase disk latency by 1.0; 2. Decrease disk latency by 3.0. Please only provide a list of actions, and the format of the action list is in Markdown format...". It should be understood that the above response instructions are only exemplary. According to the actual application requirements, the response instructions in the prompt template 312 can have other forms and may include more or less content.
[0048] The prompt 322 can be provided to the language model 330. The language model 330 can utilize its semantic understanding ability, logical reasoning ability, language expression ability, big data support ability, etc., to analyze the relationship between a set of previous states and a set of previous rewards included in the prompt 322, and then generate an action 332 in natural language form. As an example, the generated action 332 in natural language form may be "1. Decrease disk latency by 1.0; 2. Increase IOPS by 3.0".
[0049] It should be understood that as described above in connection with Figure 3The described process for generating actions in natural language form for a target application is merely exemplary. According to actual application requirements, the steps in the process for generating actions in natural language form can be replaced or modified in any way, and the process can include more or fewer steps. For example, in process 300, the prompt 322 is created based on the task description 302, previous state, previous reward, and prompt template 312, but the embodiments of the present disclosure are not limited thereto. In some embodiments, only one or more of the task description 302, previous state, previous reward, and prompt template 312 may be considered. In the absence of the prompt template 312, the prompt 322 can be created by combining one or more of the task description 302, previous state, previous reward, and response instruction.
[0050] Figure 4 An exemplary process 400 for calculating a reward according to an embodiment of the present disclosure is shown. Process 400 may correspond to Figure 1 the operations at the reward evaluation component 190 in. In process 400, the reward 432 corresponding to the recent action 404 for the target application can be calculated by the language model 430. The language model 430 can be a language model capable of scoring the input, such as the GPT-4 model, Vicuna model, etc. Prior to this, the prompt 422 to be provided to the language model 430 can be created by the prompt creator 420. Returning to reference Figure 1 , the recent action 404 can correspond to Figure 1 the action 142 in.
[0051] The task description 402 of the target application can be obtained. The task description 402 can correspond to Figure 1 the task description 112 in. Considering the task description of the target application when calculating the reward associated with the target application can make the calculated reward more accurate.
[0052] The recent state 406 corresponding to the recent action 404 can be obtained. Returning to reference Figure 1 , the recent state 406 can correspond to Figure 1 the state 172 in, which can be obtained after the target application 102 executes the action code 162 corresponding to the action 142. A set of historical states 408 associated with the target application can be obtained. The set of historical states 408 can be stored in long-term memory. The long-term memory is, for example, Figure 1 the long-term memory 180 in. The historical state 408 can correspond to Figure 1The historical state 182 therein. The recent state 406 and the historical state 408 can be combined into a previous state. A previous reward can be obtained. The previous reward can include a set of historical rewards 410 corresponding to a set of historical states 408. The set of historical rewards 410 can be stored in long-term memory. The long-term memory is, for example, Figure 1 the long-term memory 180 therein. The historical reward 410 can correspond to Figure 1 the historical reward 184 therein. Initially, there may be no historical state 408 and historical reward 410.
[0053] The prompt creator 420 can create a prompt 422 based on one or more of the task description 402 of the target application, the recent action 404, the previous state including the recent state 406 and the historical state 408, and the previous reward including the historical reward 410. Preferably, a prompt template 412 for the prompt creator 420 can be pre-designed. The prompt template 412 can include multiple parts for loading the task description, the recent action, the previous state, and the previous reward respectively. When creating the prompt 422, the task description 402 of the target application, the recent action 406, the previous state including the recent state 406 and the historical state 408, and the previous reward including the historical reward 410 can be loaded into the corresponding parts in the prompt template 412 respectively. Additionally, the prompt template 412 can include response instructions for guiding how the language model 430 should respond. Taking the target application being a workload management application as an example, the response instructions can be "I will provide the task description, the recent action, the previous state, and the previous reward. Please provide the reward corresponding to this action. Remember: 1. Only return the reward; 2. The value of the reward should be between 1 and 10...". It should be understood that the above response instructions are only exemplary. According to actual application requirements, the response instructions in the prompt template 412 can have other forms and can include more or less content.
[0054] The prompt 422 can be provided to the language model 430. The language model 430 can utilize its semantic understanding ability, logical reasoning ability, language expression ability, big data support ability, etc., to analyze the relationship between a set of previous states and a set of previous rewards included in the prompt 422, and then generate a reward 432 for the recent action 404. Optionally, the language model 430 can provide an explanation for the reward 432 to facilitate the improvement of subsequent actions.
[0055] It should be understood that in combination with the above Figure 4The described process for calculating rewards is merely exemplary. According to actual application requirements, the steps in the process for calculating rewards can be replaced or modified in any way, and the process can include more or fewer steps. For example, in process 400, the prompt 422 is created based on the task description 402, the recent action 404, the previous state, the previous reward, and the prompt template 412, but the embodiments of the present disclosure are not limited thereto. In some embodiments, only one or more of the task description 402, the recent action 404, the previous state, the previous reward, and the prompt template 412 may be considered. In the absence of the prompt template 412, the prompt 422 can be created by combining one or more of the task description 402, the recent action 404, the previous state, the previous reward, and the response instruction.
[0056] Figure 5 is a flowchart of an exemplary method 500 for action decision-making for a target application according to an embodiment of the present disclosure.
[0057] At 510, a set of previous states associated with the target application and a set of previous rewards corresponding to the set of previous states can be obtained.
[0058] At 520, based on the set of previous states and the set of previous rewards, an action in natural language form for the target application can be generated.
[0059] At 530, a set of predefined rules can be used to verify whether the action is reasonable.
[0060] At 540, in response to verifying that the action is reasonable, an action code in computer language form corresponding to the action can be generated.
[0061] At 550, the target application can be made to execute the action code.
[0062] In one implementation, the set of predefined rules may include: a rule for evaluating whether the action is included in a predetermined set of allowable actions, and / or a rule for evaluating whether the action is meaningless or untrustworthy.
[0063] The set of allowable actions can be determined by: obtaining a task design specification and / or an indicator description associated with the target application; preprocessing the task design specification and / or the indicator description to generate a task description and / or a set of candidate indicators for the target application; and generating the set of allowable actions based on the task description and / or the set of candidate indicators.
[0064] In one embodiment, the action may be generated by a language model. The method 500 may further include: in response to verifying that the action is unreasonable, adjusting the hyperparameters of the language model and / or modifying the prompt words provided to the language model; and regenerating, by the language model, a natural language form action for the target application based on the set of previous states and the set of previous rewards.
[0065] In one embodiment, the method 500 may further include: collecting the state generated by the target application executing the action code; calculating a reward corresponding to the action based on the action, the state, the set of previous states, and the set of previous rewards; and generating a subsequent natural language form action for the target application based on the state, the reward, the set of previous states, and the set of previous states.
[0066] The collecting the state generated by the target application executing the action code may include: determining observation metrics to be concerned during running the target application; collecting operation and maintenance data generated after the target application executes the action code; and extracting data corresponding to the observation metrics from the operation and maintenance data as the state.
[0067] The determining observation metrics to be concerned during running the target application may include: obtaining a task design specification and / or metric description associated with the target application; preprocessing the task design specification and / or the metric description to generate a task description and / or a set of candidate metrics for the target application; and selecting the observation metrics from the set of candidate metrics based on the task description.
[0068] The method 500 may further include: storing the state and / or the reward in a long-term memory associated with the target application.
[0069] It should be understood that the method 500 may further include any other steps / processes for action decision-making for a target application according to the embodiments of the present disclosure as described above.
[0070] Figure 6 An exemplary apparatus 600 for action decision-making for a target application according to an embodiment of the present disclosure is shown.
[0071] The apparatus 600 may include: a status and reward acquisition module 610, configured to acquire a set of previous states associated with a target application and a set of previous rewards corresponding to the set of previous states; an action generation module 620, configured to generate, based on the set of previous states and the set of previous rewards, an action in natural language form for the target application; an action verification module 630, configured to verify whether the action is reasonable by using a set of predefined rules; an action code generation module 640, configured to generate, in response to verifying that the action is reasonable, an action code in computer language form corresponding to the action; and an action code execution module 650, configured to cause the target application to execute the action code. In addition, the apparatus 600 may further include any other module configured for action decision-making for a target application according to an embodiment of the present disclosure as described above.
[0072] Figure 7 Another exemplary apparatus 700 for action decision-making for a target application according to an embodiment of the present disclosure is shown.
[0073] The apparatus 700 may include: a processor 710; and a memory 720 storing computer-executable instructions. When the computer-executable instructions are executed, they may cause the processor 710 to: acquire a set of previous states associated with a target application and a set of previous rewards corresponding to the set of previous states, generate, based on the set of previous states and the set of previous rewards, an action in natural language form for the target application, verify whether the action is reasonable by using a set of predefined rules, generate, in response to verifying that the action is reasonable, an action code in computer language form corresponding to the action, and cause the target application to execute the action code.
[0074] In one implementation, the set of predefined rules may include: a rule for evaluating whether the action is included in a pre-determined set of allowable actions, and / or a rule for evaluating whether the action is meaningless or untrustworthy.
[0075] The set of allowable actions may be determined by: acquiring a task design specification and / or a metric description associated with the target application; preprocessing the task design specification and / or the metric description to generate a task description and / or a set of candidate metrics for the target application; and generating the set of allowable actions based on the task description and / or the set of candidate metrics.
[0076] In one embodiment, the action may be generated by a language model. When the computer-executable instructions are executed, they may further cause the processor 710 to: in response to verifying that the action is unreasonable, adjust the hyperparameters of the language model and / or modify the prompt words provided to the language model; and, through the language model, regenerate an action in natural language form for the target application based on the set of previous states and the set of previous rewards.
[0077] In one embodiment, when the computer-executable instructions are executed, they may further cause the processor 710 to: collect the state generated by the target application executing the action code; calculate a reward corresponding to the action based on the action, the state, the set of previous states, and the set of previous rewards; and generate a subsequent action in natural language form for the target application based on the state, the reward, the set of previous states, and the set of previous states.
[0078] The collecting of the state generated by the target application executing the action code may include: determining observation metrics to be concerned during the running of the target application; collecting operation and maintenance data generated after the target application executes the action code; and extracting data corresponding to the observation metrics from the operation and maintenance data as the state.
[0079] The determining of the observation metrics to be concerned during the running of the target application may include: obtaining task design specifications and / or metric descriptions associated with the target application; preprocessing the task design specifications and / or the metric descriptions to generate a task description and / or a set of candidate metrics for the target application; and selecting the observation metrics from the set of candidate metrics based on the task description.
[0080] When the computer-executable instructions are executed, they may further cause the processor 710 to: store the state and / or the reward in a long-term memory associated with the target application.
[0081] It should be understood that the processor 710 may also execute any other steps / processes of the method for action decision for a target application according to the embodiments of the present disclosure as described above.
[0082] Embodiments of the present disclosure propose a computer program product for action decision-making for a target application, including a computer program, the computer program being executed by a processor for: obtaining a set of previous states associated with the target application and a set of previous rewards corresponding to the set of previous states; generating, based on the set of previous states and the set of previous rewards, an action in natural language form for the target application; using a set of predefined rules to verify whether the action is reasonable; in response to verifying that the action is reasonable, generating an action code in computer language form corresponding to the action; and causing the target application to execute the action code. In addition, the computer program may also be executed for implementing any other steps / processes of the method for action decision-making for a target application according to the embodiments of the present disclosure as described above.
[0083] Embodiments of the present disclosure may be embodied in a computer-readable medium. The computer-readable medium may include instructions that, when executed, cause a processor: obtain a set of previous states associated with the target application and a set of previous rewards corresponding to the set of previous states; generate, based on the set of previous states and the set of previous rewards, an action in natural language form for the target application; use a set of predefined rules to verify whether the action is reasonable; in response to verifying that the action is reasonable, generate an action code in computer language form corresponding to the action; and cause the target application to execute the action code. In addition, when executed, the instructions may also cause the processor to perform any other steps / processes of the method for action decision-making for a target application according to the embodiments of the present disclosure as described above.
[0084] It should be understood that all operations in the methods described above are merely exemplary, and the present disclosure is not limited to any operation in the methods or the order of these operations, but should cover all other equivalent transformations under the same or similar concepts. Additionally, unless otherwise specified or clearly understood from the context for the singular form, the articles "a" and "an" as used in this specification and the appended claims should generally be construed to mean "one" or "one or more".
[0085] It should also be understood that all modules in the devices described above may be implemented in various ways. These modules may be implemented as hardware, software, or a combination thereof. In addition, any of these modules may be further functionally divided into sub-modules or combined together.
[0086] Processors have been described in connection with various apparatuses and methods. These processors may be implemented using electronic hardware, computer software, or any combination thereof. Whether a processor is implemented as hardware or software will depend upon the particular application and the overall design constraints imposed on the system. As an example, a processor, any part of a processor, or any combination of processors given in this disclosure may be implemented using a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic device (PLD), a state machine, a gated logic unit, discrete hardware circuitry, and other suitable processing components configured to perform the various functions described in this disclosure. The functions of a processor, any part of a processor, or any combination of processors given in this disclosure may be implemented using software executed by a microprocessor, a microcontroller, a DSP, or other suitable platform.
[0087] Software should be broadly construed to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, threads of execution, procedures, functions, etc. Software may reside on a computer-readable medium. A computer-readable medium may include, for example, a memory, which may be, for example, a magnetic storage device (e.g., a hard disk, a floppy disk, a magnetic strip), an optical disk, a smart card, a flash memory device, a random access memory (RAM), a read only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), a register, or a removable disk. Although the memory is shown as being separate from the processor in several aspects given in this disclosure, the memory may also be located within the processor, such as a cache or a register.
[0088] The foregoing description has been provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein. All structural and functional equivalents to the elements of the various aspects described in this disclosure that are known or later become known to those of ordinary skill in the art are expressly incorporated herein and are covered by the claims.
Claims
1. A method for action decision-making for a target application, comprising: obtaining a set of previous states associated with the target application and a set of previous rewards corresponding to the set of previous states; generating, based on the set of previous states and the set of previous rewards, an action in natural language form for the target application; using a set of predefined rules to verify whether the action is reasonable; in response to verifying that the action is reasonable, generating an action code in computer language form corresponding to the action; and causing the target application to execute the action code.
2. The method according to claim 1, wherein, the set of predefined rules includes: rules for evaluating whether the action is included in a predefined set of allowable actions, and / or rules for evaluating whether the action is meaningless or untrustworthy.
3. The method according to claim 2, wherein, the set of allowable actions is determined by: obtaining a task design specification and / or metric description associated with the target application; preprocessing the task design specification and / or the metric description to generate a task description and / or a set of candidate metrics for the target application; and generating the set of allowable actions based on the task description and / or the set of candidate metrics.
4. The method according to claim 1, wherein, the action is generated by a language model, and the method further includes: in response to verifying that the action is unreasonable, adjusting hyperparameters of the language model and / or modifying the prompt words provided to the language model; and regenerating, by the language model, an action in natural language form for the target application based on the set of previous states and the set of previous rewards.
5. The method according to claim 1, further comprising: collecting the state generated by the target application executing the action code; calculating a reward corresponding to the action based on the action, the state, the set of previous states, and the set of previous rewards; and generating a subsequent action in natural language form for the target application based on the state, the reward, the set of previous states, and the set of previous states.
6. The method according to claim 5, wherein, the collecting the state generated by the target application executing the action code includes: determining observation metrics to be concerned during running the target application; collecting operation and maintenance data generated after the target application executes the action code; and extracting data corresponding to the observation metrics from the operation and maintenance data as the state.
7. The method according to claim 6, wherein, the determining observation metrics to be concerned during running the target application includes: obtaining a task design specification and / or metric description associated with the target application; preprocessing the task design specification and / or the metric description to generate a task description and / or a set of candidate metrics for the target application; and selecting the observation metrics from the set of candidate metrics based on the task description.
8. The method according to claim 5, further comprising: Store the state and / or the reward in a long-term memory associated with the target application.
9. An apparatus for action decision-making for a target application, comprising: a processor; and a memory storing computer-executable instructions that, when executed, cause the processor to: Obtain a set of previous states associated with the target application and a set of previous rewards corresponding to the set of previous states, Based on the set of previous states and the set of previous rewards, generate an action in natural language form for the target application, Use a set of predefined rules to verify whether the action is reasonable, In response to verifying that the action is reasonable, generate an action code in computer language form corresponding to the action, and Cause the target application to execute the action code.
10. The apparatus according to claim 9, wherein, The set of predefined rules includes: Rules for evaluating whether the action is included in a pre-determined set of allowable actions, and / or Rules for evaluating whether the action is meaningless or untrustworthy.
11. The apparatus according to claim 10, wherein, The set of allowable actions is determined by: Obtaining a task design specification and / or metric description associated with the target application; Preprocessing the task design specification and / or the metric description to generate a task description and / or a set of candidate metrics for the target application; and Generating the set of allowable actions based on the task description and / or the set of candidate metrics.
12. The apparatus according to claim 9, wherein, The action is generated by a language model, and the computer-executable instructions, when executed, further cause the processor to: In response to verifying that the action is unreasonable, adjust the hyperparameters of the language model and / or modify the prompt words provided to the language model; and Through the language model, regenerate an action in natural language form for the target application based on the set of previous states and the set of previous rewards.
13. The apparatus according to claim 9, wherein, The computer-executable instructions, when executed, further cause the processor to: Collect the state generated by the target application executing the action code; Based on the action, the state, the set of previous states and the set of previous rewards, calculate the reward corresponding to the action; and Based on the state, the reward, the set of previous states and the set of previous states, generate a subsequent action in natural language form for the target application.
14. The apparatus according to claim 13, wherein, The collecting the state generated by the target application executing the action code includes: Determining the observation metrics to be concerned during the operation of the target application; Collecting the operation and maintenance data generated after the target application executes the action code; and Extracting the data corresponding to the observation metrics from the operation and maintenance data as the state.
15. The apparatus according to claim 14, wherein, The determining the observation metrics to be concerned during the operation of the target application includes: Obtain a task design specification and / or metric description associated with the target application; Preprocess the task design specification and / or the metric description to generate a task description and / or a set of candidate metrics for the target application; and Select the observation metric from the set of candidate metrics based on the task description.
16. The apparatus according to claim 13, wherein, when the computer-executable instructions are executed, they further cause the processor to: Store the state and / or the reward in a long-term memory associated with the target application.
17. A computer-readable medium for action decision-making for a target application, comprising instructions that, when executed, cause a processor to: Obtain a set of previous states associated with the target application and a set of previous rewards corresponding to the set of previous states; Generate an action in natural language form for the target application based on the set of previous states and the set of previous rewards; Verify whether the action is reasonable using a set of predefined rules; In response to verifying that the action is reasonable, generate an action code in computer language form corresponding to the action; and Cause the target application to execute the action code.
18. The computer-readable medium according to claim 17, wherein, the set of predefined rules includes: Rules for evaluating whether the action is included in a pre-determined set of allowable actions, and / or Rules for evaluating whether the action is meaningless or untrustworthy.
19. The computer-readable medium according to claim 17, wherein, the action is generated by a language model, and when the instructions are executed, they further cause the processor to: In response to verifying that the action is unreasonable, adjust the hyperparameters of the language model and / or modify the prompt words provided to the language model; and Regenerate an action in natural language form for the target application based on the set of previous states and the set of previous rewards through the language model.
20. The computer-readable medium according to claim 17, wherein, when the instructions are executed, they further cause the processor to: Collect the state generated by the target application executing the action code; Calculate the reward corresponding to the action based on the action, the state, the set of previous states, and the set of previous rewards; and Generate a subsequent action in natural language form for the target application based on the state, the reward, the set of previous states, and the set of previous states.
Citation Information
Cited By
Vehicle underwater driving control method, vehicle and computer equipment
CN121934571A