A service execution method and device, a storage medium and an electronic device
By combining conservative strategy optimization and local strategy convexity, the problem of inaccurate resource consumption estimation is solved, and maximum benefit optimization under constraints is achieved, thereby improving the accuracy and stability of business execution.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING SANKUAI NETWORK TECH CO LTD
- Filing Date
- 2023-10-17
- Publication Date
- 2026-07-24
AI Technical Summary
In constrained optimization problems, existing technologies often fail to accurately estimate resource consumption, leading to biases in strategy optimization and making it difficult to achieve maximum benefits with limited resources.
We employ a combination of conservative strategy optimization and local strategy convexification. By introducing upper confidence values and Lagrange multipliers, we optimize business execution strategies, reduce the uncertainty of resource estimation, and improve algorithm stability.
It improves the accuracy and stability of business execution strategies, ensures maximum benefits under constraints, and reduces the uncertainty of resource consumption.
Smart Images

Figure CN117391624B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to a business execution method, apparatus, storage medium, and electronic device. Background Technology
[0002] In scenarios involving the execution of specific business tasks, the executor needs to expend certain resources to achieve the desired outcome. The most crucial question in executing this type of task is how to maximize benefits with limited resources. This type of problem is typically referred to as a constrained optimization problem.
[0003] For example, in a scenario where a platform distributes multimedia content, merchants can bid to send multimedia information to users and generate revenue. For the platform, given that the total bid amount for multimedia distribution is fixed, how to maximize the revenue for merchants using limited bid space becomes a constrained optimization problem.
[0004] Currently, these algorithms typically transform the primal problem into a dual problem by introducing Lagrange multipliers. However, existing algorithms are highly sensitive to resource consumption estimates; inaccurate resource estimates can lead to errors in the Lagrange multipliers, and even small errors can cause significant deviations in policy optimization.
[0005] Therefore, how to optimize strategies more efficiently and stably in constrained optimization problems to achieve the best business performance is an urgent problem to be solved. Summary of the Invention
[0006] This specification provides a business execution method, apparatus, storage medium, and electronic device to at least partially solve the aforementioned problems existing in the prior art.
[0007] The following technical solution is adopted in this specification:
[0008] This specification provides a business execution method, including:
[0009] Obtain the current status of the target business;
[0010] Based on the current state and the predetermined business execution strategy, a target action for executing the target business is determined in the current state. The target action represents the resources required to execute the target business. The business execution strategy is obtained by optimizing the strategy to be optimized with the upper confidence value not being greater than the constraint value and the estimated benefit being maximized as the optimization objective. The upper confidence value is obtained based on the mean and variance of the estimated resources. The estimated resources and the estimated benefit are obtained based on each state of the target business and the corresponding actions of each state. The corresponding actions of each state are obtained based on each state and the strategy to be optimized. Each state is contained in the state space of the target business.
[0011] The target action is used to execute the target service.
[0012] Optionally, the target service is a multimedia delivery service, the status includes at least a user profile and a merchant profile, and the action is used to represent the bid made by the merchant when delivering multimedia information to the user.
[0013] Optionally, the estimated resources and estimated revenue are determined based on each state and the corresponding action, specifically including:
[0014] The states and their corresponding actions are input into a pre-trained analysis model to obtain the estimated resources and estimated profits output by the analysis model.
[0015] Optionally, there are at least two analytical models, each with a different structure and / or parameters;
[0016] Determining the mean and variance of the estimated resources specifically includes:
[0017] For each analysis model, the states and the corresponding actions are input into the analysis model to obtain the independent estimated resources output by the analysis model.
[0018] The average of each independent estimated resource is determined as the mean of the estimated resource, and the variance of each independent estimated resource is determined as the variance of the estimated resource.
[0019] Optionally, the analysis model is pre-trained, specifically including:
[0020] Obtain the sample state and the corresponding sample action, and obtain the annotation resources and annotation revenue of the sample state;
[0021] The sample state and the corresponding action are input into the analysis model to be trained to obtain the resource to be optimized and the benefit to be optimized output by the analysis model.
[0022] The analysis model is trained with the optimization objective of minimizing the difference between the resource to be optimized and the labeled resource, and minimizing the difference between the revenue to be optimized and the labeled revenue.
[0023] Optionally, the strategy to be optimized is optimized with the goal of ensuring that the upper confidence value is not greater than the constraint value and that the estimated return is maximized. Specifically, this includes:
[0024] The difference between the upper confidence value and the constraint value is defined as the dual difference;
[0025] Initialize the Lagrange multipliers and adjust the dual difference using the Lagrange multipliers to obtain the dual resources;
[0026] The optimization objective is to maximize the difference between the estimated revenue and the dual resource, and to optimize the strategy to be optimized and the Lagrange multiplier.
[0027] Optionally, before optimizing the policy to be optimized and the Lagrange multipliers, the method further includes:
[0028] The dominant term is determined based on the dual difference, and the dominant term is positively correlated with the square of the dual difference;
[0029] The optimization objective is to maximize the difference between the estimated revenue and the dual resource. The optimization of the strategy to be optimized and the Lagrange multiplier includes:
[0030] The optimization objective is to maximize the difference between the estimated revenue and the sum of the dual resources and the dominant term, and to optimize the strategy to be optimized and the Lagrange multiplier.
[0031] This specification provides a business execution device, the device comprising:
[0032] The acquisition module is used to acquire the current status of the target business.
[0033] The determination module is used to determine the target action for executing the target business in the current state based on the current state and a pre-determined business execution strategy. The target action is used to characterize the resources required to execute the target business. The business execution strategy is obtained by optimizing the strategy to be optimized with the upper confidence value not being greater than the constraint value and the estimated benefit being maximized as the optimization objective. The upper confidence value is obtained based on the mean and variance of the estimated resources. The estimated resources and the estimated benefit are obtained based on each state of the target business and the actions corresponding to each state. The actions corresponding to each state are obtained based on each state and the strategy to be optimized. Each state is contained in the state space of the target business.
[0034] An execution module is used to execute the target service using the target action.
[0035] This specification provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described business execution method.
[0036] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described business execution method.
[0037] The above-mentioned technical solutions adopted in this specification can achieve the following beneficial effects:
[0038] In the business execution method provided in this specification, the current state of the target business is obtained; based on the current state and a pre-determined business execution strategy, a target action for executing the target business in the current state is determined. The target action represents the resources required to execute the target business. The business execution strategy is obtained by optimizing the strategy to be optimized with the upper confidence value not exceeding the constraint value and the estimated benefit being maximized as the optimization objective. The upper confidence value is obtained based on the mean and variance of the estimated resources. The estimated resources and the estimated benefit are obtained based on each state of the target business and the corresponding action for each state. The corresponding action for each state is obtained based on each state and the strategy to be optimized. Each state is contained in the state space of the target business; the target business is executed using the target action.
[0039] When executing a target business with a constrained optimization problem using the business execution method provided in this specification, the target business can be executed based on the current state and the target action determined by the pre-determined business execution strategy. The business execution strategy can be updated and optimized by combining conservative strategy optimization and local strategy convexity proposed in this method, with the introduction of Lagrange multipliers to transform the primal problem into a dual problem. This method addresses the underestimation of predicted resource consumption in the dual problem through conservative strategy optimization. Subsequently, local strategy convexity is combined with augmented Lagrange multipliers to modify the original target, thereby convexifying the neighborhood of the locally optimal strategy. This gradually reduces the uncertainty of resource estimation within this region, ultimately improving the overall performance of the algorithm and obtaining a better business execution strategy. Attached Figure Description
[0040] The accompanying drawings, which are included to provide a further understanding of this specification and form part of this specification, illustrate exemplary embodiments and are used to explain this specification, but do not constitute an undue limitation thereof. In the drawings:
[0041] Figure 1 This is a flowchart illustrating one of the business execution methods described in this specification.
[0042] Figure 2 This specification provides a schematic diagram of a business execution device.
[0043] Figure 3 The corresponding information provided in this specification Figure 1 A schematic diagram of an electronic device. Detailed Implementation
[0044] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments in this specification without creative effort are within the scope of protection of this application.
[0045] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0046] Figure 1 This is a flowchart illustrating one business execution method described in this specification, which specifically includes the following steps:
[0047] S100: Obtain the current status of the target service.
[0048] All steps in the business execution method provided in this specification can be implemented by any electronic device with computing capabilities, such as a terminal, server, etc.
[0049] The purpose of this method is to determine the target action based on the state and execution strategy of the target business, and then execute the target business using the target action. Based on this, the current state of the target business can be determined first in this step. The target business can be any business in a scenario with constrained optimization problems.
[0050] There are various common target businesses that can be used as examples. For instance, taking multimedia advertising as an example, in a multimedia advertising scenario, merchants can spend costs to deliver multimedia messages to users through a platform to attract them to buy their products and generate corresponding revenue. Each multimedia ad delivery requires a certain cost from the merchant; this cost is the bid for the ad placement. Generally, merchants only need to set the goals they want to achieve with their multimedia ad delivery on the platform, such as return on investment (ROI), clicks, and conversions. Within a subsequent preset time period, i.e., an ad delivery cycle, the platform will intelligently help merchants plan their costs and complete the multimedia ad delivery through its multimedia advertising system. The upper limit of the cost that can be used for multimedia ad delivery within an ad delivery cycle is predetermined by the merchant. For the platform, if the cost of delivering multimedia information reaches the merchant's preset cost limit but fails to achieve the merchant's expected goals, it will result in a poor merchant experience and cause dissatisfaction, which is something the platform needs to avoid. Therefore, how the platform can help merchants maximize their revenue with limited resources is a constrained optimization problem.
[0051] For example, taking cloud service as the target business, in a cloud service scenario, the platform provides users with cloud computing resources to meet their computing needs in exchange for revenue. In this scenario, the total amount of resources the platform can provide has an upper limit, while the computing needs made by users each time are fixed. Therefore, how the platform can utilize limited cloud computing resources to maximize its own revenue is a constrained optimization problem.
[0052] The status of a target business can refer to relevant information during its execution. Any target business can exist in multiple different statuses, and the relevant information contained in these statuses varies. Taking multimedia advertising as an example, the status of multimedia advertising can include, but is not limited to, time factors, user profiles, and merchant profiles.
[0053] S102: Based on the current state and the pre-determined business execution strategy, determine the target action for executing the target business in the current state. The target action represents the resources required to execute the target business. The business execution strategy is obtained by optimizing the strategy to be optimized with the upper confidence value not being greater than the constraint value and the estimated benefit being maximized as the optimization objective. The upper confidence value is obtained based on the mean and variance of the estimated resources. The estimated resources and the estimated benefit are obtained based on each state of the target business and the actions corresponding to each state. The actions corresponding to each state are obtained based on each state and the strategy to be optimized. Each state is contained in the state space of the target business.
[0054] In this step, the target action for executing the target service can be determined based on the current state of the target service as determined in step S100 and the pre-determined service execution strategy. The target action represents the resources required to execute the target service; the service execution strategy represents the rules or methods for determining the target action in different states of the target service. In other words, the service execution strategy determines how many resources should be consumed to execute the target service in each state.
[0055] For example, suppose the target service has five states: S1, S2, S3, S4, and S5. Then the service execution strategy can be represented as {S1: C1, S2: C2, S3: C3, S4: C4, S5: C5}, which means that the target service is executed by consuming resource C1 in state S1, resource C2 in state S2, and so on.
[0056] Taking multimedia advertising as an example, the number of slots available for platforms to display multimedia information to users is limited. Therefore, merchants often need to compete for these slots. For each multimedia slot, the merchant who consumes more resources—that is, the one with the higher bid—is able to advertise in that slot. Of course, during the competition phase, the bid only represents the cost a merchant is willing to pay; only the merchant who wins the slot actually incurs the corresponding cost. Merchants who don't win the slot don't actually incur any cost. For merchants, the cost they can consume within a campaign period is limited. If they consistently use high bids for multimedia advertising, they may end up failing to achieve their goals due to insufficient campaign runs. Therefore, having a reasonable business execution strategy—that is, how to consume costs for multimedia advertising in different states—is crucial. For example, if the preferences of active users match the multimedia information offered by the merchant within a certain period of time, then a higher bid can be considered to compete for the multimedia placement slots; conversely, if the preferences of active users do not match the multimedia information offered by the merchant within a certain period of time, then a lower bid can be considered to compete or to give up the competition.
[0057] Similarly, in cloud service scenarios, platforms can choose to consume different cloud computing resources to obtain different benefits when meeting user needs. In this scenario, business status may include, but is not limited to, user profiles, user needs, and resource information; the business execution strategy is to determine how to consume computing resources to meet user needs in each status.
[0058] In constrained optimization problems where business execution strategy is crucial, obtaining a better business execution strategy is the most critical aspect of this method. The business execution method provided in this specification specifically involves determining the business execution strategy by: acquiring the state space and constraint values of the target business; determining the actions corresponding to each state based on the states contained in the state space and the strategy to be optimized; determining the estimated resources and estimated revenue based on each state and its corresponding actions; determining the upper confidence value based on the mean and variance of the estimated resources; and optimizing the strategy to be optimized with the upper confidence value not exceeding the constraint value and the estimated revenue being maximized as the optimization objective, thereby obtaining the business execution strategy.
[0059] Since the business execution strategy determines how many resources to consume to execute the target business in each state, the first step in determining the business execution strategy is to obtain the state space of the target business. The state space of the target business contains all possible states of the target business. Constraints can be the constraints on task execution, such as the total amount of resources.
[0060] In this method, the business execution strategy is not generated directly from scratch in one step, but is ultimately determined after continuous updates and optimizations. Therefore, when determining the business execution strategy, there is initially a strategy to be optimized. The business execution strategy can be optimized multiple times; that is, in this method, the process of determining the business execution strategy can be executed multiple times, continuously updating and optimizing the strategy to be optimized until a satisfactory business execution strategy is finally obtained. In each round of executing the above process, the strategy to be optimized can be the business execution strategy obtained after the previous round. In the initial round, the strategy to be optimized can be randomly initialized.
[0061] Based on the states and optimization strategies within the obtained state space, the actions corresponding to each state during this optimization process can be determined, which represents the resources required to execute the target business in each state. Based on this, the estimated resources and estimated revenue can be determined according to each state and its corresponding actions. Estimated resources represent the total resources expected to be consumed after completing a preset cycle of target business execution, and estimated revenue represents the total revenue expected after completing a preset cycle of target business execution. In optimizing the business execution strategy, the preset cycle can be represented by the number of times the target business is completed, or by a specific time period, depending on the requirements. In actual target business execution scenarios, the cycle is generally set by the business executor. For example, in a multimedia advertising scenario, the merchant can set the number of multimedia campaigns to be conducted or the duration of the multimedia campaign.
[0062] Taking multimedia advertising as an example, in this scenario, the business execution strategy is a bidding strategy that helps merchants intelligently bid during multimedia advertising campaigns. When the multimedia advertising system is not yet online and is in the adjustment phase, the bidding strategy can be continuously optimized through testing. After each round, or a preset period of multimedia advertising, the bidding strategy can be optimized and adjusted based on the final results. In a preset period of multimedia advertising, the constraint value is the total cost that the merchant has pre-set to consume, the estimated resources are the total cost the merchant expects to consume after completing all multimedia advertising, and the estimated revenue is the total revenue the merchant expects to receive after completing all multimedia advertising. In each round of multimedia advertising, the corresponding action for each state can be determined based on the current bidding strategy, that is, the single bid for each state, thus obtaining the expected total cost and total revenue. The bidding strategy is optimized and adjusted with the optimization goal of ensuring that the expected total cost does not exceed the merchant's preset total cost and maximizing the expected total revenue.
[0063] In cloud service scenarios, the business execution strategy refers to the platform's resource consumption strategy when meeting user needs. In a cloud service provision scenario with a preset period, the constraint is the total amount of cloud computing resources the platform can provide, the estimated resources are the total amount of resources the platform expects to consume after meeting all user needs, and the estimated revenue is the total revenue the platform expects to obtain after meeting all user needs. When providing cloud services, the amount of cloud computing resources provided for a single user's needs in each state can be determined based on the current resource consumption strategy, thus obtaining the expected total resource consumption and total revenue. The resource consumption strategy is optimized and adjusted with the optimization objective of ensuring that the total resource consumption does not exceed the total amount of cloud computing resources the platform can provide and maximizing the expected total revenue.
[0064] There are various methods for determining the estimated resources and estimated revenues. This specification provides one specific implementation method for reference. Since the state of the target business is not fixed each time it is executed within a preset cycle, a neural network model can be used to learn the transition relationships between the states in the target business scenario, and the estimated resources and estimated revenues can be obtained by combining the actions corresponding to each state. Specifically, the states and their corresponding actions can be input into a pre-trained analysis model to obtain the estimated resources and estimated revenues output by the analysis model.
[0065] The structure of the analysis model can be configured according to requirements, as long as it can determine the estimated resources and estimated revenue. This manual does not impose specific restrictions on this. Additionally, depending on the structure or training method of the analysis model, a preset period can be input into the analysis model, allowing it to predict the estimated resources and estimated revenue based on each state, its corresponding action, and the preset period.
[0066] Based on the data identified above, a more specific constrained optimization problem can be derived. Here, S represents the state space, S = {s0, s1, s2, ...}, A represents the action space, A = {a0, a1, a2, ...}, with s and a corresponding to the same label; π represents the strategy to be optimized. The constrained optimization problem can be defined using the following formula:
[0067]
[0068]
[0069] The above formula will be referred to as Formula (1) in the following sections of this specification. Wherein, ρ represents the stationary distribution generated by the strategy π to be optimized. Indicates estimated resources, d represents the estimated revenue, and d represents the constraint value.
[0070] For this type of problem, the common approach is to introduce the Lagrange multiplier λ to transform the primal problem into a dual problem for solution. Specifically, it can be done as follows:
[0071]
[0072] In the following sections of this manual, the above formula will be referred to as Formula (2). During multiple rounds of update and optimization, the above formula is solved by alternately optimizing the strategy to be optimized π and updating the Lagrange multiplier λ, thus obtaining the final business execution strategy.
[0073] Generally, if the estimated resources are no greater than the constraint value and the optimization objective is to maximize the estimated revenue, the optimization strategy can be optimized using the formula given above. However, it is important to consider that in constrained optimization problems, using dual methods can easily lead to an underestimation of the resources required to execute the target business. With low estimated resources, the feasibility boundary of the optimization strategy can easily exceed the feasible space during the optimization process, resulting in a very aggressive business execution strategy and ultimately violating the constraints. To address this issue, this method proposes using an upper confidence boundary for resources to encourage overestimation, thereby generating a more conservative feasibility boundary for the optimization strategy, narrowing the feasible space of the strategy, and thus improving constraint satisfaction.
[0074] In this method, the upper confidence value of the resource can be determined based on the mean and variance of the estimated resource. When solving the dual form of the constrained optimization problem using formula (2), when the Lagrange multiplier is greater than 0, the optimization strategy will estimate the estimated resource by minimizing resource consumption. Assuming that the resource consumption distribution is a Gaussian distribution with noise and zero mean each time the target business is executed, the zero mean attribute is usually not maintained under the operation of the minimization function, which will lead to the estimated resource being less than the actual resource consumed.
[0075] This method proposes a conservative optimization strategy to address the aforementioned problem. In this method, a determined upper confidence value for the resource is used. To replace the estimated resources in the dual problem In this process, the upper confidence value can be calculated using the following formula:
[0076]
[0077] In the following sections of this manual, the above formula will be referred to as formula (3). This represents the upper confidence value, and N represents the number of estimated resources identified. Let represent the i-th estimated resource, and k represent the weight of the variance.
[0078] Obviously, under the above method, multiple estimated resources need to be determined. There are various ways to determine multiple estimated resources; this specification provides a specific embodiment for reference. In the embodiment given above, which uses analytical models to determine estimated resources and estimated revenues, multiple analytical models can be used to predict multiple independent estimated resources. Specifically, for each analytical model, the states and their corresponding actions are input into the model to obtain the independent estimated resources output by the model; the average of the independent estimated resources is determined as the mean of the estimated resources, and the variance of the independent estimated resources is determined as the variance of the estimated resources. Each analytical model can be designed or trained differently to obtain analytical models with different structures and / or parameters, thus yielding different prediction results.
[0079] Each analysis model can be pre-trained. Specifically, sample states and corresponding sample actions can be obtained, along with labeled resources and labeled rewards for those sample states. The sample states and their corresponding actions are then input into the analysis model to be trained, yielding the resources and rewards to be optimized output by the model. The analysis model is trained with the optimization objective of minimizing the difference between the resources to be optimized and the labeled resources, and minimizing the difference between the rewards to be optimized and the labeled rewards.
[0080] Labeled resources and labeled revenue refer to the actual resources consumed and the actual revenue gained after executing the target business using sample states and corresponding sample actions within a preset period. Sample states, corresponding sample actions, and labeled resources and revenues for each sample state can all be obtained from historical execution data of the target business. When multiple analysis models exist, each model can be trained independently using the above method. When analysis models with the same structure exist, different training samples can be used to train different analysis models.
[0081] After obtaining the upper confidence value of the resources, the upper confidence value can be used to replace the estimated resources in the constrained optimization problem. That is, the upper confidence value is no greater than the constraint value, and the optimization objective is to maximize the estimated profit, and then optimize the optimization strategy. It should be noted that when solving the dual problem, the Lagrange multipliers also need to be updated during the optimization of the optimization strategy.
[0082] At this point, the dual problem corresponding to the constrained optimization problem is solved. Specifically, when optimizing the strategy to be optimized, the difference between the upper confidence value and the constraint value can be determined as the dual difference; the Lagrange multipliers are initialized, and the dual difference is adjusted using the Lagrange multipliers to obtain the dual resources; the optimization objective is to maximize the difference between the estimated revenue and the dual resources, and the strategy to be optimized and the Lagrange multipliers are optimized. The above can be expressed by the following formula:
[0083]
[0084] The above formula will be referred to as formula (4) in the following sections of this specification. The estimated resources in formula (2) Replace with the determined upper confidence value This allows us to obtain the above formula (4).
[0085] Furthermore, during the optimization of the target strategy, the dual method is highly sensitive to resource estimation. Even a small error in the resource estimation can cause significant deviations in subsequent updates and optimizations, even leading to optimizations in the completely opposite direction. For example, when the actual resource consumption is lower than the constraint value, but the estimated resource exceeds the constraint value, the Lagrange multipliers will be updated in a direction completely opposite to the actual direction, becoming a misleading penalty term in the original objective.
[0086] This method addresses the aforementioned problem by employing local policy convexity. Specifically, an augmented Lagrangian method is used to modify the original objective, thereby convexifying the neighborhood of the locally optimal policy. By correcting the policy gradient, local policy convexity can stabilize policy learning within a local region, gradually reducing the uncertainty of resource prediction in that region.
[0087] Specifically, before optimizing the strategy and Lagrange multipliers, a dominant term can be determined based on the duality difference, where the dominant term is positively correlated with the square of the duality difference. In this case, maximizing the difference between the estimated return and the sum of the dual resources and the dominant term can be used as the optimization objective to optimize the strategy and the Lagrange multipliers. The above method can be expressed by the following formula:
[0088] In the following sections of this specification, the above formula will be referred to as Formula (5). Here, c is a hyperparameter greater than zero. When updating and optimizing the strategy according to Formula (5), for sufficiently large c, the objective function will be dominated by the quadratic penalty term, thus becoming convex in the neighborhood of the boundary solution. Here, the boundary solution can be understood as... π at that time.
[0089] S104: Execute the target service using the target action.
[0090] After obtaining the target action in step S102, the target business can be executed in this step using the obtained target action.
[0091] For example, in multimedia advertising, the target action is for the merchant to bid for the multimedia advertising campaign. The merchant can use the obtained bid to compete for multimedia advertising slots. If the bid is successful, the merchant can incur the corresponding costs to place multimedia information in the slot; if the bid is unsuccessful, no costs are incurred.
[0092] For example, in a cloud service scenario, the target action is the amount of cloud computing resources consumed by the platform when fulfilling a user's request. Unlike multimedia delivery scenarios, there is no competition in cloud service scenarios; the platform will inevitably consume corresponding computing resources when fulfilling a user's request.
[0093] In this method, conservative policy optimization and local policy convexity can be used in combination. Local policy convexity can stabilize and concentrate the policy to be optimized, bringing it closer to a local optimum. This allows the collected samples to be concentrated in a region that conforms to the distribution of the local optimum policy, which in turn eliminates cognitive uncertainty within the convexity region. As uncertainty decreases, the independent resource estimates obtained by each analytical model become more accurate, and the inconsistency between the independent resource estimates predicted by each analytical model, which causes cognitive uncertainty, also decreases, meaning the variance of the estimated resources gradually approaches zero. In this way, the conservatism in resource estimation caused by prediction inconsistency is also reduced, thus the conservative prediction boundary can be gradually pushed towards the true boundary.
[0094] When executing a target business with a constrained optimization problem using the business execution method provided in this specification, the target business can be executed based on the current state and the target action determined by the pre-determined business execution strategy. The business execution strategy can be updated and optimized by combining conservative strategy optimization and local strategy convexity proposed in this method, with the introduction of Lagrange multipliers to transform the primal problem into a dual problem. This method addresses the underestimation of predicted resource consumption in the dual problem through conservative strategy optimization. Subsequently, local strategy convexity is combined with augmented Lagrange multipliers to modify the original target, thereby convexifying the neighborhood of the locally optimal strategy. This gradually reduces the uncertainty of resource estimation within this region, ultimately improving the overall performance of the algorithm and obtaining a better business execution strategy.
[0095] The above describes the business execution method provided in this manual. Based on the same approach, this manual also provides corresponding business execution devices, such as... Figure 2 As shown.
[0096] Figure 2 A schematic diagram of a business execution device provided in this specification specifically includes:
[0097] The acquisition module 200 is used to acquire the current status of the target business.
[0098] The determining module 202 is used to determine, based on the current state and a pre-determined business execution strategy, a target action for executing the target business in the current state. The target action represents the resources required to execute the target business. The business execution strategy is obtained by optimizing the strategy to be optimized with the upper confidence value not being greater than the constraint value and the estimated benefit being maximized as the optimization objective. The upper confidence value is obtained based on the mean and variance of the estimated resources. The estimated resources and the estimated benefit are obtained based on each state of the target business and the actions corresponding to each state. The actions corresponding to each state are obtained based on each state and the strategy to be optimized. Each state is contained in the state space of the target business.
[0099] The execution module 204 is used to execute the target service using the target action.
[0100] Optionally, the target service is a multimedia delivery service, the status includes at least a user profile and a merchant profile, and the action is used to represent the bid made by the merchant when delivering multimedia information to the user.
[0101] Optionally, the device further includes a pre-determining module 206, specifically used to input the states and the actions corresponding to the states into a pre-trained analysis model to obtain the estimated resources and estimated revenues output by the analysis model.
[0102] Optionally, there are at least two analytical models, each with a different structure and / or parameters;
[0103] The pre-determining module 206 is specifically used to input each state and the corresponding action into the analysis model for each analysis model to obtain the independent estimated resources output by the analysis model; to determine the average of each independent estimated resource as the mean of the estimated resources, and to determine the variance of each independent estimated resource as the variance of the estimated resources.
[0104] Optionally, the pre-determining module 206 is specifically used to obtain the sample state and the sample action corresponding to the sample state, and to obtain the labeled resources and labeled benefits of the sample state; input the sample state and the action corresponding to the sample state into the analysis model to be trained, and obtain the resources to be optimized and the benefits to be optimized output by the analysis model; and train the analysis model with the optimization objective of minimizing the difference between the resources to be optimized and the labeled resources, and minimizing the difference between the benefits to be optimized and the labeled benefits.
[0105] Optionally, the pre-determining module 206 is specifically used to determine the difference between the upper confidence value and the constraint value as the dual difference; initialize the Lagrange multiplier and adjust the dual difference using the Lagrange multiplier to obtain the dual resource; and optimize the strategy to be optimized and the Lagrange multiplier with the maximum difference between the estimated revenue and the dual resource as the optimization objective.
[0106] Optionally, the pre-determining module 206 is specifically used to determine the dominant term based on the dual difference, wherein the dominant term is positively correlated with the square of the dual difference; and to optimize the strategy to be optimized and the Lagrange multiplier with the optimization objective of maximizing the difference between the estimated revenue and the sum of the dual resources and the dominant term.
[0107] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above-described... Figure 1 The business execution methods provided.
[0108] This instruction manual also provides Figure 3 The diagram shows a schematic structural representation of the electronic device. Figure 3 At the hardware level, the electronic device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for the business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to achieve the above-mentioned functions. Figure 1 The described business execution method. Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. In other words, the execution subject of the following processing flow is not limited to individual logic units, but can also be hardware or logic devices.
[0109] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0110] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0111] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0112] For ease of description, the above devices are described in terms of function, divided into various units. Of course, in implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware components.
[0113] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0114] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0115] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0116] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0117] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0118] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0119] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0120] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0121] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0122] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0123] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0124] The above description is merely an embodiment of this specification and is not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this application.
Claims
1. A business execution method, characterized in that, include: Obtain the current status of the target business; Based on the current state and the predetermined business execution strategy, a target action for executing the target business is determined in the current state. The target action represents the resources required to execute the target business. The business execution strategy is obtained by optimizing the strategy to be optimized with the upper confidence value not being greater than the constraint value and the estimated benefit being maximized as the optimization objective. The upper confidence value is obtained based on the mean and variance of the estimated resources. The estimated resources and the estimated benefit are obtained based on each state of the target business and the corresponding actions of each state. The corresponding actions of each state are obtained based on each state and the strategy to be optimized. Each state is contained in the state space of the target business. Based on each state and the corresponding action, the estimated resources and estimated revenue are determined, specifically including: The states and their corresponding actions are input into a pre-trained analysis model to obtain the estimated resources and estimated profits output by the analysis model. There are at least two analytical models, and each analytical model has a different structure and / or parameters; Determining the mean and variance of the estimated resources specifically includes: For each analysis model, the states and the corresponding actions are input into the analysis model to obtain the independent estimated resources output by the analysis model. The average of each independent estimated resource is determined as the mean of the estimated resource, and the variance of each independent estimated resource is determined as the variance of the estimated resource. The target action is used to execute the target service.
2. The method as described in claim 1, characterized in that, The target business is a multimedia delivery business, and the state includes at least a user profile and a merchant profile. The action is used to represent the bid made by the merchant when delivering multimedia information to the user.
3. The method as described in claim 1, characterized in that, Pre-training the analysis model specifically includes: Obtain the sample state and the corresponding sample action, and obtain the annotation resources and annotation revenue of the sample state; The sample state and the corresponding action are input into the analysis model to be trained to obtain the resource to be optimized and the benefit to be optimized output by the analysis model. The analysis model is trained with the optimization objective of minimizing the difference between the resource to be optimized and the labeled resource, and minimizing the difference between the revenue to be optimized and the labeled revenue.
4. The method as described in claim 1, characterized in that, The optimization objective is to maximize the estimated return while ensuring that the confidence value is not greater than the constraint value. The optimization process includes: The difference between the upper confidence value and the constraint value is defined as the dual difference; Initialize the Lagrange multipliers and adjust the dual difference using the Lagrange multipliers to obtain the dual resources; The optimization objective is to maximize the difference between the estimated revenue and the dual resource, and to optimize the strategy to be optimized and the Lagrange multiplier.
5. The method as described in claim 4, characterized in that, Before optimizing the policy to be optimized and the Lagrange multipliers, the method further includes: The dominant term is determined based on the dual difference, and the dominant term is positively correlated with the square of the dual difference; The optimization objective is to maximize the difference between the estimated revenue and the dual resource. The optimization of the strategy to be optimized and the Lagrange multiplier includes: The optimization objective is to maximize the difference between the estimated revenue and the sum of the dual resources and the dominant term, and to optimize the strategy to be optimized and the Lagrange multiplier.
6. A business execution device, characterized in that, include: The acquisition module is used to acquire the current status of the target business. The determination module is used to determine the target action for executing the target business in the current state based on the current state and a pre-determined business execution strategy. The target action is used to characterize the resources required to execute the target business. The business execution strategy is obtained by optimizing the strategy to be optimized with the upper confidence value not being greater than the constraint value and the estimated benefit being maximized as the optimization objective. The upper confidence value is obtained based on the mean and variance of the estimated resources. The estimated resources and the estimated benefit are obtained based on each state of the target business and the actions corresponding to each state. The actions corresponding to each state are obtained based on each state and the strategy to be optimized. Each state is contained in the state space of the target business. An analysis module is used to determine estimated resources and estimated profits based on each state and the corresponding actions. The analysis module inputs each state and the corresponding actions into a pre-trained analysis model to obtain the estimated resources and estimated profits output by the analysis model. At least two analysis models exist, each with a different structure and / or parameters. For each analysis model, the analysis module inputs each state and the corresponding actions into that model to obtain independent estimated resources output by that model. The average of these independent estimated resources is determined as the mean of the estimated resources, and the variance of these independent estimated resources is determined as the variance of the estimated resources. An execution module is used to execute the target service using the target action.
7. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the method described in any one of claims 1 to 5.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method described in any one of claims 1 to 5.