System and method for implementing a sequential decision-making agent considering uncertain states

By constructing a sequential decision-making intelligent agent system that considers uncertain states, and combining prior and subsequent information with reinforcement learning methods, the decision-making difficulties caused by feedback delays are solved, achieving more efficient optimization and stable transformation results, and reducing the risk of exceeding limits.

CN115983321BActive Publication Date: 2026-04-24SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI JIAOTONG UNIV
Filing Date
2022-12-30
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing sequential decision-making agents struggle to achieve efficient optimization when faced with feedback delays and uncertainties in real-world environments, resulting in significantly reduced decision-making effectiveness and increased risk of exceeding limits.

Method used

By constructing a sequential decision-making agent system that considers uncertain states, combining prior and subsequent information with reinforcement learning methods, and utilizing a transformation delay distribution model and a deep neural network for action-state functions, a decision-making model for the agent is built to handle uncertain states and optimize the decision-making process.

Benefits of technology

It significantly improved the optimization performance of the decision-making agent, reduced the probability of exceeding limits, and improved the stability and efficiency of the transformation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115983321B_ABST
    Figure CN115983321B_ABST
Patent Text Reader

Abstract

A system and method for implementing a sequential decision-making agent considering uncertain states, comprising: a prior-posterior information combination processing module, an input-distribution decision-making agent module, wherein the prior-posterior information combination processing module utilizes the prior estimated information and the posterior real feedback information to obtain the distribution of the conversion amount and the unit conversion cost parameter; the input-distribution decision-making agent module samples the distribution information of the unit conversion cost parameter to obtain the corresponding discrete distribution, and inputs the distribution into the parallel action state neural network to obtain the optimal decision under the reference uncertain state. The present application utilizes the feature distribution and the reinforcement learning method when making sequential decisions, and significantly improves the optimization effect of the agent when making sequential decisions by constructing the agent with low complexity cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a technology in the field of neural network applications, specifically a system and method for implementing a sequential decision-making intelligent agent that considers uncertain states. Background Technology

[0002] In the current context of big data and information technology, limited by massive amounts of data and users' limited manual control capabilities, sequential decision-making agents are often needed to assist users in achieving various optimization goals. For example, in industrial automation, industrial internet, and autonomous driving, users utilize sequential decision-making agents for real-time control to achieve goals such as industrial parameter regulation, traffic allocation, and autonomous driving. The efficiency and optimization of sequential decision-making agents are therefore of paramount importance.

[0003] The problem of constructing a sequential decision-making agent is to improve the final optimization effect by adjusting the strategy in real time based on observed feedback information, given an unknown future environment. In recent years, many sequential decision-making methods based on this idea have been proposed. Most of these methods assume that the decision-making agent can observe real feedback in real time, thereby rationally adjusting the real-time decision-making strategy. However, existing decision-making agents ignore the uncertainty caused by the feedback delay in the real environment. Regardless of the theoretical effectiveness of these strategy adjustment methods, if they cannot obtain real feedback, the decision-making effect will be greatly reduced.

[0004] The uncertainty of feedback characteristics is a key property that distinguishes the online real-world environment from the offline ideal environment. It is determined by the complexity and randomness of feedback in the real-world environment. Taking the internet industry as an example, independent operators use automated decision-making agents to formulate strategies to compete for traffic and achieve conversions. However, whether the acquired traffic will lead to conversions, and when those conversions will occur, are highly uncertain. Conversion delays can last for hours or even days, which significantly impacts the decision-making performance of sequential decision-making agents with high real-time feedback requirements, leading to losses in optimization effectiveness and increasing the risk of exceeding limits.

[0005] The aforementioned independent operators refer to special users on the Internet who have traffic needs. They compete with other independent operators for traffic in order to achieve the effect of attracting traffic and thus achieve their own conversion goals.

[0006] The competition process refers to the process by which the platform, after receiving the competition strategy from each independent operator, allocates traffic according to certain rules and deducts the corresponding fees from the winner.

[0007] The conversion mentioned refers to the traffic acquisition needs of independent operators. For example, the conversion behavior defined by independent operators such as "live-streaming influencers" is the purchase behavior of traffic.

[0008] The aforementioned risk of exceeding limits refers to the independent operator setting a unit conversion cost for the decision-making agent, i.e., the cost required to win a conversion. If, for various reasons, the bidding results of the serialized decision-making agent cause the actual unit conversion cost to exceed the independent operator's expected value, this is called the "limit exceeding phenomenon." The limit exceeding phenomenon indicates that the serialized decision-making agent has not performed well in achieving its target task and is a key consideration in its design. Summary of the Invention

[0009] To address the aforementioned shortcomings of existing technologies, this invention proposes a serialization decision-making agent implementation system and method that considers uncertain states. By utilizing feature distribution and reinforcement learning methods during serialization decision-making and constructing an agent, the optimization performance of the agent during serialization decision-making is significantly improved with lower complexity and cost.

[0010] This invention is achieved through the following technical solution:

[0011] This invention relates to a sequential decision-making agent implementation system considering uncertain states, comprising: a priori and a posteriori information combination processing module and a decision-making agent module with distributed input. The priori and a posteriori information combination processing module performs comprehensive processing of prior estimated information and posterior real feedback information to obtain the distribution of conversion amount and unit conversion cost parameter. The decision-making agent module with distributed input samples the unit conversion cost parameter distribution information to obtain the corresponding discrete distribution, and inputs the distribution into a parallel action-state neural network to obtain the optimal decision under the reference uncertain state.

[0012] This invention relates to a method for implementing a serialized decision-making agent based on the above-described system, considering uncertain states, comprising:

[0013] Step 1: Combining the prior conversion information transmitted from the front link, the posterior conversion information observed by the agent, and the real-time feedback information, the distribution of the unit conversion cost of the traffic won by the agent at present is obtained using the conversion delay distribution model.

[0014] The aforementioned prior conversion information transmitted in the front-link refers to the estimated conversion rate (pcVR) of a specific traffic flow i provided by the platform to the agent before the agent makes a decision in the Internet industrial field. i For your reference.

[0015] The posterior transformation information of the real observation refers to the following: after an agent wins a traffic stream, the transformation result of that traffic stream is observed at a certain moment. If the traffic stream transformation can be observed, it is called positive posterior information; if the traffic stream transformation has not yet been observed, it is called negative posterior information.

[0016] The aforementioned real-time feedback information refers to information such as the agent's real-time expenditure and the remaining time in the decision-making cycle. This information is characterized by its immediacy and determinism.

[0017] The aforementioned conversion delay distribution model refers to the distribution of traffic conversion delay under the premise that traffic is ultimately converted; that is, when traffic is ultimately converted, its conversion delay is less than τ. i The probability of H is specifically: T (τ i )=P(T≤τ i ), where: H T (τ i τ is the conversion delay function, representing the distribution of possible conversion delays in the flow undergoing conversion; i This represents the time elapsed from the moment traffic i clicks to the current observation.

[0018] Step 2: Formalize the sequential decision problem and use reinforcement learning to obtain the solution under a given state.

[0019] The formal modeling mentioned refers to: modeling the decision problem with known traffic attributes as a linear programming problem in the offline phase. Under the premise of budget constraints and unit conversion cost constraints, the agent selects as many cost-effective traffic sources as possible. Specifically, the optimization objective is: Restrictions: Where: N is the total traffic during the competition cycle, x i To decide whether to bid for a specific traffic stream i, v i The value of traffic i (v) i =pctr i *pcvr i c i The cost of acquiring traffic i, B is the independent operator's budget, and tCPA is the independent operator's pre-set target cost per unit conversion.

[0020] The linear programming problem described has an optimal solution in an offline environment, and the optimal decision form is as follows: Where λ1 and λ2 are the dual factors corresponding to the two constraints in the dual problem of the linear programming problem.

[0021] The phrase "all traffic attributes are known" refers to a simulated environment constructed using historical data, in which the agent can obtain the value information v of all traffic throughout the day. i .

[0022] The reinforcement learning approach refers to further modeling the formally modeled linear programming problem into a Markov Decision Process (MDP) problem, and then using a policy deep neural network and an action-state function (Q-value) deep neural network to approximate the decision process. Specifically, for the state s at time t+1... t+1 Its possible distribution depends only on the state s at time t. t and the decision-making actions of the intelligent agent a t It is related to all states or agent decisions before time t, and is unrelated to the state s at time t. t It refers to an intelligent agent's perception of the current environment and its own properties.

[0023] The aforementioned deep neural network obtains the agent's policy action based on the agent's current state. When the agent learns the current state s, it refers to the deep neural network to determine the decision action.

[0024] The deep neural network for action state function obtains the action state function value based on the agent's current state and decision action. When the agent obtains the current state s, it can know the remaining time reward corresponding to each decision a, which is Q(s,a), where Q(s,a) is the action state function under the corresponding state s and decision action a.

[0025] Step 3: Considering the uncertainty of the current state, referencing the discrete distribution of the current state, and utilizing the theory of uncertain states, combined with the action-state function deep neural network in the reinforcement learning model, construct a sequential decision-making agent to assist independent traffic operators in making resource allocation decisions in the traffic allocation environment carried out on the platform.

[0026] The aforementioned state uncertainty refers to the fact that in a real environment, there is a delay in the feedback of certain physical quantities (such as unit conversion cost information), preventing the decision-making agent from obtaining the current accurate state. The decision-making agent relies on the real-time state to make decisions. The uncertainty of this physical quantity can be expressed using corresponding distribution information.

[0027] The discrete distribution refers to the fact that, given a finite number of N possible states, the agent can know its current specific state s. i The probability b(s) i ), i∈{1,2,…,N}.

[0028] The aforementioned uncertainty theory refers to: Where: a * Let b(s) be the current optimal decision, and b(s) be the discrete distribution of possible states. This optimal decision is responsible for all possible states.

[0029] Technical effect

[0030] This invention utilizes conversion delay to link prior and subsequent information; the input is a distributed decision model capable of handling uncertain states. Compared to existing technologies, this invention achieves more accurate estimation of relevant parameters and constructs possible parameter distributions; it achieves a more stable and efficient decision-making method, enabling independent operators to win traffic conversion effects and reducing the phenomenon of independent operators exceeding limits. Attached Figure Description

[0031] Figure 1 This is a flowchart illustrating the implementation of the present invention.

[0032] Figure 2 A schematic diagram for implementing an industrial scenario.

[0033] Figure 3 This is a schematic diagram illustrating the effect of an embodiment. Detailed Implementation

[0034] like Figure 1 As shown, this embodiment relates to a serialized decision-making agent implementation system and method that considers uncertain states, including:

[0035] The first step is to construct the distribution of unit conversion cost by combining prior and posterior information: The sequential decision-making agent needs to refer to real-time conversion volume and unit conversion cost information to make decisions. In the agent's estimation of the final conversion volume of acquired traffic, effective information includes prior and posterior information. Prior information is the estimated conversion rate, which is relatively certain, while posterior information evolves over time and specifically includes:

[0036] 1.1) For traffic that has already been clicked, introduce an event. For an agent to observe flow i at time t, whether the flow has already undergone transformation. For the agent to observe the transformation, Transformation has not yet been observed; introduce event p. i Whether traffic i ultimately converts, For traffic i to ultimately convert, Ultimately, no transformation occurred.

[0037] 1.2) The prediction of the total conversion amount of traffic at time t is redefined as... in: Let be the probability that flow i will eventually transform under posterior observation conditions. If flow transformation has already been observed, then flow transformation is inevitable. Then we get The total amount of this flow is denoted as For flows where no transformation was observed, a Bayesian variational process can be used to obtain... in: This refers to the probability that the agent has not yet observed the transformation of the flow under the condition that the flow will eventually transform. Its physical meaning is equivalent to: under the condition that the flow will eventually transform, T ≥ τ. i Where T is the conversion delay of the traffic, and τ i P(p) represents the time elapsed from the moment the click occurred to the current observation; while P(p) represents the time elapsed from the moment the click occurred to the current observation. i This represents the prior probability of traffic conversion, which is pcvr. i .

[0038] 1.3) Under the equivalent physical meaning described above, the following can be calculated: We obtain an unbiased estimate of the total amount of transformation at time t: Wherein: In the unbiased estimate of the total conversion at time t, each flow is independent and in Bernoulli variable form; Bernoulli variable form means that flow i will change with probability pcvr i Taking 1 represents traffic conversion; it will proceed with a probability (1-pcvr) i If the value is 0, it means the traffic has not been converted.

[0039] 1.4) After the above modeling, the total conversion amount is expressed as the sum of heterogeneous independent Bernoulli variables, which can be reduced to a distribution form by the central limit theorem. This distribution is the distribution of the possible results of the final total conversion amount. The total conversion amount distribution takes the form of a Gaussian distribution, which expresses the agent's perception of environmental uncertainty. The distribution of unit conversion cost is obtained by dividing the real-time expenditure by the above conversion amount distribution.

[0040] The second step is to model the decision-making process of converting the reference unit cost into a continuous decision-making problem under uncertainty. Specifically, after the above modeling, the online decision-making problem is remodeled into a sequential decision-making problem with state uncertainty. Based on this, a sequential decision-making agent with distributed input states can be designed.

[0041] This embodiment includes the sequential decision algorithm GQOUS, which specifically refers to: constructing a decision model under deterministic conditions using reinforcement learning based on the constraints of the optimization problem; this model uses the current state as a reference to make decisions, and the specific steps include:

[0042] i) Using the training process and method proposed by Yue He et al. in "A Unified Solution to Constrained Bidding in Online Display Advertising" (KDD'2020), an action-state function model under deterministic conditions is constructed. Model parameters are initialized, and using past samples, the model is continuously trained by interacting with the environment in a simulated environment. A penalty term is set to regulate model behavior. The decision model and the corresponding action-state function model Q(s,a) under deterministic conditions are obtained and kept for later use.

[0043] ii) In an online uncertain environment, the sequential decision-making agent observes the current state distribution information. The agent performs quantile sampling from this continuous distribution to obtain the discrete distribution b(s) corresponding to this state distribution.

[0044] iii) The agent makes decisions based on a discrete distribution, given N possible states s in the discrete distribution. i Let i∈{1,2,…,N}, and set N action-state function models Q(s). i Their inputs 'a' are all left blank. The N action-state function model Q(s) i a) Connect them in parallel, and they share the same decision action a. For each input a, sum the outputs of all the parallel deep neural networks to obtain ∑ i Q(s i ,a), is the global action state function considering the discrete state distribution.

[0045] iv) Obtain the decision action a that maximizes the global action state function. * This is the optimal solution in this decision-making process.

[0046] The aforementioned past samples refer to traffic records for a specific day, including the estimated click-through rate, estimated conversion rate, and the decisions required to win that traffic for each traffic item.

[0047] The simulated environment refers to the process of using past samples within a day to simulate the online decision-making and traffic competition of intelligent agents.

[0048] The aforementioned penalty refers to the reduction of the agent's gains if the simulation results in the simulation environment exceed the limits, which is for the purpose of exceeding the limits control.

[0049] The deterministic condition refers to the ideal condition where there is no feedback delay, under which the agent can accurately obtain all the information on which its decision depends.

[0050] Through specific practical experiments, in a simulated industrial scenario mimicking an online environment, this scenario provides millions of traffic data points per day, including the time, estimated click-through rate, estimated conversion rate, and cost required to acquire each traffic item. In the experiment, to reflect the conversion behavior characteristics of an online environment, the simulated conversions were given sparsity and latency. The USCB algorithm's policy model and Q-value model were used as comparison items, and the GQOUS algorithm was used as the experimental item. Two types of comparative experiments were constructed:

[0051] Experiment 1: Statistically analyze the average effect of 1000 different traffic allocation scenarios, running each traffic allocation scenario once; Experiment 2: Repeat the same traffic allocation scenario 1000 times, examine the average effect of 1000 times, and take 5 scenarios for comparison.

[0052] The sparsity in the experiment refers to the fact that the click and conversion results of the simulated traffic were obtained through 0-1 sampling.

[0053] The parameters used for sampling here are the estimated click-through rate and estimated conversion rate of the traffic.

[0054] The latency in the experiment refers to the fact that for traffic determined to be converted, a conversion delay is randomly assigned according to the delay distribution, and the agent can only observe its conversion if the time held by the traffic exceeds the delay.

[0055] Table 1 shows a comparison of the average results of 1000 repetitions of five traffic allocation scenarios in the simulation experiment. The performance of the USCB algorithm's policy deep neural network, action-state function deep neural network, and GQOUS algorithm are compared. Here, parameter G represents the relative conversion rate, and P represents the over-limit rate. The table shows that, compared to the other two methods, the decision agent constructed by the GQOUS algorithm has a higher average conversion rate in a single traffic allocation scenario and can effectively reduce the over-limit rate.

[0056] Table 1

[0057]

[0058] Table 2 shows a comparison of the average performance across 1000 different traffic allocation scenarios in the simulation experiment. Here, parameter G represents the relative conversion rate, and P represents the over-limit rate. The figure shows that the agent constructed using the GQOUS algorithm improves the daily conversion rate across different advertising campaigns and significantly reduces the campaign over-limit rate. This leads to the conclusion that the sequential decision-making agent constructed using the GQOUS algorithm has strong generalization and versatility.

[0059] Table 2

[0060]

[0061] like Figure 3 The figure shows the final conversion results of the GQOUS algorithm and two other algorithms in 1000 repeated experiments in a single scenario. As can be seen from the figure, the GQOUS algorithm generally achieves a higher number of conversions, with a mode of 2, indicating greater stability. Therefore, it can be concluded that GQOUS has a higher optimization effect and better robustness.

[0062] Compared with existing technologies, this invention constructs the current state distribution by combining a priori and post-a priori information, making the agent's understanding of the current state clearer. Furthermore, by using a decision agent with a distributed input, it can significantly improve the optimization effect of the decision agent in environments with delayed feedback and reduce the probability of exceeding the limit (meaning that the final unit conversion cost exceeds the value preset by the independent operator). It achieves decoupling of offline training and online inference processes with very low additional computational cost. It is simple to implement and can be deployed in existing reinforcement learning decision frameworks.

[0063] The above-described specific implementations can be partially adjusted by those skilled in the art in different ways without departing from the principles and purpose of the present invention. The scope of protection of the present invention is defined by the claims and is not limited to the above-described specific implementations. All implementation schemes within the scope of the claims are bound by the present invention.

Claims

1. A serialized decision-making agent implementation system considering uncertain states, characterized in that, include: The system consists of a combined prior and subsequent information processing module and a decision-making agent module with distributed input. The former and subsequent information processing module integrates prior prediction information and subsequent real feedback information to obtain the distribution of conversion volume and unit conversion cost parameters. The decision-making agent module with distributed input samples the unit conversion cost parameter distribution information to obtain the corresponding discrete distribution, and inputs the distribution into a parallel action state neural network to obtain the optimal decision under the reference uncertainty state. The optimal decision is obtained as follows: combining prior conversion information transmitted from the previous link, posterior conversion information observed by the agent, and real-time feedback information, the distribution of unit conversion cost of the traffic won by the agent is obtained using a conversion delay distribution model; the serial decision problem is formally modeled, and the solution under a deterministic state is obtained using reinforcement learning; considering the uncertainty of the current state, referring to the discrete distribution of the current state, the theory of uncertain states is used, combined with the action-state function deep neural network in the reinforcement learning model, to construct a serial decision agent to assist independent traffic operators in making resource allocation decisions in the traffic allocation environment carried out on the platform; The aforementioned prior transformation information transmitted in the front-link refers to the information provided by the platform to an agent in the Internet industrial field before the agent makes a decision, specifically a certain traffic flow. Estimated conversion rate For reference only; The posterior transformation information of the real observation refers to: after the agent wins a traffic, the transformation result of the traffic is observed at a certain moment. If the traffic transformation is observed, it is called positive posterior information; if the traffic transformation has not been observed, it is called negative posterior information. The aforementioned real-time feedback information refers to the agent's real-time expenditure and the remaining time in the decision-making cycle, which are characterized by their immediacy and determinism.

2. A method for implementing a serialized decision-making agent considering uncertain states based on the system of claim 1, characterized in that, include: Step 1: Combining the prior conversion information transmitted from the front link, the posterior conversion information observed by the agent, and the real-time feedback information, the distribution of the unit conversion cost of the traffic won by the agent at present is obtained using the conversion delay distribution model. Step 2: Formalize the sequential decision problem and use reinforcement learning to obtain the solution under a given state; Step 3: Considering the uncertainty of the current state, referencing the discrete distribution of the current state, and utilizing the theory of uncertain states, combined with the action-state function deep neural network in the reinforcement learning model, construct a sequential decision-making agent to assist independent traffic operators in making resource allocation decisions in the traffic allocation environment carried out on the platform.

3. The method for implementing a serialized decision-making agent considering uncertain states according to claim 2, characterized in that, The aforementioned conversion delay distribution model refers to the distribution of traffic conversion delay under the premise that traffic is ultimately converted; that is, when traffic is ultimately converted, its conversion delay is less than [a certain value]. The probability is as follows: ,in: Let be the conversion delay function, representing the distribution of possible conversion delays in the traffic that undergoes conversion; To obtain traffic Click to see the elapsed time for the current observation.

4. The method for implementing a serialized decision-making agent considering uncertain states according to claim 2, characterized in that, The formal modeling mentioned refers to: modeling the decision problem with known traffic attributes as a linear programming problem in the offline phase. Under the premise of budget constraints and unit conversion cost constraints, the agent selects the most cost-effective traffic. Specifically, the optimization objective is: Restrictions: ,in: This represents the total traffic during the competition cycle. To decide whether to win a certain traffic stream , For traffic value , To win traffic Expenses, For the budget of independent operators, The target unit conversion cost preset for independent operators; The linear programming problem described has an optimal solution in an offline environment, and the optimal decision form is bid = ,in and In the dual problem of a linear programming problem, the dual factors corresponding to the two constraints are: The phrase "all traffic attributes are known" refers to a simulated environment constructed using historical data, in which the agent can obtain the value information of all traffic throughout the day. .

5. The method for implementing a serialized decision-making agent considering uncertain states according to claim 2, characterized in that, The reinforcement learning approach refers to further modeling the formally modeled linear programming problem into a Markov Decision Process (MDP) problem, and then using a deep neural network for the policy and a deep neural network for the action-state function Q-value to approximate the decision process. Specifically, for state of time Its possible distribution is only related to state of time and the decision-making actions of intelligent agents Related to All states or agent decisions prior to that time are irrelevant. state of time It refers to an intelligent agent's perception of the current environment and its own properties.

6. The method for implementing a serialized decision-making agent considering uncertain states according to claim 5, characterized in that, The aforementioned deep neural network derives the agent's policy action based on the agent's current state. This action occurs when the agent learns of its current state. At that time, it refers to this deep neural network to determine the decision action; The deep neural network describing the action state function derives the action state function value based on the agent's current state and its decision action. When the agent obtains the current state... At that time, it learns about every decision The corresponding remaining time revenue is ,in That is, the corresponding state and decision-making actions The action state function below.

7. The method for implementing a serialized decision-making agent considering uncertain states according to claim 2, characterized in that, The uncertainty of the state refers to the fact that in a real environment, there is a delay in the feedback of unit conversion cost information, and the decision-making agent cannot obtain the current accurate state; while the decision-making agent relies on the real-time state to make decisions; and the uncertainty of unit conversion cost information is expressed using the corresponding distribution information.

8. The method for implementing a serialized decision-making agent considering uncertain states according to claim 2, characterized in that, The discrete distribution mentioned above refers to: for a finite The agent knows its current specific state from among several possible states. probability , .

9. The method for implementing a serialized decision-making agent considering uncertain states according to claim 2, characterized in that, The aforementioned uncertainty theory refers to: ,in: The optimal decision to make at this time is... Given a discrete distribution of possible states, the optimal decision is responsible for all possible states.

Citation Information

Patent Citations

  • Collaborative optimization decision-making method, system and device of energy internet and storage medium

    CN114977326A

  • Bicirculating application method and system for partially observable Markov decision problem

    CN115356923A