Multi-model fusion decision-making method and system based on deep reinforcement learning
Through the multi-model fusion decision-making method of deep reinforcement learning, the model weights are adaptively allocated by Actor and Critic networks, the problems of high computing complexity and low real-time performance in the existing technology are solved, and efficient and accurate multi-model fusion decisions are achieved, which are suitable for financial, production and traffic flow scenarios.
Patent Information
- Application Number
- CN202510498793.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-18
AI Technical Summary
The existing multi-model fusion decision-making method has high computational complexity and low operational efficiency in continuous decision space, and the allocation of weights by manual experience leads to low real-time and automation of the system.
The multi-model fusion decision-making method of deep reinforcement learning is adopted, and the model weights are adaptively allocated through the Actor network and the Critic network, the fusion strategy is dynamically updated, and the reward benefits are evaluated in combination with the backtest system, and the network parameters are updated.
It realizes efficient and accurate multi-model fusion decision-making in a dynamic environment, improves the real-time and automation of the system, and is suitable for complex decision-making problems in multiple scenarios.
Smart Images

Figure CN120338034A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-model fusion decision-making, and specifically provides a multi-model fusion decision-making method and system based on deep reinforcement learning. Background Art
[0002] Existing multi-model fusion decision-making methods can be divided into end-to-end methods and two-stage methods.
[0003] The end-to-end method means that input data can be directly input into a large integrated model of multiple models to directly output the final decision-making strategy. The advantage of this fusion method is that it is intuitive and simple, and only requires end-to-end training of the integrated large model during training. However, if the decision space is a continuous space (such as cases with continuous numerical ranges like asset value, risk probability, etc.), the end-to-end method will have problems of high computational complexity and low operating efficiency when searching for the optimal solution in the continuous space. The two-stage method is to input the input data into each model respectively, first obtain the independent output results of each model, and then perform weighted fusion on the output results of each model to obtain the final decision-making strategy. Compared with the end-to-end method, the two-stage method has higher operating efficiency. However, the disadvantage of this method is that the weight of each model in the fusion process often depends on manual allocation based on experience, and when the external environmental conditions change, it is necessary to manually intervene to re-adjust the weights of model fusion, which greatly reduces the real-time performance and automation degree of the system.
[0004] Therefore, we propose a multi-model fusion decision-making method and system based on deep reinforcement learning. Summary of the Invention
[0005] The purpose of the present invention is to provide a multi-model fusion decision-making method and system based on deep reinforcement learning, which solves the problems proposed in the background art.
[0006] To achieve the above purpose, the present invention provides the following technical solution: A multi-model fusion decision-making method based on deep reinforcement learning, including the following method steps:
[0007] Step 1: Initialization: At the starting moment, initialize the weights of the reinforcement learning agent, where the agent includes an Actor network and a Critic network, which are respectively used to generate the model weight allocation strategy and evaluate the strategy value.
[0008] Step 2: Obtain the current system state.
[0009] Step 3: The Actor network outputs the weight allocation of the model according to the current system state.
[0010] Step 4: Perform weighted fusion on the outputs of multiple models according to the weight assignment to obtain the final output strategy;
[0011] Step 5: Input the fused output strategy into the backtest system to evaluate the strategy performance and obtain the reward income;
[0012] Step 6: Update the total resource amount according to the reward income and the current total resource amount;
[0013] Step 7: Store the current system state, fusion weights, and reward income in the buffer to form a training dataset;
[0014] Step 8: Sample training data from the buffer, update the parameters of the Actor network and the Critic network, and repeat Steps 2 to 8 until the preset termination condition is met.
[0015] As a preferred embodiment of the present invention, the parameter initialization method of the Actor network and the Critic network is Gaussian distribution initialization or Kaiming initialization.
[0016] As a preferred embodiment of the present invention, the model historical performance data is a performance index vector of each model in the historical time period, and the environmental variable factor is an external environmental parameter data vector.
[0017] As a preferred embodiment of the present invention, in the resource update step, the calculation rule of the total resource amount at the next moment is: weighted summation based on the current total resource amount and the reward income.
[0018] As a preferred embodiment of the present invention, the parameter update of the Actor network is achieved by the gradient ascent method of maximizing the cumulative reward; the parameter update of the Critic network is achieved by the gradient descent method of minimizing the value function error.
[0019] The present invention also relates to a multi-model fusion decision-making system based on deep reinforcement learning, including:
[0020] Initialization module: used to construct and initialize the Actor network, Critic network, and system state;
[0021] Weight generation module: dynamically generate model fusion weights through the Actor network;
[0022] Fusion decision module: perform weighted fusion of multi-model strategies;
[0023] Backtest evaluation module: configure resources and evaluate strategy performance to generate reward income;
[0024] Resource management module: update the total resource amount and record state data;
[0025] Data storage module: manages the training data set in the buffer;
[0026] Parameter update module: updates the parameters of the Actor network and the Critic network based on the buffer data.
[0027] Furthermore, the model library module includes multiple models for handling different tasks.
[0028] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0029] The present invention adaptively fuses the output strategies of multiple models by the reinforcement learning agent intelligently allocating fusion weights for multiple models. Compared with the traditional method of manually assigning fusion weights based on experience, the proposed intelligent multi-model fusion strategy can achieve more accurate decision-making;
[0030] The implemented system can update the fusion strategy of multi-model fusion in real time according to the external environmental changes (environmental variable factors, system allocable resources, historical performance of individual models), and has the adaptive ability to the dynamic external environment;
[0031] The proposed method can adapt to complex decision-making problems in multiple scenarios, support different types of data in different scenarios (such as financial data, production data, traffic flow data), and has a wide range of application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Other features, objects, and advantages of the present invention will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings:
[0033] Figure 1 It is a flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] To make the technical means, creative features, achieved purposes, and functions of the present invention easy to understand, the present invention will be further described below in conjunction with specific embodiments.
[0035] A multi-model fusion decision-making system based on deep reinforcement learning includes:
[0036] Initialization module: used to construct and initialize the Actor network, the Critic network, and the system state;
[0037] Weight generation module: dynamically generates model fusion weights through the Actor network;
[0038] Fusion decision module: performs weighted fusion of multi-model strategies;
[0039] Backtest evaluation module: configures resources and evaluates the strategy performance to generate reward benefits;
[0040] Resource management module: Update the total amount of resources and record status data;
[0041] Data storage module: Manage the training data set in the buffer;
[0042] Parameter update module: Update the parameters of the Actor network and the Critic network based on the buffer data.
[0043] A multi-model fusion decision-making method based on deep reinforcement learning is specifically implemented as follows:
[0044] (Step 0001) Initialization: At the starting time t = 0, initialize the weights of the reinforcement learning agent. In the present invention, the reinforcement learning agent is implemented through two networks, namely the Critic network and the Actor network. Among them, the Actor network is responsible for generating the weight allocation strategy of the model, and the Actor network can be implemented by a parameterized neural network π θ with its parameters denoted as θ. The Critic network is responsible for evaluating the value of the policy generated by the current Actor network, and the Critic network can be implemented by a parameterized neural network V φ (S t , E t , G t ) with its parameters denoted as φ. Initializing the weights of the reinforcement learning agent means initializing the parameters θ and φ of these two networks, and the specific implementation can be achieved through Gaussian initialization, kai-ming initialization, etc. At the same time, it is necessary to initialize the system state, and the system state includes: the total amount of allocable resources G t , the historical performance S t of the model, and the environmental variable factor E t . The historical performance S t of the model is a vector containing multiple features S t = [s 1,t , s 2,t ,..., s χ,t , where χ is the total number of models. The environmental variable factor E t is also a vector containing multiple environmental variables E t = [e 1,t , e 2,t ,..., e m,t , where m is the total number of environmental variable factors.
[0045] (Step 0002) Generate fusion weights: Obtain the historical performance S t of the model and the environmental variable factor E t at time t. The Actor network, according to the current resource G t, the historical performance S of the model t and the environmental variable factor E t , sample a set of fusion weights w of the model according to the policy network i,t :
[0046] w i,t ~π θ (w i |S t ,E t ,G t )
[0047] (Step 0003) Fusion decision: Let each model be represented as M i , i = 1, 2,..., χ, where χ is the total number of models. The output policy y of each model according to the current system state i,t can be expressed as:
[0048] y i,t = M i (E t , G t )
[0049] Use the current weight w i,t to perform weighted fusion on the outputs of each model to obtain the final output policy Y t :
[0050]
[0051] Output the weight vector w t = [w 1,t , w 2,t ,..., w k,t , and satisfy
[0052] (Step 0004) Backtest system: Input the fused output policy Y t into the backtest system, and evaluate the performance of the policy Y t by configuring the current resource G t to obtain the reward return R t :
[0053] R t = Backtest(Y t , G t )
[0054] where Backtest() is the evaluation reward function of the backtest system.
[0055] (Step 0005) Update the total amount of resources: Update the total amount of resources G according to the reward return R t and the current total amount of resources G t t+1 .
[0056] (Step 0006) Storage: Store information such as the current state, weight, and reward in buffer B as a set of training data for subsequent learning and updating:
[0057] B = B ∪ {(S t , S t+1 , E t , E t+1 , G t , G t+1 , w i,t , R t )}
[0058] By continuously repeating Step 0002 - Step 0006 until a predetermined stop condition is reached (such as N sets of training data have been stored in buffer B). After the buffer reaches the predetermined stop condition, the next step is to update the parameters of the reinforcement learning agent based on the existing buffer B.
[0059] (Step 0007) Update agent weight: Sample K sets of training data (S t , S t+1 , E t , E t+1 , G t , G t+1 , w i,t , R t ) from the buffer to update the Actor network and the Critic network. The implementation method of the update is as follows. The goal of the Actor network is to maximize the expected cumulative reward, and the parameters θ of the Actor network are updated by gradient ascent:
[0060] A t = R t + γV φ (S t+1 , E t+1 , G t+1 ) - V φ (S t , E t , G t )
[0061]
[0062] where a θ is the learning rate of the Actor network.
[0063] The goal of the Critic network is to minimize the mean squared error of the value function:
[0064]
[0065] where γ is the discount factor, R t is the current reward, and V φ (S t+1 , E t+1 ) is the value of the next state. Update the parameters φ of the Critic network by gradient descent:
[0066]
[0067] where a φ is the learning rate of the Critic network.
[0068] Delete the first Z samples in the buffer B, and repeat steps 0002 - 0007 until the system stops. The update of the agent can be realized under dynamic external conditions, so as to realize the function of adaptively outputting the multi - model fusion strategy.
[0069] In summary, through reinforcement learning, the agent intelligently allocates fusion weights for multiple models, so as to adaptively fuse the output strategies of multiple models. Compared with the traditional method of manually allocating fusion weights based on experience, the proposed intelligent multi - model fusion strategy can make more accurate decisions;
[0070] The implemented system can update the fusion strategy of multi - model fusion in real time according to the changes in the external environment (environmental variable factors, system - allocable resources, historical performance of individual models), and has the ability to adapt to the dynamic external environment;
[0071] The proposed method can adapt to complex decision - making problems in multiple scenarios, support different types of data in different scenarios (such as financial data, production data, traffic flow data), and has a wide range of application scenarios.
[0072] The above shows and describes the basic principles, main features and advantages of the present invention. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above - mentioned exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic features of the present invention. Therefore, in any aspect, the embodiments should be regarded as exemplary and non - restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention.
[0073] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A multi-model fusion decision-making method based on deep reinforcement learning, characterized in that: It includes the following method steps: Step 1: Initialization: At the starting moment, initialize the weights of the reinforcement learning agent. The agent includes an Actor network and a Critic network, which are respectively used to generate the model weight allocation policy and evaluate the policy value; Step 2: Obtain the current system state; Step 3: The Actor network outputs the weight allocation of the model according to the current system state; Step 4: Perform weighted fusion on the outputs of multiple models according to the weight allocation to obtain the final output policy; Step 5: Input the fused output policy into the backtest system to evaluate the policy performance and obtain the reward income; Step 6: Update the total resource amount according to the reward income and the current total resource amount; Step 7: Store the current system state, fusion weights and reward income in the buffer to form a training data set; Step 8: Sample training data from the buffer, update the parameters of the Actor network and the Critic network, and repeat Steps 2 to 8 until the preset termination condition is met.
2. The multi-model fusion decision-making method based on deep reinforcement learning according to claim 1, characterized in that: The parameter initialization methods of the Actor network and the Critic network are Gaussian distribution initialization or Kaiming initialization.
3. A multi-model fusion decision-making method based on deep reinforcement learning according to claim 1, characterized in that: The model historical performance data is the performance metric vector of each model in the historical time period, and the environmental variable factor is the external environmental parameter data vector.
4. A multi-model fusion decision-making method based on deep reinforcement learning according to claim 1, characterized in that: In the resource update step, the calculation rule of the total resource amount at the next moment is: based on the weighted sum of the current total resource amount and the reward income.
5. A multi-model fusion decision-making method based on deep reinforcement learning according to claim 1, characterized in that: The parameter update of the Actor network is achieved by the gradient ascent method of maximizing the cumulative reward; the parameter update of the Critic network is achieved by the gradient descent method of minimizing the value function error.
6. A multi-model fusion decision-making system based on deep reinforcement learning, applicable to the multi-model fusion decision-making method based on deep reinforcement learning described in claims 1-3, characterized in that: It includes: Initialization module: used to construct and initialize the Actor network, Critic network and system state; Weight generation module: dynamically generate model fusion weights through the Actor network; Fusion decision module: perform weighted fusion of multi-model policies; Backtest evaluation module: configure resources and evaluate policy performance to generate reward income; Resource management module: update the total resource amount and record status data; Data storage module: manage the training data set in the buffer; Parameter update module: update the parameters of the Actor network and the Critic network based on the buffer data.