Offline reinforcement learning method, device and apparatus for target control

By constructing a policy optimization objective function based on the constraint terms and policy performance improvement terms based on the maximum likelihood estimation in autonomous driving, the problem of limited offset between optimization strategies and behavioral strategies in the existing technology is solved, and a more flexible and effective optimization strategy in autonomous driving is achieved.

CN114186474BActive Publication Date: 2025-05-09TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111256006.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-27
Publication Date
2025-05-09
Estimated Expiration
2041-10-27

AI Technical Summary

Technical Problem

In the prior art, KL divergence is used to constrain the optimization strategy, which limits the deviation between the optimization strategy and the behavior strategy, making it difficult to find an effective optimization strategy in autonomous driving.

Method used

By constructing constraint terms based on maximum likelihood estimation and policy performance improvement terms related to behavioral policy reward expectations, a polynomial policy optimization objective function is formed, allowing the optimization strategy to generate a large offset in a high confidence state, and the effectiveness of the optimization strategy is evaluated through the policy performance improvement terms.

Benefits of technology

It realizes that the optimization strategy is allowed to generate a large deviation in a high confidence state, improves the effectiveness and flexibility of the optimization strategy in autonomous driving, and cooperates with the agent to determine the optimization strategy to support better autonomous driving control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114186474B_ABST
    Figure CN114186474B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of deep learning technology, and specifically provides an offline reinforcement learning method, device and equipment for target control. Among them, the offline reinforcement learning method for target control includes: obtaining historical data; based on the historical data, updating a preset behavior strategy simulator, determining the behavior strategy and the reward expectation of the behavior strategy; based on the historical data, the behavior strategy and the strategy optimization objective function, the behavior optimization is performed through a preset intelligent agent to obtain an optimization strategy; wherein the strategy optimization objective function is constructed based on constraints and strategy performance improvement items; the constraints are constructed based on the maximum likelihood estimation method; the strategy performance improvement item is constructed based on the reward expectation of the behavior strategy. In this way, the constraints constructed based on the maximum likelihood estimation method constrain the maximized probability distribution of the optimization strategy to be the behavior strategy, allowing the optimization strategy to produce a large deviation under a high confidence state, thereby improving the expressiveness of the optimization strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning technology, and in particular to an offline reinforcement learning method, device and equipment for target control. Background Art

[0002] With the advancement of technology and the development of society, autonomous driving has begun to enter people's lives.

[0003] In order to realize autonomous driving, it is necessary to obtain the vehicle's driving environment information and the corresponding driver's operation information. Then, reinforcement learning is carried out based on this information to obtain the behavior strategy and optimization strategy, and the optimization strategy is used to support the vehicle's autonomous driving.

[0004] However, in existing solutions, the offline reinforcement learning used is generally based on the KL divergence to constrain the optimization strategy, which does not allow the optimization strategy to deviate significantly from the behavioral strategy. The restriction is very strict and is not conducive to seeking optimization strategies to control vehicle autonomous driving. Summary of the invention

[0005] The present invention provides an offline reinforcement learning method, device and equipment for target control, which are used to solve the problem that the prior art uses KL divergence to constrain the optimization strategy, does not allow the optimization strategy to have a large deviation compared to the behavior strategy, and the restriction is very strict, which is not conducive to seeking optimization strategies to control the defects of vehicle automatic driving.

[0006] In a first aspect, the present invention provides an offline reinforcement learning method for target control, comprising:

[0007] Get historical data;

[0008] Based on the historical data, updating a preset behavior strategy simulator, determining a behavior strategy and a reward expectation of the behavior strategy;

[0009] Based on the historical data, the behavior strategy and the strategy optimization objective function, the behavior optimization is performed through a preset intelligent agent to obtain an optimization strategy; wherein the strategy optimization objective function is constructed based on constraints and strategy performance improvement items; the constraints are constructed based on the maximum likelihood estimation method with the maximization probability distribution of the constraint optimization strategy as the goal of the behavior strategy; the strategy performance improvement item is constructed to be related to the reward expectation of the behavior strategy.

[0010] Optionally, the method further includes: controlling the target based on the optimization strategy.

[0011] Optionally, the construction process of the constraint item includes:

[0012] Based on the maximum likelihood estimation method, determining a determinant for indicating the degree of support of the behavior strategy for the optimization strategy;

[0013] The determinant is used as the constraint item.

[0014] Optionally, the construction process of the policy performance improvement item includes:

[0015] Determine the importance sampling factor;

[0016] Based on the importance sampling coefficient and the behavior strategy reward expectation, a strategy performance improvement item is determined.

[0017] Optionally, determining the importance sampling coefficient includes:

[0018] Determining a target average deviation; the target average deviation is a maximized average deviation of the importance sampling coefficient and the inverse importance sampling coefficient;

[0019] The importance sampling coefficient is determined by minimizing the target mean deviation.

[0020] Optionally, determining the target average deviation includes:

[0021] Determine the kernel function;

[0022] A target mean deviation is constructed based on the kernel function.

[0023] Optionally, the construction process of the strategy optimization objective function includes:

[0024] Add the constraint term and the policy performance improvement term to get a polynomial;

[0025] Based on the goal of maximizing the value corresponding to the polynomial, a policy optimization objective function is constructed.

[0026] Optionally, the historical data includes: vehicle driving environment information and vehicle control behavior information.

[0027] In a second aspect, an embodiment of the present invention provides an offline reinforcement learning device for target control, comprising:

[0028] An acquisition unit, used for acquiring historical data;

[0029] A determination unit, configured to update a preset behavior strategy simulator based on the historical data, and determine a behavior strategy and a reward expectation of the behavior strategy;

[0030] An optimization unit is used to perform behavior optimization through a preset intelligent agent based on the historical data, the behavior strategy and the strategy optimization objective function to obtain an optimization strategy; wherein the strategy optimization objective function is constructed based on constraints and strategy performance improvement items; the constraints are constructed based on the maximum likelihood estimation method with the goal of maximizing the probability distribution of the constraint optimization strategy as the behavior strategy; the strategy performance improvement item is constructed to be related to the reward expectation of the behavior strategy.

[0031] In a third aspect, an embodiment of the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the offline reinforcement learning method for target control as provided in the first aspect are implemented.

[0032] In a fourth aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the offline reinforcement learning method for target control provided in the first aspect.

[0033] The offline reinforcement learning method for target control provided by the present invention first obtains historical data; then based on the historical data, updates a preset behavior strategy simulator to determine the behavior strategy and the reward expectation corresponding to the behavior strategy; then, based on the constructed strategy optimization objective function, the behavior optimization is performed through a preset intelligent agent to obtain an optimized strategy. Specifically, the strategy optimization objective function includes: a constraint term constructed based on the maximum likelihood estimation method and a strategy performance improvement term constructed based on the behavior strategy reward expectation. Compared with the prior art that constructs constraint terms based on KL divergence, the constraint terms constructed based on the maximum likelihood estimation method constrain the maximized probability distribution of the optimization strategy to be the behavior strategy, allowing the optimization strategy to produce a large deviation under a high confidence state. At the same time, in order to determine the optimization effect of the optimization strategy after the constraint terms are changed, in the solution provided by the embodiment of the present invention, the optimization effect of the optimization strategy is reflected by a strategy performance improvement term constructed based on the expected reward of the behavior strategy. In this way, the strategy optimization objective function can allow the optimization strategy to produce a large deviation under a high confidence state. At the same time, based on the strategy performance improvement term, the optimization effect of the optimization strategy is judged, so as to realize the function of the collaborative intelligent agent determining the optimization strategy in offline reinforcement learning, so as to obtain a better optimization strategy, and control the vehicle automatic driving through the optimization strategy. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0035] Figure 1 It is one of the flow charts of the offline reinforcement learning method for target control provided by the present invention;

[0036] Figure 2 This is the second flow chart of the offline reinforcement learning method for target control provided by the present invention;

[0037] Figure 3 It is a structural schematic diagram of an offline reinforcement learning device for target control provided by the present invention;

[0038] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0039] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0040] First, the application scenario of the embodiment of the present invention is described. With the advancement of science and technology and the development of society, autonomous driving has begun to enter people's lives. In order to achieve autonomous driving, it is necessary to obtain the vehicle's driving environment information and the corresponding driver's operation information. Then, reinforcement learning is performed based on this information to obtain the behavior strategy and optimization strategy, and the optimization strategy is used to support the vehicle's autonomous driving.

[0041] Offline reinforcement learning and online reinforcement learning are two major branches of reinforcement learning. Compared with online reinforcement learning, offline reinforcement learning does not require online interaction between the agent and the environment. Instead, the agent learns the optimal strategy from a historical data set B that records the transfer information of "state-action-reward-state {s, a, r, s'}" so that the strategy can obtain the maximum cumulative reward.

[0042] However, in the offline reinforcement learning adopted by the existing solutions, the optimization strategy is generally constrained based on the KL divergence, which does not allow the optimization strategy to deviate greatly from the behavior strategy. The restriction is very strict, which is not conducive to seeking the optimization strategy to control the vehicle's automatic driving. This application proposes a corresponding solution to this problem. Figure 1-Figure 4 The present invention describes an offline reinforcement learning method, device and apparatus for target control.

[0043] Figure 1 This is one of the flow charts of the offline reinforcement learning method for target control provided by the present invention, which can be executed by the offline reinforcement learning method provided by the embodiment of the present invention. Figure 1 , the method may specifically include the following steps:

[0044] Step 110, obtaining historical data.

[0045] Exemplarily, when the target of control is a vehicle that needs to be autonomously driven, the historical data may be, but is not limited to: vehicle driving environment information and vehicle control behavior information. It should be noted that the historical data may be different based on different control targets. In an embodiment of the present invention, the vehicle driving environment information is the state; the vehicle control behavior information is the action; the reward can be determined after comparing the vehicle control behavior information based on preset rules, or it can be added by relevant personnel. Of course, the historical data may not include rewards, and the rewards are determined by updating the simulator in the preset behavior strategy simulator in the subsequent steps.

[0046] Step 120: based on the historical data, update the preset behavior strategy simulator to determine the behavior strategy and the reward expectation of the behavior strategy.

[0047] Among them, the updated policy simulator is a simulator of the behavior strategy μ reflected in the corresponding historical data. When observing the same state s, the policy simulator will simulate the behavior of the behavior strategy μ as much as possible, that is, make an action a similar to the behavior strategy as much as possible. Exemplarily, the behavior strategy is the correspondence between the vehicle driving environment information and the vehicle control behavior information summarized based on the historical data, and the policy simulator can simulate this behavior strategy. Furthermore, the policy simulator can also calculate the reward expectations corresponding to various strategies in the behavior strategy.

[0048] Step 130, based on the historical data, the behavior strategy and the strategy optimization objective function, the behavior optimization is performed through a preset intelligent agent to obtain an optimization strategy; wherein the strategy optimization objective function is constructed based on constraints and strategy performance improvement items; the constraints are constructed based on the maximum likelihood estimation method with the goal of maximizing the probability distribution of the constraint optimization strategy as the behavior strategy; the strategy performance improvement item is constructed to be related to the reward expectation of the behavior strategy.

[0049] With such a setting, compared with the prior art in which the constraint terms are constructed based on KL divergence, the constraint terms constructed based on the maximum likelihood estimation method constrain the maximized probability distribution of the optimization strategy to be the behavior strategy, allowing the optimization strategy to produce a large deviation under a high confidence state. At the same time, in order to determine the optimization effect of the optimization strategy after the constraint terms are changed, in the solution provided by the embodiment of the present invention, the optimization effect of the optimization strategy is reflected by a strategy performance improvement term constructed based on the expected reward of the behavior strategy. In this way, the strategy optimization objective function can allow the optimization strategy to produce a large deviation under a high confidence state. At the same time, based on the strategy performance improvement term, the optimization effect of the optimization strategy is judged, so as to realize the function of the collaborative intelligent agent determining the optimization strategy in offline reinforcement learning, so as to obtain a better optimization strategy, and control the vehicle automatic driving through the optimization strategy.

[0050] It should be noted that the process of constructing the constraint term includes: determining a determinant for indicating the degree of support of the behavior strategy for the optimization strategy based on the maximum likelihood estimation method.

[0051] Specifically, when the state s is observed t Then, the optimization strategy π is used to select action a t , and then determine the state s observed by the simulator according to the preset behavior strategy simulator t Will make an action a t The probability, specifically, the probability here refers to the mean value μ t , the variance is σ t 2 The conditional probability distribution of the Gaussian distribution is shown in formula (1):

[0052]

[0053] Based on formula (1), the constraint term αlogμ(a t |s t ), where α is a manually adjusted hyperparameter, μ t and σ t 2 is the learnable parameter of the behavior policy simulator. When the action a selected by the optimization policy π t When the action selected by the simulator μ is significantly different, αlogμ(at |s t ) will become smaller. On the contrary, when the action a selected by the optimization strategy π t When the difference between the action selected by the simulator μ is small, αlogμ(a t |s t ) will become larger. Therefore, by maximizing αlogμ(a t |s t ) can force the optimization strategy π to choose actions similar to the simulator μ.

[0054] Specifically, refer to Figure 2 The construction process of the strategy performance improvement item includes the following steps:

[0055] Step 210, determining a kernel function;

[0056] Step 220: construct a target average deviation based on the kernel function.

[0057] Among them, the target average deviation is the maximized average deviation of the importance sampling coefficient and the inverse importance sampling coefficient; through steps 210 and 220, the target average deviation can be constructed to facilitate determining the importance sampling coefficient based on the target average deviation and improve the accuracy of the importance sampling coefficient.

[0058] Step 230, determining the importance sampling coefficient by minimizing the target average deviation.

[0059] Step 240: Determine a strategy performance improvement item based on the importance sampling coefficient and the behavior strategy reward expectation.

[0060] In this way, the strategy performance improvement item can determine the reward expectation of the optimization strategy. The higher the reward expectation, the better the effect of the optimization strategy.

[0061] It should be noted that the performance improvement of the optimization strategy π compared with the behavior strategy μ is Positive correlation. μ (s,a) is an indicator used to measure whether an action is good or bad, which is related to the expected cumulative reward Q μ (s, a) is positively correlated, so the performance improvement of the optimization strategy π compared to the behavior strategy μ is Positive correlation. However, the calculation When you need to constantly distribute π The state s is collected in the π The calculation of requires the agent to interact with the environment, so this cannot be achieved in offline reinforcement learning, which is an application scenario where the agent cannot interact with the environment. That is: It is usually difficult to calculate. To solve this problem, the method adopted in the embodiment of the present invention is to calculate By estimate

[0062] calculate The specific steps are as follows: First, determine or estimate the importance sampling coefficient ω. After obtaining the importance sampling coefficient ω, the importance sampling method can be used to more conveniently calculate When the importance sampling coefficient ω is more accurate, It can estimate the difficult-to-calculate

[0063] Therefore, in the solution provided in the embodiment of the present invention, π (s)Q μ (s,a) is used as a strategy performance improvement item to indicate the performance of the optimization strategy. At the same time, in order to ensure that the indication effect of the strategy performance improvement item is more accurate, it is necessary to accurately estimate the importance sampling coefficient ω.

[0064] Specifically, the specific meaning of the importance sampling coefficient ω is as follows:

[0065]

[0066] In formula (2), d μ is the state distribution of the behavior strategy, which can be directly calculated from historical data; d π To accurately estimate the importance sampling coefficient ω, the present invention uses the importance sampling coefficient ω and the inverse importance sampling coefficient The average deviation between By using the method of , we can get the estimated value of the importance sampling coefficient ω, that is, determine the importance sampling coefficient ω. At the same time, as the average deviation Minimization can also gradually improve the estimation accuracy of the importance sampling coefficient ω and reduce the variance of the estimated value of the importance sampling coefficient ω, thereby making the training more stable without large fluctuations.

[0067] Specifically, calculate Previously, it was necessary to select the type of kernel function K(·,·), and the kernel function K(·,·) may be, but is not limited to, a Gaussian kernel function or a Laplace kernel function.

[0068] The Gaussian kernel function is as follows:

[0069]

[0070] The Laplace kernel function is as follows:

[0071]

[0072] The calculation process is divided into two steps.

[0073] Step 1: From historical data Randomly extract independent state-action-state transition pairs from

[0074] Step 2: Refer to formula (5) to calculate the importance sampling coefficient ω and the inverse importance sampling coefficient The maximum average deviation between Formula (5) is as follows:

[0075]

[0076] So far, the maximum average deviation is calculated By minimizing The calculation accuracy of the importance sampling coefficient ω can be gradually improved.

[0077] In the solution provided by the embodiment of the present invention, the strategy optimization objective function is constructed based on the constraint term and the strategy performance improvement term; based on the above-mentioned related embodiments, the embodiment of the present invention has provided a method for determining the strategy performance improvement term and the constraint term. Among them, the strategy performance improvement term is ω π (s)Q μ (s,a). By continuously maximizing the policy performance improvement term, the policy performance can be improved. The constraint term is αlogμ(a|s). By maximizing the constraint term, the optimization policy π can be guaranteed to be within the distribution support range of the behavior policy μ. Based on this, the constraint term and the policy performance improvement term are added to obtain a polynomial; based on the goal of maximizing the value corresponding to the polynomial, the policy optimization objective function is constructed. Specifically, the policy optimization objective function is shown in formula (6):

[0078]

[0079] The offline reinforcement learning device for target control provided by the present invention is described, and the offline reinforcement learning device for target control described below and the offline reinforcement learning method for target control described above can refer to each other.

[0080] Figure 3 Schematic diagram of the structure of the target controlled offline reinforcement learning device provided by the present invention; Figure 3 , the target controlled offline reinforcement learning device provided by the embodiment of the present invention comprises:

[0081] An acquisition unit 31, used for acquiring historical data;

[0082] A determination unit 32, configured to update a preset behavior strategy simulator based on the historical data, and determine a behavior strategy and a reward expectation of the behavior strategy;

[0083] The optimization unit 33 is used to perform behavior optimization through a preset intelligent agent based on the historical data, the behavior strategy and the strategy optimization objective function to obtain an optimization strategy; wherein the strategy optimization objective function is constructed based on constraints and strategy performance improvement items; the constraints are constructed based on the maximum likelihood estimation method with the goal of maximizing the probability distribution of the constraint optimization strategy as the behavior strategy; the strategy performance improvement item is constructed to be related to the reward expectation of the behavior strategy.

[0084] In the device provided by the embodiment of the present invention, the constraint terms constructed based on the maximum likelihood estimation method constrain the maximized probability distribution of the optimization strategy to be the behavior strategy, allowing the optimization strategy to produce a large deviation under a high confidence state. Furthermore, in order to determine the optimization effect of the optimization strategy after the constraint terms are changed, in the device provided by the embodiment of the present invention, the optimization effect of the optimization strategy is reflected by a strategy performance improvement term constructed based on the expected reward of the behavior strategy. In this way, the strategy optimization objective function can allow the optimization strategy to produce a large deviation under a high confidence state. At the same time, based on the strategy performance improvement term, the optimization effect of the optimization strategy is judged to realize the function of the collaborative intelligent agent determining the optimization strategy in offline reinforcement learning, so as to obtain a better optimization strategy, and control the vehicle automatic driving through the optimization strategy.

[0085] Optionally, the construction process of the constraint item includes:

[0086] Based on the maximum likelihood estimation method, a determinant indicating the degree of support of the behavior strategy for the optimization strategy is determined.

[0087] Optionally, the construction process of the policy performance improvement item includes:

[0088] Determine the importance sampling factor;

[0089] Based on the importance sampling coefficient and the behavior strategy reward expectation, a strategy performance improvement item is determined.

[0090] Optionally, determining the importance sampling coefficient includes:

[0091] Determining a target average deviation; the target average deviation is a maximized average deviation of the importance sampling coefficient and the inverse importance sampling coefficient;

[0092] The accuracy of the importance sampling coefficient is determined by minimizing the target mean deviation.

[0093] Optionally, determining the target average deviation includes:

[0094] Determine the kernel function;

[0095] A target mean deviation is constructed based on the kernel function.

[0096] Optionally, the construction process of the strategy optimization objective function includes:

[0097] Add the constraint term and the policy performance improvement term to get a polynomial;

[0098] Based on the goal of maximizing the value corresponding to the polynomial, a policy optimization objective function is constructed.

[0099] Optionally, the historical data includes: vehicle driving environment information and vehicle control behavior information.

[0100] Figure 4 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 4 As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430 and a communication bus 440, wherein the processor 410, the communication interface 420 and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute an offline reinforcement learning method for target control, the method comprising: obtaining historical data; updating a preset behavior strategy simulator based on the historical data, determining a behavior strategy and a reward expectation of the behavior strategy; performing behavior optimization through a preset intelligent agent based on the historical data, the behavior strategy and the strategy optimization objective function to obtain an optimization strategy; wherein the strategy optimization objective function is constructed based on constraint items and strategy performance improvement items; the constraint items are constructed based on a method of maximum likelihood estimation with the maximum probability distribution of the constraint optimization strategy as the goal of the behavior strategy; the strategy performance improvement item is constructed to be related to the reward expectation of the behavior strategy.

[0101] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0102] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the offline reinforcement learning method for target control provided by the above-mentioned methods, and the method includes: obtaining historical data; based on the historical data, updating a preset behavior strategy simulator to determine the behavior strategy and the reward expectation of the behavior strategy; based on the historical data, the behavior strategy and the strategy optimization objective function, performing behavior optimization through a preset intelligent agent to obtain an optimized strategy; wherein the strategy optimization objective function is constructed based on constraints and strategy performance improvement items; the constraints are constructed based on a maximum likelihood estimation method with the maximization probability distribution of the constraint optimization strategy as the goal of the behavior strategy; the strategy performance improvement item is constructed to be related to the reward expectation of the behavior strategy.

[0103] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the above-mentioned offline reinforcement learning methods for target control, the methods comprising: obtaining historical data; updating a preset behavior strategy simulator based on the historical data to determine the behavior strategy and the reward expectation of the behavior strategy; performing behavior optimization through a preset intelligent agent based on the historical data, the behavior strategy and the strategy optimization objective function to obtain an optimized strategy; wherein the strategy optimization objective function is constructed based on constraints and strategy performance improvement items; the constraints are constructed based on a maximum likelihood estimation method with the goal of maximizing the probability distribution of the constrained optimization strategy as the behavior strategy; the strategy performance improvement item is constructed to be related to the reward expectation of the behavior strategy.

[0104] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0105] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An offline reinforcement learning method for target control, characterized in that: Applied to vehicle autonomous driving, including: Acquire historical data, wherein the historical data includes vehicle driving environment information and vehicle control behavior information, wherein the vehicle driving environment information is a state, and the vehicle control behavior information is an action; Based on the historical data, a preset behavior strategy simulator is updated to determine a behavior strategy and a reward expectation of the behavior strategy, wherein the behavior strategy is a correspondence between vehicle driving environment information and vehicle control behavior information summarized based on the historical data, and the behavior strategy simulator is used to simulate the behavior strategy and calculate the reward expectations corresponding to various strategies in the behavior strategy; Based on the historical data, the behavior strategy and the strategy optimization objective function, the behavior optimization is performed through a preset intelligent agent to obtain an optimization strategy; wherein the strategy optimization objective function is constructed based on constraint items and strategy performance improvement items; the constraint items are constructed based on the maximum likelihood estimation method with the maximum probability distribution of the constraint optimization strategy as the behavior strategy as the goal; the strategy performance improvement item is constructed to be related to the reward expectation of the behavior strategy; The construction process of the constraint item includes: Based on the maximum likelihood estimation method, determining a determinant for indicating the degree of support of the behavior strategy for the optimization strategy; Using the determinant as the constraint item; When the state is observed Then, the optimization strategy is adopted Select Action , and then determine the state observed by the simulator according to the preset behavior strategy simulator Will take action The probability that the mean is , the variance is The conditional probability distribution of the Gaussian distribution is calculated as: (1); Based on formula (1), the constraint term is calculated , where α is a manually adjusted hyperparameter, and is the learnable parameter of the behavior policy simulator. When the action selected by the optimization policy is With simulator When the selected actions differ greatly, becomes smaller; on the contrary, when the optimization strategy Selected Action With simulator When the difference between the selected actions is small, Get bigger by maximizing Promote optimization strategy Selection and Simulator Similar actions.

2. The offline reinforcement learning method for target control according to claim 1, characterized in that: The construction process of the strategy performance improvement item includes: Determine the importance sampling factor; Based on the importance sampling coefficient and the behavior strategy reward expectation, a strategy performance improvement item is determined.

3. The offline reinforcement learning method for target control according to claim 2, characterized in that: The determining of the importance sampling coefficient comprises: Determining a target average deviation; the target average deviation is a maximized average deviation of the importance sampling coefficient and the inverse importance sampling coefficient; The importance sampling coefficient is determined by minimizing the target mean deviation.

4. The offline reinforcement learning method for target control according to claim 3, characterized in that: Determining the target average deviation includes: Determine the kernel function; A target mean deviation is constructed based on the kernel function.

5. The offline reinforcement learning method for target control according to claim 1, characterized in that: The construction process of the strategy optimization objective function includes: Add the constraint term and the policy performance improvement term to get a polynomial; Based on the goal of maximizing the value corresponding to the polynomial, a policy optimization objective function is constructed.

6. An offline reinforcement learning device for target control, characterized in that: Applied to vehicle autonomous driving, including: An acquisition unit, used to acquire historical data, wherein the historical data includes vehicle driving environment information and vehicle control behavior information, wherein the vehicle driving environment information is a state, and the vehicle control behavior information is an action; A determination unit, used to update a preset behavior strategy simulator based on the historical data, determine a behavior strategy and a reward expectation of the behavior strategy, wherein the behavior strategy is a correspondence between vehicle driving environment information and vehicle control behavior information summarized from the historical data, and the behavior strategy simulator is used to simulate the behavior strategy and calculate the reward expectations corresponding to various strategies in the behavior strategy; An optimization unit, configured to perform behavior optimization through a preset intelligent agent based on the historical data, the behavior strategy and the strategy optimization objective function to obtain an optimization strategy; wherein the strategy optimization objective function is constructed based on a constraint term and a strategy performance improvement term; the constraint term is constructed based on a maximum likelihood estimation method with the maximum probability distribution of the constraint optimization strategy as the behavior strategy as the goal; the strategy performance improvement term is constructed to be related to the reward expectation of the behavior strategy; The construction process of the constraint item includes: Based on the maximum likelihood estimation method, determining a determinant for indicating the degree of support of the behavior strategy for the optimization strategy; Using the determinant as the constraint item; When the state is observed Then, the optimization strategy is adopted Select Action , and then determine the state observed by the simulator according to the preset behavior strategy simulator Will take action The probability that the mean is , the variance is The conditional probability distribution of the Gaussian distribution is calculated as: (1); Based on formula (1), the constraint term is calculated , where α is a manually adjusted hyperparameter, and is the learnable parameter of the behavior policy simulator. Selected Action With simulator When the selected actions differ greatly, becomes smaller; on the contrary, when the optimization strategy Selected Action With simulator When the difference between the selected actions is small, Get bigger by maximizing Enabling optimization strategy selection and simulator Similar actions.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the offline reinforcement learning method for target control according to any one of claims 1 to 5 are implemented.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the offline reinforcement learning method for target control according to any one of claims 1 to 5 are implemented.