Intelligent Device Online Integrated Learning and Decision-making Method Based on Incomplete Information

By designing intelligent equipment for multi-base and integrated equipment, using online gradient descent algorithms and integration methods, online learning and decision-making problems under incomplete information are solved, and effective decision-making in a dynamic environment is achieved.

CN115293281BActive Publication Date: 2025-07-25NANJING UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210990653.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-18
Publication Date
2025-07-25
Estimated Expiration
2042-08-18

AI Technical Summary

Technical Problem

Existing online learning and decision-making methods for intelligent devices usually require the acquisition of complete information, while the information is often incomplete in actual scenarios, which limits its scope of application.

Method used

Design an intelligent device, including multiple base devices and an integrated device, through an online gradient descent algorithm and an out-of-synchronous base device, combined with parallel or serial integration methods, to deal with incomplete environmental observations, degree of change and evaluation results.

Benefits of technology

Under the conditions of incomplete information, intelligent equipment can effectively conduct online learning and decision-making, expanding its scope of application, especially improving the effectiveness of decision-making in a dynamic environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115293281B_ABST
    Figure CN115293281B_ABST
Patent Text Reader

Abstract

The present invention discloses an online integrated learning and decision-making method for intelligent devices based on incomplete information, enabling intelligent devices to make full use of incomplete information for online learning and decision-making tasks. The intelligent device consists of multiple base devices and an integrated device. First, various base devices are designed to explore the environment, and the integrated device is used to combine the prediction and decision results of each base device, thereby achieving the effective utilization of incomplete environmental information. For the design of base devices, a decision-making method adaptive to the cumulative noise of the environment is used to effectively utilize incomplete environmental observations, and gradient estimation techniques in the field of stochastic optimization are adopted to cope with incomplete decision evaluation results. In addition, two different types of online integration methods, including parallel and serial types, are proposed to effectively address the information incompleteness caused by environmental changes. Compared with existing intelligent devices, the present invention can make more full use of incomplete information to complete online learning and decision-making tasks in complex scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an online integrated learning and decision-making method for intelligent devices based on incomplete information, and belongs to the technical field of online learning of intelligent devices. Background Art

[0002] With the rapid development of information technology, people increasingly rely on intelligent devices with certain information collection and processing capabilities. From mobile phones and computers closely related to daily use to temperature control devices, fault detection devices, and logistics scheduling devices in the manufacturing industry, etc., all belong to the category of intelligent devices. At the same time, modern society generates tens of thousands of pieces of information every day, and the large amount of information increasing in a short time puts higher requirements on intelligent devices - being able to collect and process information in real time and give responses. Moreover, intelligent devices need to interact with the environment, that is, their current predictions and decisions may affect the future environment, and the current environment will also be affected by past predictions and decisions, thus raising the need for online learning and decision-making methods for intelligent devices.

[0003] Traditional online learning and decision-making methods usually provide good theoretical guarantees in static environments. Specifically, existing methods require that the online average performance of intelligent devices can be comparable to their average performance in the offline state. However, in actual applications, the environment may change drastically and rapidly, and at this time, traditional online decision-making methods usually perform poorly. Recently, there have also been some online decision-making methods for dynamic environments, but these methods all require complete information about the degree of environmental change, and this requirement greatly limits the scope of application of the algorithms. At the same time, existing online decision-making methods usually assume that intelligent devices can obtain complete observations and evaluation results of the environment. However, in actual applications, we usually need to face incomplete observations and evaluation results. For example, a smart phone can only obtain environmental information through the sensors it is equipped with; the information that a smart fighter can obtain is limited by radar, the visual range of the pilot, etc., and the information in such scenarios is extremely limited. Summary of the Invention

[0004] Object of the Invention: Existing online learning and online decision-making methods for intelligent devices usually need to be able to obtain various complete information about the environment, while many information in real-world scenarios is naturally incomplete, which greatly limits the scope of application of existing intelligent devices. Aiming at the problems and deficiencies in the prior art, the present invention considers using integrated learning to design an online decision-making method that can handle incomplete information, and an intelligent device that can make online decisions under incomplete environmental change degree, observations, and evaluation results.

[0005] Technical solution: An online integrated learning and decision-making method for intelligent devices based on incomplete information. The intelligent device includes multiple base devices and one integrated device. Each base device runs the online gradient descent algorithm at each decision moment. The difference between different base devices lies in the step size of the online gradient descent algorithm, and different step sizes correspond to different degrees of environmental changes. On this basis, the integrated device is used to synthesize the decisions of the base devices and generate the final decision signal.

[0006] First, initialize the integrated device and the base devices of the intelligent device. Next, the intelligent device obtains the observation of the environment, and in this process, it is necessary to deal with the problem of incomplete environmental observations. Then, the intelligent device gives the decision result, and in this process, it is necessary to deal with the problem of incomplete degrees of environmental changes. Secondly, the intelligent device obtains the evaluation result of the environment for the decision, and in this process, it is necessary to deal with the problem of incomplete evaluation results; finally, the intelligent device updates its internal state, and the environment changes due to the influence of the decision made by the intelligent device.

[0007] An online integrated learning and decision-making method for intelligent devices based on incomplete information, which needs to deal with incomplete environmental observations, incomplete degrees of environmental changes, and incomplete evaluation reward results.

[0008] Initialize the integrated device and the base devices of the intelligent device:

[0009] Set the number of base devices N, where N can be set to the order of log2T, and T is the total number of decision moments;

[0010] The intelligent device initializes the decision parameter vectors M of N base devices 1,1 ,…,M 1,N , representing the decision parameters of the Nth base device at the first moment, representing the decision parameters of the ith base device at the tth moment;

[0011] The intelligent device initializes the weight parameter vector p1=(p 1,1 ,…,p 1,N ) of the integrated device for each base device, where p 1,i >0 and p t,i represents the weight of the integrated device for the ith base device at the tth moment.

[0012] The intelligent device obtains the observation of the environment, and the specific steps to deal with incomplete environmental observations are as follows:

[0013] The intelligent device executes the following steps at each decision moment t = m + 1, m + 2,…, T:

[0014] The intelligent device obtains the incomplete observation vector y of the environmentt .

[0015] The intelligent device extracts the cumulative environmental noise vector where G is a known inherent variable based on the environment, and G [i] is the i-th component of G, and u t-i is the decision signal given by the intelligent device at the (t - i)-th decision moment.

[0016] The intelligent device gives a decision that adapts to the environmental noise in the past m moments where is the cumulative environmental noise at the (t - i)-th decision moment, is the i-th component of the decision parameter vector M t .

[0017] The specific steps for the intelligent device to give a decision result to cope with the degree of incomplete environmental change are as follows:

[0018] The intelligent device executes the following steps at each decision moment t = m + 1, m + 2,..., T:

[0019] Obtain the decision parameter vectors M t,1 , …, M t,N .

[0020] The present invention designs two online integration methods to cope with the degree of incomplete environmental change, including a parallel integration method and a serial integration method. The intelligent device can arbitrarily select one online integration method to cope with the degree of incomplete environmental change.

[0021] (1) The specific steps of the parallel integration method are as follows:

[0022] The intelligent device performs a weighted sum of the decision parameters of N base devices and the weight parameters of the integration device to obtain the final decision parameter vector

[0023] The intelligent device obtains the evaluation reward result of the environment for the decision.

[0024] Each base device runs the online gradient descent algorithm to update its decision parameter from M t,i to M t+1,i .

[0025] (2) The specific steps of the serial integration method are as follows:

[0026] The intelligent device selects the decision parameter vector M t,i of the i-th base device with a probability of p t,i , and obtains the final decision parameter vector M t = M t,i .

[0027] The intelligent device obtains the evaluation reward result of the environment for the decision.

[0028] The intelligent device updates the i-th base device, that is, the i-th base device runs the online gradient descent algorithm, and updates its decision parameter from M t,i to M t+1,i . The gradient required for running the online gradient descent is given in the next part.

[0029] The intelligent device evaluates the decision quality of each base device and updates the weight of the integrated device for each base device according to its performance.

[0030] The specific steps for the intelligent device to obtain the evaluation result of the environment for the decision and handle the incomplete evaluation reward result are as follows:

[0031] At each decision moment t = m + 1, m + 2, …, T, the intelligent device executes the following steps:

[0032] The intelligent device obtains the evaluation reward result c t (y t , u t ) of the environment for the decision, where c t is the evaluation function, y t is the incomplete environment observation vector, and u t is the decision signal;

[0033] The intelligent device performs gradient estimation, and the gradient g t = d / δ · c t (y t , u t ) · s t , where δ is the hyperparameter controlling the exploration range, d is the dimension of the decision parameter, and s t is a random sampling on the d-dimensional unit sphere. The base device to be updated uses the estimated gradient to run the online gradient descent algorithm to update its decision parameter.

[0034] A computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the above computer program, it implements the online integrated learning and decision-making method of the intelligent device based on incomplete information as described above.

[0035] A computer-readable storage medium stores a computer program for executing the online integrated learning and decision-making method of the intelligent device based on incomplete information as described above.

[0036] Beneficial effects: Compared with the prior art, the present invention designs an online integrated learning and decision-making method for intelligent devices, which is applicable to online decision-making tasks with incomplete information. Specifically, the intelligent device designed by the present invention can cope with incomplete environmental observations, incomplete evaluation reward results, and incomplete environmental change degrees, greatly expanding the potential applicable scope of intelligent devices. Description of the Drawings

[0037] Figure 1 is the online decision-making flowchart of the intelligent device of the present invention;

[0038] Figure 2 is the method flowchart for the present invention to cope with incomplete environmental observations;

[0039] Figure 3 is the method flowchart for the present invention to cope with incomplete decision evaluation reward results;

[0040] Figure 4 is the method flowchart for the present invention to cope with the degree of incomplete environmental change. Detailed Embodiments

[0041] The following further clarifies the present invention in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, those skilled in the art's various equivalent modifications of the present invention fall within the scope defined by the appended claims of this application.

[0042] Taking the intelligent temperature control device in the manufacturing industry as an example. Existing intelligent temperature control devices often need to collect as much environmental information as possible, require complete evaluation reward results of the environment when making adjustments, and are unable to cope with potential environmental changes. However, in the real manufacturing industry, the interaction between intelligent devices and the environment often only obtains incomplete information, resulting in limited application scenarios for existing intelligent temperature control devices.

[0043] The online integrated learning and decision-making method for intelligent devices based on incomplete information disclosed by the present invention can effectively address the above problems, thereby greatly expanding the applicable scope of intelligent devices.

[0044] An online integrated learning and decision-making method for intelligent devices based on incomplete information, wherein the intelligent device includes a plurality of base devices and an integrated device. The base devices run the online gradient descent algorithm at each decision moment. The difference between different base devices lies in the step size of the online gradient descent algorithm, and different step sizes are for different degrees of environmental change. On this basis, the integrated device is used to synthesize the decisions of the base devices and generate the final decision signal.

[0045] First, initialize the integrated device and the base device of the intelligent device. Next, the intelligent device obtains observations of the environment, and in this process, it needs to deal with the problem of incomplete environmental observations. Then, the intelligent device gives a decision result, and in this process, it needs to deal with the problem of incomplete environmental change degree. Secondly, the intelligent device obtains the evaluation result of the environment on the decision, and in this process, it needs to deal with the problem of incomplete evaluation results; finally, the intelligent device updates its internal state, and the environment changes due to the decision made by the intelligent device.

[0046] An online integrated learning and decision-making method for intelligent devices based on incomplete information, which needs to deal with incomplete environmental observations, incomplete environmental change degree, and incomplete evaluation reward results.

[0047] Initialize the integrated device and the base device of the intelligent device:

[0048] Set the number of base devices N, where N can be set to the order of log2T, and T is the total number of decision-making moments;

[0049] The intelligent device initializes the decision parameter vectors M of N base devices 1,1 ,…,M 1,N , representing the decision parameters of the Nth base device at the first moment, representing the decision parameters of the ith base device at the tth moment;

[0050] The intelligent device initializes the weight parameter vector p1=(p 1,1 ,…,p 1,N ) of the integrated device for each base device, where p 1,i >0 and p t,i represents the weight of the integrated device for the ith base device at the tth moment.

[0051] The intelligent device obtains observations of the environment. The specific steps to deal with incomplete environmental observations are as follows:

[0052] The intelligent device executes the following steps at each decision-making moment t = m + 1, m + 2, …, T:

[0053] The intelligent device obtains the incomplete observation vector y of the environment t .

[0054] The intelligent device extracts the cumulative environmental noise vector where G is a known inherent variable based on the environment, and G [i] is the ith component of G, and u t-i is the decision signal given by the intelligent device at the (t - i)th decision-making moment.

[0055] The intelligent device gives a decision adapted to the environmental noise in the past m moments where is the cumulative environmental noise at the (t - i)-th decision moment, is the i-th component of the decision parameter vector M t .

[0056] The specific steps for the intelligent device to give a decision result to cope with the degree of incomplete environmental change are as follows:

[0057] The intelligent device executes the following steps at each decision moment t = m + 1, m + 2, …, T:

[0058] Obtain the decision parameter vectors M t,1 , …, M t,N .

[0059] Two online integration methods are designed to cope with the degree of incomplete environmental change, including the parallel integration method and the serial integration method. The intelligent device can arbitrarily select one online integration method to cope with the degree of incomplete environmental change.

[0060] (I) The specific steps of the parallel integration method are as follows:

[0061] The intelligent device performs weighted summation of the decision parameters of N base devices and the weight parameters of the integration device to obtain the final decision parameter vector

[0062] The intelligent device obtains the evaluation reward result of the environment for the decision.

[0063] Each base device runs the online gradient descent algorithm to update its decision parameter from M t,i to M t+1,i .

[0064] (II) The specific steps of the serial integration method are as follows:

[0065] The intelligent device selects the decision parameter vector M t,i of the i-th base device with a probability of p t,i , and obtains the final decision parameter vector M t = M t,i .

[0066] The intelligent device obtains the evaluation reward result of the environment for the decision.

[0067] The intelligent device updates the i-th base device, that is, the i-th base device runs the online gradient descent algorithm to update its decision parameter from M t,i to M t+1,i . The gradient required to run the online gradient descent is given in the next part.

[0068] The intelligent device evaluates the decision-making quality of each basic device and updates the weight of the integrated device for each basic device according to its performance.

[0069] The intelligent device obtains the evaluation result of the environment for the decision-making. The specific steps to handle the incomplete evaluation reward result are as follows:

[0070] The intelligent device executes the following steps at each decision-making moment \(t = m + 1, m + 2, \ldots, T\):

[0071] The intelligent device obtains the evaluation reward result \(c\) of the environment for the decision-making t (y t ,u t ), where \(c\) t is the evaluation function, \(y\) t is the incomplete environment observation vector, and \(u\) t is the decision signal;

[0072] The intelligent device performs gradient estimation, and the gradient \(g\) t = \(\frac{d}{δ} \cdot c\) t (y t ,u t ) \cdot s t , where \(δ\) is the hyperparameter controlling the exploration range, \(d\) is the dimension of the decision parameter, and \(s\) t is a random sampling on the \(d\)-dimensional unit sphere. The basic device to be updated uses the estimated gradient to run the online gradient descent algorithm to update its decision parameter.

[0073] Next, taking the intelligent temperature control device in the manufacturing industry as an example, the specific implementation manner of the present invention will be described in detail.

[0074] Figure 1 is the workflow diagram of the intelligent temperature control device proposed by the present invention. This device consists of an integrated control device and multiple basic control devices. Each basic control device gives its own temperature control signal at each control moment, and the integrated control device synthesizes the control signals of each basic control device to give the final temperature control signal. Step 1 Initializes the intelligent temperature control device, sets its internal counter to 1, and determines the total number of control moments \(T\). Step 3 Starts the online operation at each control moment. At each control moment, Step 4 The intelligent temperature control device obtains the current temperature, humidity, the number of operating devices, etc. environmental information. Step 5 The intelligent temperature control device gives a temperature control signal according to this incomplete observation of the environment. Step 6 The intelligent temperature control device receives the evaluation reward result of the environment for the control signal, such as the power consumption, etc. Step 7 The intelligent temperature control device updates its internal state. Specifically, the integrated control device and the basic control devices update the internal variables they maintain according to the evaluation reward result. Step 8 The environment changes under the influence of the control signal, so as to achieve the effect of temperature control.

[0075] Figure 2 is the workflow of the intelligent temperature control device for coping with incomplete environmental observations. Taking a certain control moment t as an example. Step 41: The intelligent temperature control device obtains the incomplete observation vector y of the environment t =(y t,1 , y t,2 , y t,3 , …), where the components such as y t,1 , y t,2 , yt,3 respectively represent information such as the temperature, humidity, and the number of running devices in the environment at the current control moment t. Step 42: The intelligent temperature control device, based on this part of the observation and the past control signals u1, …, u t-1 , u t is a real number representing temperature control information, u t >0 indicates increasing the temperature, u t <0 indicates decreasing the temperature, |u t | represents the amplitude of temperature control. Then, based on the formalized modeling of the environment, the cumulative environmental noise is extracted, that is, the observation result at the t-th moment when no control is performed on the environment, denoted as The calculation formula is where G is the known inherent variable based on the environment, and u t-k is the control signal given by the intelligent device at the (t - k)-th control moment. Step 43: The intelligent temperature control device gives a control signal that adapts to the cumulative environmental noise in the past m moments where is the j-th component of the decision parameter M t , and m is the hyperparameter of the algorithm. M t is given by the integrated temperature control device by synthesizing the control signals of each basic control device. The specific process is shown in Figure 4 .

[0076] Figure 3 is the workflow of the intelligent temperature control device for coping with incomplete evaluation reward results. Taking a certain control moment t as an example. Step 61: The intelligent temperature control device receives the evaluation reward result of the environment for the control signal, such as the power consumption at the current control moment, denoted as c t (y t , u t ), where c t is the evaluation function, y t is the incomplete observation vector of the environment, and u t is the temperature control signal. Step 62: The intelligent temperature control device performs gradient estimation g t = d / δ·c t (y t , u t)·s t , where δ is a hyperparameter that controls the exploration range, d is the dimension of the decision parameter, and s t is a random sample on the d-dimensional unit sphere. The random perturbation applied to the control signal is reflected in the parameter M t of the control signal u t . Through this method, the intelligent temperature control device can obtain an unbiased estimate of the gradient of the evaluation function at the decision point. The estimated gradient can be used for Figure 4 the online gradient descent algorithm of the base control device to be updated in

[0077] Figure 4 is the workflow diagram of the intelligent temperature control device to cope with the degree of incomplete environmental changes. First, initialize the intelligent temperature control device and determine the total number of control times T. Step 52 initializes the control signal parameters M 1,1 ,…,M 1,N (there are N base control devices in total, and N is of the order of log2T) and the weight parameter vector p1=(p 1,1 ,…,p 1,N ) of the integrated control device, where p 1,i >0 and represents the weight of the integrated control device for the Nth base control device at the first moment; p t,i represents the weight of the integrated control device for the ith base control device at the tth moment. Step 53 starts the online operation at each control time. At each control time, step 54 the integrated control device obtains the control signal parameters M t,1 ,…,M t,N of all base control devices. Next, the two online integration methods designed for the present invention will be described separately (taking a certain control time t as an example).

[0078] For the parallel integration method, step 56 performs a weighted sum of the control signal parameters of the base control device and the weight parameters of the integrated control device to obtain the final control signal parameter Step 57 After obtaining the evaluation reward result of the environment for the control signal, step 58 each base control device runs the online gradient descent algorithm and updates its decision parameter from M t,i to M t+1,i .

[0079] For the serial integration method, step 59 selects the control signal parameter M t,i of the ith base control device with a probability of p t,i according to the weight parameter of the integrated control device to obtain the final control signal parameter M t = M t,i, after step 510 obtains the evaluation reward result of the environment for the regulation signal, in step 511, the i-th basic regulation device runs the online gradient descent algorithm, and updates its decision parameter from M t,i to M t+1,i .

[0080] Next, after the intelligent temperature regulation device gives a regulation signal according to the final regulation signal parameter, it obtains the evaluation reward result of the environment for the regulation signal. In step 512, the intelligent temperature regulation device evaluates the regulation signal quality of each basic regulation device, and in step 513, it updates the weight of the integrated regulation device according to its performance.

[0081] Obviously, those skilled in the art should understand that each step of the above-mentioned intelligent device online integrated learning and decision-making method based on incomplete information in the embodiments of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, the embodiments of the present invention are not limited to any specific combination of hardware and software.

Claims

1. An online integrated learning and decision-making method for intelligent devices based on incomplete information, characterized in that, The intelligent device includes multiple base devices and one integrated device; each base device runs the online gradient descent algorithm at each decision moment, and the difference between different base devices lies in the step size of the online gradient descent algorithm, and different step sizes are for different degrees of environmental change; the integrated device is used to synthesize the decisions of the base devices and generate the final decision signal; First, initialize the integrated device and the base devices of the intelligent device; next, the intelligent device obtains the observation of the environment, and the problem of incomplete environmental observation needs to be addressed in this process; then, the intelligent device gives the decision result, and the problem of incomplete degree of environmental change needs to be addressed in this process; secondly, the intelligent device obtains the evaluation result of the environment on the decision, and the problem of incomplete evaluation result needs to be addressed in this process; finally, the intelligent device updates its internal state, and the environment changes under the influence of the decision made by the intelligent device. The intelligent device is an intelligent temperature control device, which consists of one integrated control device and multiple base control devices. The base control devices are the base devices, and each gives its own temperature control signal at each control moment. The integrated control device is the integrated device, which synthesizes the control signals of each base control device to give the final temperature control signal. At each control moment, the intelligent temperature control device obtains environmental information, and the intelligent temperature control device gives the temperature control signal according to the environmental information. The intelligent temperature control device receives the evaluation reward result of the environment on the control signal. The integrated control device and the base control devices update the internal variables they maintain according to the evaluation reward result, and the environment changes under the influence of the control signal, so as to achieve the effect of temperature control. The intelligent device obtains the observation of the environment. The specific steps to address the incomplete environmental observation are as follows: The intelligent device performs the following steps at each decision moment t = m + 1, m + 2, …, T: Obtain an incomplete observation vector y of the environment t ; Extract the cumulative ambient noise vector where G is a known inherent variable based on the environment, G [k] is the k-th component of G, u t-k is the decision signal given by the intelligent device at the (t - k)-th decision moment; Make a decision that adapts to the cumulative ambient noise in the past m time instants where is the cumulative ambient noise at the (t - j)-th decision instant, is the j-th component of the decision parameter vector M t of.

2. The intelligent device online integrated learning and decision-making method based on incomplete information according to claim 1, characterized in that The specific steps to initialize the integrated device and the base devices of the intelligent device are as follows: Set the number of base devices N, where N is set to the order of log2T, and T is the total number of decision moments; Initialize the decision parameter vectors M of N base devices 1,1 ,…,M 1,N , denotes the decision parameter of the Nth base device at the first moment. Similarly, denotes the decision parameter of the ith base device at the tth moment; Initialize the weight parameter vector p1 = (p 1,1 , …, p 1,N ) for each base device of the integrated device, where p t,i > 0 and p t,i represents the weight of the integrated device for the i-th base device at the t-th moment.

3. The intelligent device online integrated learning and decision-making method based on incomplete information according to claim 1, characterized in that The intelligent device gives the decision result. The specific steps to address the incomplete degree of environmental change are as follows: The intelligent device performs the following steps at each decision moment t = m + 1, m + 2, …, T: Obtain the decision parameter vector M of all base devices t,1 ,…,M t,N ; Two online integration methods are designed to address the incomplete degree of environmental change, including the parallel integration method and the serial integration method. The intelligent device arbitrarily selects one online integration method to address the incomplete degree of environmental change; (1) The specific steps of the parallel integration method are as follows: The intelligent device performs weighted summation on the decision parameters of N basic devices and the weight parameters of the integrated device to obtain the final decision parameter vector The intelligent device obtains the evaluation reward result of the environment on the decision; Each base device runs the online gradient descent algorithm to update its decision parameters from M t,i to M t+1,i ; (2) The specific steps of the serial integration method are as follows: The intelligent device selects the decision parameter vector M of the i-th base device with a probability of p t,i to obtain the final decision parameter vector M t,i such that t M t, i ; The intelligent device obtains the evaluation reward result of the environment on the decision; The intelligent device updates the i-th base device, that is, the i-th base device runs the online gradient descent algorithm and updates its decision parameters from M t,i to M t+1,i ; The intelligent device evaluates the decision quality of each base device and updates the weight of the integrated device for each base device according to its performance.

4. The intelligent device online integrated learning and decision-making method based on incomplete information according to claim 1, characterized in that, The intelligent device obtains the evaluation result of the environment on the decision. The specific steps to address the incomplete evaluation reward result are as follows: The intelligent device performs the following steps at each decision moment t = m + 1, m + 2, …, T: The intelligent device obtains the evaluation reward result c of the environment for decision-making t (y t ,u t ), where c t is the evaluation function, y t is the incomplete environment observation vector, and u t is the decision signal; The intelligent device performs gradient estimation, and the gradient g t = d / δ·c t (y t , u t )·s t , where δ is a hyperparameter that controls the exploration range, d is the dimension of the decision parameter, and s t is a random sample on the d-dimensional unit sphere; the base device to be updated uses the estimated gradient to run the online gradient descent algorithm to update its decision parameter.

5. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the online integrated learning and decision-making method for intelligent devices based on incomplete information as described in any one of claims 1-4.

6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program that executes the online integrated learning and decision-making method for intelligent devices based on incomplete information as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Multi-machine collaborative air combat planning method and system based on deep reinforcement learning

    CN112861442A