Interpretability Monitoring Data Recognition Method Based on Improved DQN
By improving the DQN algorithm and Markov chain model, combined with the mutual attention mechanism, the problem of ininterpretation in medical artificial intelligence algorithms is solved, real-time online learning of medical diagnosis and treatment data is realized, and the credibility and accuracy of the warning is improved.
Patent Information
- Application Number
- CN202211680701.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-07
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-12-07
AI Technical Summary
Existing medical artificial intelligence algorithms are unexplainable in the early warning of medical diagnosis and treatment data, and it is difficult to provide credible decision-making support to doctors.
Using an interpretability monitoring data recognition method based on improved DQN, a state transition probability is generated through real-time online learning and Markov chain model, a reinforcement learning network and mutual attention mechanism are constructed to generate interpretability text to explain the prediction results.
Real-time online learning of monitoring data is realized, and interpretable text is provided while generating accurate prediction results, which improves the credibility and accuracy of early warning of medical diagnosis and treatment data.
Smart Images

Figure CN116304855B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an interpretable monitoring data recognition method based on improved DQN, belonging to data mining, and is particularly applicable to the interpretable monitoring data recognition method based on improved DQN. Background Art
[0002] At present, the use of artificial intelligence technology to analyze medical diagnosis and treatment data for early warning research has become a new field of interdisciplinary integration. Previous research reports and our previous research found that using machine learning to mine intraoperative monitoring data has the value of early diagnosis and early warning of postoperative cardiovascular diseases. By delving deeper into "early warning algorithm improvement and interpretable text", it is expected to significantly improve the accuracy and credibility of early warning of perioperative cardiovascular adverse events.
[0003] The Deep Q-Learning (DQN) algorithm is different from the previously used supervised learning. Based on reinforcement deep learning, for the sequential decision-making problem of dynamic direct monitoring data, in the case of no labeled data, the agent will interact with the environment and obtain information, learn the mapping between states and actions, and guide the agent to make the best decision according to the state. The prediction of cardiovascular critical adverse events based on dynamic direct monitoring vital sign data is a continuous process, which is similar to the scenario in the DQN algorithm where the agent interacts with the environment and receives feedback based on the actions taken. In addition, the DQN algorithm can make sequential decisions without a large number of manually labeled training samples, which is conducive to modeling the prediction of time series data and has been used to formulate treatment strategies for epilepsy and lung cancer, treatment strategies for sepsis, and anemia treatment by controlling erythropoietin-stimulating agents (ESA) during renal failure dialysis.
[0004] Interpretable artificial intelligence requires that while the model evaluates data to draw conclusions, it provides decision-making data for doctors to understand how the conclusions are reached, so as to achieve quality control and assist doctors in making correct medical decisions. One of the bottleneck problems of current medical artificial intelligence algorithms for prediction and early warning lies in the unexplainability of model conclusions; that is, it is difficult for medical decision-makers to determine whether the conclusions given by the algorithm model are correct and their credibility. Therefore, it is urgent to carry out research that can explain the decisions, predictions of machines and prove their reliability. That is, it requires traditional diagnostic prediction models to provide interpretable text that can be read by medical decision-makers. At present, the interpretability of model conclusions is also a newly focused issue in the field of AI. Some scholars have tried to build models through interpretable representations, but this method is only applicable to specific classifiers and cannot be generalized. Other scholars have tried to visualize the influence of hidden elements on prediction results using heatmaps. These exploratory studies have certain interpretive effects, but do not use fine-grained information that can be used to explain model behavior. Conducting research on high-quality text to explain medical prediction results can make AI model conclusions more accurate and trustworthy, and has important research significance. Summary of the Invention
[0005] In view of this, the present invention provides an interpretable monitoring data recognition method based on improved DQN, aiming to directly perform real-time online learning on monitoring data, generate accurate prediction results, and generate interpretable text to explain the prediction results at the same time.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] An interpretable monitoring data recognition method based on improved DQN, combined with Figure 1 , which is characterized by including the following steps:
[0008] S1: After preprocessing the monitoring data, sample it according to time series and period T and use it as input;
[0009] S2: Use the Markov chain model to generate the state transition probability P and construct a reinforcement learning network (DQN);
[0010] S3: Use the reinforcement learning network described in step S2 as the predictor of the mutual attention mechanism and construct an improved DQN model;
[0011] S4: Use the historical monitoring data and the corresponding explanatory text as input to train the improved DQN model;
[0012] S5: Real-time collect monitoring data and use the improved DQN model to analyze it, identify the state therein and output the corresponding interpretable text;
[0013] The reinforcement learning network consists of a tuple containing five elements (S, A, R, P, γ), where R is the reward function, P is the state transition probability, and γ is the discount factor; S is the state space, which is the input monitoring data; A is the action space, including two types of actions: waiting to monitor more data and making a timely choice corresponding to an explanatory label;
[0014] The mutual attention mechanism consists of an encoder in series with a pair of parallel generators and predictors, and then in series with a classifier; all explanatory label categories are preset in the classifier.
[0015] Further, the preprocessing of the monitoring data described in step S1 includes filling in the gaps and normalizing the monitoring data; the selection of the period T should be much smaller than the total monitoring duration, and at the same time, the accuracy and computing power should be taken into account; the input monitoring data is a q×T-dimensional matrix, where q is the category of the monitoring data.
[0016] Further, the generation of the state transition probability P described in step S2 is specifically as follows: The historical monitoring data and the corresponding explanatory texts are statistically analyzed, and a state transition matrix is established using a Markov chain model, which is the state transition probability P. The action at this moment is determined based on the probability.
[0017] Optionally, step S2 can be implemented using a recommendation system based on a reinforcement learning network. By sorting the similarities of historical cycle data, the action corresponding to the cycle with the highest similarity is recommended.
[0018] Further, the reward function R of the reinforcement learning network at time t corresponds to
[0019]
[0020] where s t ∈S is the state at time t, and a t ∈A is the action at time t; p>0 are the compromise parameters for accuracy and early predictability, respectively, and can be obtained through training in step S4.
[0021] Further, the encoder and the generator are convolutional neural networks (CNNs); the classifier is a multi-class classifier, such as a random forest classifier, a naive Bayes classifier, a convolutional neural network (CNN), etc.
[0022] Further, the training of the improved DQN model described in step S4 specifically includes two training processes:
[0023] (1) For the labeled test data set D, the parameters of both the reinforcement learning network and the mutual attention mechanism network are trained simultaneously; its performance evaluation mechanism is the accuracy of the classifier C:
[0024]
[0025] where i = 1,..., n is the number of the labeled test data set D, # is the operation of finding the number of data in the set; s i ∈S, l i is the corresponding state and the true explanatory label of the i-th test data; is the explanatory label corresponding to the i-th test data predicted by the reinforcement learning network; C(s i ) is the explanatory label corresponding to the i-th test data predicted by the classifier;
[0026] (2) For the unlabeled test data set, the parameters of the reinforcement learning network are tuned using the trained mutual attention mechanism network.
[0027] Furthermore, the improved DQN model described in step S5 can be adjusted according to specific requirements: for cases with not very high accuracy requirements but high real-time requirements, the generator can be trimmed, and only the encoder is used to concatenate a reinforcement learning network as a predictor, and then a classifier is concatenated to achieve the goal.
[0028] An electronic device applied to the interpretable monitoring data recognition method based on the improved DQN includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor , The instructions are executed by the at least one processor, so that the at least one processor can execute the steps of the interpretable monitoring data recognition method based on the improved DQN according to any one of claims 1 to 8.
[0029] A readable storage medium applied to the interpretable monitoring data recognition method based on the improved DQN, the readable storage medium is a computer-readable storage medium, and a program for implementing the interpretable monitoring data recognition method based on the improved DQN is stored on the computer-readable storage medium. The program for implementing the interpretable monitoring data recognition method based on the improved DQN is executed by a processor to implement the steps of the interpretable monitoring data recognition method based on the improved DQN according to any one of claims 1 to 8.
[0030] The beneficial effects of the present invention are as follows: The present invention provides an interpretable monitoring data recognition method based on the improved DQN, which directly performs real-time online learning on the monitoring data, uses the improved DQN model of the temporal convolutional network based on the attention mechanism to perceive the monitoring state, generates accurate prediction results, and at the same time generates interpretable text to explain the prediction results, which can assist early diagnosis and interpretable early warning. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to illustrate the purpose and technical solutions of the present invention, the following drawings are provided for description:
[0032] Figure 1 It is a flowchart of the method of the present invention;
[0033] Figure 2 It is an improved DQN architecture diagram of Embodiment 1 of the present invention;
[0034] Figure 3 It is a simplified improved DQN architecture diagram of Embodiment 2 of the present invention;
[0035] Figure 4 It is a schematic diagram of functional modules of Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0036] Example 1: The historical monitoring and surveillance data of patients provided by a certain hospital from 2014 to 2019 contains common indicators of 8 different critical illnesses (diastolic blood pressure, systolic blood pressure, heart rate, body temperature, respiratory rate, arterial oxygen partial pressure, central venous pressure, blood pH value, etc.), and some of the monitoring data has been annotated with interpretable text. Further, for model training, all data is divided into a training set (from 2014 to 2018) and a test set (2019) by year. To help accurately predict the patient's status and give interpretable text, the present invention proposes an "interpretable monitoring data recognition method based on improved DQN".
[0037] The following will be combined with the attached Figure 1 , and the preferred examples of the present invention will be described in detail.
[0038] S1: After preprocessing the monitoring data, sample it according to time series and period and use it as input.
[0039] The preprocessing of the monitoring data includes filling in the gaps and normalizing the monitoring data; the selection of the period T should be much smaller than the total monitoring duration, and at the same time, the accuracy and computing power should be taken into account; the monitoring data used as input is a q×T-dimensional matrix, where q is the category of the monitoring data.
[0040] S2: Combine Figure 2 , the agent uses the Markov chain model to generate the state transition probability P and constructs a reinforcement learning network (DQN); under the action of the action a t at time t, the environmental state changes from s t to the state s t ' at the next moment.
[0041] The reinforcement learning network is used to describe and solve the problem that the agent maximizes the reward or achieves a specific goal by learning strategies during the interaction with the environment. It consists of a tuple containing five elements (S, A, R, P, γ) of two convolutional neural network structures. Among them, R is the reward function, P is the state transition probability, and γ is the discount factor; S is the state space, which is the input monitoring data; A is the action space, including two types of actions: waiting to monitor more data and making a timely choice corresponding to an explanation label.
[0042] The specific method for generating the state transition probability P is: statistically analyze the historical monitoring data and the corresponding explanation text, and use the Markov chain model to establish a state transition matrix, which is the state transition probability P, and determine the action at this moment according to the probability.
[0043] The reward function R of the reinforcement learning network at time t corresponds to:
[0044]
[0045] Among them, s t is the state at time t, and a t is the action at time t; p > 0 are the compromise parameters for accuracy and early predictability, respectively, which can be obtained through training in step S4.
[0046] S3: Use the reinforcement learning network described in step S2 as the predictor of the mutual attention mechanism to construct an improved DQN model.
[0047] The mutual attention mechanism consists of an encoder in series with a pair of parallel generators and predictors, and then in series with a classifier; all the interpretation label categories are preset in the classifier.
[0048] The encoder and generator are convolutional neural networks; the classifier is a multi-class classifier.
[0049] S4: Use the historical monitoring data and the corresponding interpretation text as inputs to train the improved DQN model.
[0050] The training of the improved DQN model specifically includes two training processes:
[0051] (1) For the labeled test data set D, train the parameters of both the reinforcement learning network and the mutual attention mechanism network simultaneously; its performance evaluation mechanism is: the accuracy of the classifier C
[0052]
[0053] Among them, i = 1,..., n is the number of the labeled test data set D, # is the operation of finding the number of data in the set; s i , l i are the corresponding state and the true interpretation label of the i-th test data; is the interpretation label corresponding to the i-th test data predicted by the reinforcement learning network; C(s i ) is the interpretation label corresponding to the i-th test data predicted by the classifier;
[0054] (2) For the unlabeled test data set, use the trained mutual attention mechanism network to optimize the parameters of the reinforcement learning network.
[0055] S5: Collect monitoring data in real time and use the improved DQN model to analyze it, identify the state therein and output the corresponding interpretable text.
[0056] Example 2: Movie-Lens is a movie recommendation system based on ratings, created by the Group-Lens research group at the University of Minnesota in the United States. It includes three datasets of sizes 100KB, 1MB, and 10MB. For the historical behavior data of users, movies are recommended to users with reasons for recommendation (interpretability text) stated. The present invention provides an "Interpretability Monitoring Data Identification Method Based on Improved DQN".
[0057] The following will be combined with the attached Figure 1 , content that is the same as or similar to the above Example 1 can be referred to the above introduction and will not be elaborated hereinafter. Specifically, it includes the following steps:
[0058] S1: After preprocessing the monitoring data, it is sampled according to time series and period and used as input.
[0059] S2: Construct a recommendation system for the reinforcement learning network, and use the similarity ranking of historical period data to recommend the action corresponding to the period with the highest similarity.
[0060] The reinforcement learning network consists of a tuple containing five elements (S, A, R, P, γ). Among them, R is the reward function, P is the state transition probability, and γ is the discount factor; S is the state space, which is the input monitoring data; A is the action space, including two types of actions: waiting to monitor more data and making a timely choice corresponding to a certain explanation label.
[0061] S3: Use the reinforcement learning network described in step S2 as the predictor of the mutual attention mechanism to construct an improved DQN model.
[0062] The mutual attention mechanism consists of an encoder in series with a pair of parallel generators and predictors, and then in series with a classifier; all explanation label categories are preset in the classifier.
[0063] S4: Use the historical monitoring data and the corresponding explanation text as input to train the improved DQN model.
[0064] S5: Combine Figure 3 , after trimming the generator, the mutual attention mechanism consists of an encoder in series with a predictor, and then in series with a classifier to form a simplified improved DQN model; all explanation label categories are preset in the classifier. Real-time collect monitoring data and use the simplified improved DQN model to analyze it, identify the states therein and output the corresponding interpretability text.
[0065] Embodiment 3: An embodiment of the present invention provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the interpretable monitoring data recognition method based on the improved DQN in the above Embodiment 1 or Embodiment 2.
[0066] Reference is made below Figure 4 , which shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure. Figure 4 The electronic device shown is only an example and should not impose any limitation on the functions and scope of use of the embodiments of the present disclosure.
[0067] As Figure 4 shown, the electronic device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM 1002) or a program loaded from a storage device into a random access memory (RAM 1004). In the RAM 1004, various programs and data required for the operation of the electronic device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface is also connected to the bus 1005.
[0068] Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an electronic device having various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be alternatively implemented or had.
[0069] Specifically, according to the embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above functions defined in the method of the embodiments of the present disclosure are executed.
[0070] The electronic device provided by the present invention adopts the interpretable monitoring data recognition method based on the improved DQN in the above-mentioned Embodiment 1 or Embodiment 2, which improves the efficiency of prediction results and interpretable texts. Compared with the prior art, the beneficial effects of the electronic device provided by the embodiments of the present invention are the same as those of the interpretable monitoring data recognition method based on the improved DQN provided in the above-mentioned Embodiment 1, and other technical features in this electronic device are the same as the features disclosed in the method of Embodiment 1 or Embodiment 2, and will not be elaborated herein.
[0071] It should be understood that each part of the present disclosure can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0072] As described above, only the specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
[0073] Embodiment 4: The embodiments of the present invention provide a readable storage medium, which is a computer-readable storage medium. The computer-readable storage medium has computer-readable program instructions stored thereon, and the computer-readable program instructions are used to execute the interpretable monitoring data recognition method based on the improved DQN in the above-mentioned Embodiment 1 or Embodiment 2.
[0074] The computer-readable storage medium provided by the embodiments of the present invention can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0075] The above computer-readable storage medium may be included in an electronic device; or it may exist independently without being assembled into the electronic device.
[0076] The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by an electronic device, the electronic device is enabled to: obtain monitoring data, directly perform real-time online learning on the monitoring data, use an improved DQN model based on an attention mechanism-based temporal convolutional network to perceive the monitoring state, and generate an interpretable text while generating an accurate prediction result to explain the prediction result.
[0077] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the computer of the person to be detected, partially on the computer of the person to be detected, executed as an independent software package, partially on the computer of the person to be detected and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the computer of the person to be detected through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by connecting through an Internet service provider using the Internet).
[0078] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0079] The modules involved in the embodiments described in the present disclosure may be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.
[0080] The computer-readable storage medium provided by the present invention stores computer-readable program instructions for executing the above-mentioned interpretability monitoring data recognition method based on improved DQN, which improves the efficiency and accuracy of prediction results and interpretability texts. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the embodiments of the present invention are the same as those of the interpretability monitoring data recognition method based on improved DQN provided in the above-mentioned Embodiment 1 or Embodiment 2, and will not be elaborated here.
[0081] Embodiment 5: The embodiments of the present invention further provide a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the interpretability monitoring data recognition method based on improved DQN as described above are implemented.
[0082] The computer program product provided by the present application improves the efficiency and accuracy of prediction results and interpretability texts. Compared with the prior art, the beneficial effects of the computer program product provided by the embodiments of the present invention are the same as those of the interpretability monitoring data recognition method based on improved DQN provided in the above-mentioned Embodiment 1 or Embodiment 2, and will not be elaborated here.
[0083] Finally, it should be noted that the above preferred embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made in form and details without departing from the scope defined by the claims of the present invention.
Claims
1. An interpretable monitoring data recognition method based on improved DQN, characterized in that, it includes the following steps: S1: After preprocessing the monitoring data, sample it according to time sequence and period T and use it as input; S2: Use the Markov chain model to generate the state transition probability P and construct a reinforcement learning network (DQN); S3: Use the reinforcement learning network described in step S2 as the predictor of the mutual attention mechanism to construct an improved DQN model; S4: Use historical monitoring data and corresponding explanatory texts as input to train the improved DQN model; S5: Real-time collect monitoring data and use the improved DQN model to analyze it, identify the status therein and output the corresponding interpretable text; The reinforcement learning network consists of a tuple containing five elements (S, A, R, P, γ), where R is the reward function, P is the state transition probability, and γ is the discount factor; S is the state space, which is the input monitoring data; A is the action space, including two types of actions: waiting to monitor more data and making a timely choice corresponding to an explanatory label; The mutual attention mechanism consists of an encoder in series with a pair of parallel generators and predictors, and then in series with a classifier; all explanatory label categories are preset in the classifier; The training of the improved DQN model described in step S4 specifically includes two training processes: (1) For the labeled test data set D, train the parameters of both the reinforcement learning network and the mutual attention mechanism network simultaneously; its performance evaluation mechanism is: the accuracy of the classifier C where \(i = 1,\ldots,n\) is the number of the already marked test data set \(D\), \(\#\) is the operation of finding the number of data in the set; \(s\) i , \(l\) i are the corresponding state and the true interpretation label of the \(i\)-th test data; is the interpretation label corresponding to the \(i\)-th test data predicted by the reinforcement learning network; \(C(s\) i ) is the interpretation label corresponding to the \(i\)-th test data predicted by the classifier; (2) For the unlabeled test data set, use the trained mutual attention mechanism network to optimize the parameters of the reinforcement learning network.
2. The interpretable monitoring data recognition method based on improved DQN according to claim 1, characterized in that, The monitoring data preprocessing described in step S1 includes filling in the gaps and normalizing the monitoring data; the selection of the period T should be much smaller than the total monitoring duration, and at the same time, the accuracy and computing power should be taken into account; the input monitoring data is a q×T-dimensional matrix, where q is the category of the monitoring data.
3. The interpretable monitoring data recognition method based on improved DQN according to claim 1, characterized in that, The specific method for generating the state transition probability P described in step S2 is: count the historical monitoring data and corresponding explanatory texts, use the Markov chain model to establish the state transition matrix, which is the state transition probability P, and determine the action at the current moment according to the probability.
4. The interpretable monitoring data recognition method based on improved DQN according to claim 1, characterized in that, The step S2 is implemented by using a recommendation system based on the reinforcement learning network, and uses the similarity ranking of historical period data to recommend the action corresponding to the period with the highest similarity.
5. The interpretable monitoring data recognition method based on improved DQN according to claim 1, characterized in that, The reward function R of the reinforcement learning network at time t corresponds to: Among them, s t is the state at time t, and a t is the action at time t; p > 0 are the compromise parameters of accuracy and early predictability, respectively, and are obtained by training in step S4.
6. The interpretable monitoring data recognition method based on improved DQN according to claim 1, It is characterized in that the encoder and the generator are convolutional neural networks; the classifier is a multi-class classifier.
7. The interpretable monitoring data recognition method based on improved DQN according to claim 1 It is characterized in that the improved DQN model described in step S5 is adjusted according to specific requirements: for the case where the accuracy requirement is not very high and the real-time requirement is high, the generator is cut off, and only the encoder is used to connect a reinforcement learning network in series as a predictor, and then a classifier is connected in series to implement.
8. An electronic device It is characterized in that the electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the interpretable monitoring data recognition method based on improved DQN according to any one of claims 1 to 7.
9. A readable storage medium It is characterized in that the readable storage medium is a computer-readable storage medium, and a program for implementing the interpretable monitoring data recognition method based on improved DQN is stored on the computer-readable storage medium, and the program for implementing the interpretable monitoring data recognition method based on improved DQN is executed by a processor to implement the steps of the interpretable monitoring data recognition method based on improved DQN according to any one of claims 1 to 7.