Embedded software timing anomaly detection method based on value function reinforcement learning

Through the method of reinforcement learning based on value functions, combined with long and short-term memory networks and deep Q learning algorithms, the adaptability and timing dependence problems of timing abnormality detection in embedded software in the prior art are solved, and the high-accurate abnormality detection effect is achieved.

CN119759797BActive Publication Date: 2025-05-16ZHEJIANG SHUXIN NETWORK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510267696.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-05-16
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

The existing embedded software timing anomaly detection technology lacks adaptive learning ability and is difficult to deal with complex and changeable anomaly scenarios. The traditional methods ignore the long-term dependence between timing data, resulting in limited detection accuracy.

Method used

Using a method of reinforcement learning based on value functions, the timing data is obtained and the state space matrix is ​​constructed, the timing feature vector is extracted using a long and short-term memory network, and the Q-value function network is trained in combination with a deep Q-learning algorithm to construct a reward function to evaluate the degree of abnormality, and the exception score of the timing fragment is calculated by optimizing the Q-value function network.

Benefits of technology

It realizes automatic learning and precise identification of timing exceptions of embedded software, improves detection accuracy and efficiency, and has stronger timing dependency modeling and generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119759797B_ABST
    Figure CN119759797B_ABST
Patent Text Reader

Abstract

The present invention provides an embedded software time series anomaly detection method based on value function reinforcement learning, which relates to the field of anomaly detection technology, including obtaining time series data and dividing time series segments, constructing a state space matrix, extracting feature vectors using a long short-term memory network, generating a Q value function network based on a deep neural network, constructing a reward function based on Mahalanobis distance, using a deep Q learning algorithm to train the network and optimize parameters, and finally realizing anomaly detection. The present invention can effectively capture the long-term dependency of time series data, improve the accuracy of anomaly detection, reduce the false alarm rate, and has a strong generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to anomaly detection technology, and in particular to an embedded software timing anomaly detection method based on value function reinforcement learning. Background Art

[0002] Embedded software may have various abnormal behaviors during operation, such as dead loops, memory leaks, execution timing anomalies, etc. These anomalies may cause system performance degradation or functional failure. Therefore, it is of great significance to effectively detect abnormal behaviors during the operation of embedded software.

[0003] Time series anomaly detection is an important research direction for anomaly detection in embedded software. Traditional time series anomaly detection methods are mainly based on statistical analysis and machine learning algorithms, which identify abnormal patterns by establishing normal behavior models. With the development of deep learning technology, anomaly detection methods based on deep neural networks have shown good performance and can automatically learn complex time series feature representations. As an important branch of artificial intelligence, reinforcement learning learns the optimal strategy through the interaction between the intelligent agent and the environment, and has broad application prospects in the field of anomaly detection.

[0004] The existing embedded software timing anomaly detection technology has the following shortcomings:

[0005] Existing methods mainly rely on pre-defined abnormal patterns and rules, lack adaptive learning capabilities, and are difficult to cope with complex and changing abnormal scenarios. This rule-based approach requires a lot of expert experience and is prone to false positives and false negatives.

[0006] Traditional anomaly detection algorithms often treat time series data as independent samples, ignoring the long-term dependencies between time series data and failing to effectively capture the dynamic changes of time series features, resulting in limited detection accuracy. Summary of the invention

[0007] The embodiment of the present invention provides an embedded software timing anomaly detection method based on value function reinforcement learning, which can solve the problems in the prior art.

[0008] A first aspect of an embodiment of the present invention provides an embedded software timing anomaly detection method based on value function reinforcement learning, comprising:

[0009] Acquire the time series data during the operation of the embedded software, and divide the time series data into a plurality of time series segments according to a preset time window; construct a state space matrix for each of the time series segments, wherein the row vector of the state space matrix represents the time series feature, and the column vector represents the time step;

[0010] Based on the state space matrix, a long short-term memory network is used to extract a time series feature vector; the time series feature vector is input into a deep neural network to generate a Q value function network, and the Q value function network is used to evaluate the action value function under the time series state; a reward function is constructed, and the reward function is calculated based on the abnormality degree of the time series segment, and the abnormality degree is obtained by measuring the Mahalanobis distance between the current time series segment and the historical normal time series segment;

[0011] The Q-value function network is trained using a deep Q-learning algorithm. During the training process, an action is selected based on an ε-greedy strategy, and the selected action and the current state are input into the Q-value function network to obtain a Q-value; the Q-value is compared with a target Q-value calculated based on the reward function, and the parameters of the Q-value function network are optimized by minimizing a loss function;

[0012] Based on the optimized Q-value function network, the anomaly score of the time series segment is calculated. The anomaly score is obtained by comparing the Q-value of the optimal action in the current state with the preset threshold. When the anomaly score exceeds the preset threshold, it is determined that the current time series segment has an anomaly, and the anomaly detection result is output.

[0013] Inputting the time series feature vector into a deep neural network to generate a Q value function network, wherein the Q value function network is used to evaluate the action value function under the time series state, including:

[0014] Acquire a time series feature vector, input the time series feature vector into a multi-layer perceptron structure, wherein the multi-layer perceptron structure includes a first fully connected layer, a second fully connected layer, and an output layer, wherein the first fully connected layer includes 256 neurons and uses a ReLU activation function, the second fully connected layer includes 128 neurons and uses a ReLU activation function, and the output layer uses a hyperbolic tangent function to limit the output value to a range of [-1, 1], thereby obtaining a depth mapping feature;

[0015] A Q-value function network of a dual-stream network structure is constructed based on the deep mapping features, wherein the Q-value function network includes a state value function branch and an advantage function branch. The output value of the state value function branch is combined with the output value of the advantage function branch, and the average value of the advantage function in the action space is subtracted to obtain a state-action value function.

[0016] The Q-value function network is trained using a deep Q-learning algorithm. During the training process, an action is selected based on an ε-greedy strategy, and the selected action and the current state are input into the Q-value function network to obtain a Q-value; the Q-value is compared with a target Q-value calculated based on the reward function, and the parameters of the Q-value function network are optimized by minimizing a loss function, including:

[0017] Construct an action selection probability distribution model, determine the optimal action according to the state-action value output by the Q-value function network, assign a selection probability of 1-ε+ε / |A| to the optimal action, and assign a selection probability of ε / |A| to other actions, where |A| is the size of the action space; update the exploration rate ε based on the training rounds, and the initial value of the exploration rate is the product of the difference between the maximum exploration rate and the minimum exploration rate and the decay exponent;

[0018] Calculate state uncertainty, obtain multiple Q value samples corresponding to the state-action, calculate the variance of the multiple Q value samples and their mean, use the variance as a measure of state uncertainty, and use the product of the exploration rate and the state uncertainty as an adaptive exploration rate; select an action based on the adaptive exploration rate;

[0019] Construct a multi-step temporal difference target based on the selected action, calculate the discounted accumulation of the immediate reward and the maximum Q value of the terminal state to obtain the n-step cumulative reward, perform a weighted combination of the n-step cumulative rewards to obtain the multi-step temporal difference target value, and take the difference between the multi-step temporal difference target value and the current Q value as the temporal difference error;

[0020] A priority experience replay mechanism is established according to the temporal difference error, the sum of the absolute value of the temporal difference error and the preset priority bias is used as the sample priority, the α-power of the sample priority is divided by the sum of the α-powers of all sample priorities to obtain a sampling probability distribution, samples in the experience replay pool are sampled based on the sampling probability distribution, and the β-power of the inverse of the product of the sampling probability and the size of the experience pool is used as the importance weight;

[0021] Construct a parameter optimization target. When the absolute value of the time series difference error is less than a threshold parameter, the square error is used. When the absolute value of the time series difference error is greater than the threshold parameter, the linear error minus the fixed bias is used to obtain the Huber loss. The product of the Huber loss and the importance weight is used as the final optimization target. Based on the optimization target, the parameters of the Q-value function network are gradient updated.

[0022] Based on the selected action, a multi-step TD target is constructed, the discounted accumulation of the immediate reward and the maximum Q value of the terminal state are calculated to obtain the n-step cumulative reward, the n-step cumulative reward is weightedly combined to obtain the multi-step TD target value, and the difference between the multi-step TD target value and the current Q value is taken as the TD error, including:

[0023] The formula for calculating the cumulative return in n steps is as follows:

[0024] ;

[0025] in, : The cumulative return of n steps starting at time t;

[0026] γ: discount factor, range [0,1], used to balance immediate rewards and future rewards;

[0027] r t+k : The instant reward obtained at time t+k;

[0028] : Q value of all actions in the state at time t+n; s t+n : state at time t+n; a represents action; n: cumulative number of steps;

[0029] The formula for calculating the multi-step timing difference target value is as follows:

[0030] ;

[0031] y t : Multi-step time difference target value at time t;

[0032] λ: weight parameter used to control the rewards of different step lengths;

[0033] N: maximum number of steps considered;

[0034] The timing differential error calculation formula is as follows:

[0035] ;

[0036] δ t : The timing difference error at time t.

[0037] Based on the optimized Q-value function network, the anomaly score of the time series segment is calculated. The anomaly score is obtained by comparing the Q-value of the optimal action in the current state with a preset threshold, including:

[0038] Extract the optimal action of the current state from the Q-value function network, obtain the corresponding optimal Q-value based on the optimal action, calculate the mean and standard deviation of the historical optimal Q-values ​​within a preset time window, and use the weighted sum of the mean and standard deviation as the dynamic threshold;

[0039] The difference is calculated based on the average of the optimal Q value and the historical optimal Q value, and the ratio of the difference to the standard deviation is used as the basic anomaly score. The time decay function is introduced to calculate the time series weight, and the time series weight is calculated by the time difference between the current moment and the reference moment. The product of the basic anomaly score and the time series weight is used as the weighted anomaly score;

[0040] The transition probability between adjacent states is calculated, and the negative logarithm of the transition probability is used as the transition anomaly score. The state feature vector is extracted to calculate the Euclidean distance between adjacent states to obtain the state representation distance. The weighted anomaly score, the transition anomaly score, and the state representation distance are weighted and combined by a preset weight coefficient to obtain a comprehensive anomaly score.

[0041] The difference is calculated based on the average of the optimal Q value and the historical optimal Q value, and the ratio of the difference to the standard deviation is used as the basic anomaly score. The time decay function is introduced to calculate the time series weight, and the time series weight is calculated by the time difference between the current moment and the reference moment. It includes:

[0042] Obtaining the optimal Q value at the current moment, calculating the difference between the optimal Q value at the current moment and the mean to obtain an abnormal deviation, and dividing the abnormal deviation by the standard deviation to obtain a normalized abnormal score;

[0043] Constructing a time decay function based on a Gaussian kernel function, substituting the time difference between the current time and the reference time into the time decay function, and calculating a timing weight coefficient according to a preset time decay coefficient;

[0044] The normalized anomaly score is multiplied by the time series weight coefficient to obtain a weighted anomaly score, and the weighted anomaly score decays as the time difference increases.

[0045] A second aspect of the embodiments of the present invention

[0046] An electronic device is provided, comprising:

[0047] processor;

[0048] a memory for storing processor-executable instructions;

[0049] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0050] According to a third aspect of the embodiments of the present invention,

[0051] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0052] The beneficial effects of this application are as follows:

[0053] The present invention realizes embedded software time series anomaly detection by combining reinforcement learning with deep neural network, which can automatically learn the feature representation of time series data, avoid the complexity of artificial feature engineering, and improve the accuracy and efficiency of anomaly detection.

[0054] The present invention adopts long short-term memory network to extract timing features, and combines it with deep Q learning algorithm to train Q value function network. The reward function is used to guide the model to learn to distinguish normal and abnormal timing patterns, so that the model has stronger timing dependency modeling and generalization capabilities.

[0055] The present invention constructs a reward function based on the Mahalanobis distance metric, which can effectively capture the degree of abnormality of time series data, and evaluates the state action value through the Q-value function network, thereby realizing accurate identification of time series anomalies, and at the same time having good real-time and interpretability. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 It is a flow chart of an embedded software timing anomaly detection method based on value function reinforcement learning according to an embodiment of the present invention;

[0057] Figure 2 The accuracy trend of different methods during training is shown in Figure 2.

[0058] Figure 3 is the average reward value obtained by different methods in different training rounds;

[0059] Figure 4 Uncertainty distribution of this technical solution and the traditional ε-greedy strategy in the state space. DETAILED DESCRIPTION

[0060] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0061] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0062] Figure 1 FIG. 1 is a flow chart of an embedded software timing anomaly detection method based on value function reinforcement learning according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0063] Acquire the time series data during the operation of the embedded software, and divide the time series data into a plurality of time series segments according to a preset time window; construct a state space matrix for each of the time series segments, wherein the row vector of the state space matrix represents the time series feature, and the column vector represents the time step;

[0064] Based on the state space matrix, a long short-term memory network is used to extract a time series feature vector; the time series feature vector is input into a deep neural network to generate a Q value function network, and the Q value function network is used to evaluate the action value function under the time series state; a reward function is constructed, and the reward function is calculated based on the abnormality degree of the time series segment, and the abnormality degree is obtained by measuring the Mahalanobis distance between the current time series segment and the historical normal time series segment;

[0065] The Q-value function network is trained using a deep Q-learning algorithm. During the training process, an action is selected based on an ε-greedy strategy, and the selected action and the current state are input into the Q-value function network to obtain a Q-value; the Q-value is compared with a target Q-value calculated based on the reward function, and the parameters of the Q-value function network are optimized by minimizing a loss function;

[0066] Based on the optimized Q-value function network, the anomaly score of the time series segment is calculated. The anomaly score is obtained by comparing the Q-value of the optimal action in the current state with the preset threshold. When the anomaly score exceeds the preset threshold, it is determined that the current time series segment has an anomaly, and the anomaly detection result is output.

[0067] In an optional implementation, the time series feature vector is input into a deep neural network to generate a Q value function network, and the Q value function network is used to evaluate the action value function under the time series state, including:

[0068] Acquire a time series feature vector, input the time series feature vector into a multi-layer perceptron structure, wherein the multi-layer perceptron structure includes a first fully connected layer, a second fully connected layer, and an output layer, wherein the first fully connected layer includes 256 neurons and uses a ReLU activation function, the second fully connected layer includes 128 neurons and uses a ReLU activation function, and the output layer uses a hyperbolic tangent function to limit the output value to a range of [-1, 1], thereby obtaining a depth mapping feature;

[0069] A Q-value function network of a dual-stream network structure is constructed based on the deep mapping features, wherein the Q-value function network includes a state value function branch and an advantage function branch. The output value of the state value function branch is combined with the output value of the advantage function branch, and the average value of the advantage function in the action space is subtracted to obtain a state-action value function.

[0070] The time series feature vector is obtained by preprocessing and extracting features from the input data. It usually includes information in multiple dimensions and can reflect the state of the system at different time points. These features can be obtained from sensor data, historical records, or other data sources.

[0071] The multi-layer perceptron structure consists of multiple fully connected layers, including the first fully connected layer, the second fully connected layer, and the output layer. The first fully connected layer contains 256 neurons and uses the ReLU activation function to introduce nonlinear features and enhance the expressiveness of the model. The second fully connected layer contains 128 neurons and also uses the ReLU activation function. The output layer uses the hyperbolic tangent function to limit the output value to the range of -1 to 1 to ensure the stability and controllability of the output. Through this process, deep mapping features are generated, which can effectively capture the complex relationship of the input features.

[0072] Based on the deep mapping features, a Q-value function network with a dual-stream network structure is constructed. The network structure includes a state value function branch and an advantage function branch. The state value function branch is responsible for evaluating the value of the current state, while the advantage function branch evaluates the advantage of taking a specific action in this state. The output values ​​of the two are weighted and combined to form the final state-action value function. During the combination process, the output value of the advantage function is subtracted from the average value in the action space to eliminate unnecessary deviations and ensure that the final Q-value function can accurately reflect the value of the state and action.

[0073] This application can effectively extract and map complex time series features through the design of a multi-layer perceptron structure, improving the model's learning ability and prediction accuracy. The introduction of a dual-stream network structure makes the evaluation of the state value function and the advantage function more independent and accurate, and enhances the model's flexibility in action selection under different states. The output value restrictions and combination methods ensure the stability and interpretability of the Q-value function, facilitating subsequent decision-making and strategy optimization.

[0074] In an optional implementation, a deep Q learning algorithm is used to train the Q value function network. During the training process, an action is selected based on an ε-greedy strategy, and the selected action and the current state are input into the Q value function network to obtain a Q value; the Q value is compared with a target Q value calculated based on the reward function, and the parameters of the Q value function network are optimized by minimizing the loss function, including:

[0075] Construct an action selection probability distribution model, determine the optimal action according to the state-action value output by the Q-value function network, assign a selection probability of 1-ε+ε / |A| to the optimal action, and assign a selection probability of ε / |A| to other actions, where |A| is the size of the action space; update the exploration rate ε based on the training rounds, and the initial value of the exploration rate is the product of the difference between the maximum exploration rate and the minimum exploration rate and the decay exponent;

[0076] Calculate state uncertainty, obtain multiple Q value samples corresponding to the state-action, calculate the variance of the multiple Q value samples and their mean, use the variance as a measure of state uncertainty, and use the product of the exploration rate and the state uncertainty as an adaptive exploration rate; select an action based on the adaptive exploration rate;

[0077] Construct a multi-step temporal difference target based on the selected action, calculate the discounted accumulation of the immediate reward and the maximum Q value of the terminal state to obtain the n-step cumulative reward, perform a weighted combination of the n-step cumulative rewards to obtain the multi-step temporal difference target value, and take the difference between the multi-step temporal difference target value and the current Q value as the temporal difference error;

[0078] A priority experience replay mechanism is established according to the temporal difference error, the sum of the absolute value of the temporal difference error and the preset priority bias is used as the sample priority, the α-power of the sample priority is divided by the sum of the α-powers of all sample priorities to obtain a sampling probability distribution, samples in the experience replay pool are sampled based on the sampling probability distribution, and the β-power of the inverse of the product of the sampling probability and the size of the experience pool is used as the importance weight;

[0079] Construct a parameter optimization target. When the absolute value of the time series difference error is less than a threshold parameter, the square error is used. When the absolute value of the time series difference error is greater than the threshold parameter, the linear error minus the fixed bias is used to obtain the Huber loss. The product of the Huber loss and the importance weight is used as the final optimization target. Based on the optimization target, the parameters of the Q-value function network are gradient updated.

[0080] The Q-value function network is trained using a deep Q-learning algorithm. During training, actions are selected based on the ε-greedy strategy. Specifically, the probability distribution model for action selection is determined based on the state-action value output by the Q-value function network. A higher selection probability is assigned to the optimal action, while a lower selection probability is assigned to other actions. The selection probability is calculated by subtracting ε from 1 and adding ε divided by the size of the action space. This ensures that the optimal action is selected in most cases while retaining a certain amount of exploration space.

[0081] In the early stages of training, the exploration rate ε is set to a maximum value to encourage exploration. As training progresses, the exploration rate gradually decays until it reaches a minimum value. The decay process is achieved by setting a decay exponent, which allows more exploration in the early stages of training and more use of learned knowledge in the later stages.

[0082] Calculate the uncertainty of the state. By taking multiple Q-value samples corresponding to the state-action, calculate the variance between these samples and their mean. This variance is used as a measure of state uncertainty. Based on the state uncertainty, an adaptive exploration rate can be calculated, that is, the product of the exploration rate and the state uncertainty is used as the new exploration rate. This process ensures that the system can perform more exploration in states with higher uncertainty.

[0083] When selecting an action, a multi-step TD target is constructed based on the selected action. The discounted sum of the immediate reward is calculated and combined with the maximum Q value of the terminal state to obtain the cumulative reward for n steps. These cumulative rewards are then weighted and combined to form the multi-step TD target value. Finally, the difference between the multi-step TD target value and the current Q value is calculated as the TD error.

[0084] According to the time difference error, a priority experience replay mechanism is established. The absolute value of the time difference error is calculated and added to the preset priority bias to obtain the priority of the sample. The α power of the sample priority is divided by the sum of the α powers of all sample priorities to form a sampling probability distribution. Based on this distribution, samples are sampled from the experience replay pool, and the β power of the inverse of the product of the sampling probability and the size of the experience pool is used as the importance weight.

[0085] When the absolute value of the time series difference error is less than the set threshold, the square error is used for optimization; when the absolute value is greater than the threshold, the linear error minus the fixed bias is used to form the Huber loss. The product of the Huber loss and the importance weight is used as the final optimization target, and the parameters of the Q value function network are gradient updated based on this target.

[0086] Figure 2 The accuracy trends of different methods during training. This technical solution showed obvious advantages in the early stage of training (0-20,000 steps), with the fastest increase in accuracy, reaching an accuracy of more than 0.85 at 40,000 steps. When the number of training steps reached 60,000 steps, the accuracy of this technical solution stabilized at around 0.95, while other methods performed poorly at the same number of steps: the accuracy of the DQN baseline method ultimately only reached 0.85, the accuracy of Double DQN was 0.88, and the accuracy of Dueling DQN was 0.90. It is particularly noteworthy that in the critical learning stage of 20,000-40,000 steps, the slope of the learning curve of this technical solution was significantly greater than that of other methods, indicating that it has faster learning efficiency. In the later stage of training (after 80,000 steps), the performance fluctuation of this technical solution was also significantly smaller than that of other methods, and the standard deviation was reduced to within 0.01, showing excellent stability.

[0087] Figure 3The average reward values ​​obtained by the four methods at different training rounds were compared. This technical solution showed obvious advantages within the first 10,000 rounds of training, and the reward growth rate was the fastest. When the training round reached 30,000, the average reward of this technical solution had exceeded 350, while the other methods were still hovering around 300. By the end of the training (about 40,000 rounds), the average reward of this technical solution stabilized at around 400, which was about 30 points higher than the second best Dueling DQN method (370) and about 80 points higher than the worst DQN baseline method (320). It can also be observed in the figure that this technical solution showed the steepest rising curve during the 15,000-25,000 rounds, indicating that this is the stage where the algorithm learning effect is most significant. At the same time, after obtaining a high reward, the fluctuation range is controlled within ±10, reflecting good stability.

[0088] In an optional implementation, a multi-step TD target is constructed based on the selected action, the discounted accumulation of the immediate reward and the maximum Q value of the terminal state are calculated to obtain an n-step cumulative reward, the n-step cumulative rewards are weightedly combined to obtain a multi-step TD target value, and the difference between the multi-step TD target value and the current Q value is used as a TD error, including:

[0089] The formula for calculating the cumulative return in n steps is as follows:

[0090] ;

[0091] in, : The cumulative return of n steps starting at time t;

[0092] γ: discount factor, range [0,1], used to balance immediate rewards and future rewards;

[0093] r t+k : The instant reward obtained at time t+k;

[0094] : Q value of all actions in the state at time t+n; s t+n : state at time t+n; a represents action; n: cumulative number of steps;

[0095] The formula for calculating the multi-step timing difference target value is as follows:

[0096] ;

[0097] y t : Multi-step time difference target value at time t;

[0098] λ: weight parameter used to control the rewards of different step lengths;

[0099] N: maximum number of steps considered;

[0100] The timing differential error calculation formula is as follows:

[0101] ;

[0102] δ t : The timing difference error at time t.

[0103] The technical solution of constructing a multi-step TD target based on the selected actions can be achieved by the following steps:

[0104] The state space represents the state of the system at any time, while the action space is all the possible actions that the system can take in that state. By defining the state and action, we can provide a basis for the subsequent decision-making process.

[0105] The discount factor is used to balance the relationship between immediate rewards and future rewards. The selected discount factor should be between 0 and 1. The closer the value is to 1, the greater the impact of future rewards. By choosing the discount factor reasonably, the learning process can be effectively guided, making the model more cautious when considering future rewards.

[0106] Immediate rewards refer to the feedback that the system receives immediately after performing an action. By recording the immediate rewards for each action, important information can be provided for subsequent learning.

[0107] On this basis, the n-step cumulative reward is calculated. The n-step cumulative reward refers to the weighted sum of all immediate rewards obtained after n time steps from the current moment. By weighting the immediate rewards, the impact of different time steps on the current decision can be better reflected.

[0108] Get the Q value of all actions in the state after n steps. The Q value represents the expected return of taking an action in a specific state. By evaluating the Q value of all possible actions in the state after n steps, more comprehensive information can be provided for the current decision.

[0109] The multi-step temporal difference target value is obtained by weighted combination of the n-step cumulative return and the maximum Q value of the state after n steps. In this way, the current learning goal can be combined with the expected return in the future, thereby improving the effectiveness of learning.

[0110] The TD error is the difference between the multi-step TD target value and the current Q value. By calculating the TD error, the effectiveness of the current decision can be evaluated and feedback can be provided for the subsequent learning process.

[0111] Figure 4The uncertainty distribution of this technical solution and the traditional ε-greedy strategy in the state space is compared. The horizontal axis represents the state value (0-5.0), and the vertical axis represents the corresponding uncertainty measure (0-0.8). This technical solution shows the highest uncertainty measure near the state value of 2.5, with a peak value of 0.8, while the uncertainty measure of the traditional ε-greedy strategy at the same position is only 0.5. More importantly, the uncertainty distribution curve of this technical solution presents a narrower bell shape, with the main uncertainty concentrated in the interval of state values ​​1.5-3.5, which shows that the algorithm can more accurately identify the state area that needs to be explored. In contrast, the distribution curve of the traditional ε-greedy strategy is relatively flat, and it maintains a relatively high uncertainty in the range of state values ​​1.0-4.0, indicating that its exploration strategy is relatively blind. In addition, the uncertainty of this technical solution at both ends of the state space (0-1.0 and 4.0-5.0) is rapidly reduced to near 0, indicating that the algorithm can effectively avoid over-exploration of states with high certainty.

[0112] In an optional implementation, based on the optimized Q-value function network, the anomaly score of the time series segment is calculated, and the anomaly score is obtained by comparing the Q-value of the optimal action in the current state with a preset threshold, including:

[0113] Extract the optimal action of the current state from the Q-value function network, obtain the corresponding optimal Q-value based on the optimal action, calculate the mean and standard deviation of the historical optimal Q-values ​​within a preset time window, and use the weighted sum of the mean and standard deviation as the dynamic threshold;

[0114] The difference is calculated based on the average of the optimal Q value and the historical optimal Q value, and the ratio of the difference to the standard deviation is used as the basic anomaly score. The time decay function is introduced to calculate the time series weight, and the time series weight is calculated by the time difference between the current moment and the reference moment. The product of the basic anomaly score and the time series weight is used as the weighted anomaly score;

[0115] The transition probability between adjacent states is calculated, and the negative logarithm of the transition probability is used as the transition anomaly score. The state feature vector is extracted to calculate the Euclidean distance between adjacent states to obtain the state representation distance. The weighted anomaly score, the transition anomaly score, and the state representation distance are weighted and combined by a preset weight coefficient to obtain a comprehensive anomaly score.

[0116] Extract the optimal action for the current state from the Q-value function network. This process involves analyzing the characteristics of the current state and using the trained Q-value function network to identify the action that can achieve the highest Q value in this state. At this point, the Q value of the optimal action is recorded as the basis for subsequent calculations.

[0117] Based on the extracted optimal action, the corresponding optimal Q value is obtained. This Q value reflects the expected return that can be obtained by taking this action in the current state. In order to further analyze the abnormality of the Q value, it is necessary to calculate the mean and standard deviation of the historical optimal Q value within the preset time window. The historical optimal Q value refers to the set of optimal Q values ​​obtained for the same state in the past period of time. The mean provides a benchmark, while the standard deviation reflects the volatility of these Q values.

[0118] After obtaining the mean and standard deviation, the weighted sum of the two is used as the dynamic threshold. The setting of the dynamic threshold enables the calculation of the anomaly score to adapt to changes in the environment and ensures the consistency and accuracy of anomaly detection in different time periods.

[0119] The difference is calculated based on the average of the current optimal Q value and the historical optimal Q value. This difference reflects the deviation between the optimal Q value in the current state and the historical performance. Then, the ratio of this difference to the standard deviation is used as the basic anomaly score. The calculation of the basic anomaly score provides a preliminary quantitative indicator for subsequent anomaly detection.

[0120] The calculation of the timing weight is based on the time difference between the current moment and the reference moment, ensuring that states closer in time have a greater impact on the anomaly score. Finally, the product of the basic anomaly score and the timing weight is used as the weighted anomaly score, which reflects the sensitivity to anomalies in the time dimension.

[0121] The transition probability refers to the possibility of transitioning to the next state in the current state. By taking the negative logarithm of the transition probability, the transition anomaly score is obtained, which reflects the abnormality of the state transition.

[0122] The state feature vector is extracted to calculate the Euclidean distance between adjacent states. The Euclidean distance is used to quantify the similarity between adjacent states. The larger the distance, the more significant the difference between the states. Finally, the weighted anomaly score, the transition anomaly score and the state representation distance are weighted and combined by the preset weight coefficient to obtain the comprehensive anomaly score.

[0123] This application can effectively identify and quantify abnormal behaviors in time series data, and improve the accuracy and reliability of anomaly detection. The introduction of dynamic thresholds enables anomaly detection to adapt to environmental changes, enhancing the flexibility and adaptability of the system. By comprehensively considering weighted anomaly scores, transfer anomaly scores, and state representation distances, a more comprehensive anomaly assessment is provided, which can better reflect the actual operating status of the system.

[0124] In an optional implementation, the difference is calculated based on the average of the optimal Q value and the historical optimal Q value, the ratio of the difference to the standard deviation is used as the basic anomaly score, and the time decay function is introduced to calculate the time series weight, and the time series weight is calculated by the time difference between the current moment and the reference moment, including:

[0125] Obtaining the optimal Q value at the current moment, calculating the difference between the optimal Q value at the current moment and the mean to obtain an abnormal deviation, and dividing the abnormal deviation by the standard deviation to obtain a normalized abnormal score;

[0126] Constructing a time decay function based on a Gaussian kernel function, substituting the time difference between the current time and the reference time into the time decay function, and calculating a timing weight coefficient according to a preset time decay coefficient;

[0127] The normalized anomaly score is multiplied by the time series weight coefficient to obtain a weighted anomaly score, and the weighted anomaly score decays as the time difference increases.

[0128] The optimal Q value is obtained by evaluating all possible actions in the current state, reflecting the expected benefits of taking a specific action in that state. By comparing the optimal Q value at the current moment with the mean calculated from historical data, the abnormal deviation at the current moment can be determined. The abnormal deviation refers to the difference between the optimal Q value at the current moment and the mean, reflecting the degree of abnormality of the current state.

[0129] The standard deviation is a quantitative measure of the fluctuation of the historical optimal Q value, which can reflect the stability of historical data. The abnormal deviation is divided by the standard deviation to obtain the normalized anomaly score. The calculation of the normalized anomaly score allows the degree of anomaly in different time periods to be compared, which is convenient for subsequent analysis.

[0130] The function of the time decay function is to calculate a time series weight coefficient based on the time difference between the current moment and the reference moment. The Gaussian kernel function can effectively reflect the impact of the time difference on the anomaly score. As the time difference increases, the weight coefficient gradually decreases, reflecting the attenuation effect of time on the anomaly score.

[0131] The normalized anomaly score is multiplied by the time series weight coefficient to obtain the weighted anomaly score. The weighted anomaly score takes into account the degree of anomaly and the time factor at the current moment, and can more accurately reflect the anomaly of the current state. As the time difference increases, the weighted anomaly score will gradually decrease, ensuring the importance of the time factor in anomaly detection.

[0132] Based on the above steps, a closed-loop anomaly detection system is formed. By continuously updating the optimal Q value and historical mean at the current moment, the system can monitor anomalies in real time and adjust the anomaly score according to the principle of time decay to ensure the accuracy and timeliness of detection.

[0133] According to a second aspect of the embodiments of the present invention,

[0134] An electronic device is provided, comprising:

[0135] processor;

[0136] a memory for storing processor-executable instructions;

[0137] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0138] According to a third aspect of the embodiments of the present invention,

[0139] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0140] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An embedded software timing anomaly detection method based on value function reinforcement learning, characterized in that: include: Acquire time series data during the operation of the embedded software, and divide the time series data into multiple time series segments according to a preset time window; For each of the time series segments, a state space matrix is ​​constructed, wherein the row vector of the state space matrix represents the time series feature, and the column vector represents the time step; Based on the state space matrix, a long short-term memory network is used to extract a time series feature vector; the time series feature vector is input into a deep neural network to generate a Q value function network, and the Q value function network is used to evaluate the action value function under the time series state; a reward function is constructed, and the reward function is calculated based on the abnormality degree of the time series segment, and the abnormality degree is obtained by measuring the Mahalanobis distance between the current time series segment and the historical normal time series segment; The Q-value function network is trained using a deep Q-learning algorithm. During the training process, an action is selected based on an ε-greedy strategy, and the selected action and the current state are input into the Q-value function network to obtain a Q-value; the Q-value is compared with a target Q-value calculated based on the reward function, and the parameters of the Q-value function network are optimized by minimizing a loss function; Based on the optimized Q-value function network, the anomaly score of the time series segment is calculated, and the anomaly score is obtained by comparing the Q-value of the optimal action in the current state with a preset threshold; When the anomaly score exceeds a preset threshold, it is determined that an anomaly exists in the current time series segment, and an anomaly detection result is output.

2. The method according to claim 1, characterized in that Inputting the time series feature vector into a deep neural network to generate a Q value function network, wherein the Q value function network is used to evaluate the action value function under the time series state, including: Acquire a time series feature vector, input the time series feature vector into a multi-layer perceptron structure, wherein the multi-layer perceptron structure includes a first fully connected layer, a second fully connected layer, and an output layer, wherein the first fully connected layer includes 256 neurons and uses a ReLU activation function, the second fully connected layer includes 128 neurons and uses a ReLU activation function, and the output layer uses a hyperbolic tangent function to limit the output value to a range of [-1, 1], thereby obtaining a depth mapping feature; A Q-value function network of a dual-stream network structure is constructed based on the deep mapping features, wherein the Q-value function network includes a state value function branch and an advantage function branch. The output value of the state value function branch is combined with the output value of the advantage function branch, and the average value of the advantage function in the action space is subtracted to obtain a state-action value function.

3. The method according to claim 1, characterized in that The Q value function network is trained using a deep Q learning algorithm. During the training process, an action is selected based on an ε-greedy strategy, and the selected action and the current state are input into the Q value function network to obtain a Q value; Comparing the Q value with a target Q value calculated based on the reward function, and optimizing the parameters of the Q value function network by minimizing a loss function includes: Construct an action selection probability distribution model, determine the optimal action according to the state-action value output by the Q-value function network, assign a selection probability of 1-ε+ε / |A| to the optimal action, and assign a selection probability of ε / |A| to other actions, where |A| is the size of the action space; update the exploration rate ε based on the training rounds, and the initial value of the exploration rate is the product of the difference between the maximum exploration rate and the minimum exploration rate and the decay exponent; Calculate state uncertainty, obtain multiple Q value samples corresponding to the state-action, calculate the variance of the multiple Q value samples and their mean, use the variance as a measure of state uncertainty, and use the product of the exploration rate and the state uncertainty as an adaptive exploration rate; select an action based on the adaptive exploration rate; Construct a multi-step temporal difference target based on the selected action, calculate the discounted accumulation of the immediate reward and the maximum Q value of the terminal state to obtain the n-step cumulative reward, perform a weighted combination of the n-step cumulative rewards to obtain the multi-step temporal difference target value, and take the difference between the multi-step temporal difference target value and the current Q value as the temporal difference error; A priority experience replay mechanism is established according to the temporal difference error, the sum of the absolute value of the temporal difference error and the preset priority bias is used as the sample priority, the α-power of the sample priority is divided by the sum of the α-powers of all sample priorities to obtain a sampling probability distribution, samples in the experience replay pool are sampled based on the sampling probability distribution, and the β-power of the inverse of the product of the sampling probability and the size of the experience pool is used as the importance weight; Construct a parameter optimization target. When the absolute value of the time series difference error is less than a threshold parameter, the square error is used. When the absolute value of the time series difference error is greater than the threshold parameter, the linear error minus the fixed bias is used to obtain the Huber loss. The product of the Huber loss and the importance weight is used as the final optimization target. Based on the optimization target, the parameters of the Q-value function network are gradient updated.

4. The method according to claim 3, characterized in that Based on the selected action, a multi-step TD target is constructed, the discounted accumulation of the immediate reward and the maximum Q value of the terminal state are calculated to obtain the n-step cumulative reward, the n-step cumulative reward is weightedly combined to obtain the multi-step TD target value, and the difference between the multi-step TD target value and the current Q value is taken as the TD error, including: The formula for calculating the cumulative return in n steps is as follows: in, The cumulative return of n steps starting at time t; γ: discount factor, range [0,1], used to balance immediate rewards and future rewards; r t+k : The instant reward obtained at time t+k; Q(s t+n , a): Q value of all actions in the state at time t+n; s t+n : state at time t+n; a represents action; n: cumulative number of steps; The formula for calculating the multi-step timing difference target value is as follows: y t : Multi-step time difference target value at time t; λ: weight parameter used to control the rewards of different step lengths; N: maximum number of steps considered; The timing differential error calculation formula is as follows: δ t =y t -Q(s t ,a t ); δ t : The timing difference error at time t.

5. The method according to claim 1, characterized in that Based on the optimized Q-value function network, the anomaly score of the time series segment is calculated. The anomaly score is obtained by comparing the Q-value of the optimal action in the current state with a preset threshold, including: Extract the optimal action of the current state from the Q-value function network, obtain the corresponding optimal Q-value based on the optimal action, calculate the mean and standard deviation of the historical optimal Q-values ​​within a preset time window, and use the weighted sum of the mean and standard deviation as the dynamic threshold; The difference is calculated based on the average of the optimal Q value and the historical optimal Q value, and the ratio of the difference to the standard deviation is used as the basic anomaly score. The time decay function is introduced to calculate the time series weight, and the time series weight is calculated by the time difference between the current moment and the reference moment. The product of the basic anomaly score and the time series weight is used as the weighted anomaly score; The transition probability between adjacent states is calculated, and the negative logarithm of the transition probability is used as the transition anomaly score. The state feature vector is extracted to calculate the Euclidean distance between adjacent states to obtain the state representation distance. The weighted anomaly score, the transition anomaly score, and the state representation distance are weighted and combined by a preset weight coefficient to obtain a comprehensive anomaly score.

6. The method according to claim 5, characterized in that The difference is calculated based on the average of the optimal Q value and the historical optimal Q value, and the ratio of the difference to the standard deviation is used as the basic anomaly score. The time decay function is introduced to calculate the time series weight, and the time series weight is calculated by the time difference between the current moment and the reference moment. It includes: Obtaining the optimal Q value at the current moment, calculating the difference between the optimal Q value at the current moment and the mean to obtain an abnormal deviation, and dividing the abnormal deviation by the standard deviation to obtain a normalized abnormal score; Constructing a time decay function based on a Gaussian kernel function, substituting the time difference between the current time and the reference time into the time decay function, and calculating a timing weight coefficient according to a preset time decay coefficient; The normalized anomaly score is multiplied by the time series weight coefficient to obtain a weighted anomaly score, and the weighted anomaly score decays as the time difference increases.

7. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.

8. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Man-machine-object cooperation abnormal state detection method for strengthening heterogeneous graph neural network

    CN115081585A

  • Deep Q learning bearing fault diagnosis method based on Bayesian optimization

    CN117171508A