Message processing method, device, equipment, medium and product
By combining long short-term memory networks and deep reinforcement learning models, a message detection and handling decision-making system is constructed, which solves the accuracy and adaptability problems of traditional message anomaly detection methods in complex network scenarios and realizes efficient abnormal message detection and handling.
Patent Information
- Application Number
- CN202511003469.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-09-12
AI Technical Summary
Traditional message anomaly detection methods are difficult to deal with complex attacks, lack intelligence, are difficult to adapt to complex network scenarios, and have insufficient detection accuracy.
By combining the long short-term memory network model and the deep reinforcement learning model, message features and handling strategies are acquired through iterative training to build a message detection and handling decision-making system.
It improves the accuracy and adaptability of message detection, reduces false positives and missed positives, and enhances network security protection capabilities.
Smart Images

Figure CN120639474A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of message detection technology, and in particular to a message processing method, apparatus, device, medium, and product. Background Art
[0002] Packet anomaly detection analyzes network traffic to identify unusual packets that do not conform to normal behavior patterns. These unusual packets may indicate a network attack, a network failure, or a misconfiguration.
[0003] Traditional packet anomaly detection methods include rule-based and statistical approaches. These use predefined rules or statistically analyze network traffic characteristics to detect anomalies. These methods are effective for detecting packet anomalies with known attack signatures. However, these traditional methods also have limitations, including difficulty responding to complex attacks, a lack of intelligence, and difficulty adapting to complex network scenarios.
[0004] Based on this, in order to address the above problems, it is necessary to design a message anomaly detection method and disposal decision-making plan. Summary of the Invention
[0005] The embodiments of the present invention provide a message processing method, apparatus, device, medium and product to improve the accuracy and adaptability of message detection and provide protection for network security.
[0006] According to one aspect of the present invention, a message processing method is provided, comprising:
[0007] Obtaining a message to be detected and inputting the message to be detected into a first model to obtain target features and detection results; the first model is obtained by iteratively training a long short-term memory network model with a first sample set, the first sample set including message samples and annotation results corresponding to the message samples;
[0008] The target features are input into a second model to obtain a target disposal strategy; the second model is obtained by iteratively training a deep reinforcement learning model with a second sample set, and the second sample set includes feature samples and action space samples.
[0009] According to another aspect of the present invention, a message processing device is provided, the device comprising:
[0010] A first input module is configured to obtain a message to be detected and input the message to be detected into a first model to obtain target features and detection results; the first model is obtained by iteratively training a long short-term memory network model with a first sample set, the first sample set including message samples and annotation results corresponding to the message samples;
[0011] The second input module is used to input the target features into a second model to obtain a target disposal strategy; the second model is obtained by iteratively training a deep reinforcement learning model with a second sample set, and the second sample set includes feature samples and action space samples.
[0012] According to another aspect of the present invention, an electronic device is provided, comprising:
[0013] at least one processor; and
[0014] a memory communicatively connected to the at least one processor; wherein,
[0015] The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can execute the message processing method described in any embodiment of the present invention.
[0016] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the message processing method according to any embodiment of the present invention when executed.
[0017] According to another aspect of the present invention, an embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the message processing method described in any embodiment of the present invention.
[0018] The embodiment of the present invention uses message samples and the annotation results corresponding to the message samples as the first sample set, iteratively trains a long short-term memory network model through the first sample set to obtain a first model, obtains the message to be detected, and inputs the message to be detected into the first model to obtain the target features and detection results; then uses the feature samples and action space samples as the second sample set, iteratively trains a deep reinforcement learning model through the second sample set to obtain a second model, inputs the target features into the second model, and obtains the target handling strategy. The technical solution of the present invention, by constructing a long short-term memory network model to extract features in the message and determine whether it is an abnormal message, and then by constructing a deep reinforcement learning model to learn the optimal handling strategy for different features, can more accurately detect abnormal messages and quickly obtain a handling strategy, reduce false positives and false negatives, and improve detection accuracy.
[0019] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 is a flow chart of a message processing method in an embodiment of the present invention;
[0022] Figure 2 This is a flow chart of the operation of a strategy learning unit in an embodiment of the present invention;
[0023] Figure 3 It is a structural diagram of a message processing device in an embodiment of the present invention;
[0024] Figure 4 It is a structural diagram of an electronic device for implementing the message processing method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0025] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0026] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and their inclusion are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0027] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0028] Example 1
[0029] Figure 1 This is a flow chart of a message processing method in an embodiment of the present invention. This embodiment is applicable to the case of message processing. The method can be executed by the message processing device in an embodiment of the present invention. The device can be implemented in software and / or hardware. Figure 1 As shown, the method specifically includes the following steps:
[0030] S101: Obtain a message to be detected, and input the message to be detected into a first model to obtain target features and detection results.
[0031] It should be noted that the message to be detected may be a network message to be detected for anomalies.
[0032] In this embodiment, the first model may be a model for performing anomaly detection on the message data to be detected and outputting message features and detection results. Preferably, the first model may be, for example, a trained long short-term memory network model.
[0033] The first model is obtained by iteratively training a long short-term memory network model with a first sample set, and the first sample set includes message samples and labeling results corresponding to the message samples.
[0034] It should be noted that the message sample may be a network message used as a training set to train a long short-term memory network model, and the labeling result corresponding to the message sample may be labeling information indicating whether the message sample is abnormal.
[0035] The target feature may be a message feature corresponding to the message to be detected, such as a network traffic feature. The detection result may be a detection result of whether the message to be detected is abnormal. For example, the detection result may be that the message to be detected is normal or abnormal.
[0036] Specifically, the message samples and their corresponding annotation results are used as a first sample set, and a long short-term memory network model is iteratively trained using the first sample set to obtain a first model. A message to be detected is obtained and input into the trained long short-term memory network model, i.e., the first model. The message to be detected is detected for anomalies and the message features and detection results are output.
[0037] S102: Input the target features into the second model to obtain a target disposal strategy.
[0038] In this embodiment, the second model may be a model for selecting an optimal handling strategy based on the message features of the message to be detected and outputting a message processing method. Preferably, the second model may be, for example, a trained deep reinforcement learning model.
[0039] The second model is obtained by iteratively training a deep reinforcement learning model with a second sample set, and the second sample set includes feature samples and action space samples.
[0040] It should be noted that the feature samples may be message features of network messages used as a training set to train a deep reinforcement learning model. Preferably, the feature samples in the second sample set may be message features of message samples in the first sample set.
[0041] In this embodiment, the action space sample may be an action space sample constructed based on marking the message sample as normal and released (0), abnormal and intercepted (1), and observed (2).
[0042] The target handling strategy may be a processing strategy for the message to be detected. For example, the target handling strategy may be release, interception, or observation.
[0043] Specifically, the feature samples and action space samples are used as the second sample set, the deep reinforcement learning model is iteratively trained through the second sample set to obtain the second model, the target features corresponding to the message to be detected are input into the second model, and the target handling strategy corresponding to the message to be detected is obtained.
[0044] The embodiment of the present invention uses message samples and the annotation results corresponding to the message samples as the first sample set, iteratively trains a long short-term memory network model through the first sample set to obtain a first model, obtains the message to be detected, and inputs the message to be detected into the first model to obtain the target features and detection results; then uses the feature samples and action space samples as the second sample set, iteratively trains a deep reinforcement learning model through the second sample set to obtain a second model, inputs the target features into the second model, and obtains the target handling strategy. The technical solution of the present invention, by constructing a long short-term memory network model to extract features in the message and determine whether it is an abnormal message, and then by constructing a deep reinforcement learning model to learn the optimal handling strategy for different features, can more accurately detect abnormal messages and quickly obtain a handling strategy, reduce false positives and false negatives, and improve detection accuracy.
[0045] Optionally, iteratively training the long short-term memory network model using the first sample set includes:
[0046] Build a long short-term memory network model.
[0047] LSTM (Long Short-Term Memory) is a special type of RNN (Recurrent Neural Network) designed to address the long-term dependency issues of traditional RNNs. By introducing gating mechanisms (input gate, forget gate, and output gate) to selectively remember or forget information, it excels at processing long-term patterns in time series data.
[0048] Specifically, this embodiment selects the long short-term memory network (LSTM) in the deep learning model to detect the message data and output the message features and detection results.
[0049] In the specific implementation process, the constructed deep learning LSTM model includes an LSTM layer, a Dropout layer, and an output layer. The Dropout layer is a regularization technology that aims to reduce the risk of overfitting of the model by randomly "closing" some neurons in the neural network.
[0050] The message samples in the first sample set are input into the long short-term memory network model to obtain the prediction results.
[0051] It should be noted that the prediction result may be the result outputted by the long short-term memory network model after detecting the message sample data. For example, the prediction result may be that the message sample is normal or abnormal.
[0052] Specifically, the message samples in the first sample set are input into the long short-term memory network model, anomaly detection is performed on the message samples, and message features and prediction results of the message samples are output.
[0053] The parameters of the long short-term memory network model are trained based on the objective function formed by the prediction results and the labeling results.
[0054] For example, the objective function may be a loss function. Specifically, the weights and other parameters of the long short-term memory network model are trained based on the loss function formed by the prediction results and the labeling results.
[0055] Return to executing the operation of inputting the message samples in the first sample set into the long short-term memory network model to obtain the prediction results, until the first model is obtained.
[0056] Specifically, the process returns to executing the operation of inputting the message samples in the first sample set into the long short-term memory network model to obtain prediction results until a preset convergence condition is reached or the number of training iterations reaches a preset number of iterations, thereby obtaining the first model. The preset convergence condition and the preset number of iterations can both be pre-set by the user based on actual circumstances and are not specifically limited in this embodiment.
[0057] Optionally, inputting the message samples in the first sample set into the long short-term memory network model to obtain a prediction result includes:
[0058] Extract characteristic information of the message samples in the first sample set.
[0059] The characteristic information may be a message characteristic corresponding to a message sample.
[0060] Specifically, the message samples in the first sample set are input into the long short-term memory network model, and feature information of the message samples is output.
[0061] After preprocessing the feature information, it is serialized and converted to obtain sequence format data.
[0062] In this embodiment, preprocessing can be a process of performing data encoding and data standardization processing on the message samples in the first sample set after message feature extraction to generate standard data that can be used for deep learning training.
[0063] It should be noted that serialization can be the process of converting a data structure or object state into a format that can be stored or transmitted. A time series format is a standardized structure for storing and representing data points arranged in chronological order. In this embodiment, the sequence format data can be message data in a time series format.
[0064] Specifically, first start the feature extractor, extract basic features such as source IP, destination IP, protocol type, message content, and statistical features such as message rate and traffic rate from the message sample, encode the data, convert non-numerical features such as IP address and protocol type into data form, and use the Z-score standardization method (a statistical method used to measure the difference between a data point and the mean of a data set, whose core purpose is to standardize data of different scales) to normalize the numerical features. After data standardization, normal and abnormal messages are labeled, 0 for normal and 1 for abnormal, and the data set is divided into training set, validation set, and test set. The feature information of the message sample is then serialized to obtain sequence format data.
[0065] The sequence format data is input into the long short-term memory network model to obtain the prediction results.
[0066] Specifically, a deep learning LSTM model is constructed, consisting of an LSTM layer, a dropout layer, and an output layer. Training data is input to train the model. After the model is trained, new message data is passed in for message detection, and the message features and detection results are output.
[0067] Optionally, iteratively training the deep reinforcement learning model using the second sample set includes:
[0068] State space samples are constructed according to the feature samples in the second sample set.
[0069] Specifically, a state space sample can be constructed based on the feature samples corresponding to the message samples output by the long short-term memory network model to train the deep reinforcement learning model.
[0070] Build a deep reinforcement learning model.
[0071] Machine Learning (ML) is a core branch of artificial intelligence. It refers to the process by which computer systems automatically learn patterns and patterns from data and make predictions or decisions based on these patterns without explicit programming. It is primarily categorized into supervised learning, unsupervised learning, reinforcement learning, semi-supervised learning, and self-supervised learning.
[0072] Deep learning (DL) is a subfield of machine learning that automatically learns hierarchical feature representations of data using multi-layer nonlinear neural networks. Its core advantage lies in its ability to extract high-level, abstract features directly from raw data (such as images, text, and time series signals) without relying on manual feature engineering.
[0073] RL (Reinforcement Learning) is a branch of machine learning that studies how intelligent agents learn optimal strategies through interaction with their environments to maximize long-term cumulative rewards. Its core is trial-and-error learning, which does not require pre-labeled data and is particularly well-suited for sequential decision-making problems.
[0074] DRL (Deep Reinforcement Learning) is a combination of reinforcement learning (RL) and deep learning (DL). It uses neural networks to approximate the value function or policy function in reinforcement learning, enabling intelligent agents to learn optimal decision-making strategies through trial and error in complex, high-dimensional environments.
[0075] In this embodiment, the constructed deep reinforcement learning model is trained using a deep reinforcement learning algorithm (Deep Q-Network). Deep Q-Network (DQN): Deep Q-learning is a combination of the Q-learning algorithm and DNN (Deep Neural Network). It approximates the action-value function (Q function) in Q-learning through a neural network, solving the dimensionality curse problem of traditional Q-learning in high-dimensional state spaces (such as financial time series data and images). Its core innovation is the introduction of experience replay and target network to stabilize the training process.
[0076] The deep reinforcement learning model is iteratively trained according to the state-action space, action space samples, and reward function to obtain the second model.
[0077] In actual operation, a state space sample is constructed based on the feature sample, and the model is built based on the Q-Learning model, which includes three parameters: state space sample, action space sample, and reward function. The state space sample is the network traffic feature extracted by LSTM, and the action space sample is the marked message sample as normal and released (0), abnormal and intercepted (1), and observed (2). The reward function rewards or punishes several situations. The specific reward scores and scoring strategies are shown in Table 1.
[0078] The parameters of the reward evaluation function are automatically adjusted based on the accuracy, false alarm rate, and missed alarm rate, so that the model can better cope with complex network environments.
[0079] The accuracy rate refers to the proportion of correctly detected packets (normal and abnormal) to the total sample size, reflecting the model's overall judgment ability. The false positive rate refers to the proportion of normal packets that are mistakenly marked as abnormal by the system. The false negative rate refers to the proportion of abnormal packets that are mistakenly allowed by the system.
[0080] A reward mechanism for discovering new attack patterns is added to the reward function. This curiosity-driven reward encourages the model to proactively explore unknown areas, increasing its adaptability. Curiosity-driven learning is an intrinsic incentive mechanism in reinforcement learning. By instilling a desire to explore unknown states or environments with high prediction errors, it addresses the exploration efficiency issue under sparse rewards. Adaptability refers to the model's ability to dynamically adjust its behavior to respond to environmental changes.
[0081] Table 1
[0082] Condition Reward Points Score strategy Correct detection +5 When the accuracy rate is <80%, increase the correct reward False positives -3 When the false alarm rate is >15%, increase the false alarm penalty Underreporting -10 When the false negative rate is >10%, the false negative penalty will be increased. Discovering new attack patterns +0.5
[0083] Optionally, inputting the target features into the second model to obtain a target treatment strategy includes:
[0084] The target features are input into the second model to obtain the target action space.
[0085] Among them, the target action space can be the corresponding action space output by the trained deep reinforcement learning model according to the input target features.
[0086] Specifically, the target feature is input into the second model, and the target action space is output. For example, the target action space can be normal and release (0), abnormal and intercept (1), and observe (2).
[0087] Determine the target disposal strategy based on the target action space.
[0088] Specifically, a target processing strategy for the message to be detected is determined according to the target action space. For example, the target processing strategy can be release, interception, or observation.
[0089] Optionally, the feature information includes: basic features and statistical features; basic features include: source IP, destination IP, port number, protocol type, message length, timestamp, and message content; statistical features include: message rate, traffic rate, message size distribution, and message interval.
[0090] To improve the accuracy of packet anomaly detection and its adaptability to network environments, the present invention proposes a packet anomaly detection and handling decision-making system based on a combination of deep learning and deep reinforcement learning. This system can be used to accurately detect abnormal packets in network transmission and adapt to network complexity.
[0091] The system includes a data processing unit, a message detection unit, a reinforcement learning modeling unit, and a strategy learning unit. The data processing unit is used to extract message features, encode data, and standardize data on the message training set to generate standard data that can be used for deep learning training. The message detection unit converts the message data into a time series format, selects the long short-term memory network (LSTM) in the deep learning model to detect the message data, and outputs the message features and detection results. The reinforcement learning modeling unit uses the output of the message detection unit as the state, and "marking the message as normal and releasing (0), abnormal and intercepting (1), or observing (2)" as the action, and sets a reasonable reward function. The strategy learning unit uses a deep reinforcement learning algorithm (Deep Q-Network) to learn the optimal strategy, reduce the false alarm and missed alarm rate, and maximize the detection accuracy. After a new message is input, the system extracts features and makes strategy decisions on the new message data, marks the message as normal or abnormal according to the learned strategy, and decides on the message handling plan.
[0092] The data processing unit consists of a message feature extractor, a data encoder, and a data normalizer. The message feature extractor extracts basic and statistical features from messages; the data encoder converts non-numeric features such as IP addresses and protocol types into data; and the data normalizer normalizes numerical features. Specifically, the workflow of the data processing unit can be described as follows: extracting basic features of the message, such as the source IP, destination IP, port number, protocol type, message length, timestamp, message content, as well as statistical features such as message rate, traffic rate, message size distribution, and message interval time; converting non-numerical features such as IP address and protocol type into numerical form; performing one-hot encoding on the categorical feature of protocol type (one-hot encoding is a coding method that converts categorical variables into binary vectors to ensure that the machine learning model can correctly handle non-numerical features): TCP->0, UDP->1, ICMP->2; mapping the IP to a low-dimensional continuous vector space, compressing high-dimensional sparse data (high-dimensional sparse data is a common type of complex data in machine learning and data analysis, characterized by extremely high feature dimensions, but the vast majority of features in each sample are zero) into a dense vector (dense vector is a key data representation in machine learning and data analysis, characterized by low dimensionality and each dimension containing meaningful values), which can facilitate the model to capture semantics and adapt to the model input; using the Z-score normalization method to normalize data features such as message length.
[0093] The message detection unit primarily consists of four parts: message data serialization, long-term and short-term neural network training (including LSTM model construction and LSTM model training), and message anomaly detection. Message data serialization processes message data into a data type acceptable to the model. For LSTM models, message data must be time-seriesized. LSTM model construction constructs the LSTM basic model and sets parameters. LSTM model training uses serialized data to train the model, enabling it to support message anomaly detection. Message anomaly detection performs anomaly detection on newly input message data and outputs message features and detection results. The workflow can be described as follows: message data is serialized and converted into a sequence format with an output shape of [number of samples, time steps, feature dimensions]; an LSTM model is constructed, consisting of an LSTM layer, a dropout layer, and an output layer; the LSTM model is trained using the processed message data; and the trained LSTM model performs anomaly detection on new messages, outputting message features and detection results.
[0094] Among them, the reinforcement learning modeling unit mainly includes three parts: defining the state space, defining the action space, and constructing the reward evaluation function. The state space is constructed based on the message features output by the message detection unit, and the model is based on the Q-Learning model, which includes three parameters: state space, action space, and reward function. The state space is the network traffic features extracted by LSTM, and the action space is to mark the message as normal and release (0), abnormal and intercept (1), and observe (2). The reward function rewards or punishes several situations. The specific reward scores and score strategies are shown in Table 1. The parameters of the reward evaluation function are automatically adjusted according to the three parameters of accuracy, false alarm rate, and missed alarm rate, so that the model can better cope with complex network environments. A reward mechanism for discovering new attack patterns is added to the reward function. The curiosity-driven reward enables the model to actively explore unknown areas and increase the adaptability of the model.
[0095] Figure 2 This is a flow chart of the operation of a strategy learning unit in an embodiment of the present invention. The strategy learning unit uses the deep reinforcement learning algorithm Deep Q-Network (DQN) to learn the optimal strategy based on the state space, action space, and reward function constructed in the previous step, and combines the experience pool (experience replay) with the target network to disrupt the correlation. Figure 2 As shown, the workflow of the policy learning unit can be described as:
[0096] Training process:
[0097] 1. Initialize the Q network and Target network;
[0098] 2. For each episode:
[0099] - Initialize state(s) based on current message characteristics;
[0100] -For each step:
[0101] -Select an action (a) based on the current state (s);
[0102] - Perform an action (a), observe the reward (r) and the next state (s');
[0103] -Store the experience (state s, action a, reward r, next state s') in the experience replay buffer;
[0104] - Randomly sample a batch of experiences from the experience replay buffer and update the Q network;
[0105] -Update the Target network according to the loss function;
[0106] 3. Until the end of the round and traverse all rounds.
[0107] The trained model can accurately determine whether a message is abnormal based on the message status. Since a self-adjustment mechanism has been added to the reward function, it can better adapt to complex network environments and improve the accuracy of the model.
[0108] Specifically, the steps of implementing the message processing method of this embodiment based on the above system can be described as follows:
[0109] The data processing unit first activates the feature extractor, extracting basic features from the message, such as the source IP address, destination IP address, protocol type, and message content, as well as statistical features such as message rate and traffic rate. The unit then encodes the data, converting non-numeric features like IP address and protocol type into numerical form. Numerical features are then normalized using the Z-score method. After data normalization, normal and abnormal messages are labeled, with 0 representing normal and 1 representing abnormal. The dataset is then divided into training, validation, and test sets.
[0110] The packet detection unit first serializes the packet features and then builds a deep learning LSTM model, consisting of an LSTM layer, a dropout layer, and an output layer. Training data is then fed into the model to train the model. After the model is trained, new packet data is passed in for packet detection, and the packet features and detection results are output.
[0111] The reinforcement learning modeling unit first constructs a state space based on the output of the message detection unit, and then constructs an action space based on marking the message as normal and released (0), abnormal and intercepted (1), and observed (2). A reward evaluation function is constructed to give rewards or penalties for correct detection and incorrect detection.
[0112] The policy learning unit uses a deep reinforcement learning algorithm (Deep Q-Network) to train the reinforcement learning model built in the previous step. After the model is trained, it is fed with message features, selects the optimal policy, and outputs the message processing method.
[0113] The technical solution of the embodiment of the present invention automatically extracts message features through LSTM deep learning, without relying heavily on manually designed rules. It can extract features of higher dimensions, making the message features more comprehensive, providing a reliable data source for deep reinforcement learning, and improving the credibility of model training. In addition, this embodiment reduces the false alarm rate through a deep reinforcement learning model, performs strategy learning through the DQN algorithm, reduces false alarms and missed alarms, and improves detection accuracy; at the same time, reinforcement learning can also dynamically adjust the reward function strategy according to changes in the network environment, thereby improving the model's adaptability to complex network environments. In addition, this embodiment also uses curiosity-driven rewards to enable the model to explore unknown states. Through the exploration feedback in the reward function, the model can more actively detect new attack patterns, helping the model to learn the optimal strategy more quickly.
[0114] Example 2
[0115] Figure 3 This is a schematic diagram of the structure of a message processing device in an embodiment of the present invention. This embodiment is applicable to the case of message processing. The device can be implemented in software and / or hardware. The device can be integrated into any device that provides message processing functions, such as Figure 3 As shown, the message processing device specifically includes: a first input module 201 and a second input module 202.
[0116] The first input module 201 is configured to obtain a message to be detected and input the message to be detected into a first model to obtain target features and detection results; the first model is obtained by iteratively training a long short-term memory network model with a first sample set, wherein the first sample set includes message samples and annotation results corresponding to the message samples;
[0117] The second input module 202 is used to input the target features into a second model to obtain a target disposal strategy; the second model is obtained by iteratively training a deep reinforcement learning model with a second sample set, and the second sample set includes feature samples and action space samples.
[0118] Optionally, the first input module 201 includes:
[0119] The first building unit is used to build a long short-term memory network model;
[0120] A first input unit, configured to input message samples in the first sample set into the long short-term memory network model to obtain a prediction result;
[0121] A first training unit is configured to train parameters of the long short-term memory network model according to an objective function formed by the prediction result and the labeling result;
[0122] The first execution unit is used to return to executing the operation of inputting the message samples in the first sample set into the long short-term memory network model to obtain a prediction result until the first model is obtained.
[0123] Optionally, the input unit is specifically used to:
[0124] Extracting feature information of message samples in the first sample set;
[0125] The feature information is pre-processed and serialized to obtain sequence format data;
[0126] The sequence format data is input into the long short-term memory network model to obtain a prediction result.
[0127] Optionally, the second input module 202 includes:
[0128] A construction unit, configured to construct a state space sample according to the feature samples in the second sample set;
[0129] The second building unit is used to build a deep reinforcement learning model;
[0130] The second training unit iteratively trains the deep reinforcement learning model according to the state-action space, the action space samples, and the reward function to obtain a second model.
[0131] Optionally, the second input module 202 is specifically configured to:
[0132] Inputting the target features into a second model to obtain a target action space;
[0133] A target handling strategy is determined according to the target action space.
[0134] Optionally, the feature information includes: basic features and statistical features; the basic features include: source IP, destination IP, port number, protocol type, message length, timestamp, message content; the statistical features include: message rate, traffic rate, message size distribution, message interval time.
[0135] The above-mentioned product can execute the message processing method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0136] Example 3
[0137] Figure 4A schematic diagram of the structure of an electronic device 30 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0138] like Figure 4 As shown, the electronic device 30 includes at least one processor 31 and a memory, such as a read-only memory (ROM) 32, a random access memory (RAM) 33, etc., which is communicatively connected to the at least one processor 31. The memory stores a computer program that can be executed by the at least one processor. The processor 31 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 32 or the computer program loaded from the storage unit 38 into the random access memory (RAM) 33. Various programs and data required for the operation of the electronic device 30 can also be stored in the RAM 33. The processor 31, ROM 32, and RAM 33 are connected to each other via a bus 34. An input / output (I / O) interface 35 is also connected to the bus 34.
[0139] Multiple components in the electronic device 30 are connected to the I / O interface 35, including an input unit 36, such as a keyboard, a mouse, etc.; an output unit 37, such as various types of displays, speakers, etc.; a storage unit 38, such as a magnetic disk, an optical disk, etc.; and a communication unit 39, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 39 allows the electronic device 30 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0140] The processor 31 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 31 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 31 executes the various methods and processes described above, such as the message processing method:
[0141] Obtaining a message to be detected and inputting the message to be detected into a first model to obtain target features and detection results; the first model is obtained by iteratively training a long short-term memory network model with a first sample set, the first sample set including message samples and annotation results corresponding to the message samples;
[0142] The target features are input into a second model to obtain a target disposal strategy; the second model is obtained by iteratively training a deep reinforcement learning model with a second sample set, and the second sample set includes feature samples and action space samples.
[0143] In some embodiments, the message processing method can be implemented as a computer program that is tangibly contained in a computer-readable storage medium, such as the storage unit 38. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 30 via the ROM 32 and / or the communication unit 39. When the computer program is loaded into the RAM 33 and executed by the processor 31, one or more steps of the message processing method described above can be performed. Alternatively, in other embodiments, the processor 31 can be configured to perform the message processing method by any other appropriate means (e.g., by means of firmware).
[0144] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0145] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0146] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0147] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0148] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0149] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0150] In one embodiment, the present invention further includes a computer program product. The computer program product includes a computer program. When the computer program is executed by a processor, the computer program implements the message processing method of any embodiment of the present invention.
[0151] The computer program product may be implemented by writing computer program code for performing the operations of the present invention in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0152] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0153] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A message processing method, characterized in that: include: Obtaining a message to be detected, and inputting the message to be detected into a first model to obtain target features and detection results; The first model is obtained by iteratively training a long short-term memory network model with a first sample set, wherein the first sample set includes message samples and labeling results corresponding to the message samples; Inputting the target features into a second model to obtain a target treatment strategy; The second model is obtained by iteratively training a deep reinforcement learning model with a second sample set, where the second sample set includes feature samples and action space samples.
2. The method according to claim 1, characterized in that Iteratively training the long short-term memory network model through the first sample set includes: Build a long short-term memory network model; Inputting the message samples in the first sample set into the long short-term memory network model to obtain a prediction result; Training parameters of the long short-term memory network model according to an objective function formed by the prediction result and the labeling result; Return to executing the operation of inputting the message samples in the first sample set into the long short-term memory network model to obtain prediction results, until the first model is obtained.
3. The method according to claim 2, characterized in that Inputting the message samples in the first sample set into the long short-term memory network model to obtain a prediction result, including: Extracting feature information of message samples in the first sample set; The feature information is pre-processed and serialized to obtain sequence format data; The sequence format data is input into the long short-term memory network model to obtain a prediction result.
4. The method according to claim 1, wherein Iteratively training the deep reinforcement learning model through the second sample set includes: constructing a state space sample according to the feature samples in the second sample set; Building deep reinforcement learning models; The deep reinforcement learning model is iteratively trained according to the state-action space, the action space samples, and the reward function to obtain a second model.
5. The method according to claim 1, wherein Inputting the target features into the second model to obtain a target treatment strategy includes: Inputting the target features into a second model to obtain a target action space; A target handling strategy is determined according to the target action space.
6. The method according to claim 3, characterized in that The characteristic information includes: basic characteristics and statistical characteristics; the basic characteristics include: source IP, destination IP, port number, protocol type, message length, timestamp, message content; the statistical characteristics include: message rate, traffic rate, message size distribution, and message interval time.
7. A message processing device, characterized in that: include: A first input module is used to obtain a message to be detected and input the message to be detected into a first model to obtain target features and detection results; The first model is obtained by iteratively training a long short-term memory network model with a first sample set, wherein the first sample set includes message samples and labeling results corresponding to the message samples; A second input module, configured to input the target features into a second model to obtain a target disposal strategy; The second model is obtained by iteratively training a deep reinforcement learning model with a second sample set, where the second sample set includes feature samples and action space samples.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the message processing method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the message processing method according to any one of claims 1 to 6 when executed.
10. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the message processing method according to any one of claims 1 to 6.