Evidence obtaining method and device based on behavior recognition model and electronic equipment

By collecting and preprocessing data from various servers, constructing a behavior recognition model and training it using a reward function, the problem of low accuracy in identifying lightweight attacks was solved, achieving efficient identification and reliable forensics of lightweight attack behaviors.

CN120915487APending Publication Date: 2025-11-07BEIJING UNITED TRUST TECH SERVICE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510928906.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing intrusion detection systems have low accuracy in identifying lightweight attacks, struggle to adapt to evolving attack methods, and traditional methods are insufficient to meet dynamic defense requirements, thus affecting the integrity and effectiveness of the evidence collection chain.

Method used

By collecting system performance data, log data, and network connection feature data, numerical feature vectors are extracted after preprocessing, a behavior recognition model is constructed, and the model is trained using a reward function to improve the accuracy of identifying light attack behaviors. The attack process is recorded and timestamped and stored in an immutable storage medium.

Benefits of technology

It significantly improves the accuracy of identifying minor attacks, enhances the ability to detect abnormal behavior, ensures the integrity and credibility of forensic records, and supports real-time defense and subsequent forensics of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120915487A_ABST
    Figure CN120915487A_ABST
Patent Text Reader

Abstract

The invention relates to an evidence obtaining method and device based on a behavior recognition model and electronic equipment, and the method comprises the steps: collecting to-be-recognized data in a target server, the to-be-recognized data comprising one or more of system performance data, log data and network connection feature data; the method comprises the following steps: preprocessing various to-be-identified data, respectively extracting numeric feature vectors, and splicing a plurality of feature vectors to obtain a state vector; inputting the state vector into a trained behavior recognition model for recognition, wherein the behavior recognition model is obtained based on reward function training so as to improve the recognition accuracy of the light attack behavior; when the behavior recognition model recognizes that attack behaviors exist, alarm information is sent out, the attack process is recorded, and the attack behaviors include the light attack behavior and the strong attack behavior. According to the method, the attack behavior identification efficiency and accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of forensics, and in particular to a forensics method and device based on a behavior recognition model, an electronic device, and a computer program product. BACKGROUND

[0002] With the increasing complexity of information system architecture, the network attack means faced by servers are becoming diversified and concealed. In order to ensure the safe operation of server systems, server intrusion detection systems have been widely deployed in recent years, aiming to identify abnormal behavior and take response measures in time when attacks occur. The attack identification of existing intrusion detection systems mainly adopts the following two ways: First, the detection method based on rule matching. This method compares the network or system behavior with known attack features by predefining an attack feature library or a strategy template to identify attack behavior. This method highly depends on manual maintenance of the rule library and is difficult to adapt to the trend of continuous evolution of attack methods. It has weak recognition ability for new, variant or combined attacks and is difficult to meet the dynamic defense needs in actual environment.

[0003] Second, the simple machine learning identification method based on logs. This method trains an attack recognition model by collecting historical log data of servers, usually using a shallow model structure. However, due to the limited feature selection and insufficient model expression ability, the recognition accuracy is not high, the generalization ability is weak, false positives or false negatives are easy to occur, and it is difficult to adapt to cross-environment deployment requirements.

[0004] Especially in the face of lightweight attack behavior, such as low-rate penetration, disguised command injection, intermittent weak password cracking, etc., because of its hidden operation, unobvious features and high similarity to normal behavior, the recognition rate of existing detection mechanisms is generally low. This easily causes key attack behaviors to be not identified and recorded in time, thereby affecting the integrity and effectiveness of the forensics chain. SUMMARY

[0005] Therefore, the embodiments of the present application provide a forensics method and device based on a behavior recognition model, an electronic device and a storage medium, which are used to solve at least one technical problem.

[0006] The embodiment of the application provides a forensics method based on a behavior recognition model, comprising: collecting to-be-recognized data in a target server, wherein the to-be-recognized data comprises one or more of system performance data, log data and network connection feature data; pre-processing a plurality of to-be-recognized data, extracting a numerical feature vector from each of the pre-processed to-be-recognized data, and splicing a plurality of the feature vectors to obtain a state vector; inputting the state vector into a trained behavior recognition model for recognition, wherein the behavior recognition model is trained based on a reward function to improve the recognition accuracy of light attack behavior; in response to the behavior recognition model recognizing that there is an attack behavior, issuing an alarm information and recording an attack process, wherein the attack behavior comprises light attack behavior and strong attack behavior; and wherein, based on the attack process, a timestamp certificate is applied for timestamp authentication to obtain a timestamp certificate, and attack process data and the timestamp certificate are stored in an immutable storage medium for subsequent forensics.

[0007] The forensics method as described above, wherein the state vector comprises: a system performance vector formed by system performance data change rate calculation and distribution extraction; a log vector formed by field semantic coding splicing of the log data; and a network connection feature vector formed by edge attribute coding in a double-node graph constructed based on the network connection feature data.

[0008] The forensics method as described above, wherein the numerical feature vector extracted after pre-processing the system performance data comprises: collecting CPU utilization and memory occupancy in the system performance data, calculating CPU utilization change rate and memory occupancy change rate based on CPU utilization and memory occupancy in a past preset time period, respectively; collecting disk I / O data in the system performance data, and determining a disk feature vector according to disk I / O file operation distribution; collecting network traffic in a preset time period and calculating network traffic burstiness score; and determining the system performance vector according to CPU utilization change rate, memory occupancy change rate, disk feature vector and network traffic burstiness score.

[0009] The forensics method as described above, wherein the numerical feature vector extracted after pre-processing the log data comprises: collecting log data and extracting fields according to a format template; encoding the extracted fields according to semantics to obtain an encoding vector; and splicing a plurality of the encoding vectors into a log vector.

[0010] The method for extracting evidence as described above, after the network connection feature data is preprocessed to extract the numerical feature vector, includes: collecting network connection feature data, constructing a double-node graph structure based on the source IP and the target port, and binding one or more of the total number of connections, the number of successful connections, and the number of failed connections to each edge in the double-node graph structure; extracting connection edge data in the double-node graph structure, the connection edge data including one or more of connection frequency, success rate, failure rate, non-conventional port access ratio, and abnormal mode score; and encoding the connection edge data into a network connection feature vector.

[0011] The method for extracting evidence as described above, the behavior recognition model is trained based on a reward function, including: making training samples according to server historical data and inputting the trained model; counting the recognition results of the behavior recognition model in each training period and calculating the reward function value according to the recognition results output by the model; calculating the long-term return value according to the reward function value of each training period, and then calculating the advantage value; according to the advantage value, using the PPO strategy to construct the objective function for parameter optimization until the recognition result meets the optimization condition, and outputting the trained behavior recognition model.

[0012] The method for extracting evidence as described above, the reward function is represented by the following formula:

[0013] wherein, is the reward function value, is the recognition accuracy, is the recall rate of light attack samples, is the false positive rate, is the weight coefficient, and =1.

[0014] The method for extracting evidence as described above, after each training period is completed, the attack recognition accuracy, the recall rate of light attack samples, and the false positive rate data in the recognition result are counted, and the corresponding weight value is dynamically adjusted according to the data changes of the three.

[0015] The method for extracting evidence as described above, the immutable storage medium includes: a WORM hard disk and a blockchain structure.

[0016] According to another aspect of the present application, a forensics device based on a behavior recognition model is provided, comprising: a collection module configured to collect to-be-recognized data in a target server, the to-be-recognized data comprising one or more of system performance data, log data, and network connection feature data; a preprocessing module configured to pre-process the to-be-recognized data and extract numerical feature vectors therefrom, and concatenate the numerical feature vectors to obtain a state vector; an identification module configured to input the state vector into a trained behavior recognition model for identification, the behavior recognition model being trained based on a reward function to improve the identification accuracy of light attack behaviors; and a forensics module configured to, when the behavior recognition model identifies that an attack behavior exists, issue an alarm and record the attack process, the attack behavior comprising a light attack behavior and a strong attack behavior; wherein the recorded attack process is time-stamped and written into an immutable storage medium in real time for subsequent forensics.

[0017] According to another aspect of the present application, an electronic device is provided, comprising a processor and a memory, the memory storing a computer program instruction set, and the processor executing the computer program instruction set on the memory to implement the forensics method based on a behavior recognition model as described above.

[0018] According to another aspect of the present application, a computer program product is provided, comprising a computer program instruction set, the computer program instruction set being executed by a processor to implement the forensics method based on a behavior recognition model as described above.

[0019] The present application comprehensively characterizes the current behavior state of a server from multiple dimensions by introducing multiple running data including system performance data, log data, and network connection feature data, so that the behavior recognition model can simultaneously focus on system load fluctuations, user operation trajectories, and network communication features. In addition, the present application uses a reinforcement learning strategy model based on a reward function for training, the reward function guiding the model to simultaneously optimize the overall identification accuracy, light attack recall rate, and false alarm control capability, thereby greatly improving the identification accuracy of the recognition model for light attack behaviors. BRIEF DESCRIPTION OF DRAWINGS

[0020] Hereinafter, preferred embodiments of the present application will be described in further detail with reference to the accompanying drawings, in which: Figure 1 is a forensics method flowchart based on a behavior recognition model according to an embodiment of the present application.

[0021] Figure 2 is a method flowchart for extracting numerical feature vectors after pre-processing system performance data according to an embodiment of the present application.

[0022] Figure 3is a method flow chart of extracting a numerical feature vector after pre-processing log data according to an embodiment of the present application.

[0023] Figure 4 is a method flow chart of extracting a numerical feature vector after pre-processing network connection feature data according to an embodiment of the present application.

[0024] Figure 5 is a method flow chart of training a behavior recognition model according to an embodiment of the present application.

[0025] Figure 6 is a structural schematic diagram of a forensic device based on a behavior recognition model according to an embodiment of the present application.

[0026] Figure 7 is a hardware structural schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0027] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0028] In the following detailed description, reference can be made to the various drawings that form a part of the present disclosure and are used to illustrate specific embodiments of the present application. In the drawings, like numerals describe generally similar components throughout the several views. Each of the various embodiments of the present application is described in enough detail to enable those skilled in the art to implement the technical solutions of the present application. It should be understood that other embodiments can also be utilized and structural, logical or electrical changes can be made to the embodiments of the present application.

[0029] In order not to interfere with the normal work of the target server, the honeypot enhanced recording unit (i.e., the forensics device based on the behavior recognition model) of the present application is in an isolated state with the target server, which can be deployed in a physically independent host or a secure virtual environment to ensure that it is not detected or disturbed by the attacker when the attack occurs. Preferably, the honeypot enhanced recording unit can be deployed on a mirror port or a bypass switching device in the network where the target server is located, for transparent monitoring of network traffic to and from the server without invasive modification of the target server network stack. The acquisition unit in the honeypot enhanced recording unit sends the raw data collected from the target server to the behavior recognition model in the honeypot enhanced recording unit for recognition and judgment.

[0030] Figure 1 is a flowchart of a forensics method based on a behavior recognition model according to an embodiment of the present application. As shown in Figure 1 , the method comprises: S110, acquiring to-be-identified data in a target server, the to-be-identified data comprising one or more of system performance data, log data, and network connection feature data; S120, extracting a numerical feature vector after pre-processing a plurality of the to-be-identified data, and splicing a plurality of the feature vectors to obtain a state vector; S130, inputting the state vector into a trained behavior recognition model for recognition, the behavior recognition model being trained based on a reward function to improve the recognition accuracy of light attack behavior; S140, in response to the behavior recognition model recognition result being that there is an attack behavior, issuing an alarm information and recording an attack process, the attack behavior comprising a light attack behavior and a strong attack behavior; wherein the recorded attack process is time-stamped and written in real time into an immutable storage medium for subsequent forensics.

[0031] In step S110, a plurality of types of raw data generated during the operation of the target server are acquired, which include but are not limited to one or more of system performance data, log data, and network connection feature data. Among them, the system performance data such as CPU utilization, memory occupancy, etc. reflects the time series features of the server running state; the log data can record system events and operation processes; the network connection feature data includes source IP, target port, connection frequency, etc., reflecting the network interaction behavior. By introducing multi-dimensional data sources, it is helpful to build a more complete attack behavior portrait and enhance the perception ability of abnormal behavior, especially to improve the detectability of lightweight attacks. In an embodiment, the to-be-identified data in the target server is periodically acquired, and the period can be any value in 1s-30s.

[0032] In step S120, after standardizing and structuring the various types of raw data (preprocessing), numerical vector features are extracted respectively, and then a plurality of feature vectors are spliced to form a unified state vector. For example, the CPU utilization rate in the system performance data can be calculated and normalized; the log field is spliced into a log vector after semantic encoding; the network connection features are constructed based on the source IP and port connection graph, and the connection success rate, failure rate, etc. are extracted as vector inputs. The structured multi-modal information is uniformly encoded into a numerical vector that can be processed by the model, and each state vector is bound to the corresponding timestamp, providing input with time effectiveness and behavior feature correlation for the subsequent recognition model.

[0033] In step S130, the above state vector is input into the trained behavior recognition model, which is trained based on a reward function using a reinforcement learning algorithm (such as PPO) to construct a reward function weighted by accuracy, light attack recall rate, and false positive rate. By guiding the model to optimize the above comprehensive indicators during training, it has higher recognition accuracy and practicality in identifying attack behaviors, especially when there are light attack behaviors.

[0034] Among them, the light attack behavior refers to an intrusion behavior with low attack frequency, concealed operation, unobvious single behavior characteristics, but with potential harm, which is usually not easily identified by traditional intrusion detection systems based on rules or shallow feature matching. The strong attack behavior refers to a network attack behavior with obvious destructive, high intensity, and high visibility, which usually directly affects the availability, data integrity, or control authority of the server system, has obvious intrusion characteristics and high risk level. Further, the behavior recognition model of the present application has online updating capability, which can continuously adapt to new attack patterns and improve the system's detection capability in real combat. In step S140, when the behavior recognition model identifies that there is an attack behavior, an alarm information will be immediately sent out, and the attack process will be recorded. Based on the recorded attack process, a timestamp certificate is obtained by timestamp authentication, and the attack process data and timestamp certificate are stored in an immutable storage medium for subsequent evidence. Among them, the attack behavior includes: light attack behavior and strong attack behavior.

[0035] In one embodiment, the behavior recognition model uses the vector distance (such as cosine similarity) between the state vector and the normal behavior baseline model to determine whether there is an attack behavior, specifically: When the vector distance between the state vector and the normal behavior baseline model is higher than the light attack threshold and lower than the strong attack threshold, or the number of occurrence of the related attack event within a unit time is lower than the preset frequency threshold, it is determined that there is a light attack behavior, an alarm information is immediately sent out, and the attack process is recorded.

[0036] When the state vector is higher than the strong attack threshold from the normal behavior baseline model or the triggered network traffic burstiness score is higher than the attack threshold, or the number of high-risk system logs triggered by it within a unit time exceeds the count threshold, it is determined that there is a strong attack behavior, and an alarm information is immediately sent out, and the attack process is recorded.

[0037] According to an embodiment of the present application, recording the attack process includes one or more of recording the keyboard input sequence initiated by the attack source, the executed command and the execution result, and the terminal screen snapshot in the attack session. Through the above operation, the operation path of the attacker can be completely restored, and detailed basis is provided for security analysis. Based on the "session replay level" record, the interactive process of the attacker (such as the input command, the system response, and the screen dynamic change) can be reproduced frame by frame, helping to quickly locate the attack method, tool and intention, and improving the event investigation efficiency.

[0038] Further, the attack process content can be written into an immutable storage medium, which includes a WORM hard disk and a blockchain structure. Writing the attack process into the immutable storage medium can prevent data tampering and loss, ensure the integrity and credibility of the forensic record, and enhance the post-event review and legal support ability of the system.

[0039] In an embodiment of the present application, after identifying the attack behavior and completing the attack process recording, a timestamp authentication operation is further performed to ensure the authenticity and integrity of the data file at a specific time point. The timestamp authentication process includes the following steps: First, a hash calculation is performed on the extracted forensic fragment set (such as disk cluster blocks, memory page data, network session summaries, etc.) to generate a unique hash digest value (Hash Value) of the data set. The digest value can uniquely identify the data content and has a fixed nature.

[0040] Then, the digest value is submitted to a nationally authorized trusted timestamp service center (TSA), which records the receiving time and generates a timestamp response file (TSR) containing time information and the original digest value, i.e. a timestamp certificate. The timestamp certificate is essentially an electronic certificate issued by the service center, which is used to identify that the submitted data exists at that time and has not been changed.

[0041] Subsequently, the timestamp certificate and the original forensic fragment set are stored in an immutable storage medium, such as a WORM disk, a blockchain system or a trusted computing platform, for future audit, compliance investigation or judicial evidence use.

[0042] ​In the subsequent verification phase, the Hash value can be recalculated by the current data, and compared with the original digest value in the timestamp certificate to determine whether the data has remained unchanged since self-authentication. Among them: When the Hash value generated when verifying is consistent with the Hash value in the timestamp file, it means that the data has existed at the time of applying for the timestamp, and has not been modified or forged since then; When the Hash value is inconsistent, it means that the data has been changed or forged since the timestamp is generated, and cannot be used as original and reliable evidence.

[0043] This method effectively solves the technical problems of "time point cannot be verified" and "data tampering is suspicious" in the process of electronic evidence identification, significantly improving the feasibility and credibility of the application in electronic evidence solidification, responsibility identification and judicial application.

[0044] In the above method, by recording the attack process in real time and writing it into the immutable storage medium after applying the timestamp, complete and reliable technical evidence can be provided for the fact that the server has been attacked. When the server cannot normally provide services to the superior customer (such as Party A) due to network attacks, the cause of the failure can be proved to be an irresistible factor, not the responsibility of Party B, so as to avoid disputes or commercial responsibility caused by misjudgment.

[0045] The present application comprehensively characterizes the current behavior state of the server from multiple dimensions by introducing various running data including system performance data, log data and network connection feature data, so that the behavior recognition model can simultaneously focus on system load fluctuation, user operation trajectory and network communication features. Compared with the data structure relying only on a single log, the multi-dimensional fusion input significantly improves the recognition efficiency and attack discrimination accuracy of the model.

[0046] In addition, the present application adopts a reinforcement learning strategy model based on a reward function for training. The reward function guides the model to simultaneously optimize the overall recognition accuracy, light attack recall rate and false alarm control ability, so that the training process is more goal-oriented, and the recognition accuracy of the identification model for light attack behavior is greatly improved.

[0047] According to an embodiment of the present application, the state vector includes: a system performance vector formed by the change rate calculation and distribution extraction of the system performance data; a log vector formed by the semantic encoding splicing of the fields in the log data; and a network connection feature vector formed by the edge attribute encoding in the double-node graph constructed by the network connection feature data. The specific steps of extracting the numerical feature vector after preprocessing the system performance data, log data and network connection feature data will be described in detail below.

[0048] Figure 2is a method flowchart for extracting a numerical feature vector after pre-processing system performance data according to an embodiment of the present application. As shown in Figure 2 The method comprises the following steps: S210, collecting CPU utilization and memory occupancy in system performance data, and calculating CPU utilization rate of change and memory occupancy rate of change based on CPU utilization and memory occupancy in a past preset time period, respectively; S220, collecting disk I / O data in system performance data, and determining a disk feature vector according to disk I / O file operation distribution; S230, collecting network traffic in a preset time period and calculating network traffic burstiness score; S240, determining a system performance vector according to CPU utilization rate of change, memory occupancy rate of change, disk feature vector and network traffic burstiness score.

[0049] In step S210, the CPU utilization and memory occupancy of the server at the current time are collected in real time, and the "rate of change" is calculated based on the change in the past preset time period, which is used to depict the dynamic fluctuation behavior of system load. Specifically, let the current time be t, and the past preset time period be Δt (for example, 3 seconds), then the change rate calculation formula is:

[0050] Wherein, r is the rate of change, is the data (CPU utilization or memory occupancy) at time t, and Δt is the past preset time period.

[0051] In step S220, disk I / O operation records are collected, and the behavior categories and their proportions of file operations in a unit of time are identified, including but not limited to file writing, file deletion, file renaming, etc. The proportion of the number of each type of operation to the total number of operations is counted, and a triple vector is constructed:

[0052] Wherein, is the disk feature vector, is the number of file writes, is the number of file deletions, is the number of file renames.

[0053] In step S230, the network traffic in and out is collected in a preset time window (such as 5 seconds), and the burstiness score is calculated according to the entropy model. High burstiness score often corresponds to attack behaviors such as port scanning and instantaneous burst, which helps to assist in identifying high-variability communication patterns in lightweight attacks. The entropy value is calculated using the normalized change distribution of the traffic sampling sequence, and the change amplitude is counted as a burstiness indicator:

[0054]

[0055]

[0056] wherein H is an entropy value, the more uniform the connection distribution, the higher the entropy, is the proportion of the i-th port in all connections in the second, is the maximum entropy value, is the minimum entropy value, B is the burstiness score, ΔH represents the entropy change amplitude, and n is the number of sampling points.

[0057] In step S240, the above-mentioned various features are spliced into a system performance vector, including: CPU utilization rate change rate (1 dimension), memory occupancy rate change rate (1 dimension), disk operation distribution vector (3 dimensions), network traffic burstiness score (1 dimension). That is, the output system performance vector format is:

[0058] The above method performs fine-grained structured analysis on system performance data, introduces dynamic indicators such as change rate, distribution proportion and burstiness score, thereby sensitively modeling the implicit changes of attack behaviors at the input feature level. This vectorization method not only can capture the short-time resource abnormalities caused by light attack behaviors, but also can enhance the model's ability to distinguish fuzzy samples at the behavior boundary, thereby improving the overall recognition accuracy and the response ability to light attacks.

[0059] Figure 3 is a flow chart of a method for extracting a numerical feature vector after pre-processing log data according to an embodiment of the present application. As shown in Figure 3 , the method comprises: S310, collecting log data and extracting fields according to a format template; S320, encoding the extracted fields according to semantics to obtain an encoding vector; S330, splicing multiple encoding vectors into a log vector.

[0060] In step S310, multiple log types including system logs, security logs and application logs are collected, and the log samples are as follows:

[0061] Based on a pre-defined format template (such as a regular expression or a log parsing rule), the key fields of the log are extracted, including: timestamp: ; module name: ; operation type: ; user: , source IP: Port number: .

[0062] In step S320, for each extracted field, different encoding methods are used according to its data type and semantic type: discrete fields (such as module name, operation type, result status) use One-hot encoding or pre-trained word embedding; numerical fields (such as port number) are normalized; time fields can be discretized by hour or day of the week (time slice bucketing); IP address fields can be mapped to risk levels according to frequency classification (such as suspicious IP with high frequency mapped to high risk encoding).

[0063] In step S330, the multiple field vectors corresponding to each log are spliced to form a uniform length log vector. Assuming that there are m total extracted fields, and each field is encoded as a vector vi, then each structured log is:

[0064] The method analyzes and encodes the original log data through template parsing, not only retaining the key information of operation behavior, but also enabling the model to identify the context semantics and potential risk patterns behind the behavior. By splicing the semantic encoding of multiple fields into a fixed-length vector, the structural expressiveness of log data and the model processing efficiency are significantly improved. Especially for low-frequency abnormal log events commonly seen in light attack behavior, structured processing helps the model accurately capture their feature differences, improving the accuracy and reliability of attack recognition.

[0065] Figure 4 is a flow chart of a method for extracting a numerical feature vector from preprocessed network connection feature data according to an embodiment of the present application.

[0066] S410, network connection feature data is collected, and a double-node graph structure is constructed based on the source IP and target port, wherein each edge in the double-node graph structure is bound with one or more of the connection total number, connection success number and connection failure number; S420, connection edge data is extracted in the double-node graph structure, wherein the connection edge data includes one or more of the connection frequency, success rate, failure rate, irregular port access proportion and abnormal mode score; S430, the connection edge data is encoded into a network connection feature vector.

[0067] In step S410, network connection records are collected in a fixed time window (e.g., every minute), including the source IP, target port, connection state (success / failure), etc. of each TCP or UDP connection. Based on such connection events, a bipartite graph is constructed with "source IP-target port" as a set of nodes. This graph structure can retain the directionality and port association of connections, and is an important basis for describing port scanning, distributed connection, and other attack behaviors. Each edge represents one or more connection attempts. Each edge is bound to an attribute field, including: N total : total number of connections; N success : number of successful connections; N fail : number of failed connections.

[0068] In step S420, connection frequency represents the total number of attempted connections initiated per unit time, and the proportion of irregular port access represents the proportion of the number of access to uncommon ports. If the target port of a connection edge is p, and the port does not belong to the predefined set of common ports (such as 22, 80, 443, etc.), it is marked as "irregular", and the proportion of its occurrence in all connections is counted. A scoring function is defined in combination with multiple features (such as high failure rate, random port distribution), and the degree of deviation of the connection behavior from the normal communication mode is determined by the scoring function value.

[0069] In step S420, several feature data extracted from each connection edge are normalized and spliced into a uniform structure of network connection feature vector, for example:

[0070] Wherein, is the network connection feature vector, is the connection frequency, is the connection success rate, is the connection failure rate, is the proportion of irregular port access, is the abnormal mode score.

[0071] This method effectively mines the implicit abnormal patterns in server external connections by modeling network connection behavior as a bipartite graph structure and combining statistical and behavior feature extraction methods. Compared with methods that only rely on total number of connections or port blacklist, this method can more sensitively identify the low-frequency, high-failure-rate, and widely-spreading connection attempts commonly seen in light attack behaviors, and inject them into the model in a structured form, improving the system's perception ability and attack discrimination accuracy for complex communication behaviors.

[0072] The system performance vector, log vector, and network connection feature vector calculated by Figures 2-4 are spliced into a state vector S t , and the state vector St The trained behavior recognition model can be input to quickly identify whether there is an attack behavior. The behavior recognition model is trained by using a reinforcement learning method, preferably a proximal policy optimization (PPO) algorithm. The method guides and evaluates the model behavior by defining a reward function, optimizes the recognition strategy to improve the attack recognition accuracy, especially the recognition ability of light attack behavior.

[0073] Figure 5 is a flowchart of a behavior recognition model training method according to an embodiment of the present application. As shown in Figure 5 , the method comprises: S510, training samples are made according to server historical data and input into the model to be trained, the training samples include system performance vector, log vector and network connection feature vector, and corresponding action type and confidence. Wherein, the identified action a∈{0,1,2}, a=0: normal behavior; a=1: light attack behavior; a=2: strong attack behavior. The confidence score p∈[0,1].

[0074] S520, the recognition results of the behavior recognition model in each training period are counted, and the reward function value is calculated according to the recognition results output by the model; wherein, the reward function is represented by the following formula:

[0075] Wherein, is the reward function value, is the recognition accuracy, is the recall rate of light attack samples, is the false positive rate, is the weight coefficient and =1.

[0076] S530, the long-term return value of the reward function value in each round of training period is calculated, and the advantage value is calculated;

[0077]

[0078] Wherein, is the long-term return value of the cumulative return obtained at time step t, is the discount factor, is the reward function value at time step t; is the advantage value at time step t, is the state vector S t under the average expected value, which is predicted by the value network.

[0079] S540, according to the advantage value, using the PPO strategy to construct the target function for parameter optimization until the identification result meets the optimization condition, stopping optimization, and outputting the trained behavior identification model; the target function is represented by the following formula:

[0080] wherein, is the target function of PPO, is the parameter of the policy network, is the probability ratio of the new and old policies, is the advantage function, is the clipping function, is the policy change clipping threshold.

[0081] In the model training process, the application introduces a reward function. The reward function integrates multiple indicators such as attack recognition accuracy, light attack recall rate, and false positive rate into a unified scoring standard, so that the model can actively iterate and optimize in the direction of "high accuracy, high light attack recall rate, and low false positive rate" during training, solving the problem of single target or deviation from actual security needs in traditional training.

[0082] According to an embodiment of the application, after each training cycle is completed, the attack recognition accuracy, the recall rate of light attack samples, and the false positive rate data in the identification result are counted, and the corresponding weight values are dynamically adjusted according to the data changes of the three.

[0083] After each training cycle is completed, the weight parameters in the reward function are dynamically adjusted according to the performance change trend of the current model in the three indicators, so that they guide the model to further optimize in the direction of weak performance in the subsequent training cycle. For example: If the light attack recall rate R in the current round is lower than the expected threshold (such as <60%), the weight proportion of is appropriately increased to strengthen the model's learning motivation for light attack recognition; If the false positive rate F is high, the punishment effect of is improved to guide the model to shrink the attack boundary and reduce the false positive tendency.

[0084] By introducing the reward function weight dynamic adjustment mechanism based on training result feedback, the application realizes the "self-calibration" of the training target: the model can automatically adjust the training bias according to its performance in recognition accuracy, light attack coverage ability, and false positive control ability, thereby realizing adaptive optimization of multiple target performance.

[0085] Corresponding to the method embodiment of the application, the application also provides a forensic device based on a behavior identification model, as shown in Figure 6 The forensic device based on the behavior identification model 100 comprises: The collection module 101 is configured to collect to-be-identified data in a target server, and the to-be-identified data includes one or more of system performance data, log data, and network connection feature data; The preprocessing module 102 is configured to pre-process a plurality of to-be-identified data, respectively extract numerical feature vectors, and splice the plurality of feature vectors to obtain a state vector. The identification module 103 is configured to input the state vector into a trained behavior identification model for identification, and the behavior identification model is trained based on a reward function to improve the identification accuracy of light attack behavior. The evidence collection module 104 is configured to issue an alarm information and record an attack process when the behavior identification model identifies that there is an attack behavior, and the attack behavior includes light attack behavior and strong attack behavior; wherein the recorded attack process is time-stamped and written into an immutable storage medium in real time for subsequent evidence collection. The attack process content is written into the storage module 105, which can be a WORM hard disk or an immutable storage medium in a blockchain structure.

[0086] Figure 7 FIG. 1 is a schematic diagram of a hardware structure of an electronic device according to an embodiment of the present application. The electronic device can be implemented as a server or other various terminal devices, such as a desktop personal computer, a tablet computer, a laptop computer, a mobile phone, etc. The electronic device includes a processor 601 and a memory 602. The memory 602 stores a set of program instructions. The processor 601 executes the set of program instructions stored in the memory 602 to implement the above-described evidence collection method based on the behavior identification model.

[0087] Specifically, the processor 601 can include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present application.

[0088] The memory 602 can include a mass storage for data or instructions. For example, and without limitation, the memory 602 can include a hard disk drive (HDD), a floppy disk drive, a compact disc (CD) drive, a digital versatile disc (DVD) drive, flash memory, a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 602 can include removable or non-removable (or fixed) media. Where appropriate, the memory 602 can be internal or external to the integrated gateway disaster recovery device. In some embodiments, the memory 602 is non-volatile, solid-state memory.

[0089] The memory can include read-only memory (ROM), random-access memory (RAM), magnetic disk storage mediums, optical storage mediums, flash memory devices, electrical, optical, or other physically tangible / tangible memory storage devices. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software that, when executed (e.g., by one or more processors), is operable to perform the forensic method based on behavior recognition model provided by the present application.

[0090] In one example, the electronic device further includes a communication interface 603 and a bus 604. The processor 601, the memory 602, and the communication interface 603 are connected through the bus 604 and complete the communication between each other. The communication interface 603 is mainly used to realize the communication between the modules, devices, units and / or equipment in the embodiments of the present application. The bus 604 includes hardware, software or both to couple the components of the online data traffic billing device to each other. By way of example and not limitation, the bus can include an accelerated graphics port (AGP) or other graphics bus, an enhanced industry standard architecture (EISA) bus, a front side bus (FSB), a hyper transport (HT) interconnect, an industry standard architecture (ISA) bus, an infiniband interconnect, a low pin count (LPC) bus, a memory bus, a micro channel architecture (MCA) bus, a peripheral component interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a serial advanced technology attachment (SATA) bus, a video electronics standards association local (VLB) bus, or other suitable bus or combination of two or more of these. The bus 604 can include one or more buses as appropriate. Although specific buses are described and illustrated in the embodiments of the present application, the present application contemplates any suitable bus or interconnect.

[0091] The present application also provides a computer-readable storage medium having computer program instructions stored thereon, which can be executed by a processor to implement the forensic method based on behavior recognition model in any of the foregoing embodiments. The computer-readable storage medium can be any medium that can tangibly contain or store computer-executable instructions for use by or in connection with an instruction execution system, apparatus, and device. The storage medium can be a transitory computer-readable storage medium or a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium can include, but is not limited to, magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Embodiments corresponding to such storage devices include, for example, magnetic disks, optical disks based on CD, DVD, or Blu-ray technologies, and persistent solid-state memory such as flash memory, solid-state drives, and the like.

[0092] The application further provides a computer program product comprising a set of computer program instructions which, when executed by a processor, implement the forementioned evidence obtaining method based on the behavior recognition model. The computer program product includes but is not limited to an application installation package, an application plug-in, a small program that can run in some applications, and the like published in a website or an application store.

[0093] It should be noted that the present application is not limited to the specific configurations and processes described above and shown in the drawings. For the sake of brevity, detailed descriptions of well-known methods are omitted herein. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between steps, after understanding the spirit of the present application.

[0094] The above embodiments are only for illustrating the present application, and are not intended to limit the present application. Those skilled in the art can make various changes and modifications without departing from the scope of the present application. Therefore, all equivalent technical solutions shall also belong to the scope disclosed by the present application.

Claims

1. A forensics method based on a behavior recognition model, characterized in that, The method comprises the following steps: Collecting to-be-identified data in a target server, wherein the to-be-identified data comprises one or more of system performance data, log data and network connection feature data; After pre-processing the to-be-identified data, numerical feature vectors are extracted respectively, and a state vector is obtained by splicing the numerical feature vectors; The state vector is input into a trained behavior identification model for identification, wherein the behavior identification model is trained based on a reward function to improve the identification accuracy of light attack behavior; When the behavior identification model identifies that there is an attack behavior, an alarm information is sent and an attack process is recorded, wherein the attack behavior comprises light attack behavior and strong attack behavior; A timestamp certificate is obtained by applying timestamp authentication based on the attack process, and attack process data and the timestamp certificate are stored in an immutable storage medium for subsequent evidence collection.

2. The forensic method according to claim 1, characterized in that, The state vector comprises a system performance vector formed by rate of change calculation and distribution extraction of system performance data, a log vector formed by splicing of field semantic encoding of the log data, and a network connection feature vector formed by edge attribute encoding of a double-node graph constructed based on the network connection feature data.

3. The forensic method according to claim 1, characterized in that, After pre-processing the system performance data, the numerical feature vectors are extracted, comprising: CPU utilization and memory occupancy in the system performance data are collected, and CPU utilization rate of change and memory occupancy rate of change are calculated based on CPU utilization and memory occupancy in a past preset time period, respectively; Disk I / O data in the system performance data are collected, and a disk feature vector is determined according to disk I / O file operation distribution; Network traffic in a preset time period is collected and network traffic burstiness score is calculated; The system performance vector is determined according to the CPU utilization rate of change, the memory occupancy rate of change, the disk feature vector and the network traffic burstiness score.

4. The forensic method of claim 1, wherein, After pre-processing the log data, the numerical feature vectors are extracted, comprising: The log data are collected and fields are extracted according to a format template; The extracted fields are encoded according to semantics to obtain an encoding vector; A plurality of encoding vectors are spliced into a log vector.

5. The method of forensics according to claim 1, wherein, After pre-processing the network connection feature data, the numerical feature vectors are extracted, comprising: The network connection feature data are collected, a double-node graph structure is constructed based on source IP and target port, and one or more of connection total number, connection success number and connection failure number are bound to each edge in the double-node graph structure; Connection edge data are extracted in the double-node graph structure, wherein the connection edge data comprises one or more of connection frequency, success rate, failure rate, irregular port access proportion and abnormal mode score; The connection edge data are encoded into a network connection feature vector.

6. The forensic method according to claim 1, characterized in that, The behavior identification model is trained based on a reward function, comprising: Training samples are made according to server historical data and input into a to-be-trained model; Identification results of the behavior identification model in each training period are counted, and a reward function value is calculated according to the identification results; A long-term return value is calculated according to the reward function value of each training period, and an advantage value is calculated. According to the advantage value, a PPO strategy is used to construct a target function for parameter optimization until the identification result meets an optimization condition, and a trained behavior identification model is output.

7. The forensic method according to claim 1, characterized in that, The reward function is expressed by the following formula: wherein, is a reward function value, is an identification accuracy, is a light attack sample recall rate, is a false positive rate, is a weight coefficient and = 1.

8. The forensic method according to claim 7, characterized in that, After each training cycle, the attack identification accuracy, recall rate and false positive rate of the light attack sample in the identification result are counted, and the corresponding weight values are dynamically adjusted according to the data changes of the three.

9. The method of forensics according to claim 1, wherein, The immutable storage medium includes a WORM hard disk and a blockchain structure.

10. A forensics device based on a behavior recognition model, characterized by, It includes: The acquisition module is configured to acquire to-be-identified data in the target server, the to-be-identified data including one or more of system performance data, log data and network connection feature data; The preprocessing module is configured to preprocess the to-be-identified data and extract numerical feature vectors, and splice the feature vectors to obtain a state vector; The identification module is configured to input the state vector into the trained behavior identification model for identification, the behavior identification model being trained based on a reward function to improve the identification accuracy of light attack behavior; The evidence collection module is configured to issue an alarm information and record an attack process when the behavior identification model identifies that there is an attack behavior, the attack behavior including a light attack behavior and a strong attack behavior; wherein the recorded attack process is time-stamped and written into an immutable storage medium in real time for subsequent evidence collection.

11. An electronic device, comprising: The evidence collection method based on the behavior identification model includes a processor and a memory, and the memory stores a computer program instruction set, and the processor executes the computer program instruction set on the memory to implement the evidence collection method based on the behavior identification model.

12. A computer program product, characterised in that, The evidence collection method based on the behavior identification model includes a computer program instruction set, and the computer program instruction set is executed by the processor to implement the evidence collection method based on the behavior identification model.

Citation Information

Patent Citations

  • Collaborative network electronic evidence obtaining technology based on third-party signature

    CN102932145A

  • Full-scene network security threat association analysis method and system

    CN117478403A

  • Network attack identification method based on behavior modeling

    CN118740513A

  • Network information security analysis method and system based on data analysis

    CN118784348A

  • Foggy day driving augmented reality auxiliary method and system based on multi-mode perception

    CN120171555A