Computer network information security management method and system based on data processing
By constructing login time, space, and device models, and combining access rhythm with resource coupling coefficients, anomaly scores are generated and automatic blocking is implemented. This solves the problems of identity spoofing, advanced threat identification, and time lag in computer network information security management, and improves the accuracy and real-time performance of security management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LIHE TECHNOLOGY CO LTD
- Filing Date
- 2026-03-13
- Publication Date
- 2026-06-05
AI Technical Summary
Existing computer network information security management technologies face problems such as a surge in identity theft risks, high frequency of false alarms, difficulty in identifying advanced threats, and time lag, especially in terms of device fingerprinting and session monitoring.
By collecting current and historical security data of users, a login time probability distribution model, a spatial location cluster center model, and a dynamic device fingerprint model are constructed. Combined with access rhythm and resource coupling coefficient, anomaly scores are generated and automatic blocking is achieved.
It effectively identifies false alarms caused by device counterfeiting and hardware aging, deeply identifies automated script attacks, and achieves precise risk control and millisecond-level automatic blocking across the entire chain, thereby improving the accuracy of information security management and real-time defense capabilities.
Smart Images

Figure CN122160137A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information security technology, and specifically to a computer network information security management method and system based on data processing. Background Technology
[0002] With the rapid development of network technology, user authentication and access control have become the core defense line for ensuring computer network information security. However, existing security management technologies still face severe challenges in practical applications. On the one hand, traditional device fingerprinting relies heavily on static software features such as operating system version and browser type. This information is easily forged and tampered with by attackers through proxy tools, virtual machines, or scripts, leading to a surge in identity theft risks. Simultaneously, legitimate users' hardware devices naturally age over time, causing a slow drift in microscopic physical characteristics. If static models are used, they cannot distinguish between "natural aging" and "malicious forgery," easily resulting in high-frequency false alarms and severely disrupting normal business operations. On the other hand, session monitoring after successful login is often limited to single resource access paths or frequency detection, making it difficult to effectively identify advanced threats such as "slow-probing" or "automated scripts" used by attackers. These attacks may appear normal in a single resource access, but their operational rhythm differs fundamentally from user behavior. Existing methods lack in-depth analysis of the coupling relationship between time rhythm and spatial resources, resulting in insufficient sensitivity in detecting lateral movement behavior. In addition, traditional solutions mostly focus on post-event log auditing, lacking the ability to conduct millisecond-level real-time risk assessment and automatic blocking at the moment of login and the early stage of the session. This results in a significant time lag in information security protection, making it impossible to effectively curb data leakage and system damage during the golden window period when attacks occur. Summary of the Invention
[0003] The purpose of this invention is to address the problems existing in the background technology by proposing a computer network information security management method and system based on data processing.
[0004] The technical solution of this invention: A computer network information security management method based on data processing, comprising: S1. Collect key security data and historical security data of the user's current login session, and preprocess the key security data to form a feature vector of the current login behavior; S2. Based on historical security data, construct a login time probability distribution model, a login spatial location cluster center model, and a historical device fingerprint model for each user, and form a user behavior baseline vector; S3. Within 1 second after the user completes login, compare the current login behavior feature vector with the user behavior baseline vector, calculate the time deviation, spatial deviation and device anomaly, and generate a login anomaly score to preliminarily determine whether the account is at risk of being stolen. S4. Calculate the resource access anomaly degree and access path jump anomaly degree based on the preprocessed key security data. Calculate the access rhythm anomaly degree by introducing the access rhythm and resource coupling coefficient. Generate an access anomaly score by weighted fusion. S5. Generate a comprehensive risk value based on login anomaly scoring and access anomaly scoring, which is used to trigger automatic blocking for information security management.
[0005] As a further improvement to this technical solution, in step S1, the key security data is preprocessed to form a current login behavior feature vector, including the following steps: The critical security data is preprocessed. Based on statistical feature analysis and embedding coding methods, the current login feature vector is extracted from the preprocessed critical security data. The extracted current login feature vector includes at least the current time feature vector, spatial location feature vector, and device fingerprint feature vector. The login feature vector is normalized. All the normalized login feature vectors are combined to form the final current login behavior feature vector, which serves as a representation of the user's current login behavior.
[0006] As a further improvement to this technical solution, in step S2, a login time probability distribution model, a login spatial location cluster center model, and a historical device fingerprint model are constructed to form a user behavior baseline vector, including the following steps: S2.1 Preprocess historical security data; Historical security data includes user historical login data, historical access record data, and historical device fingerprint information; S2.2 Extract user historical login time information based on preprocessed user historical login data, calculate the historical login probability of each historical login time period, and generate a user login time probability distribution model; S2.3 Based on the preprocessed user historical login data, the K-means clustering algorithm is used to identify the user's frequently used login areas, generate spatial location cluster centers, and construct a login spatial location cluster center model; S2.4 Based on the preprocessed historical device fingerprint information, a hardware aging drift model and multifractal spectrum analysis are introduced to generate dynamic device fingerprint vectors. S2.5. Generate historical time feature vectors based on the login time probability distribution model, generate spatial feature vectors based on the login spatial location cluster center model, and concatenate the historical time feature vectors, spatial feature vectors, and dynamic device fingerprint vectors to generate a complete user behavior baseline vector.
[0007] As a further improvement to this technical solution, in step S2.4, based on the preprocessed historical device fingerprint information, a hardware aging drift model and multifractal spectrum analysis are introduced to generate a dynamic device fingerprint vector, including the following steps: S2.41. During the user's historical login process, a fixed sampling frequency is used. In a short time window The micro-fluctuation sequence of the hardware operation characteristics of the internal acquisition device includes CPU clock offset sequence, memory access latency sequence, and GPU rendering time sequence; S2.42. Perform multifractal spectral analysis on each micro-wave sequence and calculate the scaling function. And based on the scaling function Extracting the generalized Hurst exponent and the strange spectrum; S2.43. Using a fractal stability adaptive confidence screening mechanism to evaluate the generalized Hurst exponent. Make corrections and base them on the corrected generalized Hurst exponent. Reconstructing the singular spectrum to calculate the stable singular spectrum width. The generalized Hurst exponent and stable singular spectral width As a multifractal feature of the micro-fluctuation sequence of the equipment; S2.44, Replace the time variable with the login event sequence number. And based on the login event sequence number Establish an aging drift model; S2.45. Based on the aging drift model, predict the expected value of the multifractal characteristics of the current device according to the current login event sequence number. And construct the confidence interval for this multifractal feature; S2.46. For the current login number, repeat steps S2.41 to S2.43 to obtain the current measured value of the multifractal feature. And the measured values of multifractal features The final multifractal features are obtained by comparing them with the confidence intervals. ; S2.47, Based on multifractal features Generate the dynamic device fingerprint vector of the current device. .
[0008] As a further improvement to this technical solution, in S2.43, the generalized Hurst exponent is evaluated using a fractal stability adaptive confidence screening mechanism. Make corrections and base them on the corrected generalized Hurst exponent. Reconstructing the singular spectrum to calculate the stable singular spectrum width. This includes the following steps: Scale intervals are constructed based on micro-fluctuation sequences, and the scaling function is used to... The scale interval is divided into several sub-intervals, and the piecewise scaling function corresponding to each sub-interval is calculated separately. ; For the same order Statistical analysis was performed on the calculation results across different scale sub-intervals, and their variances were calculated. Through variance Measuring the order Statistical stability in multi-scale estimation forms a set of stable orders. ; and based on Further determine whether the number of its orders satisfies the minimum stability statistics requirement; Scaling function for preserving order Perform discrete second-order difference to calculate local curvature, and perform neighborhood smoothing correction when the absolute value of local curvature exceeds a set threshold; After completing the neighborhood smoothing correction, based on the variance of each order... Construct the corresponding confidence weights Based on confidence weights For the generalized Hurst exponent The weighted adjustment is performed to generate the corrected generalized Hurst exponent. ; In the stable order set Based on the modified generalized Hurst exponent within the range Calculate the corresponding singularity index And construct a stable singular spectrum Generate stable singular spectral width And multifractal features.
[0009] As a further improvement to this technical solution, in step S3, the current login behavior feature vector is compared with the user behavior baseline vector to calculate the time deviation, spatial deviation, and device anomaly degree, and a login anomaly score is generated, including the following steps: S3.1 Obtain the historical login probability of the time period to which the current login time belongs based on the login time probability distribution model. Based on historical login probability Calculate time deviation and the degree of time deviation Perform normalization to generate the normalized time deviation. ; S3.2 Extract the current spatial coordinates from the current login behavior feature vector And obtain the set of cluster centers for the user's historical login areas based on the login spatial location cluster center model. Spatial coordinates are calculated using the Haversine formula. With cluster center set Spatial distance between Traverse spatial distance Select the minimum spatial distance and the corresponding cluster centers And calculate the spatial deviation. Spatial deviation Perform normalization to generate the normalized spatial deviation. ; S3.3 Extract the device fingerprint feature vector from the current login behavior feature vector, and calculate the device fingerprint feature vector and the dynamic device fingerprint vector. cosine similarity Traversing cosine similarity Select the maximum similarity Calculate the degree of equipment anomaly ; S3.4 Normalized time deviation Normalized spatial deviation Equipment abnormality Perform weighted fusion to generate login anomaly scores. .
[0010] As a further improvement to this technical solution, in step S4, the resource access anomaly degree and access path jump anomaly degree are calculated based on the preprocessed key security data. The access rhythm anomaly degree is calculated by introducing the access rhythm and resource coupling coefficient, and an access anomaly score is generated through a weighted fusion method, including the following steps: S4.1, A monitoring time window after a user successfully logs in. Internally, it monitors user access behavior in real time, extracts user access resource sequences from preprocessed key security data, and constructs a set of users' historical resource accesses based on historical access record data. S4.2 Construct the current monitoring time window based on the user's resource access sequence The system calculates the rhythm deviation index and the access interval variation coefficient based on the current access interval sequence, and also counts the proportion of first accesses to resources within the current window. Finally, it constructs an access rhythm-resource coupling coefficient to generate the access rhythm anomaly degree. ; S4.3 Match each accessed resource in the current user's accessed resource sequence with each accessed resource in the historical resource access set, and calculate the monitoring time window. The average rarity of internally accessed resources is used to calculate the resource access anomaly score. ; S4.4 Construct a user historical access transition probability matrix based on historical access record data. For the current access sequence, calculate the path probability between adjacent access resources and construct the path anomaly degree. S4.5 Normalize the path anomaly score to obtain the access path hop anomaly score. ; S4.6, Determine the abnormality of the access rhythm Resource access anomaly Access path jump anomaly degree Perform weighted fusion to construct a horizontal movement index And use the lateral movement index as a score for access anomalies.
[0011] As a further improvement to this technical solution, in step S4.2, the rhythm deviation index and the coefficient of variation of the access interval are calculated based on the current access interval sequence, and the current monitoring time window is statistically analyzed. The proportion of first-time accesses to internal resources is calculated, and a coupling coefficient between access rhythm and resources is constructed to generate access rhythm anomalies. This includes the following steps: S4.21, in the monitoring window The system records the timestamp sequence of accessed resources and constructs the access time interval sequence within the current monitoring window based on the timestamp sequence. S4.22 Calculate the average access interval of the current window based on the access time interval sequence. Based on the average access interval of the current window Constructing a rhythm deviation index ; S4.23, Average access interval based on the current window Calculate the coefficient of variation of the current access interval sequence. ; S4.24, Statistical Monitoring Window Number of resources accessed for the first time First-time access ratio of computing resources ; S4.25, Based on rhythm deviation index Coefficient of variation and the percentage of first-time visits to resources Build access rhythm and resource coupling coefficient ; S4.26. Integrating access frequency with resource coupling coefficient Mapped to access rhythm anomaly degree .
[0012] As a further improvement to this technical solution, step S5, which generates a comprehensive risk value based on login anomaly scoring and access anomaly scoring, includes the following steps: A comprehensive risk value is formed by weighting and combining login anomaly scores and access anomaly scores. Through preset risk thresholds Perform risk assessment, and trigger automatic protection and execute automatic blocking strategies based on the risk assessment results.
[0013] On the other hand, the present invention provides a computer network information security management system based on data processing, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned computer network information security management method based on data processing.
[0014] Compared with the prior art, the above-mentioned technical solution of the present invention has the following beneficial technical effects: by introducing dynamic device fingerprinting technology based on multifractal spectrum analysis and hardware aging drift model, the problems of traditional static fingerprints being easily forged and high false alarm rate caused by natural hardware aging are effectively overcome. At the same time, by combining access rhythm and resource coupling coefficient to deeply identify automated script attacks and slow lateral movement behavior, the end-to-end accurate risk control and millisecond-level automatic blocking from the login source to the access process are realized, which significantly improves the accuracy, robustness and real-time defense capability of computer network information security management. Attached Figure Description
[0015] Figure 1 This is a flowchart of the overall method of the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0017] Example 1: Please refer to Figure 1 As shown, this embodiment provides a computer network information security management method based on data processing, including the following steps: S1. When a successful user login event is detected, collect key security data and historical security data of the user's current login session, and preprocess the key security data to form a feature vector of the current login behavior. In this embodiment, the key security data is preprocessed to form a feature vector of the current login behavior, including the following steps: Critical security data undergoes preprocessing (critical security data includes at least login timestamps, login IP addresses and spatial locations, device fingerprint information, and access resource records; preprocessing specifically includes: standardization (including at least timestamp unification, login IP address normalization, and device fingerprint field cleaning) to eliminate structural inconsistencies caused by differences in data sources; outlier identification and repair of standardized data (including at least time anomaly detection and IP spatial resolution anomaly handling); and current login feature vector extraction from the preprocessed critical security data based on statistical feature analysis and embedding coding methods. The extracted current login feature vector includes at least the current time feature vector (hour, day of the week), spatial location feature vector (latitude and longitude encoding or cluster center distance), and device fingerprint feature vector, essentially transforming the preprocessed critical security data into computable vector features (the extraction of the current login feature vector from the preprocessed critical security data specifically involves extracting...). Login time information is analyzed and statistically, converting the login time into a time feature vector (including hour and day of the week information). The login IP is geo-analyzed to obtain its corresponding spatial coordinates, and a spatial location feature vector is generated through latitude and longitude encoding or distance calculation to historical spatial cluster centers. Simultaneously, device fingerprint information (including operating system type, browser type, device hardware identifier, or browser fingerprint features) is extracted and converted into a device fingerprint feature vector through embedding encoding. Finally, the time feature vector, spatial location feature vector, and device fingerprint feature vector are normalized and concatenated in a preset order to form a current login feature vector representing the user's current login behavior. The login feature vector is then normalized (e.g., Min-Max normalization), and category features are one-hot encoded or embedded to ensure comparability in numerical calculation. All normalized login feature vectors are combined to form the final current login behavior feature vector, serving as a representation of the user's current login behavior.
[0018] S2. Based on historical security data, construct a login time probability distribution model, a login spatial location cluster center model, and a historical device fingerprint model for each user, and form a user behavior baseline vector; In this embodiment, a login time probability distribution model, a login spatial location cluster center model, and a historical device fingerprint model are constructed to form a user behavior baseline vector, including the following steps: S2.1 Preprocess historical security data (i.e., organize, remove duplicate, missing or abnormal records, and standardize fields such as timestamps, IP addresses, and device fingerprints to ensure data format uniformity). Historical security data includes user historical login data, historical access record data, and historical device fingerprint information; S2.2 Extract user historical login time information based on preprocessed user historical login data, and calculate the historical login probability for each historical login time period. ( For users in time periods The number of historical logins within the app. Generate a probability distribution model of user login time based on the total number of historical logins. ( (indexed by time period), used to characterize users' regular login habits and the degree of deviation from abnormal times; S2.3. Based on the preprocessed user historical login data, use the K-means clustering algorithm to identify frequently used login areas and generate several stable spatial cluster centers (and record the cluster center coordinates of each cluster center). Cluster radius Weighting ), and construct a login spatial location clustering center model. ( This is a login spatial location clustering center model, representing a set of multiple login region cluster centers, used to describe the spatial distribution characteristics of a user's historical login locations. The index number of the spatial location cluster center. (the total number of spatial location cluster centers), used to characterize the distribution characteristics of users' frequently logged-in areas; S2.4 Based on the preprocessed historical device fingerprint information, a hardware aging drift model and multifractal spectrum analysis are introduced to generate dynamic device fingerprint vectors, which are used to characterize the set of device features that users have been using stably for a long time. In this embodiment, the introduction of a hardware aging drift model and multifractal spectrum analysis primarily addresses the problems of traditional device fingerprints being easily forged, difficult to persist, and unable to adapt to hardware aging. In network security scenarios, attackers often use proxies, virtual machines, or forged software information (such as User-Agent) to hide their true identities, making device fingerprints that rely solely on software-layer features easily bypassed. Simultaneously, legitimate users' hardware devices undergo normal aging over time, causing drift in microscopic fluctuation features at the hardware level. If a static fingerprint model is used, it will lead to numerous false positives because it cannot distinguish between natural aging and attack behavior. This invention, on one hand, collects hardware microscopic fluctuation sequences such as CPU clock offset and memory access latency, and uses multifractal spectrum analysis to extract their inherent physical features, resulting in fingerprints with extremely high uniqueness and anti-forging capabilities. Even if attackers obtain software-layer information, they cannot simulate the same hardware fluctuation pattern. On the other hand, a hardware aging drift model is introduced, using the number of device logins as a proxy variable to describe the slow drift of features over time, and combined with a fractal stability adaptive filtering mechanism to ensure the robustness of feature extraction. This design can tolerate reasonable changes caused by the natural aging of hardware (reducing false alarms) and can also keenly capture anomalous device behavior that changes abruptly (improving the detection rate), thereby significantly improving the accuracy and reliability of device anomaly detection in network security systems. The process of generating dynamic device fingerprint vectors based on preprocessed historical device fingerprint information, incorporating a hardware aging drift model and multifractal spectrum analysis, includes the following steps: S2.41. During the user's historical login process, a fixed sampling frequency is used. (50–200Hz) in a short time window Within 5–10 seconds, the micro-fluctuation sequence of the device hardware operation characteristics in a short period of time is collected, including CPU clock offset sequence, memory access latency sequence, and GPU rendering time sequence. S2.42. Perform multifractal spectral analysis on each micro-wave sequence and calculate the scaling function. And extract the generalized Hurst exponent based on the scaling function. And the Strange Spectrum : in, ; (when hour, Obtained through limits or interpolation and Through the exist (The slope at the point is obtained by linear interpolation or limit calculation). , ; In the formula, Let be the order of the statistical moments. For time intervals, In time Collected hardware micro-fluctuation sequence values (such as CPU clock offset sequence, memory access latency sequence, GPU rendering time sequence). In time The collected hardware micro-fluctuation sequence values, The singularity index is the scaling function. For order The derivative of reflects the local scaling properties corresponding to different orders of moments. Indicates to Differentiation; to ensure the stability of multifractal spectrum estimation, the length of each micro-wave sequence should be greater than or equal to 200; if insufficient, extend the acquisition time or use interpolation methods; S2.43. Using a fractal stability adaptive confidence screening mechanism to evaluate the generalized Hurst exponent. The process involves correction (because micro-fluctuations such as CPU clock skew are easily affected by instantaneous system load or acquisition noise, directly calculating the generalized Hurst exponent may lead to unstable feature values. This fractal stability adaptive confidence method automatically selects statistically reliable orders through multi-scale sub-interval partitioning and statistical variance evaluation, and smooths the scaling function, thereby reconstructing a singular spectrum that truly reflects the inherent properties of the hardware after denoising. The final stable singular spectrum width and the corrected generalized Hurst exponent can accurately characterize the essential fractal structure of hardware fluctuations, ensuring the stability of feature extraction from the same device at different times (reducing false alarms) and amplifying the physical differences between different devices (improving the detection rate), providing a solid feature foundation for subsequent anomaly detection), and based on the corrected generalized Hurst exponent... Reconstructing the singular spectrum to calculate the stable singular spectrum width. The generalized Hurst exponent and stable singular spectral width As a multifractal feature of the micro-fluctuation sequence of the equipment; during the long-term use of the equipment, the above multifractal features will slowly drift with hardware aging and changes in the operating environment. Therefore, it is necessary to construct an aging drift model that changes over time. Furthermore, the generalized Hurst exponent is evaluated using a fractal stability adaptive confidence screening mechanism. Make corrections and base them on the corrected generalized Hurst exponent. Reconstructing the singular spectrum to calculate the stable singular spectrum width. This includes the following steps: Scale intervals are constructed based on micro-wave sequences (the range of scale parameters is set according to the length and wave characteristics of the micro-wave sequence). ,in, This indicates the minimum analytical scale, typically set to 2 or 3 sampling points. (representing the maximum analytical scale), based on the scaling function To evaluate different orders To assess the stability of the statistical results, the scale interval is divided into several sub-intervals (divided at logarithmic intervals). Multiple overlapping or non-overlapping sub-intervals ( The number of scale subintervals is usually taken as... And calculate the piecewise scaling function for each subinterval. ( For scale sub-interval indexes ); For the same order Statistical analysis was performed on the calculation results across different scale sub-intervals, and their variances were calculated. Through variance Measuring the order Statistical stability in multiscale estimation: When the order exceeds a preset threshold (e.g., 0.015), the order is determined to be unstable and removed from the valid set, and the remaining orders form the stable order set. ; and based on Further determine whether the number of orders meets the minimum stable statistical requirement: when the number of stable orders is lower than the preset minimum order threshold, in order to avoid the divergence of higher-order statistics under finite sample conditions, the original order is... The range of values is adaptively shrunk, The range of values is limited to the low to medium order interval (e.g. ), and reconstruct the stable order set (i.e., recalculate within this range) (and perform stability screening again) to ensure the statistical reliability of subsequent fractal spectrum calculations; Scaling function for preserving order Perform discrete second-order difference calculations to determine local curvature, and when the absolute value of the local curvature exceeds a set threshold (e.g., 0.05), perform neighborhood smoothing correction (specifically, for adjacent orders). Scale function value at Calculate discrete second-order difference and will As a local curvature index at that order; when When the curvature exceeds a preset curvature threshold (e.g., 0.05), it is considered that the scaling function at that order has abnormal fluctuations or estimation noise. In this case, neighborhood smoothing correction is performed on the point, that is, the scaling function values of its neighboring orders are used for weighted averaging and updating. For example, using Alternatively, a weighted smoothing method with weighted coefficients can be introduced to replace it, thereby suppressing the influence of local noise on the shape of the scaling function and ensuring... (ensuring continuity and smoothness in the order space, and maintaining the concave structure required for subsequent singular spectrum calculations) to guarantee The continuity and concave structure of the singular spectrum; After completing the neighborhood smoothing correction, based on the variance of each order... Construct the corresponding confidence weights ( To adjust the parameters, a value of 30 can be used to control the sensitivity of the weights to variance (the larger the variance, the smaller the weight). This is based on confidence weights. For the generalized Hurst exponent The weighted adjustment is performed to generate the corrected generalized Hurst exponent. (Specifically: for each order) The local generalized Hurst exponent calculated using subintervals of different scales (Depend on and Relationship export, , For the first The order of the scale subintervals Confidence weights For the first In each scale subinterval, the order scaling function The weighted average is then used to obtain the corrected generalized Hurst exponent. ); In the stable order set Based on the modified generalized Hurst exponent within the range Calculate the corresponding singularity index And construct a stable singular spectrum Generate stable singular spectral width and multifractal features; specifically: based on the modified generalized Hurst exponent Reconstructing the scaling function (Applicable to one-dimensional wave signals; adjustments can be made accordingly for other definitions.) The singularity index is calculated using numerical differentiation. And the singular spectrum is obtained using the Legendre transform. Then calculate the stable singular spectral width. As one of the multifractal characteristics of the device's micro-fluctuation sequence; and the aforementioned stable singular spectral width and the generalized Hurst index (Usually a representative order is chosen, such as...) These are collectively referred to as the measured values of the multifractal features under the current login event; S2.44, Replace the time variable with the login event sequence number. (When constructing an aging drift model, the cumulative operating time of the equipment should be used.) The time variable is used to describe the long-term trend of multifractal characteristics; however, the actual cumulative power-on time of the device cannot be obtained in a browser environment or a normal terminal environment, so the time variable is replaced by the device login event sequence number. That is, the device's first The login behavior is used as a discrete proxy variable for device usage time, and is based on the login event sequence number. Establish an aging drift model: In the formula, For the first The expected value of a certain multifractal feature (such as the generalized Hurst exponent or stable spectral width) at the time of the first login. The initial eigenvalues are, i.e. eigenvalues at time, The aging drift rate parameter controls how quickly the characteristic changes with logarithmic time. S2.45. Based on the aging drift model, according to the current login event sequence number... Predict the expected value of the multifractal features of the current device. And construct the confidence interval for this multifractal feature. : ; ; In the formula, The standard deviation of the historical fitting residuals. The quantile corresponding to the confidence level (e.g., taking 3 corresponds to the 99.7% prediction interval, used for anomaly detection); S2.46, For the current login serial number Repeat steps S2.41 to S2.43 to obtain the current measured values of the multifractal features. And the measured values of multifractal features confidence interval By comparing the results, the final multifractal features are obtained. ( (Index representing different multifractal features); the specific comparison steps are: if If the measured value is correct, then the actual value is used; otherwise, the equipment is considered to be in an abnormal state or the aging model needs to be updated, and the predicted value is used. Replace the measured values to ensure the stability of the fingerprint vector; S2.47, Based on multifractal features Generate the dynamic device fingerprint vector of the current device. Among them, dynamic device fingerprint vector It consists of four key fractal features that have been modified for stability and adjusted for aging drift: It represents the stable singular spectral width, which measures the multifractal intensity of hardware fluctuations; , , These are the modified generalized Hurst exponents at the moment order. , , The value at which, Capture subtle jitters within small fluctuations. Reflecting the overall fractal dimension properties, Emphasizing response behavior to large fluctuations; , , In fact, it is Specific examples (i.e.) , These four components (e.g., etc.) precisely cover the core information of the multifractal spectrum in a low dimension, which facilitates subsequent device identification and anomaly detection. S2.5. Generate historical time feature vectors based on the login time probability distribution model (encode the probability values of each time period in the login time probability distribution model into historical time feature vectors), generate spatial feature vectors based on the login spatial location cluster center model (encode the cluster center coordinates, cluster radius, and occurrence weight in the login spatial location cluster center model into spatial feature vectors), and concatenate the historical time feature vectors, spatial feature vectors, and dynamic device fingerprint vectors to generate a complete user behavior baseline vector.
[0019] S3. Within 1 second after the user completes login, compare the current login behavior feature vector with the user behavior baseline vector, calculate the time deviation, spatial deviation and device anomaly, and generate a login anomaly score to preliminarily determine whether the account (i.e. the account the user logged in) is at risk of being stolen. In this embodiment, the current login behavior feature vector is compared with the user behavior baseline vector to calculate the time deviation, spatial deviation, and device anomaly degree, and a login anomaly score is generated, including the following steps: S3.1 Obtain the historical login probability of the time period to which the current login time belongs based on the login time probability distribution model. (Map the current login time (e.g., 21:35) to a pre-divided time period, such as the time interval 21:00–22:00. Calculate the proportion of logins in each time period to the total number of logins based on the login time probability distribution model, thereby obtaining the historical login probability of the time period to which the current login time belongs.) Based on historical login probability Calculate time deviation (In the formula, To prevent taking a small constant whose logarithm is zero, this step uses the historical login probabilities of the time period to which the current login time belongs. As a metric for matching historical login habits, and based on this matching degree, a time deviation is calculated, and the time deviation is analyzed. Perform normalization (achieved through min-max normalization) to generate the normalized time deviation. ; S3.2 Extract the spatial coordinates of the current login IP after resolution from the current login behavior feature vector. And obtain the set of cluster centers for the user's historical login areas based on the login spatial location cluster center model. Spatial coordinates are calculated using the Haversine formula. With cluster center set Spatial distance between (In the formula, Let be the Earth's radius, and take . or ), Traversing spatial distance Select the minimum spatial distance and the corresponding cluster centers And calculate the spatial deviation. ( (where the cluster radius is used to standardize the impact of different region sizes), and spatial deviation. Perform normalization (achieved through min-max normalization) to generate the normalized spatial deviation. ; S3.3 Extract the device fingerprint feature vector from the current login behavior feature vector, and calculate the device fingerprint feature vector and the dynamic device fingerprint vector. cosine similarity (In the formula, To extract the device fingerprint feature vector from the current login behavior feature vector, The first in the historical login record (device fingerprint vector), traversing cosine similarity Select the maximum similarity Calculate the degree of equipment anomaly ; S3.4 Normalized time deviation Normalized spatial deviation Equipment abnormality Perform weighted fusion to generate login anomaly scores. (In the formula, As the weight for time deviation, For spatial deviation weight, As the weight for equipment anomaly degree, ).
[0020] S4. Calculate the resource access anomaly degree and access path jump anomaly degree based on the preprocessed key security data. Calculate the access rhythm anomaly degree by introducing the access rhythm and resource coupling coefficient. Generate an access anomaly score by weighted fusion. In this embodiment, resource access anomaly and access path jump anomaly are calculated based on preprocessed key security data. Access rhythm anomaly is calculated by introducing access rhythm and resource coupling coefficient. An access anomaly score is generated through weighted fusion. The steps include: S4.1, A monitoring time window after a user successfully logs in. Within a 30–120 second period, user access behavior is monitored in real time. User access resource sequences are extracted from preprocessed critical security data (first, the critical security data is sorted according to timestamps to filter out access records belonging to the current monitoring window; then, for each access record, access resource identifiers (such as URLs, API interfaces, file paths, or resource IDs) and access order information are obtained to form ordered access resource entries; then, all entries are arranged in order of access time, duplicate or invalid access records are removed, and user access resource sequences are generated based on statistical indicators such as access count or access duration). At the same time, a set of users' historical resource accesses is constructed based on historical access record data to characterize the range of resources frequently accessed by users in normal business behavior. S4.2 Construct an access time interval sequence within the current monitoring window based on the user access resource sequence. Calculate the rhythm deviation index and access interval variation coefficient based on the current access interval sequence, and simultaneously statistically analyze the current monitoring time window. The proportion of first-time accesses to internal resources is calculated, and a coupling coefficient between access rhythm and resources is constructed to generate access rhythm anomalies. ; In this embodiment, traditional network security monitoring typically focuses on deviations in a single dimension, such as accessing sensitive resources never accessed before (resource anomaly) or exhibiting unusual access sequences (path anomaly). However, after successfully logging in, attackers often employ slow, low-frequency stealth probing methods. Their individual resource accesses may not be high-risk, and single-step paths may occur occasionally, making detection rules based on isolated dimensions easily bypassable. Furthermore, the human-computer interaction behavior of attackers differs fundamentally from the business rhythm of normal users: the access intervals of automated scripts typically exhibit highly uniform or extremely sudden patterns, while human operations have natural fluctuations. This invention couples the behavioral rhythm in the time dimension with the resource exploration in the spatial dimension, achieving a deep characterization of user behavior patterns: by constructing an access rhythm and resource coupling coefficient (CEI), it cleverly uses the rhythm deviation index (RDI) to capture the overall speed offset, uses the coefficient of variation to quantify the natural fluctuation of operations, and introduces the proportion of first resource accesses to reflect the exploration intensity. Finally, the three are fused through a nonlinear product. This design allows attackers to mimic normal users at the resource access level (e.g., a low first-time access rate), but as long as the statistical characteristics of their access rhythm (e.g., overly uniform script behavior) do not match user habits, the coupling coefficient can still keenly amplify abnormal signals, thereby significantly improving the sensitivity and accuracy of identifying advanced threats such as slow detection and automated script attacks. Specifically, the rhythm deviation index and access interval variation coefficient are calculated based on the current access interval sequence. Simultaneously, the proportion of first-time resource accesses within the current window is statistically analyzed, and an access rhythm-resource coupling coefficient is constructed to generate the access rhythm anomaly degree. This includes the following steps: S4.21, in the monitoring window The system records the timestamp sequence of accessed resources, and constructs an access time interval sequence within the current monitoring window based on the timestamp sequence. Where, if... This indicates that the window contains only a single or no valid access records, and cannot form a time interval sequence, so let The calculation steps S4.22 to S4.26 are not executed; S4.22 Calculate the average access interval of the current window based on the access time interval sequence. ( For the first The second visit and the first (Time difference between visits), based on the average access interval of the current window. Constructing a rhythm deviation index ( This represents the average historical visit interval. (Standard deviation of historical access intervals) S4.23, Average access interval based on the current window Calculate the coefficient of variation of the current access interval sequence. ( (Standard deviation of the current window access interval). S4.24, Statistical Monitoring Window Number of resources accessed for the first time First-time access ratio of computing resources ; S4.25, Based on rhythm deviation index Coefficient of variation and the percentage of first-time visits to resources Build access rhythm and resource coupling coefficient In the formula, The modulation coefficient for rhythm instability ranges from 0.5 to 2.0 and is determined based on expert experience. The modulation coefficient for resource exploration ranges from 0.5 to 2.0 and is determined based on expert experience. S4.26. Integrating access frequency with resource coupling coefficient Mapped to access rhythm anomaly degree (In the formula, This is the threshold for anomaly detection, ranging from 1.0 to 3.0, based on historical normal access data. Distribution settings (e.g., 95th percentile). This is the slope control coefficient of the Sigmoid function, with a value ranging from 1 to 5, used to control the rate of change of the function near the threshold. S4.3 Match each accessed resource in the current user's accessed resource sequence with each accessed resource in the historical resource access set, and calculate the monitoring time window. The average rarity of internally accessed resources is used to calculate the resource access anomaly score. This is used to quantify the degree of anomaly in a user's access to new resources; specifically, it involves classifying each resource in the current user's resource access sequence as an example. Match the resource with the user's historical resource access set to obtain the number of times each resource was accessed in the user's history. (If the user's historical resource access set has never appeared) ,but ), for the current monitoring time window Calculate the rarity weight of each resource by analyzing the sequence of all user accesses within the resource. And calculate the average rarity of resources accessed within the window. (In the formula, This indicates the total number of times a user accesses a resource within the current monitoring window (i.e., the length of the user resource access sequence). Current monitoring time window (All user access resource sequences within), processed by the Sigmoid function Mapped to resource access anomaly degree (In the formula, This is the gain coefficient (or slope parameter), ranging from 1 to 5, determined through expert experience, and controls the steepness of the Sigmoid function near the threshold. The threshold parameter (with a value ranging from 0.1 to 0.5, determined using historical data) is used to quantify the degree of anomaly when a user accesses a rare new resource. S4.4 To identify abnormal access path changes caused by attackers moving laterally within the system, a user historical access transition probability matrix is constructed based on historical access record data. For the current access sequence, the path probability between adjacent accessed resources is calculated, and the path anomaly degree is constructed. Specifically, firstly, the access order of each resource accessed by the user in past sessions is statistically analyzed to construct a historical access transition frequency matrix. ,in, Indicates from resources Transfer to resources The historical number of visits; normalizing the transfer frequency matrix yields the historical access transfer probability matrix. ,in, This indicates that the user is accessing the resource. Transfer to resources Historical probability; access resource sequences extracted within the current monitoring window. Calculate adjacent access resource pairs in the sequence In the historical access transition probability matrix The corresponding transition probability Then, construct the path anomaly degree based on the transition probabilities of all adjacent resource pairs. ( To prevent taking a small constant whose logarithm is zero, if Not in the historical access transition probability matrix If it appears in the middle, then (Recorded as 0), used to quantify the degree of path deviation of the current access sequence relative to historical behavior, thereby reflecting possible lateral movement or abnormal access behavior; S4.5 Normalize the path anomaly score to obtain the access path hop anomaly score. (Using linear normalization method, ...) Mapping to interval Within this, the normalized access path jump anomaly degree is obtained. ); S4.6, Determine the abnormality of the access rhythm Resource access anomaly Access path jump anomaly degree Perform weighted fusion to construct a horizontal movement index ( As the weight for the abnormality of resource access, For path anomaly weights, Weighting for abnormal access rhythm. And use the lateral movement index as a score for access anomalies.
[0021] S5. Generate a comprehensive risk value based on login anomaly scoring and access anomaly scoring, which is used to trigger automatic blocking for information security management. In this embodiment, login anomaly scores and access anomaly scores are weighted and fused to form a comprehensive risk value. Through preset risk thresholds (e.g., 0.6, determined based on the distribution of historical risk scores) Risk assessment is performed, and automatic protection and automatic blocking strategies are triggered based on the risk assessment results to achieve real-time management of computer network information security (specifically: if...). If the current session risk is deemed acceptable, normal access is allowed; if If the system determines that the current session contains high-risk behavior, it will trigger the security protection mechanism and execute automatic blocking policies. Automatic blocking policies include session blocking: immediately terminating the current user session and forcibly logging out; access restriction: temporarily freezing access to sensitive resources or restricting operations; multi-factor authentication triggering: requiring the user to perform two-factor authentication (such as dynamic verification code or SMS / email verification); security alarm: sending a high-risk alarm to the security management system or operations and maintenance personnel, and recording detailed event logs, including login time, login IP, accessed resources, device fingerprint, and lateral movement index.
[0022] Example 2: This example provides a computer network information security management system based on data processing, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program to implement the computer network information security management method based on data processing described in Example 1 above.
[0023] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited thereto. Various changes can be made within the scope of knowledge possessed by those skilled in the art without departing from the spirit of the present invention.
Claims
1. A computer network information security management method based on data processing, characterized in that, include: S1. Collect key security data and historical security data of the user's current login session, and preprocess the key security data to form a feature vector of the current login behavior; S2. Based on historical security data, construct a login time probability distribution model, a login spatial location cluster center model, and a historical device fingerprint model for each user, and form a user behavior baseline vector; S3. Within 1 second after the user completes login, compare the current login behavior feature vector with the user behavior baseline vector, calculate the time deviation, spatial deviation and device anomaly, and generate a login anomaly score to preliminarily determine whether the account is at risk of being stolen. S4. Calculate the resource access anomaly degree and access path jump anomaly degree based on the preprocessed key security data. Calculate the access rhythm anomaly degree by introducing the access rhythm and resource coupling coefficient. Generate an access anomaly score by weighted fusion. S5. Generate a comprehensive risk value based on login anomaly scoring and access anomaly scoring, which is used to trigger automatic blocking for information security management.
2. The computer network information security management method based on data processing according to claim 1, characterized in that, In step S1, the key security data is preprocessed to form a feature vector of the current login behavior, including the following steps: The critical security data is preprocessed. Based on statistical feature analysis and embedding coding methods, the current login feature vector is extracted from the preprocessed critical security data. The extracted current login feature vector includes at least the current time feature vector, spatial location feature vector, and device fingerprint feature vector. The login feature vector is normalized. All the normalized login feature vectors are combined to form the final current login behavior feature vector, which serves as a representation of the user's current login behavior.
3. The computer network information security management method based on data processing according to claim 1, characterized in that, In step S2, a login time probability distribution model, a login spatial location cluster center model, and a historical device fingerprint model are constructed to form a user behavior baseline vector, including the following steps: S2.1 Preprocess historical security data; Historical security data includes user historical login data, historical access record data, and historical device fingerprint information; S2.2 Extract user historical login time information based on preprocessed user historical login data, calculate the historical login probability of each historical login time period, and generate a user login time probability distribution model; S2.3 Based on the preprocessed user historical login data, the K-means clustering algorithm is used to identify the user's frequently used login areas, generate spatial location cluster centers, and construct a login spatial location cluster center model; S2.4 Based on the preprocessed historical device fingerprint information, a hardware aging drift model and multifractal spectrum analysis are introduced to generate dynamic device fingerprint vectors. S2.
5. Generate historical time feature vectors based on the login time probability distribution model, generate spatial feature vectors based on the login spatial location cluster center model, and concatenate the historical time feature vectors, spatial feature vectors, and dynamic device fingerprint vectors to generate a complete user behavior baseline vector.
4. The computer network information security management method based on data processing according to claim 3, characterized in that, In step S2.4, based on the preprocessed historical device fingerprint information, a hardware aging drift model and multifractal spectrum analysis are introduced to generate a dynamic device fingerprint vector, including the following steps: S2.
41. During the user's historical login process, a fixed sampling frequency is used. In a short time window The micro-fluctuation sequence of the hardware operation characteristics of the internal acquisition device includes CPU clock offset sequence, memory access latency sequence, and GPU rendering time sequence; S2.
42. Perform multifractal spectral analysis on each micro-wave sequence and calculate the scaling function. And based on the scaling function Extracting the generalized Hurst exponent and the strange spectrum; S2.
43. Using a fractal stability adaptive confidence screening mechanism to evaluate the generalized Hurst exponent. Make corrections and base them on the corrected generalized Hurst exponent. Reconstructing the singular spectrum to calculate the stable singular spectrum width. The generalized Hurst exponent and stable singular spectral width As a multifractal feature of the micro-fluctuation sequence of the equipment; S2.44, Replace the time variable with the login event sequence number. And based on the login event sequence number Establish an aging drift model; S2.
45. Based on the aging drift model, predict the expected value of the multifractal characteristics of the current device according to the current login event sequence number. And construct the confidence interval for this multifractal feature; S2.
46. For the current login number, repeat steps S2.41 to S2.43 to obtain the current measured value of the multifractal feature. And the measured values of multifractal features The final multifractal features are obtained by comparing them with the confidence intervals. ; S2.47, Based on multifractal features Generate the dynamic device fingerprint vector of the current device. .
5. The computer network information security management method based on data processing according to claim 4, characterized in that, In S2.43, the generalized Hurst exponent is evaluated using a fractal stability adaptive confidence screening mechanism. Make corrections and base them on the corrected generalized Hurst exponent. Reconstructing the singular spectrum to calculate the stable singular spectrum width. This includes the following steps: Scale intervals are constructed based on micro-fluctuation sequences, and the scaling function is used to... The scale interval is divided into several sub-intervals, and the piecewise scaling function corresponding to each sub-interval is calculated separately. ; For the same order Statistical analysis was performed on the calculation results across different scale sub-intervals, and their variances were calculated. Through variance Measuring the order Statistical stability in multi-scale estimation forms a set of stable orders. ; and based on Further determine whether the number of its orders satisfies the minimum stability statistics requirement; Scaling function for preserving order Perform discrete second-order difference to calculate local curvature, and perform neighborhood smoothing correction when the absolute value of local curvature exceeds a set threshold; After completing the neighborhood smoothing correction, based on the variance of each order... Construct the corresponding confidence weights Based on confidence weights For the generalized Hurst exponent The weighted adjustment is performed to generate the corrected generalized Hurst exponent. ; In the stable order set Based on the modified generalized Hurst exponent within the range Calculate the corresponding singularity index And construct a stable singular spectrum Generate stable singular spectral width And multifractal features.
6. The computer network information security management method based on data processing according to claim 1, characterized in that, In step S3, the current login behavior feature vector is compared with the user behavior baseline vector to calculate the time deviation, spatial deviation, and device anomaly degree, and a login anomaly score is generated, including the following steps: S3.1 Obtain the historical login probability of the time period to which the current login time belongs based on the login time probability distribution model. Based on historical login probability Calculate time deviation and the degree of time deviation Perform normalization to generate the normalized time deviation. ; S3.2 Extract the current spatial coordinates from the current login behavior feature vector And obtain the set of cluster centers for the user's historical login areas based on the login spatial location cluster center model. Spatial coordinates are calculated using the Haversine formula. With cluster center set Spatial distance between Traverse spatial distance Select the minimum spatial distance and the corresponding cluster centers And calculate the spatial deviation. Spatial deviation Perform normalization to generate the normalized spatial deviation. ; S3.3 Extract the device fingerprint feature vector from the current login behavior feature vector, and calculate the device fingerprint feature vector and the dynamic device fingerprint vector. cosine similarity Traversing cosine similarity Select the maximum similarity Calculate the degree of equipment anomaly ; S3.4 Normalized time deviation Normalized spatial deviation Equipment abnormality Perform weighted fusion to generate login anomaly scores .
7. The computer network information security management method based on data processing according to claim 1, characterized in that, In step S4, resource access anomaly degree and access path jump anomaly degree are calculated based on preprocessed key security data. Access rhythm anomaly degree is calculated by introducing access rhythm and resource coupling coefficient. Access anomaly score is generated through weighted fusion. The steps include: S4.1, A monitoring time window after a user successfully logs in. Internally, it monitors user access behavior in real time, extracts user access resource sequences from preprocessed key security data, and constructs a set of users' historical resource accesses based on historical access record data. S4.2 Construct the current monitoring time window based on the user's resource access sequence The system calculates the rhythm deviation index and the access interval variation coefficient based on the current access interval sequence, and also counts the proportion of first accesses to resources within the current window. Finally, it constructs an access rhythm-resource coupling coefficient to generate the access rhythm anomaly degree. ; S4.3 Match each accessed resource in the current user's accessed resource sequence with each accessed resource in the historical resource access set, and calculate the monitoring time window. The average rarity of internally accessed resources is used to calculate the resource access anomaly score. ; S4.4 Construct a user historical access transition probability matrix based on historical access record data. For the current access sequence, calculate the path probability between adjacent access resources and construct the path anomaly degree. S4.5 Normalize the path anomaly score to obtain the access path hop anomaly score. ; S4.6, Determine the abnormality of the access rhythm Resource access anomaly Access path jump anomaly degree Perform weighted fusion to construct a horizontal movement index And use the lateral movement index as a score for access anomalies.
8. The computer network information security management method based on data processing according to claim 7, characterized in that, In step S4.2, the rhythm deviation index and the coefficient of variation of the access interval are calculated based on the current access interval sequence, and the current monitoring time window is statistically analyzed. The proportion of first-time accesses to internal resources is calculated, and a coupling coefficient between access rhythm and resources is constructed to generate access rhythm anomalies. This includes the following steps: S4.21, in the monitoring window The system records the timestamp sequence of accessed resources and constructs the access time interval sequence within the current monitoring window based on the timestamp sequence. S4.22 Calculate the average access interval of the current window based on the access time interval sequence. Based on the average access interval of the current window Constructing a rhythm deviation index ; S4.23, Average access interval based on the current window Calculate the coefficient of variation of the current access interval sequence. ; S4.24, Statistical Monitoring Window Number of resources accessed for the first time First-time access ratio of computing resources ; S4.25, Based on rhythm deviation index Coefficient of variation and the percentage of first-time visits to resources Build access rhythm and resource coupling coefficient ; S4.
26. Integrating access frequency with resource coupling coefficient Mapped to access rhythm anomaly degree .
9. The computer network information security management method based on data processing according to claim 1, characterized in that, In step S5, a comprehensive risk value is generated based on login anomaly scores and access anomaly scores. Includes the following steps: A comprehensive risk value is formed by weighting and combining login anomaly scores and access anomaly scores. Through preset risk thresholds Perform risk assessment, and trigger automatic protection and execute automatic blocking strategies based on the risk assessment results.
10. A computer network information security management system based on data processing, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: The processor executes a computer program to implement the computer network information security management method based on data processing as described in any one of claims 1-9.