High-frequency file reading and monitoring system based on log analysis
By introducing static and dynamic risk scoring factors into the high-frequency file reading monitoring system, the context-aware distance optimization processing is solved, and the problem of inability to effectively distinguish between benign and potential abnormal high-frequency reading behavior in the prior art is solved, and more accurate abnormal identification and monitoring effects are achieved.
Patent Information
- Application Number
- CN202510585010.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-08
AI Technical Summary
When monitoring high-frequency file reading, the prior art cannot effectively integrate the context information in risk assessment, resulting in the inability to accurately distinguish between benign and potential abnormal high-frequency reading behaviors, affecting the accuracy and reliability of the monitoring system.
By constructing structured file reading behavior examples, and introducing static risk scores and exponential adjustment factors, as well as dynamic mode risk scores and exponential amplification factors, context-aware distance optimization processing is performed, static and dynamic risk optimization factors are integrated for final optimization distance calculation, and behavior pattern recognition and abnormal analysis are used based on optimized distance.
It realizes more accurate identification of potential abnormal high-frequency reading behaviors, improves the ability to identify complex attack methods or abnormal operations of the system, improves the distinction and clustering stability of the monitoring system, and reduces false alarms and missed reports.
Smart Images

Figure CN120145376A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of log analysis and information security, and particularly to a high-frequency file reading monitoring system based on log analysis. Background Art
[0002] In various information systems, files are the carriers for storing core information such as configurations, data, and codes. Effectively monitoring the access activities to these files is a fundamental step in ensuring system security, conducting compliance checks, and diagnosing operation problems. The access logs generated during the system operation record in detail information such as the time when the file is accessed, the identity of the accessor, the program used, and the specific file path, providing the raw data for monitoring and analysis.
[0003] Analyzing these log data to identify abnormal file access patterns is a common operation in the fields of information security and system operation and maintenance. Among them, the phenomenon of high-frequency file reading has become one of the focuses of monitoring and analysis because it may correspond to normal business operations or may indicate potential risks (data theft attempts, malicious probes). Accurately distinguishing these two types of high-frequency reading behaviors with completely different natures is crucial for improving the effectiveness of the monitoring system.
[0004] To address this challenge, data clustering methods such as algorithms are often used in the prior art to perform pattern analysis on file reading behaviors. The general process of this method is: First, extract the features that can quantitatively describe the reading behaviors from the log data and construct them into a multi-dimensional feature vector; then apply the algorithm to divide the reading behavior instances into different clusters according to the similarity between the feature vectors; finally, by interpreting the formed clusters, it is expected to distinguish the pattern cluster representing normal operations from the pattern cluster indicating potential anomalies. Within the algorithm framework, the standard method for measuring the similarity between two behavior feature vectors is to calculate the standard Euclidean distance between them. This distance metric reflects the geometric proximity of the two vectors in the feature space, that is, the numerical similarity.
[0005] However, in the scenario of monitoring high-frequency file reading, directly applying the Clustering has a core technical defect: when measuring the similarity of file reading behaviors using the standard Euclidean distance, its calculation completely depends on the numerical differences of feature vectors and essentially cannot integrate context information closely related to risk assessment. This inherent lack of risk context is reflected in multiple aspects: it cannot judge the inherent sensitivity based on the file path (the access to system-critical files should be differentiated from the access to ordinary log files in terms of risk); it cannot utilize user identity or process source information to evaluate the credibility of the visitor (the behaviors of known service processes are different from those of unknown scripts in terms of risk); it is also difficult to distinguish the risk differences implied by the feature combination patterns (even for high-frequency reads, the risk of a stable single-service pattern is very different from that of a scattered multi-user probing pattern).
[0006] Since the standard Euclidean distance does not introduce these key information for determining the risk level, the calculated distance only reflects the geometric proximity of behaviors in apparent numerical features and cannot accurately reflect the true potential risk gap between two reading behaviors. The direct consequence is The clustering process cannot effectively distinguish between benign high-frequency reads and potentially abnormal or malicious high-frequency reads because behaviors with completely different risk connotations are wrongly grouped into one category due to partial similarity in feature values. This seriously affects the accuracy and reliability of subsequent anomaly detection, making the entire monitoring system prone to false alarms and missed detections and unable to meet the requirements of precise security monitoring. Therefore, there is an urgent need for an improved distance metric method that can overcome the limitations of the standard Euclidean distance and effectively incorporate key risk context information into the similarity assessment of file reading behaviors. Summary of the Invention
[0007] In view of this, the present invention aims to propose a high-frequency file reading monitoring system based on log analysis to...
[0008] To achieve the above object, the technical solution of the present invention is realized as follows: A high-frequency file reading monitoring system based on log analysis, the high-frequency file reading monitoring system includes the following steps: Step S1: Construct an access risk context through log collection and auxiliary library scoring, and perform behavior feature extraction and normalization processing to obtain a structured file reading behavior instance; Step S2: Through the introduction of static risk scores and exponential adjustment factors, perform context-aware distance optimization processing to obtain static risk optimization factors; Step S3: Through the construction of dynamic mode risk scores and exponential amplification factors, perform feature vector behavior pattern analysis to obtain dynamic risk optimization factors; Step S4: Through the fusion of static and dynamic risk optimization factors, perform final optimized distance calculation processing to obtain the similarity of behavior instances with comprehensive risk weighting; Step S5: Perform behavior pattern recognition and anomaly analysis by applying the K-Means clustering algorithm based on optimized distance to obtain risk warning information for frequent file reads.
[0009] Further, construct an access risk context through log collection and auxiliary library scoring, and perform behavior feature extraction and normalization processing to obtain structured file read behavior instances, specifically including: Continuously collect file access audit log data through a log proxy deployed in the target information system. The collected file access audit log entries include a timestamp, the file path accessed, the user identifier performing the access operation, the process identifier initiating the access, and the operation type. In the central log processing system, the log data is initially screened to retain only the log records with the operation type of "read". At the same time, a sensitive file path pattern library and a visitor credibility information library are loaded into the system. The sensitive file path pattern library defines sensitive paths and their corresponding sensitivity scores through regular expressions, and the visitor credibility information library records the credibility scores corresponding to user identifiers and process identifiers. After completing log screening and auxiliary library loading, the system divides the "read" type logs into non-overlapping time windows of a fixed length in chronological order. In each time window, the system aggregates the unique file paths that have been read. The aggregation processing includes counting the number of times the path is read, the number of independent user IDs, the number of independent process IDs, and calculating the standard deviation of the time intervals of all read operation timestamps; taking the multi-dimensional features in the aggregation result as the original feature vector of the file in the time window; normalizing the original feature vector of the file in the time window to obtain a normalized feature vector. For each file and each time window, use the normalized feature vector of the file in the time window, the file path, and the user information and process information for evaluating visitor risk as a structured file read behavior instance.
[0010] Further, perform context-aware distance optimization processing by introducing static risk scores and exponential adjustment factors to obtain a static risk-weighted behavior distance, specifically including: Perform static risk assessment on structured file read behavior instances through file sensitivity scores and visitor credibility scores to obtain a first static risk assessment factor. Perform relative risk assessment between structured file read behavior instances to obtain a static risk optimization factor.
[0011] Further, perform static risk assessment on structured file read behavior instances through file sensitivity scores and visitor credibility scores to obtain a first static risk assessment factor, specifically including: Obtain an instance of structured file reading behavior. For the instance of structured file reading behavior, read the accessed file path in the instance, the visitor who performs the operation, and the process information; obtain the set file sensitivity scoring function and visitor credibility scoring function; input the visitor and process information of the operation performed in the structured file reading behavior instance into the visitor credibility scoring function, and use the output as the visitor credibility degree of the instance; input the file path in the structured file reading behavior instance into the file sensitivity scoring function, and use the corresponding output as the file sensitivity degree of the instance; perform a multiplication operation on the calculation result of subtracting the constant 1 from the visitor credibility degree of the instance and the file sensitivity degree, and use the corresponding calculation result as the first static risk assessment factor.
[0012] Furthermore, perform relative risk assessment on the instances of structured file reading behavior to obtain a static risk optimization factor, specifically including: Obtain any two instances of structured file reading behavior, and use the larger first static risk assessment factor among the first static risk assessment factors corresponding to the two instances of structured file reading behavior as the static risk assessment of the two instances of structured file reading behavior; use the mapping result of mapping the static risk assessment with the exponential function with the natural constant e as the base as the static risk optimization factor.
[0013] Furthermore, construct a dynamic pattern risk score and an exponential amplification factor, perform eigenvector behavior pattern analysis, and obtain a dynamic risk optimization factor, specifically including: Perform dynamic pattern risk analysis on the instances of structured file reading behavior through the assessment of the dispersion of access sources and the instability of behaviors based on normalized eigenvectors to obtain a first dynamic pattern risk assessment factor; Obtain a dynamic risk optimization factor through the dynamic risk comparison assessment between the instances of structured file reading behavior.
[0014] Furthermore, perform dynamic pattern risk analysis on the instances of structured file reading behavior through the assessment of the dispersion of access sources and the instability of behaviors based on normalized eigenvectors to obtain a first dynamic pattern risk assessment factor, specifically including: Obtain the normalized eigenvector of the structured file reading behavior instance. Extract the feature dimensions preset to represent the number of independent access users and the number of independent access processes in the normalized eigenvector, and assign them as the user access degree eigenvalue and the process access degree eigenvalue respectively; obtain the weighting coefficients for calculating the dispersion of access sources, multiply the user access degree eigenvalue and the process access degree eigenvalue by the corresponding weighting coefficients respectively, and divide the sum of the products by the sum of the weighting coefficients. The obtained value is used as the dispersion of access sources for this instance; at the same time, extract the feature dimension preset to represent the standard deviation of reading frequency from the normalized eigenvector, assign it as the behavior instability eigenvalue, and use this eigenvalue as the behavior instability of this instance; perform a multiplication operation on the dispersion of access sources and the behavior instability, and use the corresponding calculation result as the first dynamic mode risk assessment factor.
[0015] Further, through the dynamic risk comparison and assessment between structured file reading behavior instances, a dynamic risk optimization factor is obtained, which specifically includes: Obtain any two structured file reading behavior instances, and use the larger first dynamic mode risk assessment factor among the first dynamic mode risk assessment factors corresponding to the two structured file reading behavior instances as the dynamic risk assessment value of the two structured file reading behavior instances; use the mapping result of the exponential function with the natural constant e as the base of the dynamic risk assessment value as the dynamic risk optimization factor.
[0016] Further, through the fusion of static and dynamic risk optimization factors for the final optimization distance calculation process, the similarity of behavior instances after comprehensive risk weighting is obtained, which specifically includes: Obtain any two structured file reading behavior instances, and use the calculation result of multiplying the static risk optimization factor and the dynamic risk optimization factor of the two structured file reading behavior instances as the distance optimization factor of the two structured file reading behavior instances; Obtain the normalized eigenvectors of the two structured file reading behavior instances respectively, and use the Euclidean distance between the normalized eigenvectors of the two structured file reading behavior instances as the basic distance of the two structured file reading behavior instances; Use the calculation result of multiplying the distance optimization factor of the two structured file reading behavior instances by the basic distance of the two structured file reading behavior instances as the similarity of behavior instances after comprehensive risk weighting.
[0017] Compared with the prior art, the present invention has the following advantages: The high-frequency file reading monitoring system based on log analysis described in the present invention can comprehensively integrate the importance of the file itself and the dynamic characteristics of access behavior by constructing structured file reading behavior instances and introducing static risk factors and dynamic risk factors to double-optimize the behavior similarity measurement on this basis, so as to achieve more accurate identification of potential abnormal high-frequency reading behaviors. The present invention first quantifies the static sensitivity of the file and the credibility of the visitor through the sensitive path scoring function and the visitor credibility scoring function, and constructs a static risk assessment factor and an exponential static risk optimization factor, making up for the defect that the traditional distance measurement ignores the static risk background; further, the present invention combines and models the dispersion of access sources and the instability of access frequencies in the behavior feature vector to form a dynamic pattern risk score and constructs a corresponding dynamic risk optimization factor, so as to effectively capture the "dynamic abnormal patterns hidden under the data structure representing normal behaviors", significantly improving the recognition ability of complex attack means or system abnormal operations. By jointly applying the static risk optimization factor and the dynamic risk optimization factor to the standard Euclidean distance calculation process, the present invention realizes a risk-aware enhanced distance measurement method, which has higher discrimination and clustering stability in K-Means clustering analysis, effectively reducing phenomena such as mis-clustering and missed judgment; at the same time, the present invention can automatically identify and alarm high-risk clusters, loose clusters and outlier instances, significantly improving the security monitoring ability for high-frequency access behaviors in large-scale systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings: Figure 1 It is a flowchart of the method of the high-frequency file reading monitoring system based on log analysis according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0020] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "upper", "lower", "inner", "back", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present invention. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0021] The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.
[0022] See Figure 1 , which is the method flow chart of the high-frequency file reading monitoring system based on log analysis provided by the first embodiment of the present invention. As Figure 1 shown, the high-frequency file reading monitoring system based on log analysis may include: S1, constructing an access risk context through log collection and auxiliary library scoring, and performing behavior feature extraction and normalization processing to obtain a structured file reading behavior instance.
[0023] Through the log agent deployed in the target information system, continuously collect file access audit log data. The collected file access audit log entries include a timestamp, the file path being accessed, the user identifier performing the access operation, the process identifier initiating the access, and the operation type; In the central log processing system, the log data is preliminarily filtered, and only the log records with the operation type of "reading" are retained. At the same time, the sensitive file path pattern library and the visitor credibility information library are loaded into the system. The sensitive file path pattern library is the sensitive paths defined by regular expressions and their corresponding sensitivity scores, and the visitor credibility information library is the credibility scores corresponding to the user identifier and the process identifier; Among them, the sensitive file path pattern library contains a series of predefined regular expressions for matching sensitive file or directory paths, and associates a fixed sensitivity score with each pattern. In this embodiment, the library contains the following example rules and scores: "^ / etc / shadow$" -> 1.0, "^ / etc / passwd$" -> 0.9, ".*\\.key$" -> 0.8, "^ / var / log / secure$" -> 0.7, " / database_secrets / .*" -> 0.9, other default paths -> 0.1. Based on this library, a file sensitivity scoring function is implemented: for the input file path, this function matches the regular expressions in the library in order of priority and returns the sensitivity score corresponding to the first successfully matched rule; if no rule matches, the default score of 0.1 is returned. Its specific path scoring is determined according to the actual usage scenario and is not required.
[0024] The visitor credibility information library integrates user trust level information and process trust level information. The user trust level information stores the mapping from user identifiers to trust scores. In this embodiment, the visitor credibility is set as follows: root -> 0.2, service_backup -> 0.9, user_admin -> 0.4, and other ordinary users -> 0.5. The process trust level information stores the mapping from process identifiers proc to trust scores ( / usr / sbin / sshd -> 0.9, / usr / bin / backup_agent -> 0.9, / bin / bash -> 0.6, / tmp / .* (matching processes executed under / tmp) -> 0.1). Based on this library, a visitor credibility scoring function is obtained: for the input user name and process name , this function separately looks up the corresponding user trust score and process trust score, and then calculates the final visitor credibility score as the minimum of the two. This strategy ensures that the combined credibility is determined by the least trustworthy part. Its specific score is also determined according to the actual usage scenario and is not required.
[0025] After completing the log filtering and auxiliary library loading, the system divides the "read" type logs into non-overlapping time windows of fixed length in chronological order. In each time window, the system aggregates the unique file paths that have been read. The aggregation process includes counting the number of times the path is read, the number of independent user IDs, the number of independent process IDs, and calculating the standard deviation of the time intervals of all read operation timestamps; taking the multi-dimensional features in the aggregation result as the original feature vector of the file in the time window; normalizing the original feature vector of the file in the time window to obtain the normalized feature vector; For each file and each time window, the normalized feature vector of the file in the time window where the file is located, the file path, and the user information and process information used to evaluate the visitor risk are used as a structured file read behavior instance.
[0026] S2. By introducing a static risk score and an exponential adjustment factor, perform context-aware distance optimization processing to obtain a static risk optimization factor.
[0027] When calculating the feature vectors (denoted as and ) of two file read behavior instances (denoted as and When calculating the distance between these vectors, only the geometric proximity of these vectors in the numerical feature space is considered. However, this calculation method completely ignores the static background information of the behavior, which is crucial for evaluating potential risks. Specifically, the standard distance fails to reflect the sensitivity of the accessed file itself and the credibility of the user and process performing the read operation. A read behavior on the system password file, even if its numerical features such as the number of read times and the number of accessing users are exactly the same as those of a read behavior on an ordinary application log file, has a significantly higher inherent security risk level. Similarly, a high-frequency read performed by a known and trusted system backup service is quite different in terms of risk connotation from a high-frequency read with the same numerical features performed by a script with an unknown source running in a temporary directory. The standard Euclidean distance cannot capture these key differences based on the static context, resulting in behaviors with very different risks being wrongly grouped together during clustering.
[0028] To address the blind spot of the standard Euclidean distance in the static risk context, this step first focuses on calculating a static risk score , which aims to quantify the potential risk level inherent in each instance of file read behavior and does not change with short-term behavior patterns. The calculation of this score requires two types of key context information, which are obtained from the log records themselves and the associated knowledge base: File path information : Used to judge the sensitivity level of the accessed file. By matching or querying with a predefined library of sensitive file path patterns (including rules for key directories of the operating system, key storage locations, directories storing user privacy data, etc.), a sensitivity score is assigned to each file path . This score reflects the importance or confidentiality level of the file content.
[0029] Visitor information (user identification and process identification ): Used to evaluate the trustworthiness of the entity performing the read operation. By querying the user permission database and the known process reputation database (such as marking known system services, trusted third-party applications, potential malware, or unknown processes), a credibility score can be assigned to the combination of the user and the process . This score reflects the reliability of the access source.
[0030] Based on the obtained file sensitivity score and visitor credibility score, a comprehensive static risk score is calculated This calculation logic follows a core principle: when the two conditions of accessing highly sensitive files and low-trust visitors are both met or tend to be met, the static risk score should increase significantly. Conversely, for accessing low-sensitivity files or operations performed by high-trust visitors, the static risk score is relatively low. This design ensures that it is able to effectively quantify the baseline risk level determined by static background factors.
[0031] After obtaining the static risk scores and for each behavior instance and respectively, the goal of this step is to calculate a static risk adjustment factor which will be used to adjust the and baseline Euclidean distance between them.
[0032] Specifically, first, a static risk assessment of the structured file reading behavior instance is performed through the file sensitivity score and the visitor credibility score to obtain the first static risk assessment factor; After that, a relative risk assessment is performed between the structured file reading behavior instances to obtain the static risk optimization factor.
[0033] Among them, performing a static risk assessment of the structured file reading behavior instance through the file sensitivity score and the visitor credibility score to obtain the first static risk assessment factor specifically includes: Obtain the structured file reading behavior instance. For the structured file reading behavior instance, obtain the accessed file path, the visitor and process information of the operation performed; obtain the set file sensitivity scoring function and the visitor credibility scoring function; take the output obtained by inputting the visitor and process information of the operation performed in the structured file reading behavior instance into the visitor credibility scoring function as the visitor credibility degree of the instance; input the file path in the structured file reading behavior instance into the file sensitivity scoring function, and take the corresponding output as the file sensitivity degree of the instance; perform a multiplication operation on the calculation result of subtracting the constant 1 from the visitor credibility degree of the instance and the file sensitivity degree, and take the corresponding calculation result as the first static risk assessment factor.
[0034] In an embodiment, assume that the visitor credibility degree of the th structured file reading behavior instance is ; the file sensitivity degree of the th structured file reading behavior instance is , then the calculation expression of the first static risk assessment factor of the th structured file reading behavior instance is: Among them, represents the first static risk assessment factor of the th structured file reading behavior instance; represents the file sensitivity of the th structured file reading behavior instance; represents the credibility of the visitor of the th structured file reading behavior instance.
[0035] It should be noted that the designed in this way (the value range is in ) can directly quantify the basic risk level determined by the static background, because its calculation logic ensures that the risk score will only increase significantly when a highly sensitive file is operated by a low-credibility visitor. It directly makes up for the lack of standard Euclidean distance in this key information.
[0036] After quantifying the static risk scores of each behavior instance , these scores can be used to construct a static risk optimization factor for adjusting the distance. Specifically, by performing a relative risk assessment between structured file reading behavior instances, a static risk optimization factor is obtained, which specifically includes: Obtain any two structured file reading behavior instances, and use the larger first static risk assessment factor among the first static risk assessment factors corresponding to the any two structured file reading behavior instances as the static risk assessment of the two structured file reading behavior instances; use the mapping result of mapping the static risk assessment with the exponential function with the natural constant e as the base as the static risk optimization factor.
[0037] In one embodiment, the calculation expression of the static risk optimization factor of the th structured file reading behavior instance and the th structured file reading behavior instance is: Among them, represents the static risk optimization factor of the th structured file reading behavior instance and the th structured file reading behavior instance; represents the first static risk assessment factor of the th structured file reading behavior instance; represents the first static risk assessment factor of the th structured file reading behavior instance.
[0038] It should be noted that the static risk adjustment factor constructed through the above logic , the two key static risk context information, file sensitivity and visitor credibility, which were originally ignored by the standard distance, are successfully integrated into the distance measurement in a quantitative, high-risk-focused and non-linear amplification manner, thus overcoming the core defects of the existing technology in this regard and laying the foundation for subsequent more accurate file reading behavior clustering analysis.
[0039] S3, by constructing dynamic model risk scores and index amplification factors, characteristic vector behavior pattern analysis is performed to obtain dynamic risk optimization factors.
[0040] Although the static risk adjustment factor introduced in the previous step Taking into account the two key static risk contexts of file sensitivity and visitor trustworthiness corrects the shortcomings of the standard Euclidean distance in this regard, but it does not completely solve the limitations of distance metrics in risk assessment. Specifically, even if two file reading behaviors and Having similar static risk backgrounds, their dynamic behavior patterns during execution are reflected in their feature vectors and There may still be significant differences in the specific value distribution and combination of the data, and these differences themselves may contain different potential risks. For example, a stable high-frequency read executed by a single process (common in benign background tasks) and a high-frequency read involving multiple different users and processes with drastic frequency fluctuations (more likely to be related to anomaly detection or system failure) have completely different risk implications revealed by their dynamic patterns, even if the static risks are similar. The adjusted distance metric is still based on the Euclidean distance, so it is still not sensitive enough to identify the risk differences indicated by the specific pattern composed of feature combinations, resulting in dynamic patterns with very different risks being mistakenly judged as similar.
[0041] In order to make up for this shortcoming, that is, to solve the problem of insufficient distinguishing ability due to failure to effectively evaluate the risk of dynamic behavior patterns themselves when static risks are similar, the present invention further optimizes the distance metric. This requires the introduction of a second adjustment factor, namely the dynamic risk adjustment factor The core function of this factor is to deeply analyze the feature vector of file reading behavior It is necessary to first identify and quantify the key dynamic features that can indicate pattern risks, such as the dispersion of access sources and the instability of behavior. High dispersion and high instability, especially when they occur simultaneously, are often associated with higher potential risks.
[0042] Based on this understanding, the feature vector of each behavior instance needs to be Calculate a dynamic pattern risk score The design of this scoring function needs to comprehensively consider the above - mentioned relevant dynamic feature dimensions, aiming to quantify the degree to which this behavior pattern deviates from the "regular, stable, concentrated" state, or rather, the abnormality presented by its pattern itself. For example, the scoring logic should ensure that when the feature vector shows both high dispersion and high instability simultaneously, the score is significantly higher than the case where only a single feature is abnormal or the pattern is stable.
[0043] To construct such an effective dynamic risk adjustment factor , the primary task is to calculate a quantitative indicator that can reflect the dynamic pattern risk for the feature vector of each behavior instance within this minute - time window, which is defined as the dynamic pattern risk score . The calculation of this score must be based on the key dimensions and their combinations in the feature vector that can indicate pattern abnormality. To eliminate the influence of different feature scale differences, these dimensions need to be normalized (mapped to the interval) before calculation. Denote the normalized feature vector as , In this embodiment, the calculation of the feature vector is based on a preset time window, and the size of the time window is set to 5 minutes. The key feature aspects that can reflect the dynamic pattern risk include: Dispersion of access sources : It can be measured by features such as the normalized number of independent users , the number of independent processes , etc. In this embodiment, it is designed as the weighted average of these features: , where is the preset weight coefficient. A high value indicates that the access sources are extensive within this minute window, and the pattern is more inclined to be dispersed.
[0044] Temporal instability of behavior : It is measured by the standard deviation of the normalized read frequency : . A high value indicates that the behavior is irregular and sudden within this minute window, and the pattern is more inclined to be unstable.
[0045] Since the risk of the dynamic pattern often lies in the co - occurrence of multiple abnormal features (for example, being both dispersed and unstable within the same minute window is usually more suspicious than being only dispersed or only unstable), the dynamic pattern risk score The calculation should be able to reflect this combined effect. By multiplying the indicators reflecting different risk aspects, the final risk score will only increase significantly when multiple indicators are all high. Based on this consideration, The calculation formula of is designed as: represents the first dynamic mode risk assessment factor of the normalized eigenvector of the th structured file reading behavior instance; represents the assessment of the dispersion of the access source of the th structured file reading behavior instance; represents the assessment of the instability of the behavior of the th structured file reading behavior instance.
[0046] It should be noted that the value calculated by this formula ( and are both within the interval, and is also within the interval) can quantify the risk tendency of the dynamic mode shown by the eigenvector within this
[0047] minute time window, especially giving a higher score to high-risk modes such as "dispersed and unstable", thus solving the problem that the existing distance metric is insensitive to the risk of such modes. After successfully quantifying the dynamic mode risk scores of the eigenvectors of each behavior instance within the corresponding minute window, these scores can be used to construct a dynamic risk adjustment factor
[0048] for further adjusting the distance.
[0049] In one embodiment, the calculation expression of the dynamic risk optimization factor of the th structured file reading behavior instance and the th structured file reading behavior instance is: where represents the The dynamic risk optimization factor between the th structured file reading behavior instance and the th structured file reading behavior instance; The dynamic risk optimization factor between the th structured file reading behavior instance and the th structured file reading behavior instance; Denote the first dynamic mode risk assessment factor of the normalized eigenvector of the th structured file reading behavior instance; Denote the first dynamic mode risk assessment factor of the normalized eigenvector of the th structured file reading behavior instance; Denote the first dynamic mode risk assessment factor of the normalized eigenvector of the th structured file reading behavior instance; Denote the first dynamic mode risk assessment factor of the normalized eigenvector of the th structured file reading behavior instance.
[0050] It should be noted that similar to the construction logic of the static risk adjustment factor, considering that it is necessary to focus on the modes with higher risks and adopt a non-linear amplification mechanism, The calculation of is based on the two behaviors participating in the comparison and (the risk scores corresponding to their respective time windows) with higher dynamic mode risk scores among them, and the natural exponential function is applied for processing. This ensures that as long as there is one behavior with a higher dynamic mode risk among the two, the adjustment effect will be enhanced, and the higher the risk, the faster the adjustment amplitude (amplification factor) grows.
[0051] S4. By fusing the static and dynamic risk optimization factors, perform the final optimized distance calculation process to obtain the behavior instance similarity after comprehensive risk weighting.
[0052] Obtain any two structured file reading behavior instances, and use the calculation result of multiplying the static risk optimization factor and the dynamic risk optimization factor of these two structured file reading behavior instances as the distance optimization factor of these two structured file reading behavior instances; Obtain the normalized eigenvectors of the two structured file reading behavior instances respectively, and use the Euclidean distance between the normalized eigenvectors of the two structured file reading behavior instances as the basic distance of these two structured file reading behavior instances; Use the calculation result of multiplying the distance optimization factor of the two structured file reading behavior instances by the basic distance of these two structured file reading behavior instances as the behavior instance similarity after comprehensive risk weighting.
[0053] S5. By applying the K-Means clustering algorithm based on the optimized distance for behavior pattern recognition and anomaly analysis, obtain the risk warning information for high-frequency file reading.
[0054] Obtain the behavior instance similarity after comprehensive risk weighting of any two structured file reading behavior instances, and use the behavior instance similarity between these two instances as the optimized distance between these two instances.
[0055] Complete the optimized distance between behavior instances After the calculation of The clustering algorithm is applied to perform pattern recognition on file reading behaviors, thereby achieving the monitoring of high-frequency readings. Specifically, the steps For a specific time window All file reading behavior instances output are used as the input data of the algorithm. During the iteration process, the centroid of each cluster is obtained by calculating the arithmetic mean of the normalized feature vectors of all data points currently assigned to that cluster (ensuring effective operation even when dealing with the average centroid without native context information).
[0056] Before applying the algorithm, the number of clusters, i.e., the value, needs to be determined. The selection of the value adopts the commonly used "elbow method", which belongs to the existing method and will not be elaborated here.
[0057] Execute the algorithm and perform iterative calculations using the optimized distance until the cluster assignment converges. Finally, behavior instances are divided into clusters . Subsequently, the clustering results are analyzed to identify potential abnormal high-frequency reading patterns. The focus of the analysis is to identify clusters whose characteristics are significantly different from known normal patterns, or outliers that do not belong to any stable cluster. Clusters with an abnormally small scale, an overly large average optimized distance between internal instances (i.e., loose within the cluster), or an average risk score of instances within the cluster ( and ) significantly higher than other clusters represent a very likely presence of abnormal behavior. Similarly, a single behavior instance far from all cluster centroids is also regarded as a potential anomaly.
[0058] When a certain cluster or an outlier instance is determined to be abnormal according to the above rules, the system generates a corresponding alarm. The alarm information includes the associated file path , visitor information and behavior characteristics to support subsequent security investigations. In this way, the present invention improves the application effect of clustering in high-frequency file reading monitoring by using the optimized distance, and improves the accuracy of distinguishing between benign and abnormal behaviors.
[0059] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A high-frequency file reading monitoring system based on log analysis, characterized in that: The high-frequency file reading monitoring system comprises the following steps: Step S1: construct access risk context through log collection and auxiliary library scoring, extract and normalize behavior features, and obtain structured file reading behavior instances; Step S2: By introducing a static risk score and an exponential adjustment factor, context-aware distance optimization processing is performed to obtain a static risk optimization factor; Step S3: By constructing a dynamic model risk score and an index amplification factor, characteristic vector behavior pattern analysis is performed to obtain a dynamic risk optimization factor; Step S4: By integrating the static and dynamic risk optimization factors, the final optimization distance calculation process is performed to obtain the behavior instance similarity after comprehensive risk weighting; Step S5: By applying the K-Means clustering algorithm based on optimized distance to perform behavior pattern recognition and anomaly analysis, risk warning information of high-frequency file reading is obtained.
2. The high-frequency file reading monitoring system based on log analysis according to claim 1 is characterized in that: According to the access risk context constructed through log collection and auxiliary library scoring and the behavior feature extraction and normalization processing, a structured file reading behavior instance is obtained, which specifically includes: Through the log agent deployed in the target information system, file access audit log data is continuously collected. The collected file access audit log entries include timestamp, accessed file path, user ID performing access operation, process ID initiating access and operation type; In the central log processing system, log data is initially screened, and only log records with the operation type of "read" are retained. At the same time, the sensitive file path pattern library and the visitor credibility information library are loaded into the system. The sensitive file path pattern library is the sensitive path defined by regular expressions and its corresponding sensitivity score, and the visitor credibility information library is the credibility score corresponding to the user ID and process ID. After completing the log screening and auxiliary library loading, the system divides the "read" type log into non-overlapping time windows of fixed length in chronological order. In each time window, the system aggregates the unique file paths that have been read, and the aggregation includes counting the number of reads of the path, the number of independent user IDs, the number of independent process IDs, and calculating the standard deviation of the time interval of all read operation timestamps; using the multidimensional features in the aggregation result as the original feature vector of the file in the time window; normalizing the original feature vector of the file in the time window to obtain the normalized feature vector; For each file and each time window, the normalized feature vector in the time window where the file is located, the file path, and the user information and process information used to assess the visitor risk are used as structured file reading behavior instances.
3. The high-frequency file reading monitoring system based on log analysis according to claim 1 is characterized in that: According to the above, by introducing static risk scores and exponential adjustment factors, context-aware distance optimization processing is performed to obtain static risk-weighted behavioral distances, specifically including: Perform static risk assessment on the structured file reading behavior instance through file sensitivity score and visitor credibility score to obtain a first static risk assessment factor; By performing relative risk assessment on structured file reading behavior instances, a static risk optimization factor is obtained.
4. The high-frequency file reading monitoring system based on log analysis according to claim 3 is characterized in that: According to the static risk assessment of the structured file reading behavior instance by using the file sensitivity score and the visitor credibility score, a first static risk assessment factor is obtained, which specifically includes: Obtain a structured file reading behavior instance, and for the structured file reading behavior instance, read the accessed file path in the instance, the visitor and process information of the operator performing the operation; obtain the set file sensitivity scoring function and visitor credibility scoring function; input the visitor and process information of the operator performing the operation in the structured file reading behavior instance into the visitor credibility scoring function to obtain the output as the visitor credibility of the instance; input the file path in the structured file reading behavior instance into the file sensitivity scoring function, and use the corresponding output as the file sensitivity of the instance; multiply the calculation result of subtracting the constant 1 from the visitor credibility of the instance by the file sensitivity, and use the corresponding calculation result as the first static risk assessment factor.
5. The high-frequency file reading monitoring system based on log analysis according to claim 1 is characterized in that: According to the relative risk assessment between the structured file reading behavior instances, a static risk optimization factor is obtained, which specifically includes: Obtain any two structured file reading behavior instances, and use the larger first static risk assessment factor among the first static risk assessment factors corresponding to the any two structured file reading behavior instances as the static risk assessment of the two structured file reading behavior instances; and use the mapping result of the static risk assessment using an exponential function with a natural constant e as the base as the static risk optimization factor.
6. The high-frequency file reading monitoring system based on log analysis according to claim 1 is characterized in that: According to the above, by constructing the dynamic model risk score and the index amplification factor, the characteristic vector behavior pattern analysis is performed to obtain the dynamic risk optimization factor, which specifically includes: By evaluating the access source dispersion and behavior instability based on the normalized feature vector, a dynamic pattern risk analysis is performed on the structured file reading behavior instance to obtain a first dynamic pattern risk assessment factor; By performing dynamic risk comparison and evaluation between structured file reading behavior instances, a dynamic risk optimization factor is obtained.
7. The high-frequency file reading monitoring system based on log analysis according to claim 6 is characterized in that: According to the access source dispersion and behavior instability assessment based on the normalized feature vector, a dynamic mode risk analysis is performed on the structured file reading behavior instance to obtain a first dynamic mode risk assessment factor, which specifically includes: A normalized feature vector of a structured file reading behavior instance is obtained, and a feature dimension preset to represent the number of independent access users and a feature dimension representing the number of independent access processes in the normalized feature vector are extracted, and they are assigned as a user access degree feature value and a process access degree feature value, respectively; a weighting coefficient for calculating the access source dispersion is obtained, and the user access degree feature value and the process access degree feature value are multiplied by the corresponding weighting coefficients, and the sum of the products is divided by the sum of the weighting coefficients, and the obtained value is used as the access source dispersion of the instance; at the same time, a feature dimension preset to represent the standard deviation of the reading frequency is extracted from the normalized feature vector, and it is assigned as a behavioral instability feature value, and the feature value is used as the behavioral instability of the instance; the access source dispersion and the behavioral instability are multiplied, and the corresponding calculation result is used as the first dynamic mode risk assessment factor.
8. The high-frequency file reading monitoring system based on log analysis according to claim 6 is characterized in that: According to the above, by performing dynamic risk comparison and evaluation between structured file reading behavior instances, a dynamic risk optimization factor is obtained, which specifically includes: Obtain any two structured file reading behavior instances, and use the larger first dynamic mode risk assessment factor among the first dynamic mode risk assessment factors corresponding to the two structured file reading behavior instances as the dynamic risk assessment value of the two structured file reading behavior instances; and use the mapping result of the dynamic risk assessment value using an exponential function with a natural constant e as the base as the dynamic risk optimization factor.
9. The high-frequency file reading monitoring system based on log analysis according to claim 1 is characterized in that: According to the above, the final optimization distance calculation process is performed by fusing the static and dynamic risk optimization factors to obtain the behavior instance similarity after comprehensive risk weighting, which specifically includes: Obtaining any two structured file reading behavior instances, and using the calculation result of multiplying the static risk optimization factor and the dynamic risk optimization factor of the two structured file reading behavior instances as the distance optimization factor of the two structured file reading behavior instances; Obtaining normalized feature vectors of each of the two structured file reading behavior instances, and using the Euclidean distance between the normalized feature vectors of each of the two structured file reading behavior instances as a basic distance of the two structured file reading behavior instances; The calculation result of multiplying the distance optimization factor of the two structured file reading behavior instances by the basic distance of the two structured file reading behavior instances is used as the behavior instance similarity after comprehensive risk weighting.
Citation Information
Patent Citations
System and method for sharing atrusted platform module
CN101535957A
Sensitive data risk analysis method and system
CN118395453A
Power information system supply chain security risk static analysis method and system
CN118427828A
Power system security event log management method
CN118468342A
Ciphertext execution exception handling method and system for privacy computing algorithm
CN119670156A
Cited By
File co-processing method and system based on cloud computing
CN120561626A