Full-process automatic safety operation system and method based on AI drive

The AI-driven, fully automated security operations system addresses the issues of limited anomaly detection, inaccurate risk assessment, and reliance on human experience in existing technologies. It improves the accuracy of anomaly detection, enables quantitative risk assessment, and allows for precise classification of warning levels, thereby optimizing security operations efficiency.

CN121644191APending Publication Date: 2026-03-10LANZHOU AIKOS INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-09
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing security operation technologies suffer from limited anomaly detection dimensions, high false alarm rates, lack of quantitative standards for risk assessment, fragmented security analysis processes that fail to form a closed loop, and reliance on human experience in early warning mechanisms, which prevents accurate and tiered early warnings.

Method used

The system employs an AI-driven, fully automated security operations system, including modules for correlation feature extraction, anomaly detection, classification of similar feature data, sequence construction, and hierarchical early warning output. Through dual-model cross-validation, time-series feature analysis, behavioral feature analysis, security knowledge graph, and attack chain analysis, it improves the accuracy of anomaly detection and quantifies risk assessment, enabling multi-dimensional early warning assessment and precise early warning.

Benefits of technology

It improved the accuracy of anomaly detection, reduced false alarm and false negative rates, enabled objective quantitative assessment and scientific ranking of security risks, enhanced proactive security protection and accurate classification of early warning levels, and optimized the allocation of security operation resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121644191A_ABST
    Figure CN121644191A_ABST
Patent Text Reader

Abstract

The invention discloses a whole-process automatic safety operation system and method based on AI driving. The method comprises an associated feature extraction module, an anomaly judgment module, a similar feature data classification module, a sequence construction module and a grading early warning output module. The association feature extraction module is used for collecting security process data, establishing an anomaly analysis model and inputting the security process data into the anomaly analysis model to obtain association features; and the abnormity judgment module is used for marking the associated features as to-be-verified abnormal data, calling historical normal data, establishing an AI prediction model according to the to-be-verified abnormal data and the historical normal data, and judging the to-be-verified abnormal data through the AI prediction model. According to the method, a double-model anomaly detection mechanism of time sequence feature analysis and behavior feature analysis is established, and feature fusion is performed by adopting intersection and difference set operation, so that the anomaly recognition accuracy is effectively improved, and the false alarm rate and the missing report rate are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of network security, and in particular to an AI-driven full-process automated security operation system and method. BACKGROUND

[0002] With the deepening of the digital transformation of enterprises, the threat of network security is also gradually highlighted with the increasingly complex "invisible" and "intelligent" development trend. However, the traditional security operation mode has always been mainly based on manual review and judgment of massive alarm logs, and the security events are roughly identified and responded through pre-set static rules. However, with the increasing complexity of the network, the increasing diversification of security events and the increasing harm, simple manual review and pre-set static rules cannot meet the security needs. Based on the increasing amount of security data, various new threats are emerging, and the dependence on traditional security operation mode has become increasingly difficult to meet the security protection needs of modern enterprises. It is necessary to introduce artificial intelligence technology into security operation to realize the intelligent upgrading of security.

[0003] However, the existing security operation technology has a series of shortcomings, such as: the abnormal detection method is single, most of which only stay on a single dimensional analysis model, resulting in that the abnormal identification cannot achieve high accuracy, but causes a large number of false positive rates, and also causes a great workload and burden to the security analysts; at the same time, due to the lack of quantitative standards for risk assessment, it is difficult to objectively prioritize different types of security threats, and the security response also affects the timeliness and effectiveness of security, etc. In addition, the security analysis process also has a fragmentation phenomenon, such as the lack of effective linkage between abnormal detection, risk assessment, early warning output and other links, which cannot form a complete security operation closed loop; in addition, the early warning mechanism is not intelligent enough, the determination of the early warning level still depends on manual experience, lacks comprehensive assessment of the completeness of the attack path and the possibility of threat diffusion, and it is difficult to realize accurate hierarchical early warning, etc. SUMMARY

[0004] In view of the above existing problems, the present application is proposed.

[0005] Therefore, the present application provides an AI-driven full-process automated security operation system and method to solve the problems of single abnormal detection dimension, high false positive rate, lack of quantitative standards for risk assessment, fragmented security analysis process, inability to form a closed loop, and dependence on manual experience for early warning mechanism, and inability to realize accurate hierarchical early warning in the existing security operation technology.

[0006] To solve the above technical problems, the present application provides the following technical solutions:

[0007] In a first aspect, this invention provides an AI-driven, fully automated security operation system, including a correlation feature extraction module, an anomaly judgment module, a similar feature data classification module, a sequence construction module, and a hierarchical early warning output module. The correlation feature extraction module is used to collect security process data, establish an anomaly analysis model, and input the security process data into the anomaly analysis model to obtain correlation features. The anomaly judgment module is used to mark the correlation features as anomaly data to be verified, retrieve historical normal data, establish an AI prediction model based on the anomaly data to be verified and the historical normal data, and use the AI ​​prediction model to judge the anomaly data to be verified. The similar feature data classification module is used to retrieve feature data from the security process data. The system includes a feature data component for identifying abnormal data to be verified, classifying security process data with the same features into categories, marking them as similar data, counting the occurrence frequency of similar data, and calculating the probability of occurrence of similar data based on the frequency of occurrence of similar data and the total number of security process data; a sequence construction module for performing risk quantification calculation on feature data based on the probability of occurrence of similar data, obtaining multiple risk features, sorting multiple risk features according to the risk quantification value of the risk features, and establishing a risk feature sequence; and a hierarchical early warning output module for performing security analysis on the risk features in the risk feature sequence, obtaining security analysis results, determining the early warning level based on the security analysis results, generating security early warning information, and outputting it.

[0008] As a preferred embodiment of the AI-driven, fully automated safety operation system of the present invention, the step of obtaining associated features includes: establishing a first anomaly analysis model for safety process data, analyzing the temporal characteristics of the safety process data through the first anomaly analysis model, and obtaining a first set of anomaly features; establishing a second anomaly analysis model for safety process data, analyzing the behavioral characteristics of the safety process data through the second anomaly analysis model, and obtaining a second set of anomaly features; performing an intersection operation on the first set of anomaly features and the second set of anomaly features to obtain a set of common anomaly features, and marking the anomaly features in the set of common anomaly features as high-confidence anomaly features; performing a difference operation on the first set of anomaly features and the second set of anomaly features to obtain a set of difference anomaly features, and marking the anomaly features in the set of difference anomaly features as anomaly features to be confirmed; performing a union operation on the high-confidence anomaly features and the anomaly features to be confirmed to obtain a set of associated features, and marking the features in the set of associated features as associated features.

[0009] The beneficial effects of this preferred technical solution are: by using a dual-model cross-validation mechanism, the accuracy of anomaly detection is effectively improved, the false alarm rate caused by a single model is reduced, and potential anomaly clues are retained to avoid missed detections.

[0010] As a preferred embodiment of the AI-driven, fully automated safety operation system of the present invention, the analysis method of the first anomaly analysis model includes: performing time series decomposition on the safety process data, decomposing the safety process data into trend components, periodic components, and residual components; when analyzing the trend components, setting a trend deviation threshold, and marking the corresponding safety process data as trend anomaly data when the rate of change of the trend component exceeds the trend deviation threshold; when analyzing the periodic components, constructing a periodic baseline model, and marking the corresponding safety process data as periodic anomaly data when the deviation between the periodic component and the periodic baseline model exceeds a preset periodic deviation threshold; when analyzing the residual components, calculating the standard deviation of the residual components, and marking the corresponding safety process data as residual anomaly data when the residual components exceed a preset standard deviation multiple; and summarizing the trend anomaly data, periodic anomaly data, and residual anomaly data to generate a first anomaly feature set.

[0011] As a preferred embodiment of the AI-driven, fully automated safety operation system of the present invention, the analysis method of the second anomaly analysis model includes: extracting behavioral sequences from safety process data and establishing a normal behavior baseline library; calculating the similarity between the behavioral feature vector of the current safety process data and the behavioral pattern feature vector in the normal behavior baseline library, and marking the corresponding safety process data as behavioral deviation anomaly data when the similarity is lower than a preset similarity threshold; analyzing the behavioral state transition matrix to identify low-probability transition paths, and marking the corresponding safety process data as transition anomaly data when a low-probability transition path appears in the safety process data; and summarizing the behavioral deviation anomaly data and the transition anomaly data to generate a second anomaly feature set.

[0012] As a preferred embodiment of the AI-driven, fully automated safety operation system described in this invention, the step of judging the abnormal data to be verified using an AI prediction model includes: constructing a multi-level classification neural network as the AI ​​prediction model, wherein the multi-level classification neural network includes a feature extraction layer, an attention mechanism layer, and a classification output layer; inputting the abnormal data to be verified into the feature extraction layer, performing deep feature extraction on the abnormal data to be verified through the feature extraction layer to obtain a deep feature vector; inputting the deep feature vector into the attention mechanism layer, performing weighted processing on the key features in the deep feature vector through the attention mechanism layer to obtain a weighted feature vector; processing historical normal data through the same feature extraction layer and attention mechanism layer to obtain a weighted feature vector of normal data; in the classification output layer, calculating the distance metric between the weighted feature vector of the abnormal data to be verified and the weighted feature vector of normal data, and judging whether the abnormal data to be verified is real abnormal data based on the comparison result of the distance metric with a preset discrimination threshold.

[0013] As a preferred embodiment of the AI-driven, fully automated security operation system described in this invention, the step of performing risk quantification calculation on feature data based on the occurrence probability of similar data includes: obtaining the occurrence probability P of similar data, and calculating the information entropy value S1 of similar data according to the formula S1 = -log2(P); obtaining the impact degree value D of historical security events corresponding to similar data, and determining the impact weight coefficient W based on the impact degree value D; obtaining the timeliness parameter T of similar data, wherein the timeliness parameter T is determined based on the interval between the most recent occurrence time of similar data and the current time; and calculating the risk quantification calculation based on the formula R = S1 × W × e (-λT) Calculate the risk quantification value R, where λ is the time decay coefficient; mark the feature data whose risk quantification value R is greater than the preset risk threshold as risk features.

[0014] The beneficial effects of this preferred technical solution are: to achieve an objective quantitative assessment of security risks, and to comprehensively consider the rarity of anomalies, the degree of historical impact, and timeliness, making the risk ranking more scientific and reasonable.

[0015] As a preferred embodiment of the AI-driven, fully automated security operation system described in this invention, the step of performing security analysis on risk features in a risk feature sequence includes: constructing a security knowledge graph, which contains security threat entities, attack method entities, protective measure entities, and their relationships; matching the risk features in the risk feature sequence with the security knowledge graph to obtain security threat entities and attack method entities associated with the risk features; performing attack link analysis on multiple risk features based on the relationships in the security knowledge graph to identify potential attack paths; calculating the completeness score of the attack path, which is determined based on the ratio of the number of risk features appearing in the attack path to the number of features required for a complete attack path; and generating a comprehensive security analysis result based on the completeness score of the attack path and the risk quantification value of the risk features.

[0016] The beneficial effects of this preferred technical solution are: it enables the correlation reasoning from isolated risk characteristics to complete attack intent, and can identify potential threats before the attack is completed, thereby improving the initiative and foresight of security protection.

[0017] As a preferred embodiment of the AI-driven, fully automated security operation system described in this invention, the step of determining the warning level based on security analysis results includes: establishing a multi-dimensional warning assessment model, which includes a threat urgency dimension, an asset importance dimension, and a propagation probability dimension; determining a threat urgency score based on the completeness score of the attack path, with a higher completeness score resulting in a higher threat urgency score; determining an asset importance score based on the business level of the system assets associated with the risk characteristics; determining a propagation probability score based on the propagation characteristics of the risk characteristics and the system network topology; weighted summing of the threat urgency score, asset importance score, and propagation probability score to obtain a comprehensive warning score; and determining the corresponding warning level based on the preset score range of the comprehensive warning score.

[0018] The beneficial effects of this preferred technical solution are: it breaks through the limitations of single-indicator early warning, comprehensively assesses security threats from multiple dimensions, achieves accurate classification of early warning levels, and optimizes the allocation of security response resources.

[0019] As a preferred embodiment of the AI-driven, fully automated safety operation system described in this invention, the step of generating and outputting safety warning information includes: calling the corresponding warning template according to the warning level; filling in the structured fields in the warning template according to the safety analysis results to generate safety warning information.

[0020] Secondly, this invention provides an AI-driven, fully automated safety operation method, comprising the following steps: collecting safety process data, establishing an anomaly analysis model, inputting the safety process data into the anomaly analysis model to obtain associated features; marking the associated features as anomaly data to be verified, retrieving historical normal data, establishing an AI prediction model based on the anomaly data to be verified and the historical normal data, and judging the anomaly data to be verified through the AI ​​prediction model; retrieving feature data from the safety process data and feature data from the anomaly data to be verified that has been judged as anomalies, classifying the safety process data with the same feature data, marking them as similar data, counting the occurrence frequency of similar data, and calculating the occurrence probability of similar data based on the occurrence frequency of similar data and the total number of safety process data; performing risk quantification calculation on the feature data based on the occurrence probability of similar data, obtaining multiple risk features, sorting the multiple risk features according to the risk quantification value of the risk features, and establishing a risk feature sequence; performing safety analysis on the risk features in the risk feature sequence, obtaining safety analysis results, determining the warning level based on the safety analysis results, generating safety warning information and outputting it.

[0021] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0022] By establishing a dual-model anomaly detection mechanism that combines temporal feature analysis and behavioral feature analysis, and employing intersection and difference operations for feature fusion, the accuracy of anomaly identification is effectively improved, while the false alarm rate and false negative rate are reduced.

[0023] By constructing a risk quantification model that integrates information entropy, influence weight, and time decay, we can achieve objective quantitative assessment and scientific ranking of security risks, providing a priority basis for security responses.

[0024] By combining security knowledge graphs with attack chain analysis, we can achieve correlation reasoning from discrete risk characteristics to complete attack paths, enabling us to predict attack intentions in the early stages of an attack and improve the initiative of security protection.

[0025] By establishing a multi-dimensional early warning assessment model that considers threat urgency, asset importance, and potential for spread, we can achieve precise classification of early warning levels and optimize the rational allocation of security operation resources.

[0026] It achieves full automation of the process from data collection, anomaly detection, risk quantification, security analysis to early warning output, reducing manual intervention and improving the efficiency of security operations. Attached Figure Description

[0027] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1 This is a schematic diagram of the overall process of an AI-driven, fully automated safety operation system according to an embodiment of the present invention. Detailed Implementation

[0029] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0030] Example 1, referring to Figure 1 As an embodiment of the present invention, an AI-driven fully automated security operation system is provided, including an associated feature extraction module, an anomaly judgment module, a similar feature data classification module, a sequence construction module, and a hierarchical early warning output module;

[0031] The correlation feature extraction module is used to collect security process data, establish an anomaly analysis model, and input the security process data into the anomaly analysis model to obtain correlation features.

[0032] In one specific embodiment, security process data includes, but is not limited to, user login logs, network access logs, file operation logs, system call logs, and API request logs. Taking a company's security operations center as an example, the system collects approximately 5 million security process data entries daily, with data fields including timestamps, user IDs, source IP addresses, target IP addresses, operation types, operation objects, and operation results.

[0033] The steps for obtaining the associated features include A1 to A5:

[0034] A1. Establish a first anomaly analysis model for safety process data, and analyze the temporal characteristics of the safety process data through the first anomaly analysis model to obtain a first set of anomaly features.

[0035] In one specific embodiment, the first anomaly analysis model employs a time-series-based anomaly detection method. Taking user login behavior as an example, the system counts the number of logins per hour as time-series data. Assuming a user's average hourly login frequency over the past 30 days is 5, the system inputs this time-series data into the first anomaly analysis model for analysis. The analysis methods of the first anomaly analysis model include A1.1 to A1.5:

[0036] A1.1 Perform time series decomposition on the safety process data, decomposing the safety process data into trend components, periodic components, and residual components.

[0037] Taking the hourly login request volume of a company's intranet as an example, the system collected login request data for the past 30 days, totaling 720 hours, and used the STL decomposition method to decompose this time-series data. The original data is the hourly login request volume sequence X(t), where t ranges from 1 to 720. After decomposition, three components are obtained: the trend component T(t) reflects the long-term trend of login request volume, for example, as the number of employees in the company increases, the login request volume shows a slow upward trend, increasing from an average of 100 times per hour on day 1 to an average of 120 times per hour on day 30; the periodic component S(t) reflects the periodic fluctuation of login request volume, for example, the login request volume is higher during working hours on weekdays and lower on weekends and at night, with a period of 24 hours; the residual component R(t) reflects the random fluctuation after removing the trend and periodicity, and under normal circumstances, the residual component should fluctuate within a small range. The relationship between the three components satisfies the decomposition formula X(t) = T(t) + S(t) + R(t).

[0038] A1.2 When analyzing trend components, set a trend deviation threshold. When the rate of change of a trend component exceeds the trend deviation threshold, mark the corresponding safety process data as abnormal trend data.

[0039] The system performs sliding window analysis on the trend component T(t), with the window size set to 24 hours. It calculates the rate of change of the trend component within each window. The formula for the rate of change is the absolute value of the difference between the current trend component and the trend component at the same time the previous day, divided by the trend component at the same time the previous day, and then multiplied by 100%. The system sets a trend deviation threshold of 15%. Under normal circumstances, the daily change rate of login requests is usually within 5%. Taking the 10th hour of day 25 as an example, the trend component at that time is 180 times per hour, while the trend component at the same time the previous day was 115 times per hour, resulting in a change rate of 56.5%. Since 56.5% exceeds the 15% trend deviation threshold, the system marks the login request data for the 10th hour of day 25 as trend anomaly data, with the anomaly feature code TREND_ANO_001.

[0040] A1.3 When analyzing the periodic components, a periodic baseline model is constructed. When the deviation between the periodic component and the periodic baseline model exceeds the preset periodic deviation threshold, the corresponding safety process data is marked as periodic abnormal data.

[0041] A cyclical baseline model is constructed using historical data from the past 30 days. The average cyclical component for each hour is calculated as the baseline value over a 24-hour period. For example, the baseline value at 9 AM is +50 cycles per hour with a standard deviation of 8; at 2 PM, it is +30 cycles per hour with a standard deviation of 6; and at 2 AM, it is -40 cycles per hour with a standard deviation of 5. The system sets a cyclical deviation threshold of 2 standard deviations. Taking 2 AM on day 28 as an example, the cyclical component at that time is +25 cycles per hour, while the baseline value is -40 cycles per hour, resulting in a deviation of 65 cycles per hour. Since the deviation threshold is only 10 cycles per hour, 65 far exceeds this threshold. Therefore, the system marks the data at 2 AM on day 28 as cyclical anomaly data, with the anomaly feature code CYCLE_ANO_001, indicating abnormally high login activity during the early morning period.

[0042] A1.4 When analyzing residual components, calculate the standard deviation of the residual components. When a residual component exceeds a preset standard deviation multiple, mark the corresponding safety process data as residual abnormal data.

[0043] The mean and standard deviation of the residual components over the past 30 days are calculated. Theoretically, the residual mean is 0, and the calculated standard deviation is 12 times per hour. The system sets a preset standard deviation multiple of 3, i.e., using the 3σ principle for anomaly detection. Taking the 15th hour of day 26 as an example, the residual component at that moment is positive 58 times per hour, while the judgment threshold is 36 times per hour. Since the absolute value of 58 exceeds 36, the system marks the data from the 15th hour of day 26 as residual anomalous data, with the anomaly feature code RESID_ANO_001, indicating that there is an anomalous fluctuation at that moment that cannot be explained by trends and cycles.

[0044] A1.5. Summarize the trend anomaly data, periodic anomaly data, and residual anomaly data to generate the first set of anomaly features.

[0045] All abnormal data detected in steps A1.2 to A1.4 are summarized. After summarization, the first abnormal feature set F1 contains 5 abnormal features: the first is TREND_ANO_001, which belongs to trend anomaly, occurring at 10:00 on day 25, manifested as a sudden increase in login request volume; the second is CYCLE_ANO_001, which belongs to periodic anomaly, occurring at 2:00 on day 28, manifested as abnormal activity in the early morning; the third is RESID_ANO_001, which belongs to residual anomaly, occurring at 15:00 on day 26, manifested as random abnormal fluctuations; the fourth is TREND_ANO_002, which belongs to trend anomaly, occurring at 9:00 on day 27, manifested as a sudden decrease in login request volume; and the fifth is CYCLE_ANO_002, which belongs to periodic anomaly, occurring at 3:00 on day 29, also manifested as abnormal activity in the early morning.

[0046] A2. Establish a second anomaly analysis model for safety process data, analyze the behavioral characteristics of safety process data through the second anomaly analysis model, and obtain a second set of anomaly features.

[0047] The analysis methods for the second anomaly analysis model include A2.1 to A2.4:

[0048] A2.1 Extract the behavioral sequences from the safety process data and establish a normal behavior baseline library.

[0049] The system extracts user behavior sequences from security process data. Behavior types include login, file access, data download, email sending, system configuration, and logout. Taking user User_A as an example, the system extracts their behavior sequences from the past 30 days and establishes a normal behavior baseline database. The construction of behavior pattern feature vectors includes statistical analysis of the frequency of each behavior type, the time distribution of each behavior type, and the transition patterns of behavior sequences. User_A's normal behavior baseline shows an average of 3 logins per day, 45 file accesses per day, an average data download volume of 120MB per day, a common login time from 8:30 AM to 6:30 PM, a common login location being the office IP range, and common file types accessed being docx, xlsx, and pdf formats. The normal behavior baseline database contains behavior pattern feature vectors for all users, totaling 2000 user behavior baseline data.

[0050] A2.2 Calculate the similarity between the behavior feature vector of the current safety process data and the behavior pattern feature vector in the normal behavior baseline library. When the similarity is lower than the preset similarity threshold, mark the corresponding safety process data as abnormal behavior deviation data.

[0051] The similarity between the current behavior feature vector and the baseline feature vector is calculated using cosine similarity. The similarity is equal to the dot product of the two vectors divided by the product of their magnitudes. The system sets a preset similarity threshold of 0.75. Taking User_A's behavior on day 28 as an example, their behavior feature vector shows 8 logins (baseline 3), 180 file accesses (baseline 45), 2.5GB of data downloads (baseline 120MB), login time from 2:00 AM to 5:00 AM (baseline 8:30 AM to 6:30 PM), login location from an external IP address (baseline office IP address range), and accessed file types in SQL, BAK, and KEY formats (baseline docx, xlsx, and pdf formats). The system calculates a similarity of 0.32. Since 0.32 is lower than the preset similarity threshold of 0.75, the system marks User_A's behavior data on day 28 as abnormal behavior data, with the abnormal feature code BEHAV_ANO_001.

[0052] A2.3 Analyze the behavior state transition matrix to identify low-probability transition paths. When a low-probability transition path appears in the security process data, mark the corresponding security process data as transition anomaly data.

[0053] A behavior state transition matrix is ​​constructed based on historical data to record the transition probabilities between different behavior states. For example, the probability of transitioning from login state to file access state is 0.60, from login state to system configuration state is 0.04, and from system configuration state to download state is 0.10. The system sets a low-probability transition threshold of 0.02. Taking User_D's behavior sequence on day 28 as an example, the sequence is: log in, directly perform batch downloads, then send emails to distribute large files, and finally log out. System analysis of this behavior sequence reveals that although the probabilities of individual transitions are within the normal range, the overall behavior pattern "login, directly perform batch downloads and distribute files" has a probability of less than 0.01 in historical data, belonging to a low-probability transition path. The system marks User_D's behavior data on day 28 as transition anomaly data, with the anomaly feature encoded as TRANS_ANO_001.

[0054] A2.4. Summarize the behavioral deviation anomaly data and the transfer anomaly data to generate a second set of anomaly features.

[0055] All abnormal data detected in steps A2.2 and A2.3 are summarized. After summarization, the second abnormal feature set F2 contains four abnormal features: the first is BEHAV_ANO_001, which belongs to behavioral deviation anomaly, occurred on day 28, associated with user_A, and is characterized by logging in from a different location in the early morning and downloading a large number of sensitive files; the second is BEHAV_ANO_002, which belongs to behavioral deviation anomaly, occurred on day 28, associated with user_E, and is characterized by accessing unauthorized directories; the third is TRANS_ANO_001, which belongs to transfer anomaly, occurred on day 28, associated with user_D, and is characterized by directly downloading and sending files in batches after logging in; the fourth is BEHAV_ANO_003, which belongs to behavioral deviation anomaly, occurred on day 26, associated with user_F, and is characterized by a single-day download volume exceeding 20 times the baseline.

[0056] A3. Perform an intersection operation on the first set of abnormal features and the second set of abnormal features to obtain a common set of abnormal features, and mark the abnormal features in the common set of abnormal features as high-confidence abnormal features.

[0057] A correlation analysis is performed on the first set of abnormal features F1 and the second set of abnormal features F2 to identify abnormal events detected by both models simultaneously. The correlation matching rules include two dimensions: time window matching and entity correlation matching. Time window matching requires that the time difference between the occurrence of the anomaly is within 2 hours, and entity correlation matching requires that the anomaly involves the same user, the same IP address, or the same resource.

[0058] The system performed an intersection operation on the two sets and found that CYCLE_ANO_001 occurred at 2:00 AM on day 28, while BEHAV_ANO_001 also occurred during the early morning period of day 28, and the associated user User_A exhibited abnormal activity in the early morning. Since both sets match in time and entity, they were determined to be intersection elements. Similarly, RESID_ANO_001 occurred at 3:00 PM on day 26, while BEHAV_ANO_003 also occurred on day 26, and the associated user User_F exhibited abnormal download behavior at the same time; these were also determined to be intersection elements. Other abnormal features such as TREND_ANO_001, TREND_ANO_002, and CYCLE_ANO_002 did not find corresponding matches in the second set of abnormal features and therefore did not belong to the intersection elements.

[0059] After intersection analysis, the common anomaly feature set contains two high-confidence anomaly features: the first, coded as HIGH_CONF_001, comes from the intersection of CYCLE_ANO_001 and BEHAV_ANO_001; the second, coded as HIGH_CONF_002, comes from the intersection of RESID_ANO_001 and BEHAV_ANO_003. The system marks these anomaly features as high-confidence anomaly features, indicating that these anomalies were detected simultaneously by two independent models—time series analysis and behavioral analysis—and have high confidence.

[0060] A4. Perform a difference operation on the first abnormal feature set and the second abnormal feature set to obtain the difference abnormal feature set, and mark the abnormal features in the difference abnormal feature set as abnormal features to be confirmed.

[0061] Calculate the difference between the first set of abnormal features F1 and the second set of abnormal features F2 to obtain the abnormal features detected only by a single model. The result of the difference operation is equal to the union of F1 minus the portion of the common abnormal feature set and F2 minus the portion of the common abnormal feature set.

[0062] After difference operations, the set of differential anomaly features contains five anomaly features to be confirmed: the first is coded as PENDING_001, the original feature is TREND_ANO_001, and it comes from the first anomaly analysis model; the second is coded as PENDING_002, the original feature is TREND_ANO_002, and it comes from the first anomaly analysis model; the third is coded as PENDING_003, the original feature is CYCLE_ANO_002, and it comes from the first anomaly analysis model; the fourth is coded as PENDING_004, the original feature is BEHAV_ANO_002, and it comes from the second anomaly analysis model; and the fifth is coded as PENDING_005, the original feature is TRANS_ANO_001, and it comes from the second anomaly analysis model. The system marks these anomaly features as anomaly features to be confirmed, indicating that these anomalies were detected by only a single model.

[0063] A5. Perform a union operation on the high-confidence abnormal features and the abnormal features to be confirmed to obtain the associated feature set, and mark the features in the associated feature set as associated features.

[0064] The system performs a union operation on the high-confidence anomaly feature set and the unconfirmed anomaly feature set to obtain a complete set of related features. The result of the union operation includes all anomaly features that require further processing.

[0065] After union operation, the associated feature set contains a total of 7 associated features: the first two are high-confidence anomaly features, coded RELATED_001 and RELATED_002 respectively, derived from HIGH_CONF_001 and HIGH_CONF_002, and these two features will directly enter the risk quantification stage; the latter five are anomaly features to be confirmed, coded RELATED_003 to RELATED_007 respectively, derived from PENDING_001 to PENDING_005, and these five features need to be verified by the AI ​​prediction model before deciding whether to enter the risk quantification stage. The system uniformly marks all features in the associated feature set as associated features.

[0066] The anomaly detection module is used to mark associated features as anomalous data to be verified, retrieve historical normal data, build an AI prediction model based on the anomalous data to be verified and the historical normal data, and use the AI ​​prediction model to judge the anomalous data to be verified.

[0067] The steps for judging abnormal data to be verified using an AI prediction model include B1 to B5:

[0068] B1. Construct a multi-level classification neural network as an AI prediction model. The multi-level classification neural network includes a feature extraction layer, an attention mechanism layer, and a classification output layer.

[0069] A multi-layered classification neural network was constructed as an AI prediction model. This neural network adopts a deep learning architecture and contains three core layers. The first layer is the feature extraction layer, which uses a multi-layer fully connected neural network structure with three hidden layers. The number of neurons in each hidden layer is 256, 128, and 64, respectively. The ReLU activation function is used to perform layer-by-layer abstraction and feature extraction from the input data. The second layer is the attention mechanism layer, which uses a self-attention mechanism structure with eight attention heads, each with a dimension of 32. It is used to identify and weight the key features of the feature vector output by the feature extraction layer. The third layer is the classification output layer, which uses a fully connected layer structure with an output dimension of 64. It is used to calculate the distance metric between the data to be verified and the normal data and make the final judgment. The input dimension of the entire neural network is 128, corresponding to the feature vector dimension of the abnormal data to be verified. The total number of network parameters is approximately 150,000. It is trained using the Adam optimizer with a learning rate set to 0.001.

[0070] B2. Input the abnormal data to be verified into the feature extraction layer. The feature extraction layer performs deep feature extraction on the abnormal data to be verified to obtain a deep feature vector.

[0071] The system first preprocesses and vectorizes the abnormal data to be verified. Taking the associated feature RELATED_001 as an example, which corresponds to User_A's abnormal behavior in the early morning of the 28th day, the system converts it into a 128-dimensional input feature vector. The feature dimensions include multiple categories such as time features, behavior frequency features, resource access features, and network location features. Specifically, the time feature occupies 16 dimensions, encoding information such as the hour, day of the week, and whether it was a holiday when the anomaly occurred; the behavior frequency feature occupies 32 dimensions, encoding statistical information such as the number of logins, file accesses, and downloads; the resource access feature occupies 48 dimensions, encoding information such as the type of file accessed, the directory path accessed, and operation permissions; and the network location feature occupies 32 dimensions, encoding information such as the source IP address, the destination IP address, and the geographical location. The system then inputs this 128-dimensional input feature vector into the feature extraction layer for processing. In the first hidden layer, the input vector undergoes linear transformation and ReLU activation function processing to obtain a 256-dimensional intermediate feature vector. In the second hidden layer, the 256-dimensional intermediate feature vector is further processed to obtain a 128-dimensional intermediate feature vector. In the third hidden layer, the 128-dimensional intermediate feature vector is processed to obtain a 64-dimensional deep feature vector. Taking RELATED_001 as an example, after processing by the feature extraction layer, the system obtains a 64-dimensional deep feature vector. The values ​​in this vector reflect the high-level semantic features of the original anomalous data after multiple layers of abstraction. For example, dimensions 15 to 20 of the vector mainly encode the degree of temporal anomaly, while dimensions 35 to 45 mainly encode the degree of behavioral deviation.

[0072] B3. Input the deep feature vector into the attention mechanism layer. The attention mechanism layer performs weighted processing on the key features in the deep feature vector to obtain the weighted feature vector.

[0073] The system inputs the 64-dimensional deep feature vector obtained in step B2 into the attention mechanism layer for processing. The attention mechanism layer employs a self-attention mechanism. First, it transforms the 64-dimensional deep feature vector through three linear transformations to generate a query vector Q, a key vector K, and a value vector V, each with 64 dimensions. Then, the system calculates the dot product of the query vector and the key vector and performs scaling to obtain an attention score matrix, with the scaling factor being the square root of the key vector's dimension. Next, the system performs Softmax normalization on the attention score matrix to obtain an attention weight matrix, whose values ​​reflect the correlation and importance between the various feature dimensions. Finally, the system multiplies the attention weight matrix by the value vector to obtain a weighted feature vector.

[0074] Taking the deep feature vector of RELATED_001 as an example, after processing by the attention mechanism layer, the system identifies the key feature dimensions in this abnormal data. The attention weight distribution shows that the time-related anomaly features corresponding to dimensions 17 to 19 received a cumulative attention weight of 0.35, indicating that abnormal logins during the early morning hours are an important criterion; the download behavior features corresponding to dimensions 38 to 42 received a cumulative attention weight of 0.28, indicating that downloading a large number of sensitive files is another important criterion; and the network location features corresponding to dimensions 55 to 58 received a cumulative attention weight of 0.22, indicating that logins from different locations are also a feature that needs attention. After weighted processing, the system obtains a 64-dimensional weighted feature vector of RELATED_001, in which the key feature dimensions are enhanced and the non-key feature dimensions are suppressed.

[0075] B4. Process historical normal data through the same feature extraction layer and attention mechanism layer to obtain the weighted feature vector of normal data.

[0076] The system retrieves 50,000 manually verified normal operation records from the historical normal database. These records cover user behavior data under various normal business scenarios. To improve computational efficiency, the system uses a clustering algorithm to preprocess the historical normal data, clustering the 50,000 normal data records into 100 representative samples, with each representative sample representing a typical normal behavior pattern.

[0077] The system vectorizes each representative sample in the same way as the abnormal data to be verified, generating a 128-dimensional input feature vector. Then, these input feature vectors are sequentially fed into a feature extraction layer and an attention mechanism layer for processing, obtaining a 64-dimensional weighted feature vector for each representative sample. Taking the representative sample NORMAL_023 representing normal behavior as an example, this sample represents normal file access behavior during working hours. After processing by the feature extraction layer, a 64-dimensional depth feature vector is obtained, which is then processed by the attention mechanism layer to obtain a 64-dimensional weighted feature vector. The attention weight distribution of this normal sample shows that the attention weights in the time feature dimension are concentrated at the encoding positions corresponding to working hours, the attention weights in the behavior feature dimension are concentrated at the encoding positions corresponding to normal access frequencies, and the attention weights in the network location feature dimension are concentrated at the encoding positions corresponding to office networks.

[0078] B5. In the classification output layer, calculate the distance metric between the weighted feature vector of the abnormal data to be verified and the weighted feature vector of the normal data. Based on the comparison between the distance metric and the preset discrimination threshold, determine whether the abnormal data to be verified is real abnormal data.

[0079] The system makes the final judgment on the anomalous data to be verified in the classification output layer. For each anomalous data point to be verified, the system calculates the distance metric between its weighted feature vector and the weighted feature vectors of 100 representative normal data samples. The distance metric uses the Euclidean distance calculation method, which is calculated as the square root of the sum of the squares of the differences between the corresponding dimensions of the two vectors. The system takes the minimum distance between the anomalous data point to be verified and all normal samples as the final distance metric for that data point. This minimum value reflects the degree of difference between the anomalous data point to be verified and the most similar normal behavior.

[0080] The system has a preset discrimination threshold of 2.5, which is determined based on statistical analysis of historical data. Under this threshold, the false alarm rate is approximately 5% and the false negative rate is approximately 3%. When the distance metric value of the abnormal data to be verified is greater than 2.5, the system determines that the data is real abnormal data; when the distance metric value is less than or equal to 2.5, the system determines that the data is normal data or false alarm data.

[0081] Taking the judgment results of 7 abnormal data to be verified as an example, the system performed distance measurement calculation and judgment on each data.

[0082] The distance metric between the weighted feature vector of RELATED_001 and the most similar normal sample is 4.82, which is greater than the preset discrimination threshold of 2.5. Therefore, the system judges it as real abnormal data.

[0083] The distance metric value of RELATED_002 is 3.95, which is greater than the preset discrimination threshold of 2.5. The system judges it to be real abnormal data.

[0084] The distance metric value of RELATED_003 is 1.87, which is less than the preset discrimination threshold of 2.5. The system judges it as false alarm data. The abnormal trend corresponding to this data may be caused by normal business fluctuations.

[0085] The distance metric value of RELATED_004 is 2.13, which is less than the preset discrimination threshold of 2.5. The system judges it as a false alarm.

[0086] The distance metric value of RELATED_005 is 1.65, which is less than the preset discrimination threshold of 2.5. The system judges it as a false alarm.

[0087] The distance metric value of RELATED_006 is 3.28, which is greater than the preset discrimination threshold of 2.5. The system judges it to be real abnormal data.

[0088] The distance metric value of RELATED_007 is 4.15, which is greater than the preset discrimination threshold of 2.5. The system judges it as real abnormal data.

[0089] The same-type feature data classification module is used to retrieve feature data from security process data and feature data from abnormal data to be verified that are judged to be abnormal. It classifies security process data with the same feature data, marks them as the same type of data, counts the occurrence frequency of the same type of data, and calculates the probability of occurrence of the same type of data based on the occurrence frequency of the same type of data and the total number of security process data.

[0090] In one specific embodiment, the similar feature data classification module receives four real anomaly data points output by the anomaly judgment module: RELATED_001, RELATED_002, RELATED_006, and RELATED_007. The system also retrieves security process data from the past 30 days as the basis for analysis, totaling 1.5 million records. The system extracts feature data from the four real anomaly data points, including multiple dimensions such as the time period of the anomaly, the type of abnormal behavior, the user roles involved, the type of accessed resources, and network location characteristics.

[0091] Taking RELATED_001 as an example, the system extracts its characteristic data, including login during the early morning hours, large-scale file downloads, access to sensitive files, and login from external IP addresses. The system searches for records with the same characteristic data among 1.5 million security process records, grouping security process data with the same combination of characteristics into the same category and marking them as similar data. After classification, the system identified eight categories of similar data: The first category is abnormal logins in the early morning, including RELATED_001 and other data with the same early morning login characteristics, which appeared 45 times in 30 days; the second category is large-scale data downloads, including RELATED_002 and other data with the same large-scale download characteristics, which appeared 128 times; the third category is sensitive file access, which appeared 89 times; the fourth category is logins from different locations, which appeared 67 times; the fifth category is privilege escalation, including RELATED_006 and related data, which appeared 23 times; the sixth category is abnormal outbound data, including RELATED_007 and related data, which appeared 56 times; the seventh category is operations outside of working hours, which appeared 234 times; and the eighth category is multiple authentication failures, which appeared 312 times.

[0092] The system calculates the probability of occurrence of similar data based on the frequency of occurrence of similar data and the total number of security process data. Taking the abnormal login at midnight as an example, it occurs 45 times, and the total number of security process data is 1.5 million. The probability of occurrence, P, is equal to 45 divided by 1,500,000, which is 0.00003. The probability of occurrence of other similar data is calculated using the same method: the probability of occurrence for large-scale data download is 0.0000853, for sensitive file access is 0.0000593, for login from a different location is 0.0000447, for privilege escalation is 0.0000153, for abnormal outbound data is 0.0000373, for operations outside of working hours is 0.000156, and for multiple authentication failures is 0.000208.

[0093] The sequence construction module is used to perform risk quantification calculation on feature data based on the occurrence probability of similar data, obtain multiple risk features, sort the multiple risk features according to the risk quantification value of the risk features, and establish a risk feature sequence.

[0094] The steps for quantifying the risk of feature data based on the probability of occurrence of similar data include C1 to C5:

[0095] C1. Obtain the probability P of the occurrence of similar data, and calculate the information entropy value S1 of similar data according to the formula S1=-log2(P).

[0096] The information entropy value of each type of data is calculated using the information entropy formula. The information entropy value reflects the rarity of this type of abnormal event. The lower the probability of occurrence, the higher the information entropy value, indicating that this type of event is rarer and the potential security risks are more worthy of attention.

[0097] Taking abnormal logins in the early morning as an example, the probability of occurrence P is 0.00003. Substituting into the formula, the information entropy value S1 is calculated to be the negative logarithm of 0.00003 to the base 2, which is approximately 15.02 bits. Taking large-scale data downloads as an example, the probability of occurrence P is 0.0000853, and the information entropy value S1 is approximately 13.52 bits. Taking sensitive file access as an example, the probability of occurrence P is 0.0000593, and the information entropy value S1 is approximately 14.04 bits. Taking logins from different locations as an example, the probability of occurrence P is 0.0000447, and the information entropy value S1 is approximately 14.45 bits. Taking privilege escalation as an example, the probability of occurrence P is 0.0000153, and the information entropy value S1 is approximately 16.00 bits. Taking abnormal outbound data transmission as an example, the probability of occurrence P is 0.0000373, and the information entropy value S1 is approximately 14.71 bits. Taking the non-working-hour operation category as an example, its occurrence probability P is 0.000156, and its information entropy value S1 is approximately 12.65 bits. Taking the multiple authentication failure category as an example, its occurrence probability P is 0.000208, and its information entropy value S1 is approximately 12.23 bits.

[0098] The calculation results show that the information entropy value of privilege escalation events is the highest, indicating that this type of event is the rarest; while the information entropy value of multiple authentication failure events is the lowest, indicating that this type of event is relatively common.

[0099] C2. Obtain the impact degree value D of historical security events corresponding to the same type of data, and determine the impact weight coefficient W based on the impact degree value D.

[0100] Historical security incident records related to each type of data are retrieved from the historical security incident database. The actual impact of these historical incidents is analyzed, and an impact weighting coefficient is determined accordingly. The assessment dimensions of the impact value include the scale of the data breach, the duration of business interruption, the amount of economic loss, and the degree of compliance violations. The impact value uses a scoring system from 1 to 10, where 1 indicates minimal impact and 10 indicates significant impact. The impact weighting coefficient W is determined based on the impact value D using a piecewise function: W equals 0.5 when D is less than or equal to 3; W equals 1.0 when D is greater than 3 and less than or equal to 6; W equals 1.5 when D is greater than 6 and less than or equal to 8; and W equals 2.0 when D is greater than 8.

[0101] Taking abnormal logins in the early morning as an example, the system's query of the historical security event database revealed 12 security incidents related to abnormal logins in the early morning over the past year, of which 3 resulted in data breaches. The average impact score was assessed at 7.5, and the impact weighting coefficient W was determined to be 1.5 based on the piecewise function. Taking large-scale data downloads as an example, there were 8 related historical security incidents, of which 2 resulted in the leakage of core data. The average impact score was assessed at 8.2, and the impact weighting coefficient W was 2.0. Taking sensitive file access as an example, there were 15 related historical security incidents, mostly due to accidental operations. The average impact score was assessed at 4.5, and the impact weighting coefficient W was 1.0. Taking logins from different locations as an example, the average impact score was assessed at 5.8, and the impact weighting coefficient W was 1.0. Taking privilege escalation as an example, although the number of related historical security incidents was small, their impact was severe. The average impact score was assessed at 9.2, and the impact weighting coefficient W was 2.0. Taking abnormal outbound data transmission as an example, the average impact score was assessed at 8.5, and the impact weighting coefficient W was 2.0. For operations performed outside of working hours, the average impact score is 3.2, and the impact weighting coefficient W is 0.5. For operations involving multiple authentication failures, the average impact score is 2.8, and the impact weighting coefficient W is 0.5.

[0102] C3. Obtain the timeliness parameter T of similar data, wherein the timeliness parameter T is determined based on the interval between the most recent occurrence time of similar data and the current time.

[0103] By obtaining the most recent occurrence time of each type of data, the interval between it and the current time is calculated as the timeliness parameter T. The timeliness parameter T is in days and reflects the freshness of the abnormal event. The shorter the time interval, the more recent the abnormal event, and the higher the attention it should receive.

[0104] Assuming the current time is day 30, the system calculates the timeliness parameters for various types of similar data. The most recent occurrence time for abnormal logins in the early morning is day 28, so the timeliness parameter T equals 30 minus 28, resulting in 2 days. The most recent occurrence time for large-scale data downloads is day 29, so the timeliness parameter T is 1 day. The most recent occurrence time for sensitive file access is day 27, so the timeliness parameter T is 3 days. The most recent occurrence time for logins from different locations is day 28, so the timeliness parameter T is 2 days. The most recent occurrence time for privilege escalation is day 30, so the timeliness parameter T is 0 days, indicating that this type of exception has just occurred. The most recent occurrence time for outbound exceptions is day 29, so the timeliness parameter T is 1 day. The most recent occurrence time for operations outside of working hours is day 25, so the timeliness parameter T is 5 days. The most recent occurrence time for multiple authentication failures is day 22, so the timeliness parameter T is 8 days.

[0105] C4. According to the formula R = S1 × W × e (-λT) Calculate the risk quantification value R, where λ is the time decay coefficient.

[0106] The system uses a risk quantification formula to calculate the risk quantification value for each type of data. The system sets the time decay coefficient λ to 0.1, which means that the risk quantification value decays by about 10% per day, reflecting the time-sensitive nature of security incidents.

[0107] Taking abnormal logins in the early morning as an example, its information entropy value S1 is 15.02, the influence weight coefficient W is 1.5, and the timeliness parameter T is 2 days. Substituting these values ​​into the formula, the calculated risk quantification value R is approximately 18.45. For large-scale data downloads, S1 is 13.52, W is 2.0, T is 1 day, and R is approximately 24.47. For sensitive file access, S1 is 14.04, W is 1.0, T is 3 days, and R is approximately 10.40. For logins from other locations, S1 is 14.45, W is 1.0, T is 2 days, and R is approximately 11.84. For privilege escalation, S1 is 16.00, W is 2.0, T is 0 days, and R equals 16.00 multiplied by 2.0 multiplied by 1, resulting in a risk quantification value of 32.00. For example, for the abnormal outbound operation category, S1 is 14.71, W is 2.0, T is 1 day, and R is approximately 26.63. For the operation category outside of working hours, S1 is 12.65, W is 0.5, T is 5 days, and R is approximately 3.84. For the multiple authentication failure category, S1 is 12.23, W is 0.5, T is 8 days, and R is approximately 2.75.

[0108] C5. Mark the feature data whose risk quantification value R is greater than the preset risk threshold as risk features.

[0109] The system sets a preset risk threshold of 10.0, which is determined based on historical security operation experience and can effectively filter out high-risk features that require special attention. The system marks the feature data corresponding to the same type of data with a risk quantification value R greater than 10.0 as risk features.

[0110] After screening, 5 out of 8 similar data categories had risk quantification values ​​exceeding the preset risk threshold. The risk quantification value for privilege escalation was 32.00, greater than 10.0, and was marked as risk feature RISK_001. The risk quantification value for abnormal outbound data transmission was 26.63, greater than 10.0, and was marked as risk feature RISK_002. The risk quantification value for large-scale data downloads was 24.47, greater than 10.0, and was marked as risk feature RISK_003. The risk quantification value for abnormal logins in the early morning was 18.45, greater than 10.0, and was marked as risk feature RISK_004. The risk quantification value for logins from different locations was 11.84, greater than 10.0, and was marked as risk feature RISK_005. The risk quantification value for sensitive file access was 10.40, greater than 10.0, and was marked as risk feature RISK_006. The risk quantification value for operations outside of working hours was 3.84, less than 10.0, and was not marked as a risk feature. The risk quantification value for multiple authentication failures is 2.75, which is less than 10.0, and therefore it is not marked as a risk feature.

[0111] The system sorts six risk features in descending order based on their risk quantification values, establishing a risk feature sequence. The ranking is as follows: RISK_001 ranks first with a risk quantification value of 32.00; RISK_002 ranks second with a risk quantification value of 26.63; RISK_003 ranks third with a risk quantification value of 24.47; RISK_004 ranks fourth with a risk quantification value of 18.45; RISK_005 ranks fifth with a risk quantification value of 11.84; and RISK_006 ranks sixth with a risk quantification value of 10.40. This risk feature sequence will be output to the hierarchical early warning output module for security analysis.

[0112] The graded early warning output module is used to perform security analysis on the risk characteristics in the risk characteristic sequence, obtain the security analysis results, determine the early warning level based on the security analysis results, generate security early warning information and output it.

[0113] The steps for security analysis of risk characteristics in the risk characteristic sequence include D1 to D5:

[0114] D1. Construct a security knowledge graph, which includes security threat entities, attack method entities, protective measure entities, and their relationships.

[0115] D2. Match the risk features in the risk feature sequence with the security knowledge graph to obtain the security threat entities and attack method entities associated with the risk features.

[0116] The system performs semantic matching between each of the six risk features in the risk feature sequence and the security knowledge graph to identify the security threat entity and attack method entity corresponding to each risk feature.

[0117] RISK_001 is a privilege escalation risk characteristic. The system matches this characteristic against the security knowledge graph and identifies a high correlation with entities involved in privilege escalation attacks. The associated security threats include APT attacks and insider threats. Specific attack techniques include exploiting vulnerabilities for privilege escalation, access token manipulation, and bypassing user account controls. RISK_002 is an abnormal outward transmission risk characteristic. The system matches this characteristic and identifies a high correlation with entities involved in data leakage attacks. The associated security threats are data breach threats. Specific attack techniques include leakage through web services, email, and cloud storage. RISK_003 is a large-scale data download risk characteristic. The system matches this characteristic and identifies a high correlation with entities involved in data collection attacks. The associated security threats are data breach threats and insider threats. Specific attack techniques include local data collection, network-shared data collection, and data temporary storage. RISK_004 is a risk characteristic associated with abnormal logins in the early morning. The system has identified this risk characteristic as being associated with entities using initial access attack methods. The associated security threats include account hijacking threats and APT attack threats. Specific attack techniques include effective account exploitation and external remote services. RISK_005 is a risk characteristic associated with logins from a different location. The system has also identified this risk characteristic as being associated with entities using initial access attack methods. The associated security threat is account hijacking. RISK_006 is a risk characteristic associated with sensitive file access. The system has identified this risk characteristic as being associated with entities using discovery attack methods. Specific attack techniques include file and directory discovery and permission group discovery.

[0118] D3. Based on the relationships in the security knowledge graph, perform attack chain analysis on multiple risk characteristics to identify potential attack paths.

[0119] Based on the temporal relationships between attack method entities in the security knowledge graph, the system performs attack chain analysis on six risk characteristics to identify whether these risk characteristics constitute a complete or partial attack path.

[0120] The system first analyzes the attack phases corresponding to each risk characteristic. RISK_004 (abnormal login in the early morning) and RISK_005 (login from a different location) correspond to the initial access phase; RISK_006 (access to sensitive files) corresponds to the discovery phase; RISK_001 (privilege escalation) corresponds to the privilege escalation phase; RISK_003 (large-scale data download) corresponds to the data collection phase; and RISK_002 (abnormal data transmission) corresponds to the data leakage phase. Based on the temporal relationship of the attack phases, the system performs link analysis and finds that these six risk characteristics can be linked into a complete data leakage attack path: the attacker first obtains initial access rights through abnormal login in the early morning or login from a different location, then conducts internal reconnaissance to discover target data through access to sensitive files, then obtains higher privileges to access core resources through privilege escalation, subsequently collects target data through large-scale data download, and finally leaks the data to the outside through abnormal transmission.

[0121] The system identified the attack path as conforming to a typical insider threat or APT attack pattern, with the attack chain consisting of initial access, discovery, privilege escalation, data collection, and data exfiltration, involving a total of five attack phases. The system also identified the absence of a persistence phase and a lateral movement phase in the attack path, potentially indicating that the attacker employed a rapid penetration strategy, or that the attack activities in these two phases have not yet been detected.

[0122] D4. Calculate the completeness score of the attack path. The completeness score is determined based on the ratio of the number of risk features that have appeared in the attack path to the number of features required for a complete attack path.

[0123] The system calculates the completeness score of the currently identified attack path based on the standard attack path template defined in the security knowledge graph. Taking the data breach attack path as a reference, the standard data breach attack path includes 7 stages: initial access, execution, persistence, privilege escalation, discovery, data collection, and data leakage.

[0124] The system compares the currently identified risk characteristics with standard attack paths. The initial access phase has appeared, corresponding to risk characteristics RISK_004 and RISK_005; the execution phase has not shown any obvious characteristics; the persistence phase has not shown any obvious characteristics; the privilege escalation phase has appeared, corresponding to risk characteristic RISK_001; the discovery phase has appeared, corresponding to risk characteristic RISK_006; the data collection phase has appeared, corresponding to risk characteristic RISK_003; and the data leakage phase has appeared, corresponding to risk characteristic RISK_002. Statistical results show that 5 out of the 7 phases of the standard attack path have appeared, with a completeness score of 5 / 7, resulting in a calculated completeness of approximately 71.4%.

[0125] The system also calculates critical phase coverage as a supplementary evaluation metric. The critical phases of a data breach attack include initial access, data collection, and data leakage. The current attack path covers all three critical phases, achieving 100% critical phase coverage. Based on the overall integrity score of 71.4% and the 100% critical phase coverage, the system determines that the current attack path is highly complete, and the attack may have entered its final stage.

[0126] D5. Generate comprehensive security analysis results based on the attack path integrity score and the risk quantification value of risk characteristics.

[0127] The system comprehensively considers the attack path integrity score and the risk quantification value of each risk characteristic to generate a comprehensive security analysis result. The comprehensive security analysis result includes threat rating, attack type determination, impact scope assessment, and remediation recommendations.

[0128] The system first calculates the comprehensive threat score. The formula for the comprehensive threat score is: attack path integrity score multiplied by a weighting coefficient of 0.4, plus the average risk quantification value of the risk features multiplied by a weighting coefficient of 0.6. The attack path integrity score is 71.4%, which translates to 71.4 points on a percentage scale. The risk quantification values ​​for the six risk features are 32.00, 26.63, 24.47, 18.45, 11.84, and 10.40, with an average of 20.63, which translates to approximately 68.8 points on a percentage scale. The comprehensive threat score equals 71.4 multiplied by 0.4 plus 68.8 multiplied by 0.6, resulting in a score of approximately 69.8 points.

[0129] The system determines the alert level based on a comprehensive threat score. There are four alert levels: 0-30 for low alert, 30-50 for medium alert, 50-70 for high alert, and 70 and above for emergency alert. The current comprehensive threat score of 69.8 is at the upper end of the high alert range; however, considering that the critical phase coverage has reached 100% and the attack path has entered the data infiltration phase, the system upgrades the alert level to emergency alert.

[0130] The comprehensive security analysis indicates that the current security incident has been classified as a suspected data breach attack, classified as an insider threat or APT attack. The attack has progressed to the data leakage stage, affecting core business data and sensitive files, and the alert level is emergency alert. Recommended actions include immediately freezing the user accounts corresponding to the affected accounts RELATED_001 and RELATED_007, blocking abnormal outbound channels, initiating the data breach emergency response process, preserving relevant log evidence, and notifying the information security manager and legal department.

[0131] It is also important to know that the steps for determining the warning level based on the security analysis results include E1 to E6:

[0132] E1. Establish a multi-dimensional early warning assessment model, which includes the dimensions of threat urgency, asset importance, and diffusion probability.

[0133] In one specific embodiment, the system establishes a multi-dimensional early warning assessment model comprising three evaluation dimensions. The threat urgency dimension assesses the time urgency and attack progress of the current security threat. This dimension has a score range of 0 to 100, with higher scores indicating a more urgent threat and requiring a faster response. The asset importance dimension assesses the importance of the threatened assets to business operations. This dimension also uses a score range of 0 to 100, with higher scores indicating more important affected assets and more severe potential losses. The propagation probability dimension assesses the likelihood of the security threat spreading laterally within the system and expanding its impact. This dimension also has a score range of 0 to 100, with higher scores indicating a greater likelihood of the threat rapidly spreading to more systems and users.

[0134] The system assigns different weight coefficients to the three dimensions to reflect their importance in early warning decisions. The threat urgency dimension is weighted at 0.4 because the attack's progress is the most direct factor determining response priority. The asset importance dimension is weighted at 0.35 because protecting core business assets is the primary objective of security operations. The propagation probability dimension is weighted at 0.25 because controlling the scope of threat propagation is crucial for minimizing overall losses.

[0135] E2. Determine the threat urgency score based on the completeness score of the attack path; the higher the completeness score, the higher the threat urgency score.

[0136] In one specific embodiment, the system calculates a threat urgency score based on the completeness score of the attack path. The calculation of the threat urgency score takes into account three factors: attack path completeness, coverage of key phases, and the time of recent risk feature appearance.

[0137] The attack path integrity score was 71.4%, which the system directly converted into a basic urgency score of 71.4. Further analysis of the critical phase coverage revealed that all three critical phases of the data breach attack—initial access, data collection, and data leakage—had occurred, resulting in 100% critical phase coverage. The system added 15 urgency points for this. The system also analyzed the recent occurrence time of risk features, finding that the RISK_001 privilege escalation risk feature most recently appeared on the 30th day (the current day), with a timeliness parameter T of 0 days, indicating that the attack is still ongoing. The system added another 10 urgency points for this.

[0138] After comprehensive calculation, the threat urgency score equals the base urgency score of 71.4 plus a critical phase coverage bonus of 15 plus a sustained attack bonus of 10, totaling 96.4 points. Since the score cap is 100 points, the system sets the threat urgency score to 96.4 points, indicating that the current threat is in an extremely high urgency state.

[0139] E3. Determine the asset importance score based on the business level of the system assets associated with risk characteristics.

[0140] The system analyzes system assets involved in six risk characteristics, identifies the business level of these assets, and calculates asset importance scores. The system maintains an asset inventory database, which assigns a business level label to each system asset. The business level is divided into five levels, from low to high: general business assets, important business assets, core business assets, critical infrastructure assets, and strategic assets, with corresponding basic asset importance scores of 20, 40, 60, 80, and 100 points, respectively.

[0141] The system analyzes the assets involved in each risk characteristic one by one. The RISK_001 privilege escalation risk characteristic involves administrator privileges on the core business server. This asset is marked as a critical infrastructure asset with a base importance score of 80 points. The RISK_002 abnormal outbound risk characteristic involves customer databases and financial data. The customer database is marked as a core business asset, corresponding to 60 points, while the financial data is marked as a strategic asset, corresponding to 100 points. The RISK_003 large-scale data download risk characteristic involves R&D documents and technical materials, marked as core business assets, corresponding to 60 points. The RISK_004 abnormal login in the early morning risk characteristic involves employee accounts and intranet access permissions, marked as important business assets, corresponding to 40 points. The RISK_005 remote login risk characteristic also involves employee accounts, corresponding to 40 points. The RISK_006 sensitive file access risk characteristic involves contract documents and trade secrets, marked as core business assets, corresponding to 60 points.

[0142] The system calculates the importance score for all involved assets, taking the highest score as the primary basis for asset importance rating. In this case, the highest asset importance score is 100, corresponding to financial data as a strategic-level asset. Considering that the threat involves multiple high-level assets, the system calculates the number of assets involving core business assets or higher, finding 5 assets at or above the core business asset level. The system adds 5 importance points for these assets. The final asset importance score is determined to be 100, indicating that the threatened asset is of strategic importance to the company.

[0143] E4. Determine the diffusion probability score based on the propagation characteristics of the risk features and the system network topology.

[0144] The system calculates a diffusion probability score based on the propagation characteristics of the risk features and the enterprise's internal network topology. The system first analyzes the location and access scope of the account involved in the network. The User_A account corresponding to RELATED_001 has cross-departmental file access permissions; this account can access shared resources in the R&D, Finance, and Marketing departments. The system's query of the network topology database reveals that the User_A account's workstation has direct network connections to 32 other devices, including 15 employee workstations, 8 departmental file servers, 5 application servers, and 4 database servers.

[0145] The system analyzed the propagation characteristics of the privilege escalation risk signature RISK_001. This risk signature indicates that the attacker has obtained administrator privileges, which grant access to almost all system resources on the enterprise intranet. The system assesses the base score for the propagation potential of this privilege as 80 points. Further analysis of the network topology revealed that the network segment containing the affected account is isolated from the core business network segment by only one firewall layer, and this firewall is configured with multiple permitted business ports, making lateral movement relatively easy. Therefore, the system adds 10 points to the propagation probability score.

[0146] The system also analyzed the attacker's demonstrated technical capabilities. Based on the behavioral patterns of RISK_001 privilege escalation, RISK_003 large-scale data download, and RISK_002 abnormal outbound transmission, the attacker possesses strong technical capabilities and a clear attack target. The system assesses a high probability that the attacker will continue to expand the attack scope, therefore adding 5 points to the diffusion probability score. After comprehensive calculation, the diffusion probability score equals the base score of 80 plus a network topology bonus of 10 plus an attack capability bonus of 5, totaling 95 points, indicating a high risk of threat diffusion.

[0147] E5. The threat urgency score, asset importance score, and proliferation probability score are weighted and summed to obtain a comprehensive early warning score.

[0148] In one specific embodiment, the system calculates a weighted sum of the scores from the three dimensions according to preset weighting coefficients. The threat urgency score is 96.4 points, with a weighting coefficient of 0.4, resulting in a weighted contribution of 38.56 points. The asset importance score is 100 points, with a weighting coefficient of 0.35, resulting in a weighted contribution of 35 points. The proliferation probability score is 95 points, with a weighting coefficient of 0.25, resulting in a weighted contribution of 23.75 points.

[0149] The comprehensive early warning score is equal to the sum of the weighted contribution values ​​of the three dimensions, i.e., 38.56 + 35 + 23.75, which equals 97.31. The system rounds the comprehensive early warning score, ultimately determining it to be 97. This score falls in the highest range of the 0-100 score system, indicating that the current security incident has an extremely high comprehensive threat level.

[0150] E6. Determine the corresponding warning level based on the preset score range of the comprehensive warning score.

[0151] The warning level is determined according to the preset score range rules. There are four warning levels, and the corresponding comprehensive warning score ranges for each level are as follows: low warning corresponds to 0 to 30 points, medium warning corresponds to 31 to 55 points, high warning corresponds to 56 to 75 points, and emergency warning corresponds to 76 to 100 points.

[0152] The current comprehensive early warning score is 97 points, falling within the emergency early warning range of 76 to 100 points. The system has determined the warning level for this security incident to be an emergency warning. An emergency warning level requires the immediate activation of the highest-level emergency response procedures. Relevant security personnel must respond within 15 minutes, the emergency response team must be assembled within 30 minutes, and all involved system resources should be prioritized for protection or isolation.

[0153] Finally, the corresponding warning template is invoked according to the warning level; the structured fields in the warning template are filled in according to the security analysis results to generate security warning information.

[0154] The system invokes the corresponding alert template based on the determined emergency alert level. An emergency alert template is a structured document containing structured fields such as alert title, threat type, threat description, detection time, involved assets, scope of impact, risk assessment, attack indicators, handling recommendations, and response timeframe.

[0155] Example 2, refer to Figure 1 As an embodiment of the present invention, based on the above embodiment, an AI-driven, fully automated security operation method is provided, comprising the following steps:

[0156] Collect safety process data, establish an anomaly analysis model, and input the safety process data into the anomaly analysis model to obtain related features;

[0157] The associated features are marked as abnormal data to be verified, historical normal data is retrieved, and an AI prediction model is built based on the abnormal data to be verified and the historical normal data. The abnormal data to be verified is judged by the AI ​​prediction model.

[0158] Retrieve feature data from security process data and feature data from abnormal data to be verified that are judged to be abnormal. Classify security process data with the same feature data and mark them as similar data. Count the number of times similar data appears. Calculate the probability of similar data appearing based on the number of times similar data appears and the total number of security process data.

[0159] Risk quantification calculation is performed on feature data based on the probability of occurrence of similar data to obtain multiple risk features. The multiple risk features are then sorted according to their risk quantification values ​​to establish a risk feature sequence.

[0160] Perform security analysis on the risk characteristics in the risk characteristic sequence, obtain the security analysis results, determine the warning level based on the security analysis results, generate security warning information and output it.

[0161] In summary, this invention establishes a dual-model anomaly detection mechanism combining temporal and behavioral feature analysis, and employs intersection and difference operations for feature fusion, effectively improving the accuracy of anomaly identification and reducing false positive and false negative rates. By constructing a risk quantification model that integrates information entropy, influence weight, and time decay, it achieves objective quantitative assessment and scientific ranking of security risks, providing a priority basis for security responses. By combining security knowledge graphs with attack chain analysis, it enables associative reasoning from discrete risk features to complete attack paths, allowing for the prediction of attack intent in the early stages of an attack and enhancing proactive security protection. By establishing a multi-dimensional early warning assessment model considering threat urgency, asset importance, and diffusion probability, it achieves precise classification of early warning levels and optimizes the rational allocation of security operation resources. Finally, it automates the entire process from data collection, anomaly detection, risk quantification, security analysis to early warning output, reducing manual intervention and improving security operation efficiency.

[0162] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. An AI-driven full-process automated security operation system, characterized in that, The module comprises a correlation feature extraction module, an anomaly judgment module, a same-class feature data classification module, a sequence construction module, and a hierarchical early warning output module. The correlation feature extraction module is used for collecting safety process data, establishing an anomaly analysis model, inputting the safety process data into the anomaly analysis model to obtain correlation features. The anomaly judgment module is used for marking the correlation features as to-be-verified abnormal data, calling historical normal data, establishing an AI prediction model according to the to-be-verified abnormal data and the historical normal data, and judging the to-be-verified abnormal data through the AI prediction model. The same-class feature data classification module is used for calling feature data in the safety process data and feature data in the to-be-verified abnormal data judged as abnormal, classifying safety process data with the same feature data, marking the safety process data as same-class data, counting the number of occurrences of the same-class data, calculating the occurrence probability of the same-class data according to the number of occurrences of the same-class data and the total number of the safety process data. The sequence construction module is used for risk quantification calculation of the feature data according to the occurrence probability of the same-class data, obtaining a plurality of risk features, sorting the plurality of risk features according to risk quantification values of the risk features, and establishing a risk feature sequence. The hierarchical early warning output module is used for safety analysis of the risk features in the risk feature sequence, obtaining a safety analysis result, determining a warning level according to the safety analysis result, generating safety warning information, and outputting the safety warning information.

2. The AI-driven full-process automation security operation system of claim 1, wherein, The step of obtaining the correlation features comprises: establishing a first anomaly analysis model of the safety process data, analyzing time sequence features of the safety process data through the first anomaly analysis model, and obtaining a first abnormal feature set; establishing a second anomaly analysis model of the safety process data, analyzing behavior features of the safety process data through the second anomaly analysis model, and obtaining a second abnormal feature set; performing intersection operation on the first abnormal feature set and the second abnormal feature set to obtain a common abnormal feature set, and marking abnormal features in the common abnormal feature set as high-confidence abnormal features; performing difference set operation on the first abnormal feature set and the second abnormal feature set to obtain a difference abnormal feature set, and marking abnormal features in the difference abnormal feature set as to-be-confirmed abnormal features; performing union operation on the high-confidence abnormal features and the to-be-confirmed abnormal features to obtain a correlation feature set, and marking features in the correlation feature set as correlation features. 3.The AI-driven full-process automation security operation system of claim 2, wherein, The analysis method of the first anomaly analysis model comprises: performing time sequence decomposition on the safety process data, and decomposing the safety process data into a trend component, a periodic component, and a residual component; when analyzing the trend component, setting a trend deviation threshold, and when the change rate of the trend component exceeds the trend deviation threshold, marking the corresponding safety process data as trend abnormal data; when analyzing the periodic component, constructing a periodic baseline model, and when the deviation of the periodic component from the periodic baseline model exceeds a preset periodic deviation threshold, marking the corresponding safety process data as periodic abnormal data; when analyzing the residual component, calculating the standard deviation of the residual component, and when the residual component exceeds a preset standard deviation multiple, marking the corresponding safety process data as residual abnormal data; The trend anomaly data, the periodic anomaly data and the residual anomaly data are summarized to generate a first anomaly feature set.

4. The AI-driven full-process automation safe operation system according to claim 3, wherein, The analysis method of the second anomaly analysis model comprises: behavior sequences in the security process data are extracted, and a normal behavior baseline library is established; similarity calculation is performed on the behavior feature vector of the current security process data and the behavior pattern feature vector in the normal behavior baseline library, and when the similarity is lower than a preset similarity threshold, the corresponding security process data is marked as behavior deviation anomaly data; a behavior state transition matrix is analyzed, and a low-probability transition path is identified, and when the low-probability transition path appears in the security process data, the corresponding security process data is marked as transition anomaly data; the behavior deviation anomaly data and the transition anomaly data are summarized to generate a second anomaly feature set. 5.The AI-driven full-process automation security operation system of claim 4, wherein, The steps of judging the to-be-verified anomaly data by the AI prediction model comprise: a multi-level classification neural network is constructed as the AI prediction model, the multi-level classification neural network comprising a feature extraction layer, an attention mechanism layer and a classification output layer; the to-be-verified anomaly data is input into the feature extraction layer, and the to-be-verified anomaly data is subjected to deep feature extraction by the feature extraction layer to obtain a deep feature vector; the deep feature vector is input into the attention mechanism layer, and the key features in the deep feature vector are subjected to weighted processing by the attention mechanism layer to obtain a weighted feature vector; the historical normal data is processed by the same feature extraction layer and attention mechanism layer to obtain a weighted feature vector of the normal data; in the classification output layer, a distance measurement value between the weighted feature vector of the to-be-verified anomaly data and the weighted feature vector of the normal data is calculated, and according to a comparison result of the distance measurement value and a preset discrimination threshold, it is judged whether the to-be-verified anomaly data is real anomaly data. 6.The AI-driven full-process automation security operation system of claim 5, wherein, The steps of risk quantification calculation on the feature data according to the occurrence probability of the same type of data comprise: an occurrence probability P of the same type of data is obtained, and the information entropy value S1 of the same type of data is calculated according to the formula S1=-log2(P); an influence degree value D of a historical security event corresponding to the same type of data is obtained, and the influence weight coefficient W is determined according to the influence degree value D; a timeliness parameter T of the same type of data is obtained, and the timeliness parameter T is determined according to the interval between the latest occurrence time of the same type of data and the current time; According to the formula R = S1 x W x e (-λT) A risk quantification value R is calculated, where λ is a time decay coefficient; the feature data with a risk quantification value R greater than a preset risk threshold is marked as a risk feature.

7. The AI-driven based full-process automation security operation system of claim 6, wherein, The steps of security analysis on the risk features in the risk feature sequence comprise: a security knowledge graph is constructed, and the security knowledge graph comprises security threat entities, attack method entities, protection measure entities and their associated relationships; the risk features in the risk feature sequence are matched with the security knowledge graph to obtain the security threat entities and the attack method entities associated with the risk features; according to the associated relationships in the security knowledge graph, attack chain analysis is performed on the multiple risk features to identify potential attack paths; a completeness score of the attack path is calculated, and the completeness score is determined according to the ratio of the number of risk features that have appeared in the attack path to the number of required features of a complete attack path. According to the completeness score of the attack path and the risk quantification value of the risk feature, a comprehensive security analysis result is generated.

8. The AI-driven full-process automation security operation system of claim 7, wherein, The step of determining the warning level according to the security analysis result comprises: A multi-dimensional warning evaluation model is established, which includes threat urgency dimension, asset importance dimension and diffusion possibility dimension; A threat urgency score is determined according to the completeness score of the attack path, and the higher the completeness score, the higher the threat urgency score; An asset importance score is determined according to the business level of the system asset associated with the risk feature; A diffusion possibility score is determined according to the propagation characteristics of the risk feature and the system network topology; The threat urgency score, asset importance score and diffusion possibility score are weighted and summed to obtain a comprehensive warning score; According to the preset score interval where the comprehensive warning score is located, the corresponding warning level is determined. 9.The AI-driven full-process automation security operation system of claim 8, wherein, The step of generating and outputting security warning information comprises: According to the warning level, the corresponding warning template is called; According to the security analysis result, each structured field in the warning template is filled in to generate the security warning information.

10. An AI-driven full-process automatic security operation method, applying the system of any one of claims 1-9, characterized in that, The steps include: Collecting security process data, establishing an anomaly analysis model, inputting the security process data into the anomaly analysis model to obtain associated features; The associated features are marked as abnormal data to be verified, historical normal data is called, and an AI prediction model is established according to the abnormal data to be verified and the historical normal data. The abnormal data to be verified is judged by the AI prediction model; The feature data in the security process data and the feature data in the abnormal data to be verified judged as abnormal are called, the security process data with the same feature data are classified and marked as similar data, the occurrence frequency of the similar data is counted, and the occurrence probability of the similar data is calculated according to the occurrence frequency of the similar data and the total number of the security process data; According to the occurrence probability of the similar data, the risk quantification calculation of the feature data is performed to obtain a plurality of risk features, the risk features are sorted according to the risk quantification value of the risk features, and a risk feature sequence is established; The risk features in the risk feature sequence are subjected to security analysis to obtain a security analysis result, the warning level is determined according to the security analysis result, and the security warning information is generated and output.