A security risk assessment system and method for digital archives
By weighting and screening the characteristics of digital archives, combining the entropy value method and CRITIC method, selecting suitable algorithms for further analysis, and generating threat reports, solving the problems of high computational costs and insufficient adaptability of existing systems, and achieving more efficient security risk assessment.
Patent Information
- Application Number
- CN202510309238.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-17
AI Technical Summary
When facing a complex and changing network security environment, the existing security risk assessment system has high computing costs and lacks dynamic adaptability, which makes threat detection easy to be misjudged and difficult to adapt to the security management needs of digital archives.
By conducting preliminary analysis of archival features, combined with entropy value method and CRITIC method for weighting, further analysis is used using isolated forest algorithm or cluster analysis algorithm for further analysis, threat reports are generated and risk distribution and trends are displayed using charts or dashboards to reduce calculation costs and improve dynamic adaptability.
It achieves higher dynamic adaptability and lower computing costs, enables more accurate identification and display of security risks of digital archives, and supports real-time monitoring and analysis.
Smart Images

Figure CN119830192B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of security risk assessment, and more specifically, the present invention relates to a security risk assessment system and method for digital archives. Background Art
[0002] With the rapid development of information technology, digital archives are increasingly widely used in various fields. However, the security of digital archives faces unprecedented challenges. Problems such as cyber attacks, data leakage, malware attacks, and improper operations by internal personnel are becoming increasingly serious, bringing huge pressure to the security management of digital archives.
[0003] At present, traditional security measures such as firewalls, intrusion detection systems, and encryption technologies can, to a certain extent, prevent external threats, but they are unable to cope with the complex and ever-changing network security environment. Inability to adapt to complex environments leads to easy misjudgment in threat detection. In recent years, with the development of artificial intelligence and big data technologies, intelligent security risk assessment systems have gradually become a research hotspot. These systems can, through statistical analysis and machine learning algorithms, monitor and analyze multi-dimensional data such as user behavior, system vulnerabilities, and network traffic in real time, identify behaviors that deviate from normal patterns, and thus achieve more accurate risk assessment. However, existing systems still have some deficiencies, such as high computational costs and lack of dynamic adaptability, which limit their promotion and effectiveness in practical applications. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a security risk assessment system and method for digital archives to solve the problems raised in the above background art.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A security risk assessment method for digital archives, comprising the following steps:
[0007] Preliminarily analyze the archive features, preliminarily judge the abnormal features, and sort out the abnormal features; weight different abnormal features, and the weight of each abnormal feature is obtained by combining the entropy method and the CRITIC method;
[0008] Calculate the compliance coefficient through weighted summation according to the abnormal feature weight, timeliness, and abnormal feature data volume, and judge whether to use a single algorithm among the isolation forest algorithm and the clustering analysis algorithm to further analyze the abnormal features by comparing the compliance coefficient with the system threshold.
[0009] In a preferred embodiment, sorting out the abnormal features includes cleaning the abnormal features and counting the abnormal feature data volume.
[0010] In a preferred embodiment, the different abnormal features are weighted. Specifically, for the abnormal features, the entropy method is first used to calculate the difference coefficient, and then the CRITIC method is used to calculate the correlation; the weights of the abnormal features are calculated by weighted summation according to the difference coefficient and the correlation.
[0011] In a preferred embodiment, when the archive features are preliminarily analyzed and the abnormal features are screened out, the compliance coefficient is calculated by weighted summation according to the weights of the abnormal features, timeliness, and the amount of abnormal feature data.
[0012] In a preferred embodiment, the value of the compliance coefficient is determined. If the value of the compliance coefficient is greater than the system threshold, the Isolation Forest algorithm is used to further analyze the abnormal features. If the value of the compliance coefficient is less than the system threshold, the clustering analysis algorithm is selected to further analyze the abnormal features.
[0013] In a preferred embodiment, the evaluation system includes:
[0014] A screening module for preliminarily analyzing the archive features, preliminarily judging the weights of the abnormal features, and counting the amount of abnormal feature data;
[0015] A weighting module for weighting different abnormal features, and the weight of each abnormal feature is obtained by combining the entropy method and the CRITIC method;
[0016] An evaluation module for further analyzing the abnormal features to determine the specific threat situation.
[0017] In a preferred embodiment, the evaluation module internally includes a selection module. The selection module is used to calculate the compliance coefficient by weighted summation according to the weights of the abnormal features, timeliness, and the amount of abnormal feature data, and to judge whether to use the Isolation Forest algorithm or the clustering analysis algorithm, one of the two algorithms, to further analyze the abnormal features by comparing the compliance coefficient with the system threshold.
[0018] The technical effects and advantages of the present invention:
[0019] The present invention first collects data related to archives and preprocesses the collected data. Subsequently, a preliminary analysis of the archive features is carried out, the weights of abnormal features are preliminarily judged, and the abnormal features are sorted out; different abnormal features are weighted, and the weight of each abnormal feature is obtained by combining the entropy method and the CRITIC method. The weight of the abnormal feature, timeliness, and the amount of abnormal feature data are used to calculate the compliance coefficient by weighted summation. By comparing the compliance coefficient with the threshold, it is determined whether to use the isolation forest algorithm or the clustering analysis algorithm to further analyze the threat-related features. Finally, after determining the content of the threat report, the distribution and trend of risks are visually displayed in the form of charts or dashboards, and the data generated during the operation of the system is saved for subsequent reference and analysis, making the dynamic adaptability higher and the calculation cost lower. Description of the Drawings
[0020] For the convenience of those skilled in the art to understand, the present invention will be further described below in conjunction with the accompanying drawings;
[0021] Figure 1 It is a schematic flowchart of a method for security risk assessment of digital archives according to the present invention;
[0022] Figure 2 It is a schematic structural diagram of a security risk assessment system for digital archives according to the present invention. Detailed Embodiments
[0023] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0024] The present invention first collects data related to archives and preprocesses the collected data. Subsequently, a preliminary analysis of the archive features is carried out, the weights of abnormal features are preliminarily judged, and the abnormal features are sorted out; different abnormal features are weighted, and the weight of each abnormal feature is obtained by combining the entropy method and the CRITIC method. The weight of the abnormal feature, timeliness, and the amount of abnormal feature data are used to calculate the compliance coefficient by weighted summation. By comparing the compliance coefficient with the threshold, it is determined whether to use the isolation forest algorithm or the clustering analysis algorithm to further analyze the threat-related features. Finally, after determining the content of the threat report, the distribution and trend of risks are visually displayed in the form of charts or dashboards, and the data generated during the operation of the system is saved for subsequent reference and analysis.
[0025] Embodiment 1
[0026] A security risk assessment method for digital archives according to the present invention is as follows Figure 1 shown, and mainly includes the following steps:
[0027] Step S1: Collect archive-related data, preprocess the collected archive-related data; and extract features from the processed archive-related data;
[0028] Step S2: Use an anomaly detection method based on statistical process control to preliminarily analyze the features, preliminarily determine the anomaly feature weights; sort out the anomaly features, and weight the anomaly features;
[0029] Step S3: Determine the anomaly feature weights, timeliness, and anomaly feature data volume, calculate the compliance coefficient through weighted summation, and judge whether to use the isolation forest algorithm or the clustering analysis algorithm to further analyze the threat-related features by comparing the compliance coefficient with the threshold; generate a threat report;
[0030] Step S4: Determine the threat report, and intuitively display the risk distribution and trend using charts or dashboards; save the data generated during the operation of the system.
[0031] In step S1, the archive-related data includes but is not limited to the archive storage location, access logs, permission settings, network traffic, etc.; first, identify the data types to be collected, and each type of data has its specific format and storage location; then determine the data collection frequency and detail. For example, access logs may be collected in real time, while network traffic may be summarized hourly; set by the management personnel according to requirements; finally, define the data storage format, such as CSV, JSON, or database table structure, for subsequent processing and analysis.
[0032] Identify data sources: Server logs: Locate the location of the server log files; Database records: Confirm the database table structure and field information for storing archive information; User operation records: Obtain the log information of user logins, queries, downloads, etc.; Network traffic: Obtain relevant traffic data through network packet capture tools.
[0033] Access data sources: Use a log management tool to centrally manage server logs; Use a database connection tool to access database records; Configure a network monitoring tool to capture network traffic data.
[0034] After obtaining the archive-related data, preprocess the archive-related data, specifically including the following steps:
[0035] Step A1: Clean the archive-related data, remove noise data, and ensure the integrity and consistency of the data;
[0036] Step A2: Transform the data related to the files, unify the formatting of data from different sources to facilitate subsequent analysis.
[0037] The specific Step A1 further includes the following steps:
[0038] Step A1.1: Detect duplicate records using a hash table or the uniqueness constraint of the database; identify duplicate entries by checking fields such as timestamps, IP addresses, and operation types in the logs; then use SQL statements or programming languages to delete duplicate records, while recording the number and reasons for the deleted duplicate records for subsequent review.
[0039] Step A1.2: Identify invalid logs, check whether the log format is correct; use regular expressions to verify whether the log content conforms to the expected format; then delete log entries with incorrect formats or that cannot be parsed; at the same time, establish log quality standards and regularly check the validity of the logs.
[0040] Step A1.3: First, convert all timestamps to a unified format; then use tools such as the pytz library in Python to process and convert timestamps in different time zones to a unified time zone; finally, check whether there are jumps or overly large intervals in the timestamps to ensure the continuity of the time series.
[0041] Step A1.4: Supplement missing fields, fill in missing data with default values or use interpolation methods to fill in missing values; at the same time, mark the missing data and label the missing fields in the data.
[0042] Step A1.5: Filter out abnormal data. First, use statistical methods to identify outliers; then remove data that clearly does not conform to business logic, delete negative numbers, overly high values, or other unreasonable values; at the same time, save the removed abnormal data to a separate log file for subsequent analysis.
[0043] The specific Step A2 includes the following steps:
[0044] Step A2.1: Convert data from different sources (such as access logs, network traffic, system logs) to a unified format; and convert field names to ensure consistent field naming.
[0045] Step A2.2: First, perform standardization processing on numerical data; use the Min - Max standardization formula:
[0046] ;
[0047] Step A2.3: Encode categorical data, and perform pre - processing operations such as word segmentation and stop - word removal on text fields.
[0048] Step A2.4: Convert the timestamp to a unified time format; Aggregate data by time period, such as aggregating data by time periods of hours, days, weeks, etc.
[0049] Feature extraction is the core step in the risk identification and analysis module. Its goal is to extract key features from the preprocessed data for subsequent threat identification and vulnerability detection. The following are the detailed method steps:
[0050] Step B1: Extract key features from user behavior data. For small to medium-sized data, the Python ecosystem is the first choice; for large-scale data, Spark or Hadoop needs to be used. Identify abnormal behavior patterns; Login time distribution: Statistically analyze the login time distribution of users to check for abnormal login patterns (such as frequent logins during non-working hours). Login frequency: Calculate the number of logins of users per unit time to identify high-frequency login behaviors. Resource types accessed: Extract the types of resources accessed by users (such as documents, pictures, videos) and analyze the access patterns. IP address changes: Analyze the changes in the login IP addresses of users. Statistically analyze the access frequency of users to identify abnormal fluctuations; Extract the quantity and types of resources accessed by users.
[0051] Step B2: Extract key features from system operation data to identify potential security vulnerabilities; System monitoring tools are used to collect and store data; Security log analysis tools are used to identify abnormal behaviors; CPU usage rate: Extract the CPU usage rate data of the system and analyze whether there are abnormal fluctuations; Memory occupancy: Analyze the memory usage of the system to identify memory leaks or abnormal occupancy; Disk I / O: Extract the frequency and size of disk read and write operations and analyze whether there are abnormal activities; Extract the list of processes running in the system to identify unknown or suspicious processes; Analyze the behavior characteristics of processes.
[0052] Step B3: Extract key features from network traffic data to identify potential network threats. Obtain using the Wireshark tool; Packet size: Extract the packet size distribution in network traffic to identify abnormal patterns; Transmission protocol: Analyze the transmission protocols used in network traffic; Connection duration: Extract the duration of network connections to identify abnormal connection behaviors; Analyze the destination IP addresses and ports of network connections to identify abnormal connection behaviors; Statistically analyze the frequency and patterns of connections per unit time.
[0053] In step S2, first use the anomaly detection method based on statistical process control to conduct a preliminary analysis of the features, which specifically includes the following steps:
[0054] Step C1: Build a model based on statistical methods to identify abnormal behavior patterns; detect abnormal fluctuations in time series; calculate statistical measures such as the mean, variance, and standard deviation of time series, and use moving average method, exponential smoothing method, or ARIMA model to predict normal behavior patterns; compare the deviation between actual behavior and predicted behavior to identify abnormal fluctuations.
[0055] Identify abnormal samples through statistical hypothesis testing; set the null hypothesis and alternative hypothesis; select a test method: use methods such as t-test and chi-square test to test the hypothesis; judge whether to reject the null hypothesis according to the test results.
[0056] Establish a relationship model between variables to predict potential threats; use linear regression or logistic regression to analyze the relationship between features and threats; train a regression model based on historical data; use the model to predict potential threats and analyze the degree of influence of features on threats;
[0057] Identify abnormal behavior based on probability distribution; assume that normal behavior follows a certain probability distribution; calculate the probability density or cumulative distribution function value of the sample; detect behaviors beyond the confidence interval.
[0058] Step C2: Identify abnormal behavior according to the results of the statistical model; identify samples that deviate from the normal range; calculate the mean and variance of features to identify samples that deviate from the normal range; calculate the quantiles of features to identify samples beyond the range.
[0059] Quantify the degree of abnormality of samples; calculate the Z-score of each sample through Z-score calculation to identify behaviors that deviate from the mean. Identify outliers beyond the interquartile range through box plots.
[0060] Record abnormal samples and their reasons; mark abnormal samples according to statistical results;
[0061] Step C3: Interpret the results of the anomaly detection method based on statistical process control; classify threats according to abnormal features; classify threats according to abnormal behavior patterns.
[0062] Through the above steps, potential threats can be systematically identified using statistical modeling and anomaly detection methods. From time series analysis to hypothesis testing, then to regression analysis and probability distribution modeling, each step provides a solid theoretical basis for threat detection and preliminarily determines the weights of abnormal features.
[0063] After identifying the weights of abnormal features, collect the relevant abnormal features collected and count the amount of abnormal feature data.
[0064] For abnormal features, first, the entropy method is used to calculate the difference coefficient xs, and then the CRITIC method is used to calculate the correlation gs; the abnormal feature weights are obtained by weighted summation, and the specific formula is as follows: Q = δxs + εgs; where Q is the abnormal feature weight, and δ and ε are the weight coefficients of the difference coefficient and the correlation, respectively.
[0065] In step S3, when statistical analysis detects significant abnormalities, determine the abnormal feature weights, timeliness, and abnormal feature data volume to judge whether to use the Isolation Forest algorithm or the clustering analysis algorithm to further detect the abnormal feature weights.
[0066] Timeliness plays a crucial role in file threat detection. The Isolation Forest algorithm has more advantages in scenarios that require quick response with its efficient real-time processing ability and low resource consumption; while the clustering analysis algorithm is outstanding in discovering potential patterns in data and is suitable for offline analysis, and the timeliness is set by the staff themselves.
[0067] Normalize the abnormal feature weights, timeliness, and abnormal feature data volume; use the weighted summation formula to calculate the compliance coefficient, and the formula is as follows: F = αxn + βsx + γgm; where α, β, and γ are the weight coefficients of the abnormal feature weight, timeliness, and abnormal feature data volume; xn is the abnormal feature weight; sx is the timeliness; gm is the abnormal feature data volume.
[0068] Select the algorithm according to the value of the compliance coefficient. If the value is greater than the system threshold, adopt the Isolation Forest algorithm; otherwise, select the clustering analysis.
[0069] Isolation Forest: Suitable for scenarios directly detecting abnormalities, such as network intrusion detection, malware identification, etc.; especially suitable for anomaly detection of high-dimensional data and large datasets; suitable for real-time and efficient scenarios because of its low time complexity (O(n log n)).
[0070] Clustering analysis: Suitable for scenarios that require clear grouping, such as user behavior classification, system log grouping, etc.; can discover the natural grouping structure in data to help identify different types of threats; suitable for exploratory data analysis and can reveal potential patterns in data.
[0071] Select the algorithm according to the value of the compliance coefficient. If the value is greater than the system threshold, adopt the Isolation Forest algorithm. The steps of using the Isolation Forest for threat detection are as follows:
[0072] Step D1: Determine the threat-related features, and divide the data into a training set and a test set (such as 70% training set + 30% test set); ensure that the training set contains enough normal samples;
[0073] Step D2: Build a tree model and initialize parameters; Set the key parameters of the Isolation Forest:
[0074] n_estimators: The number of isolation trees;
[0075] max_samples: The proportion of samples used for each tree;
[0076] random_state: Random seed to ensure reproducibility of results;
[0077] contamination: The expected proportion of abnormal samples, set according to business experience;
[0078] Step D3: Train the model, use the training set data to train the Isolation Forest model; The model constructs multiple isolation trees by randomly splitting the data; The structure of each tree determines the isolation depth of the samples.
[0079] Step D4: Use the trained model to identify potential threats; Calculate the anomaly score for each sample in the test set, and the anomaly score reflects the degree of abnormality of the sample; The anomaly analysis calculation formula is as follows: ; where x represents the sample to be detected; Average Depth(x) represents the average isolation depth of sample x in all isolation trees in the Isolation Forest. The isolation depth refers to the number of splits required to isolate sample x alone. Outliers usually have a smaller isolation depth because they are easier to isolate; Expected Depth represents the expected depth when randomly splitting a complete binary tree of height h in an ideal situation, and the calculation formula is ; where H(n) is the harmonic number, defined as: ; n is the number of samples in the tree; Anomaly Score(x) represents the anomaly score of sample x. The value range of the anomaly score is [0, 1]. The closer the score is to 1, the more likely the sample is an outlier;
[0080] Set the threshold of the anomaly score according to business requirements and data distribution. Usually, samples with an anomaly score greater than 0.9 are marked as high-risk threats.
[0081] Mark abnormal samples according to the threshold; Record the reasons for anomalies; Interpret the detection results and generate a threat report.
[0082] Select the algorithm according to the value of the coincidence coefficient. If the value is less than the system threshold, adopt the clustering analysis algorithm. The steps are as follows:
[0083] Step E1: Determine the threat-related features, divide the data into a training set and a test set (e.g., 70% training set + 30% test set); Ensure that the training set contains enough normal samples;
[0084] Step E2: Select a suitable clustering algorithm according to the data characteristics and business requirements. If the data is spherically distributed, select K-means; if the data distribution is complex, select DBSCAN or OPTICS;
[0085] Step E3: Initialize the model parameters according to the selected algorithm; Example: For K-means, set the number of clusters; for DBSCAN, set epsilon and min_samples;
[0086] Step E4: Use the training set data to train the clustering model; The model divides the data into multiple clusters by calculating the similarity between samples;
[0087] The result of the clustering model is unlabeled clusters, and each cluster represents a group of behavior patterns with similar characteristics; Example: Detecting that the user behavior of a cluster is to frequently access sensitive resources may indicate unauthorized access behavior;
[0088] Step E5: Identify abnormal behavior patterns by analyzing the characteristics of the clusters; Analyze the behavior characteristics of each cluster; Identify the clusters whose behavior patterns deviate significantly from the normal patterns; Example: Detecting that the access frequency of a cluster of users is much higher than that of other clusters may indicate high-frequency attack behavior; Interpret the detection results and generate a threat report.
[0089] In step S4, receive the threat report, generate a comprehensive risk report, and display the results visually. Organize the data from different detection methods to ensure that the report content covers all identified threats and their characteristics; List the type of each threat, explain the systems, departments, or business processes that each threat may affect; Provide specific examples or cases to support the description.
[0090] Use charts or dashboards to visually display the risk distribution and trends; Example: Line chart: Show the change trend of abnormal access frequency; Bar chart: Compare the quantity distribution of different threat types; Heat map: Display the risk level distribution in different time periods.
[0091] Save the data generated by the system operation for the convenience of staff verification.
[0092] Embodiment 2
[0093] A security risk assessment system for digital archives according to the present invention, as Figure 2 shown, mainly includes the following modules:
[0094] A screening module, a weighting module, and an evaluation module;
[0095] The screening module is used to perform a preliminary analysis on the archive features, preliminarily judge the weights of abnormal features, and count the amount of abnormal feature data;
[0096] A weighting module that weights different abnormal features, and the weight of each abnormal feature is obtained by combining the entropy method and the CRITIC method;
[0097] An evaluation module for further analyzing the abnormal features to determine the specific threat situation.
[0098] The evaluation module internally includes a selection module, and the selection module is used to select an isolation forest module or a clustering analysis module to further analyze the abnormal features according to the abnormal feature weight, timeliness, and abnormal feature data volume.
[0099] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0100] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces, and the indirect coupling or communication connection of the devices or units may be in an electrical, mechanical, or other form.
[0101] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0102] In addition, the functional units in each embodiment of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0103] As described above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims described.
Claims
1. A security risk assessment method for digital archives, characterized in that It includes the following steps: Conduct a preliminary analysis of the file features, preliminarily identify abnormal features, and sort out the abnormal features; weight the difference coefficient and correlation of the abnormal features to obtain the corresponding abnormal feature weights, and each abnormal feature weight is obtained by combining the entropy method and the CRITIC method; The abnormal features include extracting key features from user behavior data, extracting key features from system operation data, and extracting key features from network traffic data; Calculate the compliance coefficient by weighted summation according to the abnormal feature weight, timeliness, and abnormal feature data volume, and judge whether to use a single algorithm among the isolation forest algorithm and the clustering analysis algorithm to further analyze the abnormal features by comparing the compliance coefficient with the system threshold.
2. The security risk assessment method for digital archives according to claim 1, characterized in that: Sort out the abnormal features, including cleaning the abnormal features and counting the abnormal feature data volume.
3. A security risk assessment method for digital archives according to claim 1, characterized in that: Regarding the weighting of the difference coefficient and correlation of the abnormal features, specifically, first use the entropy method to calculate the difference coefficient for the abnormal features, and then use the CRITIC method to calculate the correlation; calculate the weight of the abnormal features by weighted summation according to the difference coefficient and correlation.
4. A security risk assessment method for digital archives according to claim 1, characterized in that: When conducting a preliminary analysis of the file features and screening out abnormal features, calculate the compliance coefficient by weighted summation according to the abnormal feature weight, timeliness, and abnormal feature data volume.
5. The security risk assessment method for digital archives according to claim 4, characterized in that: Determine the value of the compliance coefficient. If the value of the compliance coefficient is greater than the system threshold, use the isolation forest algorithm to further analyze the abnormal features. If the value of the compliance coefficient is less than the system threshold, select the clustering analysis algorithm to further analyze the abnormal features.
6. A security risk assessment system for digital archives, which is used to implement the security risk assessment method for digital archives according to any one of the above claims 1-5, and is characterized in that: The evaluation system includes: A screening module for conducting a preliminary analysis of the file features, preliminarily identifying abnormal features, and counting the abnormal feature data volume; A weighting module for weighting the difference coefficient and correlation of the abnormal features to obtain the corresponding abnormal feature weights, and each abnormal feature weight is obtained by combining the entropy method and the CRITIC method; An evaluation module for further analyzing the abnormal features and determining the specific situation of the threat.
7. The security risk assessment system for digital archives according to claim 6, characterized in that: The evaluation module internally includes a selection module, and the selection module is used to calculate the compliance coefficient by weighted summation according to the abnormal feature weight, timeliness, and abnormal feature data volume, and judge whether to use a single algorithm among the isolation forest algorithm and the clustering analysis algorithm to further analyze the abnormal features by comparing the compliance coefficient with the system threshold.
Citation Information
Patent Citations
Abnormal battery cell detection method and device
CN116522263A
File electronization information security system
CN119377998A