A risk early warning method, device and equipment of an asset and a computer program product
Patent Information
- Application Number
- CN202611019483.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-09
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]本发明的目的是提供一种资产的风险预警方法、装置、设备及计算机程序产品,用于解决现有技术中资产热度评估不准,风险预测滞后的问题
[0057]本发明实施例,通过对访问数据提取的访问特征评估目标资产的热度值,提升了热度评估的准确性,并且结合目标资产与访问账号之间的风险关联特征组合和热度值对目标资产的风险进行预测,能够提前预判潜在风险,避免因为风险预测滞后带来的损失。
Smart Images

Figure CN122840668A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data security technology, and in particular to a method, apparatus, equipment and computer program product for risk warning of assets. Background Technology
[0002] In the field of data security, well-known technologies for asset access security management include access log statistics, static popularity ranking, and rule-based risk alerts. Among these, access log statistics are mostly single-dimensional data records, lacking in-depth feature mining; asset popularity assessments rely heavily on explicit indicators such as access frequency, without considering implicit factors such as access time, operation type, and permission matching; risk alerts mostly use preset fixed rules, which can only achieve "post-event alerts" and cannot predict potential risks in advance or prevent risks. Summary of the Invention
[0003] The purpose of this invention is to provide a method, apparatus, device, and computer program product for asset risk early warning, which solves the problems of inaccurate asset popularity assessment and delayed risk prediction in the prior art.
[0004] To achieve the above objectives, embodiments of the present invention provide a risk warning method for assets, comprising:
[0005] Feature extraction is performed on the access data of access accounts that access the target asset within the first preset time period to obtain access features;
[0006] The popularity value of the target asset within the first preset time period is evaluated based on the access characteristics.
[0007] The access data is subjected to risk association feature identification to obtain a combination of risk association features; wherein, the combination of risk association features is a set of risk association features; the risk association features are used to characterize the association risk between the target asset and the access account;
[0008] The popularity value and the access data are input into a risk prediction model to obtain the risk probability; wherein, the risk prediction model is used to predict the probability that the combination of risk-related features exists in the access data;
[0009] Based on the access characteristics and the risk probability, the target warning level of the target asset among multiple preset warning levels is determined.
[0010] Optionally, in the method, feature extraction is performed on access data of access to the target asset within a first preset time period to obtain access features, including:
[0011] The access data is input into a pre-trained model to obtain the access features.
[0012] Optionally, the method further includes:
[0013] The corresponding handling measures for the target early warning level shall be implemented for the target asset;
[0014] Continuously collect feedback data obtained after the disposal measures are performed on the target asset;
[0015] Acquire new access data of the target asset within a second preset time period; wherein the second preset time period is a time period later than the first preset time period;
[0016] The pre-trained model is incrementally trained based on the feedback data and the newly added access data.
[0017] Optionally, the method further includes:
[0018] If the handling measures are deemed effective based on feedback data, the feature weights of the access characteristics in the popularity assessment model are optimized based on the feedback data and the new access data. The feedback data is data obtained after the handling measures are applied to the target asset; the handling measures are actions corresponding to the target warning level performed on the target asset; the new access data is access data for accessing the target asset within a second preset time period; the second preset time period is a time period later than the first preset time period; and the popularity assessment model is used to assess the popularity value of the target asset within the first preset time period based on the access characteristics.
[0019] Optionally, the method further includes:
[0020] The optimization scheme for the feature weights is stored in the optimization knowledge base using a distributed storage architecture.
[0021] Optionally, in the method, the plurality of preset warning levels include one or more of the following:
[0022] High-priority risks;
[0023] Medium-priority risks;
[0024] Low priority risks;
[0025] The following measures will be taken for the target asset in accordance with the corresponding early warning level:
[0026] When the target warning level is high-priority risk, the measures taken against the target asset include freezing the access rights of abnormal accounts to the target asset and blocking the access links of the abnormal accounts; wherein, the abnormal account is an account that has abnormal access behavior to the target asset as determined by the combination of risk association characteristics.
[0027] When the target warning level is medium priority risk, the handling measures for the target asset include verifying the abnormal access behavior with the person in charge of the abnormal account of the target asset, and tracking the subsequent access behavior of the abnormal account.
[0028] If the target warning level is low priority risk, the disposal measures for the target asset are placed in the verification queue.
[0029] Optionally, the method, wherein determining the target warning level of the target asset among multiple preset warning levels based on the access characteristics and the risk probability includes:
[0030] Priority scores are calculated in real time based on the access characteristics and the risk probability;
[0031] The target warning level is determined based on the priority score.
[0032] Optionally, the method further includes:
[0033] Collect multi-source access related data for the target asset;
[0034] The multi-source access-related data is standardized, cleaned, converted, deduplicated, merged, and validated for data consistency to obtain access data with a consistent format.
[0035] The multi-source access related data includes at least two of the following:
[0036] The access log data of the target asset includes access time, account identifier, operation type, access address and access result;
[0037] The attribute data of the target asset includes the asset sensitivity level of the target asset, the system to which the target asset belongs, the department to which the target asset belongs, and the core business relevance of the target asset.
[0038] The account permission data of the target asset includes the access account level, the authorization scope of the access account, the department to which the access account belongs, and the status of the access account; the access account is the account that accesses the target asset.
[0039] Optionally, the method, wherein assessing the popularity value of the target asset within the first preset time period based on the access characteristics includes:
[0040] Obtain an access feature score based on the access features;
[0041] The popularity value is obtained by adding the product of the access feature score and the corresponding feature weight.
[0042] Optionally, the method further includes:
[0043] Obtain the original risk prediction model formed by fusing supervised learning and unsupervised learning models;
[0044] Labeled historical access data is input into the supervised learning model to obtain a first training risk probability, and unlabeled historical access data is input into the unsupervised learning model to obtain a second training risk probability; wherein, the labeled historical access data is access data of a first historical time with real labels; and the unlabeled historical access data is access data of the first historical time without labels.
[0045] The combined training risk probability is obtained by using a weighted voting mechanism based on the first training risk probability and the second training risk probability.
[0046] By comparing the combined trained risk probability with the real labels in the historical access data of the labels, the parameters of the original risk prediction model are adjusted to obtain the risk prediction model.
[0047] To achieve the above objectives, embodiments of the present invention provide an asset risk warning device, comprising:
[0048] The first acquisition module extracts features from the access data of access accounts that access the target asset within a first preset time period to obtain access features.
[0049] The first processing module is used to evaluate the popularity value of the target asset within the first preset time period based on the access characteristics.
[0050] The second acquisition module is used to identify risk association features in the access data and acquire a combination of risk association features; wherein, the combination of risk association features is a set of risk association features; the risk association features are used to characterize the association risk between the target asset and the access account;
[0051] The third acquisition module is used to input the popularity value and the access data into the risk prediction model to obtain the risk probability; wherein, the risk prediction model is used to predict the probability that the combination of risk-related features exists in the access data;
[0052] The first determining module is used to determine the target warning level of the target asset among multiple preset warning levels based on the access characteristics and the risk probability.
[0053] To achieve the above objectives, embodiments of the present invention provide an asset risk warning device, comprising: a processor, a memory, and a program or instructions stored in the memory and executable on the processor; wherein, when the processor executes the program or instructions, it implements the asset risk warning method as described above.
[0054] To achieve the above objectives, embodiments of the present invention provide a readable storage medium having a program or instructions stored thereon, wherein the program or instructions, when executed by a processor, implement the steps in the asset risk warning method described above.
[0055] To achieve the above objectives, embodiments of the present invention provide a computer program product, which includes computer instructions that, when executed by a processor, implement the steps of the asset risk warning method described above.
[0056] The beneficial effects of the above-described technical solution of the present invention are as follows:
[0057] In this embodiment of the invention, the popularity value of the target asset is evaluated by extracting access features from access data, which improves the accuracy of popularity evaluation. Furthermore, by combining the risk association features between the target asset and the access account with the popularity value, the risk of the target asset can be predicted in advance, thus avoiding losses caused by the lag in risk prediction. Attached Figure Description
[0058] Figure 1 This is a schematic diagram of the asset risk warning method according to an embodiment of the present invention;
[0059] Figure 2 This is a flowchart of the asset risk warning method described in an embodiment of the present invention;
[0060] Figure 3 This is a schematic diagram of the asset risk warning device described in an embodiment of the present invention. Detailed Implementation
[0061] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0062] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of the invention. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.
[0063] In various embodiments of the present invention, it should be understood that the sequence number of each process described below does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0064] In addition, the terms "system" and "network" are often used interchangeably in this article.
[0065] In the embodiments provided by this invention, it should be understood that "B corresponding to A" means that B is associated with A, and B can be determined based on A. However, it should also be understood that determining B based on A does not mean that B is determined solely based on A; B can also be determined based on A and / or other information.
[0066] For ease of understanding, the following describes some aspects of the embodiments of the present invention:
[0067] like Figure 1 As shown, an asset risk warning method according to an embodiment of the present invention includes:
[0068] Step A10: Extract features from the access data of the access accounts that accessed the target asset within the first preset time period to obtain access features;
[0069] It should be noted that, as Figure 2 As shown, in step S2, the large model features are extracted and the popularity is dynamically calculated (extracting explicit / latent features → multi-dimensional weighted calculation of popularity value → generating time-series change curve). By connecting to multi-source systems to collect asset access-related data, the large model's deep learning capabilities are used to automatically extract explicit and latent features of access behavior in the access data.
[0070] Step A20: Evaluate the popularity value of the target asset within the first preset time period based on the access characteristics;
[0071] It should be noted that a multi-dimensional asset access popularity assessment model is constructed based on the access characteristics, and the asset access popularity value is calculated using a weighted summation method.
[0072] Step A30: Identify risk association features in the access data to obtain a combination of risk association features; wherein, the combination of risk association features is a set of risk association features; the risk association features are used to characterize the association risk between the target asset and the access account;
[0073] It should be noted that, as Figure 2 As shown, in step S3, risk prediction model training and risk identification (learning risk features by associating historical data → hybrid training model → predicting potential risks). Before identifying risk association features, it is necessary to first use historical popularity data and security event records (such as data breaches and unauthorized access events) for association learning, thereby obtaining the ability to identify combinations of risk association features. Based on historical data, the risk prediction model is trained to achieve early prediction and priority ranking of potential security risks, link existing security protection modules to perform targeted actions, and feed the action results back to the large model for incremental training, forming a fully automated closed-loop optimization process, replacing the traditional static analysis and passive alarm mode. In this embodiment of the invention, identifying combinations of risk association features may include:
[0074] Abnormal characteristics of popularity: sudden increase in popularity (popularity value increases by more than 50% in a short period of time), sudden decrease in popularity (popularity value of core assets drops below the threshold).
[0075] Access behavior characteristics: low-privilege accounts accessing highly sensitive assets, high-frequency access outside of working hours, concentrated access from IP addresses in different locations, and high-frequency access after the activation of dormant accounts; among them, dormant accounts can be set as accounts with no login records for 90 days.
[0076] Permission matching characteristics: The scope of account authorization does not match the assets accessed; the super account accesses across departments; the super account is a management account with super privileges.
[0077] Step A40: Input the popularity value and the access data into the risk prediction model to obtain the risk probability; wherein, the risk prediction model is used to predict the probability that the combination of risk-related features exists in the access data;
[0078] It should be noted that during use, the access data and popularity value of the target asset are input into the trained risk prediction model in real time. If a combination of risk characteristics or a risk probability reaches the target, risk warning information is automatically generated, including asset identification (ID), risk type, risk probability, details of associated access behavior, and impact scope assessment.
[0079] Step A50: Based on the access characteristics and the risk probability, determine the target warning level of the target asset among multiple preset warning levels;
[0080] It should be noted that, as Figure 2 As shown, in step S4, risk warnings are generated and prioritized (generating warning information → calculating priority based on three dimensions → tiered push notifications). The warning priority score is calculated using the following mathematical formula: S = P × L × R;
[0081] Wherein, S = priority score, P = risk probability (0≤P≤1, output by the risk prediction model), L = asset sensitivity level (levels 1-4 correspond to quantitative values 1-4), and R = impact scope quantitative value (1 = single user / asset, 5 = single department, 10 = cross-departmental non-core business, 20 = core business). The asset sensitivity registration and impact scope quantitative values are obtained based on the access characteristics. The target warning level of the target asset among multiple preset warning levels is determined by the priority score.
[0082] In this embodiment, the accuracy of popularity assessment is improved by evaluating the popularity value of the target asset through the access features extracted from the access data. Furthermore, by combining the risk association features between the target asset and the access account with the popularity value, the risk of the target asset can be predicted in advance, thus avoiding losses caused by the lag in risk prediction.
[0083] Optionally, the method, wherein step A10 includes:
[0084] The access data is input into a pre-trained model to obtain the access features.
[0085] In this embodiment, the standardized fused data (i.e., the access data) is input into a pre-trained large model (a customized model based on the Transformer architecture) (i.e., the pre-trained model). The core advantage of this architecture lies in its attention mechanism, which can accurately capture the temporal correlation and feature interaction of access behavior. Compared with traditional CNN and RNN models, the accuracy of latent feature extraction is improved by more than 40%. It also supports incremental training, achieving model optimization without full retraining, and is suitable for high-frequency access data scenarios. Through a multi-head attention mechanism and deep learning algorithms, two types of access features are automatically extracted:
[0086] Explicit characteristics: access frequency, access time distribution, and percentage of operation types.
[0087] Latent features include: the degree of matching between account and asset permissions (which can be set to a degree of matching ≥0.7 for high correlation), the similarity between access behavior and historical habits, and the correlation of operation types (which can be set to a correlation of 3 or more consecutive related operations for strong correlation). Features with a large model attention weight ≥0.2 are judged as valid latent features. At the same time, the contrastive learning strategy in self-supervised learning is combined to strengthen the correlation between latent features and security risks.
[0088] Optionally, the method further includes:
[0089] The corresponding handling measures for the target early warning level shall be implemented for the target asset;
[0090] Continuously collect feedback data obtained after the disposal measures are performed on the target asset;
[0091] Acquire new access data of the target asset within a second preset time period; wherein the second preset time period is a time period later than the first preset time period;
[0092] The pre-trained model is incrementally trained based on the feedback data and the newly added access data.
[0093] In this embodiment, such as Figure 2 As shown, in step S5, early warning linkage handling and feedback collection (differentiated linkage protection → execution of handling → collection of effect feedback), and in step S6, incremental training and optimization of the large model (incremental training of feedback data → optimization of model parameters → storage of effective solutions). After the handling measures corresponding to the target early warning level are executed on the target asset, continuous collection of handling effect feedback data is carried out, including whether the risk has been eliminated, whether there is a misjudgment, handling time, optimization suggestions, terminal operation logs, access link status, etc., forming a feedback dataset which is stored in the model optimization database to provide accurate input for subsequent model optimization. Feedback data, new access data, and the latest security event data (such as new types of unauthorized access patterns) are periodically (once a week by default; if the performance indicators fail to meet the standards for two consecutive periods, this is shortened to three days) input into the large model for multi-source data fusion incremental training.
[0094] Optionally, the method further includes:
[0095] If the handling measures are deemed effective based on feedback data, the feature weights of the access characteristics in the popularity assessment model are optimized based on the feedback data and the new access data. The feedback data is data obtained after the handling measures are applied to the target asset; the handling measures are actions corresponding to the target warning level performed on the target asset; the new access data is access data for accessing the target asset within a second preset time period; the second preset time period is a time period later than the first preset time period; and the popularity assessment model is used to assess the popularity value of the target asset within the first preset time period based on the access characteristics.
[0096] In this embodiment, a process of "three-level optimization + scenario-based knowledge base construction" is used to optimize the pre-trained model:
[0097] Optimize feature weights: Adjust the feature weights in the popularity assessment model based on the handling effect (e.g., if the weight of "visit frequency" is too high in a misjudged case, automatically reduce the α value by 0.05, while proportionally increasing the weights of other features to ensure that the sum remains at 1). The accuracy of popularity assessment should be improved by ≥5% after optimization.
[0098] Update the risk feature library: Incorporate new risk patterns into the feature library of the risk prediction model to expand the scope of risk identification;
[0099] Adjust risk threshold: Dynamically adjust the risk probability threshold based on the actual false positive rate (target ≤ 5%) and false negative rate (target ≤ 3%) (if the false positive rate is too high, increase the threshold).
[0100] Optionally, the method further includes:
[0101] The optimization scheme for the feature weights is stored in the optimization knowledge base using a distributed storage architecture.
[0102] In this embodiment, effective optimization solutions are categorized and stored according to a three-dimensional index of "risk type - asset level - optimization solution". A distributed storage architecture and vector similarity matching algorithm (response time ≤ 1 second) are used to form a model optimization knowledge base, which supports retrieval and reuse by risk type and asset level, and continuously improves the accuracy of popularity assessment and risk prediction accuracy.
[0103] Optionally, in the method, the plurality of preset warning levels include one or more of the following:
[0104] High-priority risks;
[0105] Medium-priority risks;
[0106] Low priority risks;
[0107] The following measures will be taken for the target asset in accordance with the corresponding early warning level:
[0108] When the target warning level is high-priority risk, the measures taken against the target asset include freezing the access rights of abnormal accounts to the target asset and blocking the access links of the abnormal accounts; wherein, the abnormal account is an account that has abnormal access behavior to the target asset as determined by the combination of risk association characteristics.
[0109] When the target warning level is medium priority risk, the handling measures for the target asset include verifying the abnormal access behavior with the person in charge of the abnormal account of the target asset, and tracking the subsequent access behavior of the abnormal account.
[0110] If the target warning level is low priority risk, the disposal measures for the target asset are placed in the verification queue.
[0111] In this embodiment, the multiple preset warning levels are divided into three levels:
[0112] High-priority risks (priority score ≥ 80 points): Involving Level 4 sensitive assets + risk probability ≥ 90% + impact scope covering core business (Example: When P=1.0, L=4, R=20, S=1.0×4×20=80, judged as high priority).
[0113] Medium-priority risk (50 points ≤ priority score < 80 points): Involves Level 3 sensitive assets + risk probability ≥ 80% + impact limited to a single department; (Example: When P=0.95, L=4, R=20, S=0.95×4×20=76, classified as medium-priority; when P=0.98, L=4, R=20, S=0.98×4×20=78.4, classified as medium-priority).
[0114] Low-priority risks (priority score < 50 points): involve level 1-2 sensitive assets + risk probability < 80% + single impact scope.
[0115] High-priority alerts trigger pop-up windows and SMS notifications immediately, while medium- and low-priority alerts are summarized and pushed out daily.
[0116] The key improvement lies in "automated triggering of cross-module linkage + real-time feedback on the handling effect" by synchronizing risk warnings to the corresponding security protection modules and implementing targeted measures.
[0117] High-priority risks: Automatically send JSON-formatted instructions to the 4A platform via the Representational State Transfer (RESTful) interface to temporarily freeze the access permissions of abnormal accounts to target assets; transmit Internet Protocol (IP) blocking rules to the Session Description Protocol (SDP) platform via the Transmission Control Protocol (TCP) interface to block abnormal access links and force secondary authentication via cloud desktop; trigger desktop management tools to enhance monitoring and record terminal operation logs; response time ≤ 1 minute.
[0118] Medium-priority risks: Generate compliant work orders that conform to the Group's System Management Controller (SMC) work order data specifications, push them to the SMC system via the Simple Object Access Protocol (SOAP) interface, automatically match the handler based on the account's department and asset responsibility chain (dispatch accuracy ≥95%), require the account holder to verify the access's legitimacy and provide feedback within a specified time; add the work order to the key monitoring list through the situational awareness platform to track subsequent access behavior; response time ≤5 minutes.
[0119] Low-priority risks: Stored in the alarm pool for batch verification by security management personnel within 24 hours, without immediate action required.
[0120] Optionally, the method, wherein step A50 includes:
[0121] Priority scores are calculated in real time based on the access characteristics and the risk probability;
[0122] The target warning level is determined based on the priority score.
[0123] In this embodiment, the priority score is calculated using the following mathematical formula: S = P × L × R; where S = priority score, P = risk probability (0≤P≤1, output by the risk prediction model), L = asset sensitivity level (levels 1-4 correspond to quantified values 1-4), and R = quantified value of the scope of influence (1 = single user / asset, 5 = single department, 10 = cross-departmental non-core business, 20 = core business). The asset sensitivity registration and the quantified value of the scope of influence are obtained based on the access characteristics. The target warning level of the target asset among multiple preset warning levels is determined by the priority score.
[0124] The multiple preset warning levels are divided into three levels:
[0125] High-priority risks (priority score ≥ 80 points): Involving Level 4 sensitive assets + risk probability ≥ 90% + impact scope covering core business (Example: When P=1.0, L=4, R=20, S=1.0×4×20=80, judged as high priority).
[0126] Medium-priority risk (50 points ≤ priority score < 80 points): Involves Level 3 sensitive assets + risk probability ≥ 80% + impact limited to a single department; (Example: When P=0.95, L=4, R=20, S=0.95×4×20=76, classified as medium-priority; when P=0.98, L=4, R=20, S=0.98×4×20=78.4, classified as medium-priority).
[0127] Low-priority risk (priority score < 50 points): Involves level 1-2 sensitive assets + risk probability < 80% + limited impact. The target warning level is determined based on the relationship between the priority score and the threshold.
[0128] Optionally, the method further includes:
[0129] Collect multi-source access related data for the target asset;
[0130] The multi-source access-related data is standardized, cleaned, converted, deduplicated, merged, and validated for data consistency to obtain access data with a consistent format.
[0131] The multi-source access related data includes at least two of the following:
[0132] The access log data of the target asset includes access time, account identifier, operation type, access address and access result;
[0133] The attribute data of the target asset includes the asset sensitivity level of the target asset, the system to which the target asset belongs, the department to which the target asset belongs, and the core business relevance of the target asset.
[0134] The account permission data of the target asset includes the access account level, the authorization scope of the access account, the department to which the access account belongs, and the status of the access account; the access account is the account that accesses the target asset.
[0135] In this embodiment, such as Figure 2As shown, in step S1, multi-source resource access related data is collected (collecting three types of data from multiple systems → standardization processing → generating a list with version number / timestamp / MD5 → REST API synchronization → MD5 verification + version judgment →), multi-source asset access related data is collected, and standardization cleaning, transformation, deduplication and fusion, and data consistency verification are completed. If the data is complete and the format is consistent, the process continues. Interfacing with 4A platforms, data security management platforms, endpoint management tools (Beixinyuan, AsiaInfo Security, etc.), SMC systems, and other multi-source systems, three types of core data are collected:
[0136] Asset access log data: access time, account ID, asset ID, operation type (query / export / modify / delete), access IP, access result (success / failure);
[0137] Asset attribute data: asset sensitivity level (level 4 / level 3 / level 2 / level 1), system to which it belongs, department to which it belongs, and relevance to core business;
[0138] Account permission data: account level, authorization scope, department, account status (normal / silent / isolated).
[0139] The data processing module imports multi-source data through standardized interfaces, performs data cleaning (removing invalid logs and completing missing fields), format conversion (unifying field naming and data types), and deduplication to ensure data integrity and consistency.
[0140] The data processing module imports multi-source data through standardized interfaces, performs data cleaning (removing invalid logs and filling in missing fields), format conversion (unifying field naming and data types), and deduplication. It also adds asset-account permission consistency verification (verifying the matching relationship between the account's authorized scope and the accessed assets; if the account does not have access to the target asset, it is marked as abnormal data), ensuring data integrity, consistency, and validity, and providing a data foundation for subsequent popularity calculation and risk prediction.
[0141] Optionally, the method, wherein step A20 includes:
[0142] Obtain an access feature score based on the access features;
[0143] The popularity value is obtained by adding the product of the access feature score and the corresponding feature weight.
[0144] In this embodiment, the heat value = α×F1 + β×F2 + γ×F3 + δ×F4 + ε×F5; where F1 = access frequency score, F2 = permission matching score, F3 = operation risk level score, F4 = time period concentration score (which can be set to be abnormal if it deviates from the normal working hours by ±3 hours), and F5 = business relevance score. Specifically, the access frequency score F1 is calculated based on the access frequency; the permission matching score F2 is calculated based on the permission matching degree between the account and the asset; the operation risk level score F3 is calculated based on the asset sensitivity level of the target asset; the time period concentration score F4 is calculated based on the access time period distribution; and the business relevance score F5 is calculated based on the relevance of the operation type.
[0145] α, β, γ, δ, and ε are feature weights (summing up to 1), with initial values determined based on industry experience and historical data statistics: α = 0.3, β = 0.25, γ = 0.2, δ = 0.15, and ε = 0.1. These values can be dynamically adjusted through model optimization. The scores for each feature are normalized to the [0,1] interval based on the data distribution. Simultaneously, a popularity time-series curve is generated based on the popularity values at multiple time points, visually displaying the asset's popularity fluctuation trend.
[0146] Optionally, the method further includes:
[0147] Obtain the original risk prediction model formed by fusing supervised learning and unsupervised learning models;
[0148] Labeled historical access data is input into the supervised learning model to obtain a first training risk probability, and unlabeled historical access data is input into the unsupervised learning model to obtain a second training risk probability; wherein, the labeled historical access data is access data of a first historical time with real labels; and the unlabeled historical access data is access data of the first historical time without labels.
[0149] The combined training risk probability is obtained by using a weighted voting mechanism based on the first training risk probability and the second training risk probability.
[0150] By comparing the combined trained risk probability with the real labels in the historical access data of the labels, the parameters of the original risk prediction model are adjusted to obtain the risk prediction model.
[0151] In this embodiment, the training data is derived from China Mobile’s real operational data over the past two years. It covers more than 100,000 accounts, more than 50,000 assets, and more than 100 million access logs, including more than 5,000 security event samples (including real scenarios such as unauthorized access and data leakage). The samples cover the entire business line and assets of all sensitivity levels.
[0152] The original risk prediction model is trained using a combination of supervised and unsupervised learning:
[0153] Supervised learning part: A fusion model of logistic regression and gradient boosting tree (XGBoost) is adopted, using historical security event samples as labeled data to train the ability to identify known risk features;
[0154] Unsupervised learning part: The Isolation Forest algorithm is used to cluster unlabeled access data and an anomaly score threshold (≥0.8) is set to identify unknown risk patterns;
[0155] Fusion Mechanism: The two models output a final risk probability (i.e., the combined training risk probability) through a weighted voting mechanism (e.g., the supervised model weight can be set to 0.6, and the unsupervised model weight can be set to 0.4), achieving comprehensive coverage of both known and unknown risks. The combined training risk probability is compared and analyzed with the real labels in the historical access data of the labels, and the parameters of the original risk prediction model are adjusted to obtain the final risk prediction model.
[0156] During the model inference process, a risk probability threshold is set (e.g., 80%, which can be adjusted according to the actual situation). When the model predicts a risk probability greater than or equal to the risk probability threshold, the corresponding access behavior in the access data is determined to be potentially high-risk.
[0157] It should be noted that the asset risk warning method described in this embodiment of the invention has the following advantages:
[0158] 1. A multi-dimensional feature extraction and popularity calculation scheme driven by a large model is adopted, protecting the automatic extraction mechanism of "explicit features + implicit features," the feature weight optimization method based on large model learning, and the multi-dimensional asset access popularity assessment model to achieve accurate quantification of asset popularity. By automatically extracting implicit features of access behavior through a large model, a multi-dimensional popularity assessment model is constructed. The weights are dynamically optimized through model learning, and the popularity assessment results more closely reflect the actual importance and security risk level of the asset, solving the problem of "one-sided popularity assessment" in existing technologies. Furthermore, by using a hybrid training model to identify combinations of explicit and implicit risk features, potential risks can be predicted in advance, resulting in a wider risk coverage and a lower rate of missed or false positives, achieving a leap from "passive response" to "proactive prediction."
[0159] 2. Risk Prediction and Prioritization Mechanism: This mechanism protects a hybrid training model combining "supervised learning of known risk features + unsupervised mining of unknown risk features," a multi-dimensional risk feature combination identification logic, and a priority scoring algorithm based on "risk probability × asset level × impact scope," enabling accurate risk prediction and tiered risk alerts. By focusing on core risks through a priority grading mechanism and simultaneously coordinating with existing security modules to implement differentiated responses, it significantly improves security response efficiency and addresses the industry pain point of "massive redundant alarms."
[0160] 3. Early Warning and Linkage Response and Closed-Loop Optimization System: This system integrates the protection mechanism with existing security modules such as the 4A platform, SDP platform, and desktop management tools. It employs a large-scale incremental training mechanism based on response feedback, and establishes and reuses a model optimization knowledge base, forming a fully automated closed-loop process. Based on response effect feedback, it enables large-scale incremental training and automatic optimization, continuously improving system performance without manual intervention and adapting to complex and ever-changing access scenarios and security requirements.
[0161] like Figure 3 As shown, to achieve the above objectives, embodiments of the present invention provide an asset risk warning device, comprising:
[0162] The first acquisition module 301 extracts features from the access data of access accounts that access the target asset within a first preset time period to obtain access features.
[0163] The first processing module 302 is used to evaluate the popularity value of the target asset within the first preset time period based on the access characteristics.
[0164] The second acquisition module 303 is used to identify risk association features in the access data and acquire a combination of risk association features; wherein, the combination of risk association features is a set of risk association features; the risk association features are used to characterize the association risk between the target asset and the access account;
[0165] The third acquisition module 304 is used to input the popularity value and the access data into the risk prediction model to obtain the risk probability; wherein, the risk prediction model is used to predict the probability that the combination of risk-related features exists in the access data;
[0166] The first determining module 305 is used to determine the target warning level of the target asset among multiple preset warning levels based on the access characteristics and the risk probability.
[0167] Optionally, in the method, the first acquisition module 301 includes:
[0168] The first acquisition unit is used to input the access data into a pre-trained model to acquire the access features.
[0169] Optionally, the device further includes:
[0170] The second processing module is used to execute the handling measures corresponding to the target early warning level for the target asset;
[0171] The third processing module is used to continuously collect feedback data obtained after the disposal measures are performed on the target asset;
[0172] The fourth acquisition module is used to acquire new access data of the target asset within a second preset time period; wherein the second preset time period is a time period later than the first preset time period.
[0173] The third processing module is used to incrementally train the pre-trained model based on the feedback data and the newly added access data.
[0174] Optionally, the device further includes:
[0175] The fourth processing module is used to optimize the feature weights of the access characteristics in the popularity assessment model based on the feedback data and the newly added access data, when the effectiveness of the handling measures is determined based on the feedback data. The feedback data is data obtained after the handling measures are applied to the target asset; the handling measures are actions corresponding to the target warning level performed on the target asset; the newly added access data is access data of the target asset accessed within a second preset time period; the second preset time period is a time period later than the first preset time period; and the popularity assessment model is used to assess the popularity value of the target asset within the first preset time period based on the access characteristics.
[0176] Optionally, the device further includes:
[0177] The fifth processing module is used to store the optimization scheme of the feature weights into the optimization knowledge base using a distributed storage architecture.
[0178] Optionally, in the aforementioned device, the plurality of preset warning levels include one or more of the following:
[0179] High-priority risks;
[0180] Medium-priority risks;
[0181] Low priority risks;
[0182] The following measures will be taken for the target asset in accordance with the corresponding early warning level:
[0183] When the target warning level is high-priority risk, the measures taken against the target asset include freezing the access rights of abnormal accounts to the target asset and blocking the access links of the abnormal accounts; wherein, the abnormal account is an account that has abnormal access behavior to the target asset as determined by the combination of risk association characteristics.
[0184] When the target warning level is medium priority risk, the handling measures for the target asset include verifying the abnormal access behavior with the person in charge of the abnormal account of the target asset, and tracking the subsequent access behavior of the abnormal account.
[0185] If the target warning level is low priority risk, the disposal measures for the target asset are placed in the verification queue.
[0186] Optionally, in the aforementioned apparatus, the first determining module 305 includes:
[0187] The second processing unit is used to calculate the priority score in real time based on the access characteristics and the risk probability;
[0188] The first determining unit determines the target warning level based on the priority score.
[0189] Optionally, the device further includes:
[0190] The sixth processing module is used to collect multi-source access related data of the target asset;
[0191] The fifth acquisition module is used to perform standardized cleaning, format conversion, deduplication and fusion, and data consistency verification on the multi-source access-related data to obtain access data with consistent format.
[0192] The multi-source access related data includes at least two of the following:
[0193] The access log data of the target asset includes access time, account identifier, operation type, access address and access result;
[0194] The attribute data of the target asset includes the asset sensitivity level of the target asset, the system to which the target asset belongs, the department to which the target asset belongs, and the core business relevance of the target asset.
[0195] The account permission data of the target asset includes the access account level, the authorization scope of the access account, the department to which the access account belongs, and the status of the access account; the access account is the account that accesses the target asset.
[0196] Optionally, in the aforementioned apparatus, the first processing module 302 includes:
[0197] The second acquisition unit is used to acquire an access feature score based on the access features;
[0198] The third acquisition unit is used to add the product of the access feature score and the corresponding feature weight to obtain the popularity value.
[0199] Optionally, the device further includes:
[0200] The sixth acquisition module is used to acquire the original risk prediction model formed by the fusion of supervised learning model and unsupervised learning model;
[0201] The seventh acquisition module is used to input labeled historical access data into the supervised learning model to obtain a first training risk probability, and to input unlabeled historical access data into the unsupervised learning model to obtain a second training risk probability; wherein, the labeled historical access data is access data of a first historical time with real labels; and the unlabeled historical access data is access data of the first historical time without labels.
[0202] The eighth acquisition module is used to acquire the combined training risk probability based on the first training risk probability and the second training risk probability through a weighted voting mechanism;
[0203] The ninth acquisition module is used to compare and analyze the combined training risk probability with the real labels in the historical access data of the labels, adjust the parameters of the original risk prediction model, and acquire the risk prediction model.
[0204] It should be noted that the apparatus provided in this embodiment of the invention can implement all the method steps implemented in the above method embodiment and can achieve the same technical effect. Therefore, the parts and beneficial effects that are the same as those in the method embodiment will not be described in detail here.
[0205] To achieve the above objectives, embodiments of the present invention provide an asset risk warning device, comprising: a processor, a memory, and a program or instructions stored in the memory and executable on the processor; wherein, when the processor executes the program or instructions, it implements the asset risk warning method as described above.
[0206] To achieve the above objectives, embodiments of the present invention provide a readable storage medium having a program or instructions stored thereon, wherein the program or instructions, when executed by a processor, implement the steps in the asset risk warning method described above.
[0207] To achieve the above objectives, embodiments of the present invention provide a computer program product, which includes computer instructions that, when executed by a processor, implement the steps of the asset risk warning method described above.
[0208] It should be further noted that the terminals described in this specification include, but are not limited to, smartphones, tablets, etc., and many of the functional components described are referred to as modules in order to emphasize the independence of their implementation.
[0209] In this embodiment of the invention, the module can be implemented in software so that it can be executed by various types of processors. For example, an identified executable code module may include one or more physical or logical blocks of computer instructions, which may be constructed as objects, procedures, or functions. Nevertheless, the executable code of the identified module does not need to be physically located together, but may include different instructions stored in different bits, which, when logically combined, constitute the module and achieve the module's intended purpose.
[0210] In practice, an executable code module can be a single instruction or many instructions, and can even be distributed across multiple different code segments, different programs, and across multiple memory devices. Similarly, operational data can be identified within the module and can be implemented in any suitable form and organized within any suitable data structure. This operational data can be collected as a single dataset or distributed across different locations (including different storage devices), and can exist, at least in part, solely as electronic signals within the system or network.
[0211] When a module can be implemented using software, considering the current level of hardware technology, modules that can be implemented in software can be implemented using hardware circuits by those skilled in the art to achieve the corresponding functions, without considering cost. These hardware circuits include conventional very-large-scale integrated circuits (VLSI) or gate arrays, as well as existing semiconductors such as logic chips and transistors, or other discrete components. Modules can also be implemented using programmable hardware devices, such as field-programmable gate arrays, programmable array logic, and programmable logic devices.
[0212] The exemplary embodiments described above are with reference to the accompanying drawings. Many different forms and embodiments are feasible without departing from the spirit and teachings of the invention. Therefore, the invention should not be construed as limiting the exemplary embodiments set forth herein. Rather, these exemplary embodiments are provided to make the invention complete and convey the scope of the invention to those skilled in the art. In these drawings, component dimensions and relative dimensions may be exaggerated for clarity. The terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. As used herein, unless clearly indicated otherwise, the singular forms “a,” “an,” and “the” are intended to include all such forms. It will be further understood that the terms “comprising” and / or “including”, when used in this specification, indicate the presence of the stated features, integers, steps, operations, components, and / or elements, but do not exclude the presence or addition of one or more other features, integers, steps, operations, components, and / or groups thereof. Unless otherwise indicated, when stated, a range of values includes the upper and lower limits of the range and any subranges in between.
[0213] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A risk warning method for assets, characterized in that, include: Feature extraction is performed on the access data of access accounts that access the target asset within the first preset time period to obtain access features; The popularity value of the target asset within the first preset time period is evaluated based on the access characteristics. The access data is subjected to risk association feature identification to obtain a combination of risk association features; wherein, the combination of risk association features is a set of risk association features; the risk association features are used to characterize the association risk between the target asset and the access account; The popularity value and the access data are input into a risk prediction model to obtain the risk probability; wherein, the risk prediction model is used to predict the probability that the combination of risk-related features exists in the access data; Based on the access characteristics and the risk probability, the target warning level of the target asset among multiple preset warning levels is determined.
2. The method according to claim 1, characterized in that, Feature extraction is performed on access data of the target asset within a first preset time period to obtain access features, including: The access data is input into a pre-trained model to obtain the access features.
3. The method according to claim 2, characterized in that, The method further includes: The corresponding handling measures for the target early warning level shall be implemented for the target asset; Continuously collect feedback data obtained after the disposal measures are performed on the target asset; Acquire new access data of the target asset within a second preset time period; wherein the second preset time period is a time period later than the first preset time period; The pre-trained model is incrementally trained based on the feedback data and the newly added access data.
4. The method according to claim 1, characterized in that, The method further includes: If the handling measures are deemed effective based on feedback data, the feature weights of the access characteristics in the popularity assessment model are optimized based on the feedback data and the new access data. The feedback data is data obtained after the handling measures are applied to the target asset; the handling measures are actions corresponding to the target warning level performed on the target asset; the new access data is access data for accessing the target asset within a second preset time period; the second preset time period is a time period later than the first preset time period; and the popularity assessment model is used to assess the popularity value of the target asset within the first preset time period based on the access characteristics.
5. The method according to claim 4, characterized in that, The method further includes: The optimization scheme for the feature weights is stored in the optimization knowledge base using a distributed storage architecture.
6. The method according to claim 3, characterized in that, The preset warning levels include one or more of the following: High-priority risks; Medium-priority risks; Low priority risks; The following measures will be taken for the target asset in accordance with the corresponding early warning level: When the target warning level is high-priority risk, the measures taken against the target asset include freezing the access rights of abnormal accounts to the target asset and blocking the access links of the abnormal accounts; wherein, the abnormal account is an account that has abnormal access behavior to the target asset as determined by the combination of risk association characteristics. When the target warning level is medium priority risk, the handling measures for the target asset include verifying the abnormal access behavior with the person in charge of the abnormal account of the target asset, and tracking the subsequent access behavior of the abnormal account. If the target warning level is low priority risk, the disposal measures for the target asset are placed in the verification queue.
7. The method according to claim 1, characterized in that, Based on the access characteristics and the risk probability, the target warning level of the target asset among multiple preset warning levels is determined, including: Priority scores are calculated in real time based on the access characteristics and the risk probability; The target warning level is determined based on the priority score.
8. The method according to claim 1, characterized in that, The method further includes: Collect multi-source access related data for the target asset; The multi-source access-related data is standardized, cleaned, converted, deduplicated, merged, and validated for data consistency to obtain access data with a consistent format. The multi-source access related data includes at least two of the following: The access log data of the target asset includes access time, account identifier, operation type, access address and access result; The attribute data of the target asset includes the asset sensitivity level of the target asset, the system to which the target asset belongs, the department to which the target asset belongs, and the core business relevance of the target asset. The account permission data of the target asset includes the access account level, the authorization scope of the access account, the department to which the access account belongs, and the status of the access account; the access account is the account that accesses the target asset.
9. The method according to claim 1, characterized in that, Assessing the popularity of the target asset within the first preset time period based on the access characteristics includes: Obtain an access feature score based on the access features; The popularity value is obtained by adding the product of the access feature score and the corresponding feature weight.
10. The method according to claim 1, characterized in that, The method further includes: Obtain the original risk prediction model formed by fusing supervised learning and unsupervised learning models; Labeled historical access data is input into the supervised learning model to obtain a first training risk probability, and unlabeled historical access data is input into the unsupervised learning model to obtain a second training risk probability; wherein, the labeled historical access data is access data of a first historical time with real labels; and the unlabeled historical access data is access data of the first historical time without labels. The combined training risk probability is obtained by using a weighted voting mechanism based on the first training risk probability and the second training risk probability. By comparing the combined trained risk probability with the real labels in the historical access data of the labels, the parameters of the original risk prediction model are adjusted to obtain the risk prediction model.
11. A risk warning device for assets, characterized in that, include: The first acquisition module extracts features from the access data of access accounts that access the target asset within a first preset time period to obtain access features. The first processing module is used to evaluate the popularity value of the target asset within the first preset time period based on the access characteristics. The second acquisition module is used to identify risk association features in the access data and acquire a combination of risk association features; wherein, the combination of risk association features is a set of risk association features; the risk association features are used to characterize the association risk between the target asset and the access account; The third acquisition module is used to input the popularity value and the access data into the risk prediction model to obtain the risk probability; wherein, the risk prediction model is used to predict the probability that the combination of risk-related features exists in the access data; The first determining module is used to determine the target warning level of the target asset among multiple preset warning levels based on the access characteristics and the risk probability.
12. A risk warning device for an asset, comprising: A processor, a memory, and a program or instructions stored in the memory and executable on the processor; characterized in that, when the processor executes the program or instructions, it implements the asset risk warning method as described in any one of claims 1-10.
13. A readable storage medium having a program or instructions stored thereon, characterized in that, When the program or instructions are executed by the processor, they implement the steps in the asset risk warning method as described in any one of claims 1-10.
14. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the asset risk warning method as described in any one of claims 1-10.