Intelligent checking system for data assets
By analyzing dependency chains, identifying abnormal flow paths, and monitoring storage media performance and security threats through the intelligent data asset inventory system, the system solves the problems of difficult topology representation and incomplete value assessment in data asset management, and realizes the rational allocation and secure management of data assets.
Patent Information
- Application Number
- CN202511124463.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-11-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies struggle to fully represent the topological structure between data assets during data asset inventory, fail to identify abnormal flow paths, and exhibit a disconnect between storage status monitoring and security status monitoring. This results in insufficient value assessment, leading to resource misallocation and potential risks.
A data asset intelligent inventory system is provided, including a hierarchical relationship determination module, a storage status monitoring module, a security status monitoring module, an asset value calculation module, and a graded disposal execution module. By parsing the dependency relationship chain of data assets, it identifies abnormal flow paths, monitors storage media performance and security threats, and performs value quantification analysis by combining usage frequency and business relevance weights to generate retention or migration decision signals.
It enables an intuitive presentation of the network of relationships between data assets, identifies abnormal flow paths, conducts value assessments based on storage and security status, generates reasonable retention or migration decisions, ensures the comprehensiveness and security of data asset management, and makes resource allocation more aligned with business needs.
Smart Images

Figure CN120973734A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data asset inventory technology, specifically to an intelligent data asset inventory system. Background Technology
[0002] In the process of digitalization, data assets have become an important component of organizational operations. Their scale continues to expand with business development, and their types are becoming increasingly diverse, encompassing structured business data, semi-structured log information, and unstructured documents. These data assets do not exist in isolation but form complex dependencies through business processes. For example, user data supports transaction systems, transaction data is linked to risk control models, and risk control models, in turn, feed back into user profile optimization. This intertwined network of relationships significantly increases the difficulty of managing data assets.
[0003] Data asset inventory often relies on a combination of manual review and traditional tools. Manual review requires a significant time investment to verify basic information about data assets. With hundreds or even thousands of data sources and multi-layered dependencies, omissions are common, and it's difficult to fully represent the topological structure of data assets. While traditional tools can perform some basic statistics, their ability to track data flow paths is limited. When data is migrated or copied between different systems, it's difficult to identify abnormal flow nodes, resulting in some hidden data links not being included in the inventory.
[0004] Monitoring storage status is often independent of the data asset inventory process and is mostly performed by operations and maintenance personnel using storage management tools. These tools can only output basic performance parameters of storage media and cannot associate them with specific data assets. This makes it difficult to know the health status of storage media containing certain types of core business data during inventory, and also makes it impossible to adjust asset priorities based on storage status.
[0005] Security status monitoring is also disconnected from data asset inventory. Existing security tools focus on identifying individual unauthorized accesses or vulnerabilities, but fail to correlate these security threats with the business value of data assets. For example, a low-frequency document that is relevant to core business operations may have a permission configuration vulnerability, which is often overlooked because it is not included in the value assessment system of the inventory, thus creating a potential risk.
[0006] In the value assessment phase, existing methods often focus on the frequency of data asset usage, with less consideration given to the health of storage media, the level of security threats, and the depth of business relevance. This leads to a disconnect between the resulting disposal decisions and actual business needs. Some data assets that still have business value but whose storage media are aging may be misjudged as needing to be migrated, while some data assets with high security risks but low usage frequency may be retained, resulting in resource misallocation. Summary of the Invention
[0007] The purpose of this invention is to provide an intelligent data asset inventory system to solve the problems mentioned in the background art.
[0008] To achieve the above objectives, the present invention provides a data asset intelligent inventory system, the system comprising: The hierarchical relationship determination module is used to obtain the metadata structure information of the target data assets, parse the dependency relationship chain between data assets, and generate a topology relationship graph; based on the topology relationship graph, it identifies abnormal data flow paths and generates a structural anomaly alarm signal; and activates the storage status monitoring module and the security status monitoring module according to the structural anomaly alarm signal. The storage status monitoring module is used to scan the storage media performance indicators of the target data asset, analyze the read and write stability of the storage media, output the storage media health assessment results and transmit them to the asset value calculation module. The security status monitoring module is used to detect the access permission configuration and encryption policy status of the target data assets, mark unauthorized access behaviors and security vulnerabilities, generate a security threat assessment report and transmit it to the asset value calculation module. The asset value calculation module is used to receive the health assessment results of the storage medium and the security threat assessment report, combine the data asset usage frequency and business association weight, perform value quantification analysis, and generate asset retention decision signals or asset migration decision signals. The tiered disposal execution module is used to generate storage location optimization parameters and access permission reset parameters based on the asset retention decision signal or asset migration decision signal.
[0009] Preferably, the process of parsing data asset dependencies by the hierarchical relationship determination module includes: Extract metadata attribute fields from the target data assets and construct a field-level association mapping table; use the field-level association mapping table to track cross-data table reference paths and mark abnormal paths with circular references or orphaned nodes; The percentage of data table nodes in the abnormal path is counted, and a topology anomaly coefficient is generated by combining it with the reference depth threshold. When the topology anomaly coefficient exceeds the preset topology fault tolerance threshold, the structural anomaly alarm signal is triggered.
[0010] Preferably, the storage status monitoring module's analysis process of the storage medium includes: Classify and obtain the physical storage media type corresponding to the target data assets, and collect the read and write response latency of the disk array, the number of bad sectors, and the wear of solid-state storage units in real time; Based on the physical storage medium type, a baseline performance parameter is matched, and the deviation of the read / write response latency from the baseline performance parameter is calculated. By associating the number of bad sectors with the wear and tear of solid-state storage units, a storage medium health index is derived by weighting performance deviation and hardware wear values, and the storage medium health index is written into the storage medium health assessment result.
[0011] Preferably, the execution process of the safety status monitoring module includes: Iterate through the access control list of the target data asset to identify open interfaces that are not bound to an authentication protocol; simultaneously scan the key validity period of the data encryption fields and mark expired encryption fields; The number of open interfaces and expired encrypted fields are aggregated, and a security threat value is calculated by combining the frequency of historical attack events. When the security threat value reaches a preset risk threshold, the security threat value is included in the security threat assessment report.
[0012] Preferably, the value quantification analysis process of the asset value calculation module includes: Retrieve query call logs of the target data assets within a preset time period, and count the access frequency and the number of related business systems per unit time. By combining the storage media health index from the storage media health assessment results with the security threat value from the security threat assessment report, an asset value coefficient is generated using weighted access frequency, business system weight, and health and security coefficient. If the asset value coefficient is higher than the preset value retention threshold, an asset retention decision signal is generated; if the asset value coefficient is lower than the preset value retention threshold, an asset migration decision signal is generated.
[0013] Preferably, the operation logic of the tiered processing execution module includes: When the asset retention decision signal is captured, a storage strategy optimization instruction is initiated to dynamically adjust the storage medium type and backup frequency according to the data asset access pattern. When the asset migration decision signal is captured, a cold and hot data tiering instruction is initiated to migrate low-value data assets to a low-cost storage cluster based on access frequency, and to close access ports that are unrelated to the current business functions of the low-value data assets.
[0014] Preferably, the system further includes a storage strategy optimization module: Receive storage location optimization parameters from the hierarchical processing execution module and analyze the access time distribution characteristics of the target data assets; Based on the characteristics of access time distribution, cold and hot data intervals are divided, and the storage cost-effectiveness ratio of high-frequency access data is calculated. The ratio of high-speed storage resources to archive storage resources is dynamically allocated based on the aforementioned storage cost-effectiveness ratio.
[0015] Preferably, the resource allocation process of the storage policy optimization module includes: Establish a mapping relationship between the access timestamp and the data asset identifier, and generate an access heat map in the time dimension; Predict the future access peak period based on the access heat map in the time dimension, and adjust the cache space of the high-speed storage resources in advance; Compare the predicted cache space demand with the current storage resource utilization rate, and generate a resource expansion warning or a resource release instruction.
[0016] Preferably, the system further includes an abnormal handling linkage module: When the security status monitoring module generates a security threat assessment report, synchronously trigger the access log analysis engine to track the source address and operation track of the abnormal access behavior; Locate the affected data asset range according to the operation track, and generate the coordinates of the data isolation area; bind and execute the coordinates of the data isolation area with the access right reset parameters of the hierarchical handling execution module.
[0017] Preferably, the system further includes an asset portrait iteration module: Periodically obtain the asset value coefficient generated by the asset value calculation module, and associate the value change trend of the historical inventory period; Correct the business association weight and the storage medium health degree weight parameter according to the value change trend; feedback the corrected weight parameter to the topological relationship map construction process of the hierarchical relationship determination module.
[0018] Compared with the prior art, the beneficial effects of the present invention are: The intelligent inventory system for data assets obtains the metadata structure information of the target data assets through the hierarchical relationship determination module, analyzes the dependency chain and generates a topological relationship map, which can intuitively present the association network between data assets, making the originally hidden dependency paths emerge. On this basis, the module can identify abnormal data transfer paths and generate a structural anomaly warning signal, prompting the storage state monitoring module and the security state monitoring module to start synchronously, breaking the situation where structural analysis and state monitoring are separated from each other in traditional inventory, and forming a linkage feedback between the structural characteristics of data assets and their storage and security states.
[0019] The storage status monitoring module scans storage media performance metrics and analyzes read / write stability. The output health assessment results are no longer isolated technical parameters, but rather serve as crucial input for asset value calculation, establishing a link between the physical storage status of data assets and business value assessment. The security status monitoring module detects access permission configurations and encryption policy status, flags unauthorized access behaviors and security vulnerabilities, and the generated security threat assessment report, along with the storage health assessment results, participates in value quantification analysis. This integrates the security attributes of data assets into the overall value judgment system, preventing security factors from being marginalized during asset disposal.
[0020] The asset valuation module combines storage media health, security threat assessment, usage frequency, and business relevance weights to perform value quantification analysis. The resulting retention or migration decision signal no longer relies on a single dimension of judgment, but integrates the structural rationality of data assets, physical storage reliability, security risk level, and actual business needs, making the decision-making process more aligned with the multi-dimensional attributes of data assets.
[0021] The tiered disposal execution module generates storage location optimization parameters and access permission reset parameters based on decision signals, transforming asset disposal from a theoretical judgment into a concrete, directly executable operation. Storage location optimization is based on storage health and business relevance weights, ensuring that the physical storage of data assets matches business access needs; access permission resets combine security threat assessment and value assessment results, ensuring that permission configurations are adapted to the actual value and risk level of the data assets.
[0022] The collaborative operation of each module covers the entire process of data asset management, from structure analysis and status monitoring to value assessment and disposal execution. This transforms data asset inventory from a collection of scattered steps into a closed-loop intelligent management system. The identification of structural anomalies provides targeted direction for storage and security monitoring, while feedback on storage and security status enriches the dimensions of value assessment. The results of value assessment guide the specific methods of disposal execution, and the effectiveness of disposal execution, in turn, feeds back into subsequent structure analysis, forming a cycle of continuous optimization. Attached Figure Description
[0023] Figure 1 This is a schematic diagram illustrating the working principle of the intelligent data asset inventory system described in this invention. Figure 2 This is a flowchart illustrating the analysis process of the storage medium by the storage status monitoring module. Figure 3 This is a flowchart illustrating the execution process of the safety status monitoring module. Figure 4 A schematic diagram illustrating the working principle of the storage strategy optimization module. Detailed Implementation
[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] Please see Figure 1 This invention provides a data asset intelligent inventory system, the system comprising: The intelligent data asset inventory system comprises a hierarchical relationship determination module, a storage status monitoring module, a security status monitoring module, an asset value calculation module, and a tiered disposal execution module. The hierarchical relationship determination module acquires the metadata structure information of the target data assets, parses the dependency chains between data assets, and generates a topology graph. Based on the topology graph, it identifies abnormal data flow paths and generates structural anomaly alarm signals. These alarm signals activate the storage status monitoring and security status monitoring modules. The storage status monitoring module scans the storage media performance indicators of the target data assets, analyzes the read / write stability of the storage media, outputs a storage media health assessment result, and transmits it to the asset value calculation module. The security status monitoring module detects the access permission configuration and encryption policy status of the target data assets, marks unauthorized access behaviors and security vulnerabilities, generates a security threat assessment report, and transmits it to the asset value calculation module. The asset value calculation module receives the storage media health assessment result and the security threat assessment report, combines the data asset usage frequency and business relevance weights, performs value quantification analysis, and generates asset retention decision signals or asset migration decision signals. The tiered disposal execution module generates storage location optimization parameters and access permission reset parameters based on asset retention decision signals or asset migration decision signals.
[0026] Example 1: See Figure 2 The hierarchical relationship determination module begins its data asset dependency resolution process by examining the metadata attribute fields of the target data asset. These metadata attribute fields include the table name, field names, data types, constraints, and relationships with other data tables. The system constructs a field-level relationship mapping table by parsing this field information. This table records the source, destination, and reference relationships of each field with other fields. During construction, the system automatically identifies foreign key relationships, primary key dependencies, and cross-table reference paths between fields. By analyzing these reference paths, the system can depict the complete dependency chain between data assets.
[0027] During dependency chain analysis, the system focuses on detecting abnormal paths. Abnormal paths include two types: circular references and orphaned nodes. Circular references refer to closed dependencies between data tables, such as table A referencing table B, table B referencing table C, and table C referencing table A. Such circular references can lead to chaotic data processing logic or degraded system performance. Orphaned nodes refer to data tables or fields that are not referenced by any other data assets, potentially indicating that the data is obsolete or underutilized. The system automatically marks these abnormal paths by traversing the field-level relationship mapping tables and visualizes them in a topology graph.
[0028] Analyzing the number and distribution of abnormal paths is a crucial step in the hierarchical relationship determination module. The system calculates the proportion of data table nodes involved in abnormal paths relative to the total number of data tables, and generates a topology anomaly coefficient by combining this with a reference depth threshold. The reference depth threshold measures the complexity of data dependency chains; for example, multiple levels of nested references may lead to a higher risk of topology anomalies. The topology anomaly coefficient combines the proportion of abnormal paths with reference depth; a higher value indicates a more severe structural problem with the data asset. When the topology anomaly coefficient exceeds a preset topology fault tolerance threshold, the system triggers a structural anomaly alarm signal. This signal not only alerts the administrator to potential problems but also automatically activates the storage status monitoring module and the security status monitoring module to further analyze the health status of the data asset.
[0029] The topological anomaly coefficient is calculated using a hierarchical weighted method, and the specific formula is as follows: ,in, For topological anomaly coefficients, and Preset weighting coefficients (values ranging from 0 to 1 and...) ), This represents the number of abnormal paths. This represents the total number of paths. This represents the maximum reference depth of the abnormal path. This represents the average reference depth across all paths. Reference depth is defined as the number of hops from a data node to the root node; for example, the reference depth of a three-level dependency is 3. When... A structural risk warning is triggered when the threshold (e.g., 0.7) is exceeded.
[0030] The storage status monitoring module's workflow begins with the classification of physical storage media. Target data assets may be stored on different media, including hard disk drives (HDDs), solid-state drives (SSDs), or hybrid storage arrays. The system first identifies the type and performance parameters of each type of storage media, and then collects key metrics in real time. For HDDs, the system monitors read / write response latency and the number of bad sectors; for SSDs, the system focuses on wear and tear and remaining lifespan. These metrics reflect the current state and potential risks of the storage media.
[0031] Read / write response latency is a crucial metric for measuring storage performance. The system calculates the deviation by comparing actual latency with baseline performance parameters. These baseline parameters are pre-set based on the storage media's model and specifications, such as the theoretical read / write speeds of a specific solid-state drive (SSD). A higher deviation indicates a more significant performance degradation of the storage media. The number of bad sectors is particularly important for hard disk drives (HDDs), as an increase in bad sectors can lead to data loss or read / write failures. The wear and tear of SSD cells is assessed through write cycles and the percentage of remaining lifespan; high wear indicates that the storage media is nearing the end of its lifespan.
[0032] The system calculates a storage media health index by weighting performance deviation and hardware wear and tear values. Performance deviation reflects short-term performance fluctuations, while hardware wear and tear values reflect long-term wear and tear. During the weighted calculation, the system adjusts the weights based on the storage media type; for example, the number of bad sectors on mechanical hard drives has a higher weight, while the wear and tear on solid-state drives has a higher weight. The final storage media health index is a comprehensive score; the lower the value, the worse the health of the storage media. This index is written into the storage media health assessment results and transmitted to the asset valuation module for subsequent decision-making.
[0033] The collaborative operation of the hierarchical relationship determination module and the storage status monitoring module enables comprehensive monitoring of the data asset structure and storage status. Structure anomaly alarm signals not only indicate dependency issues but also trigger in-depth inspection of the storage media. This linkage mechanism ensures the system can promptly identify potential risks and provides a basis for optimized data asset management. The generation process of storage media health assessment results strictly relies on real-time data acquisition and dynamic calculation to ensure the accuracy and timeliness of the assessment results.
[0034] In the hierarchical relationship determination module, the construction of the topology graph is a continuous optimization process. The system regularly updates metadata information, re-analyzes dependencies, and adjusts the abnormal path marking strategy based on the latest data flow. This dynamic update mechanism enables the system to adapt to frequent changes in data assets, avoiding false alarms or missed alarms caused by data architecture adjustments.
[0035] The data acquisition frequency of the storage status monitoring module is dynamically adjusted based on the type of storage medium and the usage environment. The dynamic adjustment strategy is based on a dual-dimensional linkage mechanism of media type and environmental parameters: For SSD storage media, when five consecutive read / write latencies exceed 1ms, the acquisition frequency is increased from the default 1 minute / time to 10 seconds / time; if a temperature exceeding 45℃ is detected simultaneously, it is further increased to 5 seconds / time. For HDD storage media, when the vibration sensor detects an amplitude exceeding 0.5g, the acquisition frequency automatically switches to 2 seconds / time, and motor speed fluctuations are continuously monitored. Environmental parameter thresholds can be dynamically modified through the system configuration file, with default values referencing industry standard GB / T19729-2005. For high-performance storage devices, the system uses a higher acquisition frequency to capture instantaneous performance fluctuations; for archive storage devices, the system reduces the acquisition frequency to minimize resource consumption. This flexible acquisition strategy optimizes system resource utilization while ensuring monitoring effectiveness.
[0036] The marking of abnormal paths and the calculation of topology anomaly coefficients are both automated processes, reducing the need for manual intervention. The system identifies anomaly patterns through preset rules and machine learning models, and automatically triggers alarms when problems are detected. This automated processing significantly improves the efficiency of data asset inventory, and is particularly suitable for large-scale distributed storage environments.
[0037] The generation of storage media health assessment results relies not only on real-time monitoring data but also on historical trend analysis. The system records the performance change trajectory of the storage media, identifies long-term performance degradation trends, and reflects them in the health index. This trend analysis helps predict the future state of the storage media, providing a reference for preventative maintenance. The storage media health assessment employs a time-series fusion algorithm. Specifically, it first uses exponential smoothing to predict trends in historical monitoring data (such as the bad block growth rate and average response time over the past 7 days) to obtain a baseline health value. Then, the real-time monitoring data (such as the current number of bad blocks and read / write error rate) is normalized to obtain a real-time health score. Final health assessment results ,in This is a dynamic weighting coefficient (adaptively adjusted based on data fluctuations, ranging from 0.3 to 0.7). When If the score falls below the threshold (e.g., 60 points) for three consecutive times, the system will automatically trigger the data migration contingency plan.
[0038] The outputs of the hierarchical relationship determination module and the storage status monitoring module are ultimately converged into the asset value calculation module, serving as crucial inputs for the quantitative analysis of data asset value. Structural anomaly alarm signals and storage media health assessment results jointly influence the asset value calculation logic, ensuring the comprehensiveness and scientific rigor of the decision-making signals.
[0039] The entire implementation process emphasizes seamless integration between modules and real-time data flow. The structural analysis of the hierarchical relationship determination module provides triggering conditions for the storage status monitoring module, while the output of the storage status monitoring module further enriches the input dimensions of the asset value calculation module. This integrated design enables the data asset intelligent inventory system to efficiently and accurately complete the entire process management from structural analysis to health assessment and value decision-making.
[0040] Example 2: See Figure 3 The security status monitoring module's execution process begins with a comprehensive scan of the target data assets' access control configurations. The system first obtains the access control lists for all data interfaces and services, including user permission settings, role assignment rules, and access policy definitions. During the inspection, the system focuses on identifying open interfaces that lack authentication mechanisms, which may allow anonymous access or use weak authentication methods. The system determines the security level of the interfaces by analyzing their communication protocols, authentication header information, and session management mechanisms, and marks interfaces that do not meet preset security standards.
[0041] Building upon access control checks, the system simultaneously initiates a data encryption status review process. This process performs a deep scan of encrypted data fields during storage and transmission, checking aspects such as encryption algorithm strength, key management mechanisms, and key validity status. By analyzing the data storage structure and communication protocols, the system identifies fields using weak encryption algorithms or expired keys. For systems employing automatic key rotation strategies, the module verifies whether key update records comply with predetermined rotation cycle requirements. All detected expired encrypted fields and weak encryption instances are recorded and categorized by the system.
[0042] The calculation of the security threat value integrates security risk indicators from multiple dimensions. The system first counts the number of open interfaces and their risk levels. High-risk interfaces typically refer to those that allow administrator-level operations and lack multi-factor authentication. Simultaneously, it counts the number of expired encrypted fields and their impact, highlighting encryption failures involving sensitive data. These basic statistical data are correlated with a historical security event database. The system queries recent attack patterns, successful intrusion counts, and vulnerability exploitation records against similar data assets. By comprehensively analyzing the current vulnerability status and historical attack frequency, the system generates a quantified security threat value. The quantification of the security threat value is achieved through a weighted integration of risk indicators from three core dimensions. First, the risk of open interfaces is assessed, with scores assigned according to interface risk level: 5 points for each high-risk interface (allowing administrator-level operations and lacking multi-factor authentication), 3 points for each medium-risk interface (ordinary user interfaces with weak authentication mechanisms), and 1 point for each low-risk interface (internal access only). The total open interface risk score is then calculated. Secondly, the risk of expired encrypted fields is assessed, with scores assigned based on field sensitivity: 4 points are awarded for each expired encrypted field involving core business data, and 2 points for each expired encrypted field in ordinary business data. The total risk score for expired encrypted fields is then calculated. Finally, the frequency of historical attacks is evaluated, with scores assigned based on the average number of attacks per week over the past 3 months: 5 or more attacks per week = 5 points, 3-4 attacks = 3 points, 1-2 attacks = 1 point, and 0 attacks = 0 points. The total scores from these three dimensions are then combined using preset weights (e.g., 40% for open interface risk, 30% for expired encrypted field risk, and 30% for historical attack frequency) to obtain a security threat value, ranging from 0 to 10 points, with higher scores indicating a more severe security threat.
[0043] The security threat assessment report not only includes the final security threat value but also details all detected security issues and their potential impact. The report prioritizes issues by risk level, with high-risk issues typically involving unauthorized access to core business data or large-scale data breaches. The report also indicates the remediation priority of each issue, considering factors such as the difficulty of exploiting vulnerabilities, the potential business impact of an attack, and the complexity of the remediation plan. The system automatically transmits the security threat assessment report to the asset valuation module, serving as a key input for subsequent asset valuation analysis.
[0044] During security status monitoring, the system employs a layered detection mechanism to improve scanning efficiency. The first layer, rapid scanning, primarily identifies obvious security issues, such as fully open interfaces or expired encryption keys. The second layer, in-depth analysis, performs a more detailed examination of suspicious targets discovered in the initial detection, including analysis of interface call chains and tracing the actual usage scenarios of encrypted fields. This layered approach optimizes system resource utilization while ensuring comprehensive detection coverage.
[0045] The security status monitoring module maintains real-time data interaction with other components of the system. When a serious security threat is detected, the module immediately sends an alert to the anomaly handling linkage module, triggering real-time access blocking or data isolation measures. Simultaneously, all security detection results are updated to the central log database, which is used by the asset profiling iteration module for long-term security trend analysis and model optimization.
[0046] The module's execution process fully considers the varying security requirements of different data assets. For highly sensitive assets containing personal privacy data or trade secrets, the system automatically adopts stricter security check standards and more frequent scanning cycles. For publicly available data or low-sensitivity information, the detection depth and frequency are adjusted accordingly to achieve a balance between security control and system performance.
[0047] The security threat assessment algorithm continuously learns and optimizes. The system records the results of each security detection and subsequent actual security events, adjusting the parameters of the assessment model by comparing predicted threats with actual risks. This self-optimization mechanism allows the accuracy of security status monitoring to continuously improve as the system runs. The security threat assessment algorithm uses a gradient boosting decision tree model, and its acquisition process is as follows: Historical security event data from the past 36 months is collected, including 12 feature parameters such as the number of open interfaces, frequency of unauthorized access, number of expired encrypted fields, and number of successful historical attacks, as well as corresponding actual risk level labels (divided into four levels according to industry standards: low, medium, high, and extremely high). The collected feature parameters are normalized, scaling the values to the 0-1 range. Outliers are removed using box plots, and the SMOTE algorithm is used to adjust the sample ratio, ensuring that the sample size for each risk level is close to 1:1. The model was trained using a preprocessed dataset, with 100 decision trees, a maximum tree depth of 8, and a learning rate of 0.05. Five-fold cross-validation was employed, and the harmonic mean of precision and recall was used as the evaluation metric. Training was stopped when this metric consistently exceeded 0.85. The model parameters were saved and integrated into the security status monitoring module, which receives real-time feature data and outputs security threat values via an API interface. Model parameter adjustments were performed as follows: The parameters to be adjusted included the weights of each feature (range 0-1) and the risk level classification threshold (initially set to 60). The predicted threat values from the previous 24 hours were compared with the actual risk results every day at midnight. Adjustment was triggered when the absolute value of the prediction error for a single sample exceeded 15, or when the average error of 10 consecutive samples exceeded 10. If the predicted value is higher than the actual risk (false positive), the weight of the corresponding feature is reduced (by 0.05 each time, but not lower than 0.1), and the threshold for determining that risk level is increased (e.g., from 60 to 65). If the predicted value is lower than the actual risk (false negative), the weight of the corresponding feature is increased (by 0.05 each time, but not exceeding 0.9), and the threshold for determining that risk level is decreased (e.g., from 60 to 55). After each adjustment, the harmonic mean of the model must be verified. If it is lower than 0.8, the adjustment is revoked, and the parameters of the previous version are used. Full retraining on the entire dataset is performed monthly to avoid parameter drift.
[0048] At the detection technology implementation level, the module employs multiple complementary methods to improve the comprehensiveness of vulnerability identification. Static analysis primarily examines the definition files of access control policies and encryption configurations, while dynamic analysis tests the actual security protection effectiveness of the interface by simulating access requests. The combined use of these two methods can effectively reduce false positives and false negatives.
[0049] The security status monitoring module is designed with scalability and adaptability in mind. New security detection rules and vulnerability signatures can be dynamically loaded into the system, expanding detection capabilities without interrupting existing services. This design enables the system to quickly respond to emerging security threat types and attack methods.
[0050] The module's execution process is supported by a robust audit trail mechanism. All security detection operations, identified issues, and generated reports are logged with detailed information, including operation time, executing account, detection target, and result summary. This log data is encrypted, signed, and stored in a dedicated audit database to meet the needs of compliance review and post-event traceability.
[0051] The security status monitoring module and the tiered response execution module work closely together. When a security threat assessment report recommends adjusting access permissions, the relevant parameters are automatically transmitted to the tiered response execution module, triggering an immediate permission reset operation. This closed-loop processing mechanism significantly shortens the response time from threat discovery to actual remediation.
[0052] The module's exception handling mechanism can properly handle various special situations. When encountering abnormal data formats or configuration conflicts during the detection process, the system will automatically switch to safe mode to continue running, avoiding interruption of the entire monitoring process due to local problems. Detailed error diagnosis information will also be generated for technical personnel to analyze and handle.
[0053] The user interface of the security status monitoring module offers flexible report customization capabilities. Administrators can select the scope, level of detail, and presentation format of reports as needed. The system supports report filtering and aggregation by multiple dimensions, such as business department, data category, or risk level. All reports can be exported as standard format documents for further analysis or archiving.
[0054] The module's performance optimization measures include intelligent scheduling and resource allocation for detection tasks. The system dynamically adjusts the priority and execution intensity of scanning tasks based on the importance of data assets, historical security records, and real-time system load. This adaptive resource management approach ensures that security monitoring does not significantly impact the normal operation of business systems.
[0055] The security status monitoring module is implemented with full consideration of the complexity of enterprise IT environments. The module supports distributed deployment and clustered operation, effectively addressing the security management needs of large-scale data assets. It also provides standardized API interfaces for easy integration with existing enterprise security information and event management systems.
[0056] The module's configuration management adopts a layered strategy, including global default settings, business unit-level customized rules, and special configurations for specific assets. This flexible configuration system enables security policies to balance unified standards and specific needs, achieving refined management of security control.
[0057] The knowledge base of the security status monitoring module is updated regularly to incorporate the latest security threat intelligence and best practices. The update process can be automated online, ensuring that detection rules and assessment standards keep pace with evolving security landscapes. A complete update history of the knowledge base is maintained, supporting version tracking and change auditing.
[0058] Example 3: The value quantification analysis process of the asset valuation module begins with the extraction of usage characteristics of the data asset. The system retrieves recent query and call logs of the target data asset. These logs record complete usage information, including access time, access subject, and operation type. Through cleaning and parsing of the log data, the system calculates the access frequency per unit time and distinguishes the different weights of read and write operations. Simultaneously, it analyzes the number of related business systems, identifying upstream systems that directly call the data asset and downstream systems that depend on its output. Access frequency and business relevance together constitute the basic value dimensions of the data asset.
[0059] The health and safety factor is calculated using a multi-dimensional risk assessment model. The specific steps are as follows: First, basic data such as the failure rate of the storage medium, ambient temperature and humidity, and data integrity verification results are collected. Each data point corresponds to a preset risk weight coefficient (e.g., failure rate weight 0.4, temperature and humidity weight 0.3, data verification weight 0.3). Then, the data for each dimension is standardized to a score of 0-100. For example, if the failure rate exceeds a threshold, it is mapped to 80-100 points using an exponential function; if the temperature and humidity are within a safe range, it is mapped to 0-20 points. Finally, a weighted summation formula is used. Calculate the health and safety factor, where For health and safety reasons, Let i be the standardized score of the i-th indicator. The corresponding weights are denoted by , and the sum of all weight coefficients is 1; n is the number of indicators involved in the calculation. The coefficient ranges from 0 to 100, with higher values indicating lower risk.
[0060] Storage media health assessment results and security threat assessment reports were incorporated into the analysis as two additional key input dimensions. The storage media health index reflects the reliability level of data storage; a lower index indicates a more unstable storage media condition, potentially affecting data availability. The security threat value quantifies the degree of security risk faced by the data asset; a higher value indicates a more serious security vulnerability. The system integrates these indicators with usage characteristics to generate the final asset value coefficient.
[0061] The asset value coefficient is calculated using the following formula:
[0062] in, This represents the asset value coefficient. This represents the normalized access frequency. Indicates the number of associated business systems. This is a health index for storage media. It is a security threat value. , , , These are the weighting coefficients for each indicator, which are dynamically adjusted based on the data asset type and business importance. For core business data, security weighting... The weight will be increased accordingly; for frequently accessed cached data, the access frequency weight will be increased. This accounts for a larger proportion. The normalization process for access frequency is as follows: The query call logs of the target data asset for the past 30 days are retrieved, and the daily access frequency (in times / day) is calculated, resulting in 30 actual access frequency data points (all integers not less than 0). The highest access frequency within these 30 days is used as the baseline value. If there are no access records within 30 days, the baseline value is set to 1 (to avoid a denominator of 0 in the calculation); if there are access records, the baseline value is the maximum value among the 30 data points. For the daily access frequency, the actual frequency for that day is divided by the baseline value to obtain the normalized access frequency for that day (value range 0-1, 0 represents no access, 1 represents reaching the historical highest access level). The average of the normalized access frequencies over the past 30 days is taken as the final normalized access frequency. If the data asset is a newly added asset (existing for less than 30 days), the maximum access frequency within the actual number of days of existence shall be used as the benchmark value and calculated in the above manner; if there is a periodic access pattern due to business characteristics (such as monthly report data), the system administrator may manually specify the benchmark value (the source and reason for the benchmark value must be recorded in the system configuration, such as the maximum access frequency in the same period in history).
[0063] The decision signal is generated based on a comparison between the asset value coefficient and a preset threshold. The system maintains a dynamically adjusted value retention threshold, which comprehensively considers factors such as storage cost, business continuity, and data importance. When the asset value coefficient is higher than the threshold, the system generates an asset retention decision signal, indicating that the data asset is worth continuing to maintain; when the coefficient is lower than the threshold, an asset migration decision signal is generated, suggesting that the data be moved to a lower-cost storage tier. The threshold itself is also recalculated periodically to adapt to changes in the business environment. The determination and dynamic adjustment of the value retention threshold are as follows: the initial threshold is determined by comprehensively considering three factors: storage cost, business continuity, and data importance. The storage cost coefficient ranges from 0 to 1, determined by the ratio of the unit capacity cost of the current storage medium for data assets to the enterprise's average storage cost (e.g., 0.8 when the current cost is 20% higher than the average cost, and 0.3 when it is lower than the average cost). The business continuity coefficient ranges from 0 to 1, determined by the impact of data asset access interruptions on business operations (e.g., 1.0 when core transaction data interruptions cause business shutdowns, and 0.2 when non-core log data interruptions have a minor impact). The data importance coefficient ranges from 0 to 1, comprehensively assessed based on the compliance requirements of the data assets (e.g., whether they are core data) and historical contribution (the proportion of related business revenue in the past year) (e.g., 1.0 for core data, and 0.1 for ordinary archived data). The initial threshold is obtained by weighting and summing the above three coefficients according to the default weights (30% for storage cost, 40% for business continuity, and 30% for data importance). The weights can be adjusted according to the enterprise's business type (e.g., in the financial sector, the weight of business continuity can be increased to 50%). When dynamically adjusting, the threshold is automatically updated every 7 days, or adjusted immediately in the following situations: storage cost fluctuations exceed 10% (based on the average cost of the past 30 days), significant changes occur in the business system (such as core module upgrades or new business launches), or the data importance level is adjusted due to compliance requirements. During adjustment, the threshold of the previous period is used as a basis, combined with the rate of change of the storage cost coefficient, business continuity coefficient, and data importance coefficient (the difference between the current value and the previous period's value), and the adjustment magnitude is calculated using an adjustment sensitivity coefficient (default is 0.5) to ensure that each adjustment magnitude does not exceed ±20%.
[0064] The tiered disposal execution module takes corresponding actions based on the received decision signals. For asset retention decision signals, the module initiates a storage strategy optimization process. This process analyzes the access pattern characteristics of data assets, including time distribution patterns and concurrent access intensity. Based on these characteristics, the system dynamically adjusts the storage medium type, such as migrating data with significant peak access periods to higher-performance storage devices. Simultaneously, the backup strategy is optimized, increasing the backup frequency for frequently modified data and appropriately reducing the backup frequency for stable data.
[0065] Upon receiving an asset migration decision signal, the module performs a hot / cold data tiering operation. The system establishes a detailed access frequency history to identify data assets with consistently low access volume. This data is marked as cold data and migrated to a lower-cost storage cluster, such as high-density hard disk drives or cloud storage services. During the migration process, the system automatically closes unnecessary access ports, retaining only the minimum required read interfaces. For sensitive data, the encryption policy is updated synchronously during the migration process to ensure data security in the low-cost storage environment.
[0066] A closed-loop management system is formed between the asset valuation module and the tiered disposal execution module. After each migration or optimization operation, the system continuously monitors changes in data asset usage patterns and incorporates this feedback information back into the valuation model. This closed-loop mechanism enables the system to adapt to changes in business needs and adjust data asset storage strategies in a timely manner.
[0067] In terms of implementation details, the access frequency statistics employ a sliding time window algorithm. The system maintains access counters at multiple time granularities, including hourly, daily, and weekly levels, and eliminates the impact of short-term fluctuations through weighted averaging. The calculation of business system correlation considers not only direct call relationships but also analyzes the length and strength of indirect dependency chains to comprehensively assess the importance of data assets within the business architecture. Business system correlation is a quantitative indicator that measures the tightness of direct or indirect dependencies between data assets and various business systems, ranging from 0 to 1, with higher values indicating tighter correlations. The calculation method is as follows: Quantification of direct call relationships involves collecting direct interaction logs between data assets and business systems (such as API call records and database access logs), and statistically analyzing the frequency of direct calls and the proportion of key operations (the proportion of write / modify operations to total calls, ranging from 0 to 1) within a unit of time. The frequency of direct calls is normalized by dividing the actual frequency by the maximum daily call frequency within the historical period of the data asset to obtain the direct call correlation degree (value 0-1). This is then combined with the proportion of key operations to calculate the basic direct correlation value (direct call correlation degree multiplied by (0.7 + 0.3 × key operation proportion)). The length of indirect dependency chains is quantified by traversing the topological relationship graph of the data asset, identifying all indirect dependency paths (e.g., if data asset A is called by system B, and the output of system B is depended on by system C, then A and C form a second-level indirect dependency), recording the number of levels (length) of each path, and performing a length decay process (length weight is 1 divided by (length × 2), ensuring a value between 0-0.5). The quantification of indirect dependency chain strength involves calculating the importance coefficient (1 for core business systems, 0.6 for general business systems, and 0.3 for edge systems) and the interaction frequency ratio (0-1) of intermediate nodes with upstream and downstream systems for each indirect path. The strength of a single path is the product of these two values. The strengths of all indirect paths are weighted and summed according to their length to obtain the basic indirect association value (0-1). The business system association degree is the weighted sum of the basic direct association value and the basic indirect association value. The weighting coefficients are determined based on the business system type (70% direct association and 30% indirect association in core business systems, and 50% direct association and 50% indirect association in non-core business systems; this can be configured by the system administrator).
[0068] The use of storage media health indices introduces the concept of a decay factor. Considering the potential for continuous deterioration in storage device performance, the system applies a time decay to historical health indices, allowing recent test results to have a greater impact on value calculation. This approach more sensitively reflects real-time changes in storage status.
[0069] The integration of security threat values employs a risk level mapping mechanism. The original security threat values are first categorized into multiple risk level ranges, each corresponding to a different coefficient adjustment strategy. High-risk levels trigger additional value deductions, while low-risk levels may only have a slight impact. This tiered approach avoids excessive interference from security factors in value calculation.
[0070] The operation records of the tiered disposal execution module are fully saved, forming a disposal history database. This data is used to analyze the actual effects of storage strategy adjustments, including performance improvements and cost savings. The analysis results are fed back to the asset value calculation module to optimize the settings of weighting coefficients and decision thresholds.
[0071] Data exchange between modules uses a standardized message format. The decision signal output by the asset valuation module includes complete calculation basis and confidence level assessment, allowing the tiered disposal execution module to adjust the aggressiveness of its operations based on this additional information. For value coefficients near the boundary, the system may choose to observe for a period of time before making a final decision to avoid frequent strategy changes.
[0072] The anomaly handling mechanism ensures the system's robustness. When a sudden change in data access patterns or a sharp deterioration in storage status is detected, the system can temporarily override the regular decision-making process and directly trigger protective operations. These anomalies are specifically marked and given special attention in subsequent analysis.
[0073] The asset valuation model supports manual intervention and adjustments. Administrators can set mandatory retention flags for specific data assets or adjust their calculation parameters. These manual settings are explicitly recorded and displayed separately from the results of automated decisions, maintaining system transparency.
[0074] The value quantification analysis process considers the characteristics of the data lifecycle. For data assets nearing their retention period, the system automatically reduces their value coefficient to promote the timely cleanup of expiring data. Simultaneously, it identifies data with long-term preservation value and appropriately increases their weight coefficient to prevent the erroneous migration of important historical data.
[0075] The tiered processing module's resource allocation strategy balances efficiency and fairness. The system monitors the storage resource pool's usage, ensuring high-value data receives necessary performance guarantees while preventing low-value data from excessively consuming premium resources. Resource allocation records are regularly reported to help administrators understand the actual utilization efficiency of storage resources.
[0076] The module implementation employs a distributed computing architecture to handle large-scale data assets. Value computation tasks are intelligently sharded based on data classification and business units, executed in parallel, and then the results are aggregated. This architectural design significantly improves the system's response speed when processing massive amounts of data assets.
[0077] The asset valuation module's weighting coefficient setting interface provides visual aids. Administrators can intuitively adjust the relative importance of each dimension through interactive charts and view real-time changes in the value coefficients of typical data assets. This design lowers the technical barrier to parameter tuning and encourages business personnel to participate in the decision-making process.
[0078] The tiered handling process employs a gradual strategy. For large-scale data migrations, the system first moves a small sample of data to verify compatibility before proceeding with the full migration. System performance and data integrity are continuously monitored throughout the process, and any issues are immediately rolled back. This cautious approach minimizes operational risks.
[0079] Example 4: See Figure 4 The implementation process of the storage strategy optimization module can be concretely demonstrated through a user behavior data analysis case of an e-commerce platform. This platform generates approximately 2 million user behavior records daily, including browsing, searching, adding to cart, and placing orders. This data is stored in a hybrid storage environment, comprising a high-speed SSD storage pool and a traditional hard disk drive cluster. The system first analyzes the access time distribution characteristics of the target data assets. Table 1 below shows the access time distribution statistics for different types of user behavior data within a certain week:
[0080] Table 1: Statistics on the distribution of access time for different types of user behavior data within a certain week Based on this structured analysis, the storage strategy optimization module divides data into hot and cold data zones. Real-time browsing records exhibit a typical bimodal access pattern, with significant peaks at midday and evening, while access volume drops to less than 10% of the peak in the early morning. The system marks this type of data as "periodic hot data" and designs a dynamic storage strategy for it: two hours before the predicted peak, the data is automatically migrated from mechanical hard drives to an SSD storage pool; one hour after the start of the off-peak period, it is migrated back to the lower-cost mechanical hard drive environment. The migration process is executed silently in the background to ensure that normal access services are not affected.
[0081] Historical order data reveals distinct access patterns: high traffic during working hours but almost no queries at night. The system categorizes this as "weekday hot data" and implements a weekday / holiday differentiated strategy: SSD storage is maintained during weekdays, while mechanical hard drive storage is degraded during holidays and off-peak hours. This strategy, compared to an all-day SSD storage solution, saves approximately 40% on storage costs while ensuring performance requirements during peak business periods.
[0082] The access pattern of product search logs is quite unique, maintaining a relatively stable volume throughout the day without significant peaks or troughs. Further analysis at a finer time granularity revealed that while the overall volume remained stable, short-term search hotspots existed in certain product categories. Therefore, a hybrid "baseline + hotspot" strategy was adopted: basic data was stored on the hard drive, while suddenly hot products were identified through real-time monitoring, and their related search logs were temporarily moved to SSD storage. Once the hotspots subsided, the SSD storage automatically reverted to SSD storage.
[0083] Access to user profile data is concentrated during weekday working hours, but each access requires loading a large amount of data. When analyzing the access time distribution characteristics, the system considers not only the access frequency but also the amount of data per access to calculate the storage cost-effectiveness ratio. This type of data is labeled as "large-capacity periodic data," and a preloading strategy is designed for it: one hour before the start of the workday, active user profile data is preloaded into the SSD cache; during off-peak access periods, only the active user data of the most recent 3 days is retained in high-speed storage, and the rest is degraded to mechanical hard drives.
[0084] Promotional campaign data is highly time-sensitive, with access fluctuating entirely in tandem with the marketing campaign cycle. The system establishes a calendar-based mechanism for this type of data, linking storage strategies with the marketing system: during the pre-campaign period, relevant data is moved to high-performance storage; after the campaign ends, a 1-3 day observation period is set, and storage is immediately downgraded once no further access is confirmed. This precise time-sensitive management avoids the resource waste caused by traditional fixed retention periods.
[0085] The resource allocation process of the storage strategy optimization module is based on deep learning of historical access patterns. The system analyzes the mapping relationship between access timestamps and data asset identifiers over the past 12 weeks to construct a time-dimensional access heat map. This map not only reflects known periodic patterns but also identifies abnormal access patterns. For example, sudden surges in access during normal off-peak periods are recorded by the system and compared with business logs to gradually improve the predictive model.
[0086] The prediction of future peak access periods employs a multi-model fusion algorithm. The system simultaneously runs a time-series-based statistical model, a business event-based rule model, and a deep learning-based neural network model, combining the prediction results of each model to generate a cache space demand forecast. Leading up to a major e-commerce promotion, the system detected a consistent traffic surge predicted by multiple models and began gradually expanding high-speed storage resources three days in advance to ensure data access performance during the promotion period.
[0087] Resource utilization monitoring and adjustment form a closed-loop feedback loop. The system compares the predicted cache space demand with the actual usage in real time, triggering a policy review mechanism when the deviation exceeds 15%. For example, if a prediction indicates a 50% drop in nighttime access volume, but the actual drop is only 30%, the system will automatically extend the high-speed storage retention time and record the deviation to improve subsequent prediction models. Resource release commands are executed gradually, prioritizing the release of the oldest unaccessed data blocks while monitoring overall system performance metrics to avoid over-release that could negatively impact user experience.
[0088] The calculation of storage cost-effectiveness considers multiple dimensions. In addition to basic storage hardware costs, the system also quantifies the impact of differences in data access latency across different storage tiers on business metrics. For example, A / B testing revealed that migrating user profile data from SSDs to HDDs increased page load time by 200 milliseconds, resulting in a 0.8% decrease in conversion rate. These business impacts are converted into cost equivalents and incorporated into the objective function for optimizing the storage strategy.
[0089] The implementation of the module significantly improved the efficiency of storage resource utilization. Through a dynamic allocation strategy, the daily utilization rate of the high-speed SSD storage pool decreased from 92% to 68%, while the cache hit rate remained above 95%. The load distribution of the mechanical hard drive cluster was more balanced, avoiding the problem of some nodes overheating while others were idle, which is common in traditional static allocation schemes. The optimization of storage costs is directly reflected in infrastructure expenditures, saving 15-20% of storage-related expenses per month compared to the previous solution.
[0090] The abnormal situation handling mechanism ensures the reliability of the system. When a data verification error is detected during storage migration, the system automatically aborts the migration process and rolls back to the previous stable state. Simultaneously, an alarm is triggered to notify operations and maintenance personnel, and detailed error context information is recorded for subsequent analysis. For data assets that frequently fail to migrate, the system marks them as special handling objects, employing a more conservative migration strategy or requesting manual intervention.
[0091] The module's decision-making process remains highly transparent. All storage policy adjustments generate detailed execution logs, including the basis for the decision, expected effects, and actual results. Administrators can use visualization tools to trace back the storage distribution status and decision-making process at any point in time. This transparency facilitates troubleshooting and supports continuous policy optimization. The system regularly generates storage optimization reports, summarizing the policy execution status and cost savings for each data category, providing data support for infrastructure planning.
[0092] Integration with the tiered disposal execution module enables end-to-end automated management. When the asset valuation module determines a change in the value of a data asset, the storage strategy optimization module responds immediately, reassessing its storage location requirements. For example, when it detects a continuous decline in the access frequency of a certain type of historical order data, the system gradually extends its retention time on the hard drive, and may eventually remove it from the high-speed storage pool entirely. This collaborative working mode ensures that storage resources always flow to the data assets with the highest business value.
[0093] Example 5: The implementation process of the anomaly handling linkage module begins with the generation trigger mechanism of the security threat assessment report. When the security status monitoring module completes the scan and outputs the assessment report, the system automatically activates the linkage response process. This process first parses the threat level classification in the report and immediately initiates a real-time response sequence for high-risk security events. The system calls the access log analysis engine to extract access records related to the security event from massive log data, including the source IP address, access timestamp, operation type, and involved data objects. The log analysis uses multi-dimensional correlation technology to connect discrete security event points into a complete attack trajectory map, clearly showing the path evolution process of abnormal access behavior.
[0094] The location of affected data assets employs a dynamic tagging and diffusion algorithm. Starting from the initially detected anomalous access point, the system expands its analysis outward along the data association network to identify all potentially affected related data assets. This analytical method considers not only direct access paths but also indirect data dependencies, ensuring the accuracy of the isolation scope. During the location process, the system assesses the impact level of each data asset in real time, generating coordinates of the data isolation area containing physical location and logical identifiers. This coordinate information is encoded in a standardized format, including location data at the storage node location, database instance, tablespace, and field level.
[0095] The generation of access permission reset parameters follows the principle of least privilege. The system analyzes the specific operation type of abnormal access behavior and adjusts permission settings accordingly. For detected unauthorized read attempts, the system strengthens the access control list for the relevant data; for abnormal write operations, write permissions are temporarily locked or read-only mode is enabled. Permission reset schemes are implemented only after conflict detection and impact assessment to avoid affecting normal business processes due to excessive restrictions. All reset operations retain complete change records, including original permission settings, reasons for modification, execution time, and operator information.
[0096] The asset profiling iteration module operates on the basis of periodic data collection. The system is set with a fixed inventory cycle, periodically extracting the latest asset value coefficients from the asset value calculation module. During each inventory check, the module records complete value assessment data, including the raw values and weighted calculation results of each dimension's indicators. This data is compared longitudinally with historical records to analyze value change trends and the driving factors behind them. Trend analysis not only focuses on numerical changes but also identifies the rate of change and inflection point characteristics, distinguishing between normal business fluctuations and abnormal value anomalies.
[0097] The process of adjusting business association weights incorporates a dynamic feedback mechanism. The system monitors changes in the interaction patterns between various business systems and data assets. When a new call relationship is detected or the strength of an existing relationship changes, the corresponding weight coefficients are automatically adjusted. The adjustment algorithm considers the criticality level of the business system and the dependency depth of the data asset, ensuring that the weight allocation reflects the actual business value contribution. The adjustment of the storage medium health weight parameter is based on the long-term trend of hardware performance indicators, appropriately increasing the proportion of the health impact in value calculation for storage media that are continuously deteriorating.
[0098] The process of feeding back weight parameters to the hierarchical relationship determination module employs version control management. Each weight update generates a unique version identifier, which, along with metadata such as modification descriptions and effective time, is stored in the configuration repository. When constructing the topology graph, the hierarchical relationship determination module can specify the use of a specific version of the weight parameters, ensuring the repeatability of the analysis process. This mechanism also supports parameter rollback operations, allowing for rapid restoration to a previous stable state when new weights lead to abnormal results.
[0099] The exception handling linkage module and the hierarchical handling execution module are bound together using an event-driven architecture. Once the coordinates of the data isolation area are generated, the system creates standardized event messages containing information such as the handling type, target object, and operation parameters. These messages are published to the distributed event bus, and the hierarchical handling execution module, acting as a subscriber, receives and processes these events. This loosely coupled design allows the two modules to evolve independently while maintaining efficient collaborative capabilities. The event processing status is monitored in real time, and any execution failure or timeout triggers alarms and retry mechanisms.
[0100] The asset profiling iteration module employs time series forecasting technology for historical data analysis. The system establishes time-series models for the value indicators of various data assets, identifying periodic patterns and long-term trends. These forecasts are used to optimize inventory cycle settings, shortening the inventory interval for data assets with high value fluctuations and extending the cycle for stable assets to reduce system overhead. The forecasting model is periodically retrained with the latest data to adapt to changes in the business environment.
[0101] Data isolation operations during security incident handling employ a tiered control strategy. Depending on the threat level, the system adopts a progressive approach, moving from logical isolation to physical isolation. Low-risk incidents may only involve detailed access logging; medium-risk incidents trigger access restrictions and operational auditing; and high-risk incidents implement physical isolation and data encryption protection. Each isolation level defines clear recovery conditions and operational procedures to ensure an orderly restoration of normal access after the security risk is eliminated.
[0102] The weight adjustment algorithm takes into account the differences between business domains. The system maintains feature models for different business units, identifying the unique value-influencing factors of data assets in each domain. When adjusting weight parameters, the algorithm references the feature model of the relevant business domain, avoiding local irrationality caused by general rules. This differentiated processing improves the accuracy of weight allocation, making value assessment more in line with the actual needs of specific business scenarios.
[0103] The tracking and analysis of abnormal access behavior employs behavioral profiling technology. The system establishes a behavioral baseline model for each accessing entity, recording its normal access patterns. When abnormal behavior is detected, the system compares the deviation of this behavior from the entity's historical profile to help determine whether it is account theft or a normal behavioral change. This profiling-based analysis method reduces false alarms and improves the accuracy of security responses.
[0104] The asset value trend visualization analysis tool supports multi-dimensional drill-down exploration. Administrators can interactively view the value change curves of specific data assets and overlay them for comparison with relevant business indicators. The system automatically marks significant inflection points in value changes and displays related business events or system changes that occurred during the same period, helping to understand the driving factors of value fluctuations. This visualization analysis provides intuitive decision support for weight adjustments.
[0105] The automated anomaly handling process incorporates a human review node. For high-risk security incidents or actions impacting core business operations, the system pauses the automated process and transfers it to the security team for manual confirmation. The review interface presents complete analysis conclusions and handling recommendations, allowing reviewers to view original logs and related evidence. After the human decision is fed back to the system, subsequent steps continue, and the entire review process is recorded for future audits.
[0106] The asset profiling iteration module's periodic inventory tasks support flexible scheduling. In addition to fixed time-period triggers, the system also responds to special inventory requests triggered by significant business events. When a major business system upgrade, organizational restructuring, or changes in regulations and policies are detected, the module automatically initiates a temporary inventory to assess the impact of these changes on the value of data assets. This flexible mechanism ensures that value assessments reflect the latest state of the business environment in a timely manner.
[0107] Data isolation during security procedures is controlled with fine granularity. The system supports field-level isolation policies, providing special protection for sensitive fields while allowing normal access to non-sensitive fields. This fine-grained control minimizes the impact of security procedures on business continuity, and is particularly suitable for complex data objects containing mixed levels of sensitivity.
[0108] The management of the weight parameter library includes complete lifecycle control. Each weight parameter set has a clearly defined effective time and scope of application, and the system automatically handles the addition, updating, and obsolescence of parameter versions. Historical parameter versions are archived and saved, supporting point-in-time value assessment and reproduction. Changes to the parameter library are subject to strict access control; all modifications require an approval process and are recorded in the operation audit log.
[0109] The evaluation of anomaly handling effectiveness forms a closed-loop feedback loop. The system tracks and records changes in relevant indicators after each security action, including the degree of threat elimination, false positives, and the scope of business impact. These evaluation results are used to optimize handling strategies and adjust detection thresholds, forming a continuous improvement mechanism for security operations. Effectiveness evaluation reports are generated regularly to help the security team understand the actual performance of the automated handling system.
[0110] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0111] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A data asset intelligent inventory system, characterized in that, include: The hierarchical relationship determination module is used to obtain the metadata structure information of the target data assets, parse the dependency relationship chain between data assets, and generate a topological relationship graph. Based on the topological relationship map, abnormal data flow paths are identified, and structural anomaly alarm signals are generated; the storage status monitoring module and the security status monitoring module are activated based on the structural anomaly alarm signals. The storage status monitoring module is used to scan the storage media performance indicators of the target data asset, analyze the read and write stability of the storage media, output the storage media health assessment results and transmit them to the asset value calculation module. The security status monitoring module is used to detect the access permission configuration and encryption policy status of the target data assets, mark unauthorized access behaviors and security vulnerabilities, generate a security threat assessment report and transmit it to the asset value calculation module. The asset value calculation module is used to receive the health assessment results of the storage medium and the security threat assessment report, combine the data asset usage frequency and business association weight, perform value quantification analysis, and generate asset retention decision signals or asset migration decision signals. The tiered disposal execution module is used to generate storage location optimization parameters and access permission reset parameters based on the asset retention decision signal or asset migration decision signal.
2. The intelligent data asset inventory system according to claim 1, characterized in that, The process of parsing data asset dependencies by the hierarchical relationship determination module includes: Extract metadata attribute fields from the target data assets and construct a field-level association mapping table; use the field-level association mapping table to track cross-data table reference paths and mark abnormal paths with circular references or orphaned nodes; The percentage of data table nodes in the abnormal path is counted, and a topology anomaly coefficient is generated by combining it with the reference depth threshold. When the topology anomaly coefficient exceeds the preset topology fault tolerance threshold, the structural anomaly alarm signal is triggered.
3. The intelligent data asset inventory system according to claim 1, characterized in that, The storage status monitoring module's analysis process for the storage medium includes: Classify and obtain the physical storage media type corresponding to the target data assets, and collect the read and write response latency of the disk array, the number of bad sectors, and the wear of solid-state storage units in real time; Based on the physical storage medium type, a baseline performance parameter is matched, and the deviation of the read / write response latency from the baseline performance parameter is calculated. By associating the number of bad sectors with the wear and tear of solid-state storage units, a storage medium health index is derived by weighting performance deviation and hardware wear values, and the storage medium health index is written into the storage medium health assessment result.
4. The intelligent data asset inventory system according to claim 1, characterized in that, The execution process of the security status monitoring module includes: Iterate through the access control list of the target data asset to identify open interfaces that are not bound to an authentication protocol; simultaneously scan the key validity period of the data encryption fields and mark expired encryption fields; The number of open interfaces and expired encrypted fields are aggregated, and a security threat value is calculated by combining the frequency of historical attack events. When the security threat value reaches a preset risk threshold, the security threat value is included in the security threat assessment report.
5. The intelligent data asset inventory system according to claim 1, characterized in that, The asset valuation calculation module's value quantification analysis process includes: Retrieve query call logs of the target data assets within a preset time period, and count the access frequency and the number of related business systems per unit time. By combining the storage media health index from the storage media health assessment results with the security threat value from the security threat assessment report, an asset value coefficient is generated using weighted access frequency, business system weight, and health and security coefficient. If the asset value coefficient is higher than the preset value retention threshold, an asset retention decision signal is generated; if the asset value coefficient is lower than the preset value retention threshold, an asset migration decision signal is generated.
6. The intelligent data asset inventory system according to claim 1, characterized in that, The operational logic of the hierarchical processing execution module includes: When the asset retention decision signal is captured, a storage strategy optimization instruction is initiated to dynamically adjust the storage medium type and backup frequency according to the data asset access pattern. When the asset migration decision signal is captured, a cold and hot data tiering instruction is initiated to migrate low-value data assets to a low-cost storage cluster based on access frequency, and to close access ports that are unrelated to the current business functions of the low-value data assets.
7. The intelligent data asset inventory system according to claim 1, characterized in that, It also includes a storage policy optimization module: Receive storage location optimization parameters from the hierarchical processing execution module and analyze the access time distribution characteristics of the target data assets; Based on the characteristics of access time distribution, cold and hot data intervals are divided, and the storage cost-effectiveness ratio of high-frequency access data is calculated. The ratio of high-speed storage resources to archive storage resources is dynamically allocated based on the aforementioned storage cost-effectiveness ratio.
8. The intelligent inventory system for data assets according to claim 7, characterized in that, The resource allocation process of the storage strategy optimization module includes: Establish a mapping relationship between access timestamps and data asset identifiers to generate a time-dimensional access heat map; Based on the time-dimensional access heat map, predict the future access peak period and adjust the cache space of high-speed storage resources in advance; The predicted cache space requirement is compared with the current storage resource utilization rate to generate resource expansion warnings or resource release instructions.
9. The intelligent data asset inventory system according to claim 1, characterized in that, It also includes an exception handling linkage module: When the security status monitoring module generates a security threat assessment report, it simultaneously triggers the access log analysis engine to track the source address and operation trajectory of abnormal access behavior. Based on the operation trajectory, locate the affected data asset range and generate data isolation area coordinates; bind the data isolation area coordinates with the access permission reset parameters of the hierarchical disposal execution module and execute.
10. The intelligent inventory system for data assets according to claim 1, characterized in that, The system also includes an asset profiling iteration module: The asset value coefficients generated by the asset value calculation module are periodically obtained and correlated with the value change trends of historical inventory periods; The business association weight and storage medium health weight parameters are adjusted according to the value change trend; the adjusted weight parameters are then fed back to the topology graph construction process of the hierarchical relationship determination module.
Citation Information
Cited By
Data maintenance management method and system based on artificial intelligence
CN121279622A