A data asset management system based on multi-source heterogeneous data fusion

By conducting security assessments and tracing the source of multi-source heterogeneous data fusion in scientific research databases, malicious data tampering was identified and blocked, thus solving the data pollution problem after the scientific research databases were attacked and achieving data security and reliability.

CN121093393BActive Publication Date: 2026-03-13CHINA NAT INST OF STANDARDIZATION
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing technologies lack effective traceability and anomaly detection mechanisms in scientific research databases, leading to data being maliciously attacked, tampered with, and continuously transmitted, thus contaminating downstream derived data.

Method used

A database security assessment module is used to conduct regular security assessments. By obtaining database access characteristic parameters, a robustness index is calculated to identify abnormal data and shorten the assessment interval. A data recovery module is used to recover and update downstream data. A full-domain data tracing module is used to trace and isolate multi-level derived data. A computing resource management module schedules resources to optimize recovery processing.

Benefits of technology

It enables dynamic security management of scientific research databases, quickly identifies and blocks malicious data tampering, promptly isolates and repairs abnormal data, prevents the spread of data pollution, and improves the security and reliability of data sharing and reuse.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121093393B_ABST
    Figure CN121093393B_ABST
Patent Text Reader

Abstract

This invention discloses a data asset management system based on multi-source heterogeneous data fusion, belonging to the field of data management technology. It includes the following modules: a database security determination module for obtaining the database security status; a data asset protection and control module for obtaining the data asset protection level; a data recovery module for performing abnormal data recovery and downstream data update processing; a full-domain data tracing module for generating recovery determination instructions; a computing resource management module for adjusting the execution status of full-domain derived data recovery processing; and a data asset assessment secondary adjustment module for secondary adjustment of the data asset security assessment execution interval. This invention solves the problem in existing technologies where malicious attacks on research databases lead to malicious tampering of the original data, and the continuous transmission of the original data during data sharing and reuse, resulting in continuous contamination of derived data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data management technology, and in particular to a data asset management system based on the fusion of multi-source heterogeneous data. Background Technology

[0002] Existing data asset management systems use a unified data access interface to collect and access structured, semi-structured, and unstructured data. The data is then cleaned, formatted, and features extracted. By using mapping rules, semantic matching, and model-driven approaches, heterogeneous data is unified into a single logical structure, forming an applicable fused dataset. Metadata management and tagging are used to achieve unified description and traceability control. Combined with access control and auditing mechanisms, the security and compliance of the data during its use are ensured.

[0003] For example, Chinese invention patent CN118761594B discloses a method and system for assessing power grid assets, which includes: obtaining the asset operation status of each asset in the power grid based on real-time monitoring data of the power grid, and performing data fusion to obtain a dataset of the asset within a preset time period; extracting features from the dataset to obtain relevant data of the asset at the same time within the preset time period; predicting the operating environment of the asset within the preset time period to obtain at least one operating environment of the asset within the preset time period; obtaining the operating efficiency of the asset under the operating environment; obtaining the maintenance requirements of the asset under the operating environment based on the operating efficiency and a preset standard efficiency; and obtaining the asset assessment of the asset within the preset time period based on the operating efficiency and maintenance requirements.

[0004] For example, Chinese invention patent CN117350448A discloses a method for integrating multi-source data in power planning, which includes three steps: filtering data noise, clustering analysis of data, and fusion processing of multi-source heterogeneous data. This invention builds a historical measurement center based on a data platform to ensure the consistency of power grid resource and asset information. Through real-time calculation, extrapolation, and analysis fitting of the digital system, it achieves a dynamic representation of the physical power grid in the digital space.

[0005] However, in the process of implementing the inventive technical solution in the embodiments of this application, it was found that the above-mentioned technology has at least the following technical problems:

[0006] Existing technologies primarily focus on format uniformity, semantic matching, and consistency maintenance, often neglecting security risk control during data sharing and reuse. Particularly in research database scenarios, where data sources are complex and updates are frequent, existing technologies lack effective tracing and anomaly detection mechanisms if the original data is maliciously attacked or tampered with. This makes it difficult to promptly detect and block the spread of abnormal data. Consequently, there is a problem where malicious attacks on research databases lead to malicious tampering of the original data, which is then continuously transmitted during data sharing and reuse, resulting in the continuous contamination of derived data from downstream sources. Summary of the Invention

[0007] To address the technical problems existing in current technologies, such as malicious attacks on research databases leading to malicious tampering of source data and continuous contamination of derived data during data sharing and reuse, this invention provides a data asset management system based on multi-source heterogeneous data fusion. The technical solution is as follows:

[0008] A data asset management system based on multi-source heterogeneous data fusion includes: a database security assessment module, used to periodically conduct data asset security assessments on multi-source heterogeneous data within the database, obtain database access characteristic parameters, analyze and obtain a database access robustness index, thereby determining the database security status; a data asset protection and control module, used to determine the data asset protection level based on the database security status analysis; if the data asset protection level is secure, the execution interval of the data asset security assessment will not be shortened; otherwise, abnormal data is obtained, and the execution interval of the data asset security assessment will be shortened based on the database access robustness index; and a data recovery module, used to obtain data pollution load parameters of downstream data, and combine them with the database access robustness index to analyze and obtain the comprehensive duration of permission lockout, thereby performing abnormal data recovery and downstream data update processing. The data represents a dataset generated by directly reusing or processing abnormal data; the full-domain data tracing module is used to obtain the data pollution load value of the full-domain derived data, thereby generating a recovery judgment instruction. If the recovery judgment instruction is in a recovery state, the full-domain derived data is restored; otherwise, the full-domain derived data is isolated. The full-domain derived data includes downstream derived data and multi-level derived data of each downstream derived data. Downstream derived data represents a dataset generated by directly reusing or processing downstream data; the computing resource management module is used to obtain the current resource scheduling reserve of the device during the full-domain derived data recovery process, thereby adjusting the execution status of the full-domain derived data recovery processing; the data asset assessment secondary adjustment module is used to make secondary adjustments to the execution interval of the data asset security assessment based on the data pollution load value of the full-domain derived data.

[0009] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:

[0010] 1. The data asset management system based on multi-source heterogeneous data fusion provided by this invention performs periodic security assessments on multi-source heterogeneous data within the database and analyzes database access characteristic parameters in real time to calculate a database access security index. This allows for dynamic judgment of the database access status. If potential risks are detected in access behavior, the hash values ​​of each multi-source heterogeneous data are verified one by one. This enables rapid identification of tampered data at the data source level, thereby blocking the continuous transmission of maliciously tampered data from the source. This effectively solves the problem that malicious attacks on scientific research databases lead to malicious tampering of the original data, which is then continuously transmitted during data sharing and reuse, resulting in continuous contamination of derived data from downstream sources.

[0011] 2. This invention analyzes downstream data and combines it with database access security index analysis to obtain the comprehensive duration of permission lockout. Within this duration, access permissions are locked for abnormal data and its downstream data, and data recovery and update processes are carried out simultaneously. This enables timely isolation and repair of contaminated data during data sharing and reuse, ensuring the integrity and reliability of downstream data.

[0012] 3. This invention constructs a traceability link for all-domain derived data, automatically and recursively tracks multi-level reference relationships, calculates the data pollution load value of each all-domain derived data, and compares it with a preset threshold. This allows the determination of data that can be recovered and data that must be isolated, thereby achieving full coverage governance of multi-level derived data pollution and effectively preventing the continuous spread of erroneous data in complex scientific research data links.

[0013] 4. This invention introduces the computing resource scheduling margin as a constraint in the process of global data recovery. It dynamically adjusts the concurrent processing volume of data recovery based on the amount of currently schedulable resources, thereby reducing concurrency when resources are scarce and increasing concurrency when resources are abundant. This achieves a dynamic balance between the data repair process and the allocation of computing resources, effectively improving the efficiency and continuity of large-scale data recovery processing. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 A schematic diagram of the structure of a data asset management system based on multi-source heterogeneous data fusion provided in this application embodiment;

[0016] Figure 2 A flowchart of a data asset management system based on multi-source heterogeneous data fusion provided in this application embodiment;

[0017] Figure 3 A flowchart illustrating the data recovery process of a data asset management system based on multi-source heterogeneous data fusion, provided in this application embodiment. Detailed Implementation

[0018] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0019] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0020] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0021] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0022] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0023] like Figure 1The diagram shown is a structural schematic of a data asset management system based on multi-source heterogeneous data fusion provided in this application embodiment. The system includes: a database security determination module, used to periodically perform data asset security assessments on multi-source heterogeneous data within the database, obtain database access characteristic parameters, analyze and obtain a database access robustness index, thereby obtaining the database security status; a data asset protection and control module, used to obtain the data asset protection level based on the database security status analysis; if the data asset protection level is secure, the execution interval of the data asset security assessment will not be shortened; otherwise, abnormal data is obtained, and the execution interval of the data asset security assessment is shortened based on the database access robustness index; and a data recovery module, used to obtain data pollution load parameters of downstream data, and combine them with the database access robustness index to analyze and obtain the comprehensive duration of permission lockout, thereby performing abnormal data recovery and subsequent data recovery. The system comprises several modules: a downstream data update module (where downstream data represents datasets generated by directly reusing or processing abnormal data); a full-domain data tracing module (where the full-domain derived data module obtains the data pollution load value of the full-domain derived data and generates a recovery judgment instruction; if the recovery judgment instruction indicates a recovery status, the full-domain derived data is recovered; otherwise, the full-domain derived data is isolated; the full-domain derived data includes downstream derived data and multi-level derived data of each downstream derived data, where downstream derived data represents datasets generated by directly reusing or processing downstream data); a computing resource management module (where the computing resource management module obtains the current resource scheduling reserve of the device during the full-domain derived data recovery process and adjusts the execution status of the full-domain derived data recovery process); and a data asset assessment secondary adjustment module (where the data asset security assessment execution interval is adjusted based on the data pollution load value of the full-domain derived data).

[0024] In this embodiment, it should be noted that the multiplicative coupling process is specifically a multiplication process.

[0025] The data asset protection level is determined based on the database security status analysis. Specifically: if the database security status indicates potential risks, the data asset protection level is Level 1. In this case, abnormal data is analyzed, and the execution interval of the data asset security assessment is shortened based on the database access robustness index. If the database security status is secure, the data asset protection level is Level 2, and the execution interval of the data asset security assessment is not shortened.

[0026] The method for shortening the execution interval of data asset security assessment based on the database access robustness index is as follows: The database access robustness index is matched with the database to obtain an initial execution interval shortening coefficient. The shortened execution interval is then obtained by subtracting the product of the initial execution interval shortening coefficient and the data asset security assessment execution interval from the initial execution interval. Specifically, the initial execution interval shortening coefficient is obtained by matching the database access robustness index with the database. This involves first retrieving historical data from the database, including the database access robustness index at each historical time point and the corresponding initial execution interval shortening coefficient, forming historical data pairs. For the real-time measured database access robustness index, the two closest historical difference values ​​are found in the historical data, designated as the lower and higher values, and the corresponding historical initial execution interval shortening coefficients are obtained. Then, a linear interpolation method is used to proportionally map the real-time difference value between these two historical values, calculating the corresponding real-time initial execution interval shortening coefficient. Specifically, the real-time adjustment factor equals the lower historical adjustment factor plus the ratio of the real-time difference value to the historical low and high values ​​multiplied by the difference between the two historical adjustment factors. If the real-time difference value exceeds the historical range, and is lower than the historical minimum, the minimum historical adjustment factor is used directly; if it is higher than the historical maximum, the maximum historical adjustment factor is used, thus obtaining the initial execution interval shortening factor.

[0027] like Figure 2 As shown, Figure 2The flowchart of the data asset management system based on multi-source heterogeneous data fusion provided in this application embodiment obtains database access characteristic parameters, including concurrent IP connections, full table scan ratio, authentication failure count, concurrent query count, and connection request frequency. By analyzing these parameters, a database access robustness index is obtained, which is then compared with a preset security threshold to determine the database security status. If the database security status is secure, no anomaly handling is performed, maintaining data stability; if the database security status has potential risks, abnormal data is acquired and these multi-source heterogeneous data are marked. Simultaneously, the data asset security assessment execution interval is shortened based on the database access robustness index to increase the detection frequency. Data pollution load parameters of downstream data are obtained, and the comprehensive permission lock duration is calculated in conjunction with the database access robustness index to control downstream data access. A data recovery judgment instruction is generated based on the comprehensive permission lock duration. If the judgment status is recovery, recovery processing is performed on all derived data; if the judgment status is isolation, isolation processing is performed on all derived data. During the recovery process, the current device resource scheduling reserve is obtained, and the execution status of the full-domain derived data recovery processing is dynamically adjusted to ensure reasonable resource allocation and processing efficiency. Finally, a second adjustment is made to the data asset security assessment execution interval to ensure the security and integrity of the data assets, thereby completing the entire data asset management and protection process.

[0028] Furthermore, the database access robustness index is obtained through the following method: Database access characteristic parameters are obtained within a preset time period, including concurrent IP connections, full table scan ratio, authentication failure count, concurrent query count, and connection request frequency. A preset allowed set of access parameters is obtained from the database and standardized and normalized with these parameters to obtain a standardized value. Based on this standardized value, a corresponding weighting factor is introduced and coupled to obtain the database access robustness index. The allowed set of access parameters includes allowed values ​​for concurrent IP connections, full table scan ratio, authentication failure count, concurrent query count, and connection request frequency.

[0029] In this embodiment, the number of concurrent IP connections, the proportion of full table scans, the number of authentication failures, the number of concurrent queries, and the connection request frequency can be collected using the database's built-in performance monitoring tools (such as Oracle AWR reports) or the operating system's network monitoring tools (such as the netstat monitoring plugin).

[0030] IP concurrent connections represent the number of simultaneous connections from different IPs per unit of time, reflecting external access pressure and the concentration of potential abnormal requests; full table scan ratio represents the proportion of SQL statement execution involving full table scans, and an excessively high metric will increase the database I / O burden; authentication failure count measures the risk of the database encountering unauthorized access or brute-force attacks; concurrent query count represents the number of queries executed simultaneously per unit of time; connection request frequency describes the rate at which new connections are established per unit of time, reflecting the level of access activity and the intensity of instantaneous surges.

[0031] The database access robustness index is obtained using the following method:

[0032] ;

[0033] In the formula, WJ represents the database access robustness index, SL represents the number of concurrent IP connections, LJ represents the allowable number of concurrent IP connections, SM represents the proportion of full table scans, MJ represents the allowable proportion of full table scans, SC represents the number of authentication failures, CJ represents the allowable number of authentication failures, SQ represents the number of concurrent queries, QJ represents the allowable number of concurrent queries, SP represents the connection request frequency, PJ represents the allowable connection request frequency, a represents a real number, δ1 represents the weighting factor for the number of concurrent IP connections, δ2 represents the weighting factor for the proportion of full table scans, δ3 represents the weighting factor for the number of authentication failures, δ4 represents the weighting factor for the number of concurrent queries, and δ5 represents the weighting factor for the connection request frequency.

[0034] It should be noted that 'a' represents a very small real number to avoid calculation errors caused by the result being zero.

[0035] The weighting factors for concurrent IP connections, full table scan ratio, authentication failure count, concurrent query count, and connection request frequency can be obtained from the database. For example, historical concurrent IP connection data stored in the database can be retrieved and compared with the total concurrent IP connections. If a historical concurrent IP connection count is the same as the total concurrent IP connections, that historical concurrent IP connection count is used as the historical reference concurrent IP connection count. The weighting factor corresponding to that historical reference concurrent IP connection count is then used as the total concurrent IP connection count weighting factor. If there are several historical concurrent IP connection count weighting factors, their average is taken and used as the total concurrent IP connection count weighting factor. If no historical concurrent IP connection count is the same as the total concurrent IP connections, the historical concurrent IP connection count closest to that count is used as the historical reference concurrent IP connection count. Other weighting factors, such as the full table scan ratio, authentication failure count, concurrent query count, and connection request frequency, are obtained in a similar way to the IP concurrent connection count weighting factor, all by analyzing the corresponding historical reference values.

[0036] By analyzing database access characteristic parameters, including concurrent IP connections, full table scan ratio, authentication failure count, concurrent query count, and connection request frequency, a database access robustness index is derived. This index takes into account the interrelationships between these parameters. For example, an abnormally high connection request frequency directly increases the number of concurrent IP connections, leading to an increase in concurrent queries, which in turn increases the full table scan ratio, causing a decline in database performance. Conversely, if the number of authentication failures also increases when the number of concurrent queries rises, it usually indicates a batch of illegal connection attempts, further increasing the risk of database access failures. If the full table scan ratio increases abnormally, even if the connection request frequency remains stable, it will exacerbate the database load, reduce the concurrent query processing capacity, and cause a passive backlog of concurrent IP connections.

[0037] By acquiring and standardizing database access characteristic parameters, and then combining this with weighted factor coupling calculations, a database access robustness index is obtained, comprehensively reflecting the stability and security of database access. This robustness index can identify abnormal access patterns, potential risk points, and access pressure overload situations, providing a reliable quantitative basis for subsequent data asset security assessments and abnormal data identification. It ensures that data is not affected by unauthorized or abnormal access during sharing, reuse, and derivation, contributing to the integrity, authenticity, and security of research databases. Simultaneously, it provides actionable data support for dynamically adjusting the frequency of security assessments and access control strategies, thereby effectively improving the refinement and controllability of overall data asset management.

[0038] Furthermore, the database security status is obtained by: obtaining the preset database access security threshold in the database and comparing it with the database access robustness index to obtain the database security status. If the database access robustness index is above the database access security threshold, the database security status is secure; otherwise, the database security status is that there is a potential risk.

[0039] In this embodiment, the database security status is obtained by comparing the database access robustness index with a preset database access security threshold. This allows for real-time quantification of the overall database security level and a clear assessment of any potential risks. Analyzing the database security status enables scientific decisions on whether to initiate abnormal data identification and processing procedures, adjust data access permissions, or shorten the data asset security assessment execution interval. This ensures timely response when the database is maliciously or abnormally accessed, preventing the propagation of potentially abnormal data to downstream data. It provides a basis for decision-making in subsequent data recovery, downstream data updates, and comprehensive data traceability, enabling dynamic security management of multi-source heterogeneous scientific research data and effectively guaranteeing the integrity of the scientific research database and the reliability of data during data sharing and reuse.

[0040] Furthermore, abnormal data is obtained through the following method: Based on database security status analysis, if the database security status is secure, hash value judgment is not performed; if the database security status indicates potential risks, the hash values ​​of each multi-source heterogeneous data in the database are traversed and compared with preset hash values ​​to obtain a data asset security judgment result. If the hash values ​​match, the data asset security judgment result is secure; if the hash values ​​do not match, the data asset security judgment result is insecure, and the multi-source heterogeneous data is marked as abnormal data.

[0041] In this embodiment, by traversing the hash values ​​of various heterogeneous data from multiple sources within the database and comparing them with preset hash values, the integrity and consistency of each data entry can be accurately determined. This allows for the rapid identification of tampered or corrupted data, thereby marking abnormal data. Hash value verification, which does not rely on the format or type of the data content itself, achieves efficient and unambiguous security determination, ensuring the authenticity and integrity of data assets. Once abnormal data is identified, its transmission downstream can be promptly blocked, preventing erroneous or tampered data from contaminating downstream datasets and achieving dynamic security protection for the research database.

[0042] Furthermore, the comprehensive duration of access control is obtained through the following methods: First, data pollution load parameters of downstream data are obtained, including the amount of referenced data, the reference ratio, and the reference frequency. Second, a preset data pollution impact benchmark set is obtained from the database and standardized and normalized with the data pollution load parameters to obtain a standardized normalized value. Third, based on the standardized normalized value, a corresponding weighting factor is introduced for coupling processing to obtain the data pollution load value. Fourth, the initial duration of access control is obtained by matching the data pollution load value with the database. Fifth, the access control impact coefficient is obtained by matching the database access robustness index with the database. Sixth, the comprehensive duration of access control is obtained by multiplicatively coupling the initial duration of access control with the access control impact coefficient. The data pollution impact benchmark set includes the referenced data amount benchmark value, the reference ratio benchmark value, and the reference frequency benchmark value.

[0043] In this embodiment, the initial duration of permission lock is obtained by matching the data pollution load value with the database. Specifically, the method involves: first, retrieving historical records from the database, including the data pollution load value at each historical time point and the corresponding initial duration of permission lock, forming historical data pairs. For the real-time measured data pollution load value, the two closest historical difference values ​​are first found in the historical data, denoted as the lower and higher values, respectively, and the corresponding historical initial durations of permission lock are obtained. Then, using a linear interpolation method, the real-time difference value is proportionally mapped between these two historical values ​​to calculate the corresponding real-time initial duration of permission lock. Specifically, the real-time adjustment coefficient is equal to the lower historical adjustment coefficient plus the ratio of the real-time difference value to the historical low and high values ​​multiplied by the difference between the two historical adjustment coefficients. If the real-time difference value exceeds the historical range, if it is lower than the historical minimum, the minimum historical adjustment coefficient is directly used; if it is higher than the historical maximum, the maximum historical adjustment coefficient is used, thus obtaining the initial duration of permission lock.

[0044] The method for obtaining the access control impact coefficient by matching the database access robustness index with the database is as follows: First, historical records are retrieved from the database, including the database access robustness index and the corresponding access control impact coefficient at each historical time point, forming historical data pairs. For the real-time measured database access robustness index, the two closest historical difference values ​​are found in the historical data, denoted as the lower and higher values, respectively, and the historical access control impact coefficients corresponding to these two difference values ​​are obtained. Then, using a linear interpolation method, the real-time difference value is proportionally mapped between these two historical values ​​to calculate the corresponding real-time access control impact coefficient. Specifically, the real-time adjustment coefficient equals the lower historical adjustment coefficient plus the ratio of the real-time difference value to the historical low and high values ​​multiplied by the difference between the two historical adjustment coefficients. If the real-time difference value exceeds the historical range, if it is lower than the historical minimum value, the minimum historical adjustment coefficient is directly used; if it is higher than the historical maximum value, the maximum historical adjustment coefficient is used, thus obtaining the access control impact coefficient.

[0045] The overall duration of permission locking is obtained by multiplicatively coupling the initial duration of permission locking with the permission locking influence coefficient. Specifically, the overall duration of permission locking is obtained by multiplying the initial duration of permission locking with the permission locking influence coefficient and then adding the initial duration of permission locking.

[0046] Abnormal citation frequency indicates how frequently downstream data cites abnormal data, i.e., the number of times the data cites the abnormal data within a preset time period. Citation ratio indicates the proportion of abnormal data cited by downstream data out of all its cited data. The data pollution load parameters of downstream data (including the amount of cited data, citation ratio, and citation frequency) can all be obtained by querying system logs.

[0047] The specific method for obtaining the data pollution load value is as follows:

[0048] ;

[0049] In the formula, FH represents the data pollution load value, FM represents the amount of referenced data, MH represents the baseline value of the amount of referenced data, FL represents the reference ratio, LH represents the baseline value of the reference ratio, FV represents the reference frequency, and VH represents the baseline value of the reference frequency. This represents the weighting factor for the amount of referenced data. This indicates the weighting factor based on the citation ratio. This represents the citation frequency weighting factor.

[0050] The citation volume weighting factor, citation ratio weighting factor, and citation frequency weighting factor can be obtained from a database. For example, historical citation volume data stored in the database can be retrieved and compared with the citation volume. If a historical citation volume is identical to the citation volume, it is used as the historical reference citation volume, and its corresponding historical citation volume weighting factor is used as the citation volume weighting factor. If several historical citation volume weighting factors exist, they are averaged, and the average is used as the citation volume weighting factor. If no historical citation volume is identical to the citation volume, the historical citation volume closest to it is used as the historical reference citation volume. Other weighting factors, such as the citation ratio weighting factor and the citation frequency weighting factor, are obtained in a similar way to the citation volume weighting factor, by analyzing the corresponding historical reference values.

[0051] The data pollution load value is obtained by analyzing parameters including the amount of cited data, the citation ratio, and the citation frequency. This analysis considers the interrelationships between these parameters. For example, the amount of cited data represents the total amount of abnormal data directly reused or processed downstream. As the citation amount increases, the absolute degree of pollution in downstream data rises, directly affecting the citation ratio—that is, the proportion of abnormal data to the total downstream data increases, leading to a decrease in the overall reliability of downstream data. Changes in the citation ratio, however, can amplify or mitigate the risk impact of the amount of cited data: for the same amount of abnormal data, if the total downstream data volume is large and the citation ratio is low, the pollution impact is limited; conversely, a high ratio significantly increases the pollution risk. The citation frequency reflects the degree to which downstream data reuses abnormal data within a certain period. High-frequency citations accelerate the cumulative effect of data pollution, causing a multiplicative effect between the amount of cited data and the citation ratio, further exacerbating the unreliability of downstream data.

[0052] By quantifying the degree of contamination in downstream data, we can identify which data references anomalous data and carries a high risk of contamination. This allows us to rationally determine the initial duration of access control locks, avoiding resource waste or risk propagation caused by excessively long or short locks. Coupled with a database access robustness index, the access control lock duration can be dynamically adjusted based on the database's current access stability and security. This ensures stricter protection measures are implemented when the database is under heavy load or at high security risk, while resources are appropriately released when the database is secure and stable, improving overall system efficiency. Through precise calculation and dynamic adjustment of the combined access control lock duration, we can effectively prevent the propagation of anomalous data to downstream and derived data, reducing the risk of contamination of downstream derived data, ensuring the authenticity and reusability of research data. This approach considers both data contamination load and database security status while also taking into account system resource utilization efficiency, avoiding computational resource waste caused by excessive access control, and achieving a balance between security and performance.

[0053] Furthermore, a recovery determination instruction is generated. The specific method is as follows: obtain the preset data pollution impact threshold in the database and compare it with the data pollution load value. If the data pollution load value is above the data pollution impact threshold, the recovery determination instruction is in the isolation state. If the data pollution load value is less than the data pollution impact threshold, the recovery determination instruction is in the recovery state.

[0054] In this embodiment, by comparing the quantified data contamination load value with a threshold, it is possible to accurately determine whether all derived data needs to be restored or isolated, avoiding blind restoration or isolation and ensuring that processing decisions are scientific and reliable. When the data contamination load value exceeds the threshold, downstream data is placed in an isolated state, which can promptly block the continued propagation of abnormal data, thereby preventing downstream derived data from being continuously contaminated and improving the overall data credibility. Based on the real-time calculated data contamination load value and the preset threshold, the processing method can be flexibly selected according to the actual degree of data contamination, achieving dynamic response to data anomalies. By performing restoration operations only on downstream data that truly needs to be restored, while isolating severely contaminated data, resource waste can be avoided, and the utilization efficiency of data recovery and computing resources can be improved.

[0055] Furthermore, the derivation data across the entire domain is restored using the following method: The derivation data is sorted in ascending order based on its pollution load value to obtain the data restoration sorting order; the current resource scheduling capacity of the device is obtained, and the single concurrent processing capacity for data restoration is obtained by matching the resource scheduling capacity with the database; the derivation data is then restored according to the data restoration sorting order based on the single concurrent processing capacity; the derivation level of the derivation data is obtained and compared with a preset derivation level threshold in the database. If the derivation level of any derivation data exceeds the derivation level threshold, the derivation data is isolated; otherwise, the derivation data is restored.

[0056] In this embodiment, the current resource scheduling capacity of the device can be obtained by querying the device's resource manager.

[0057] The method for obtaining the single-time concurrent processing capacity for data recovery based on resource scheduling margin and database matching is as follows: First, historical records are retrieved from the database, including the resource scheduling margin and the corresponding single-time concurrent processing capacity for data recovery at each historical time point, forming historical data pairs. For the real-time measured resource scheduling margin, the two closest historical difference values ​​are found in the historical data, denoted as the lower and higher values, respectively, and the corresponding historical single-time concurrent processing capacity for data recovery is obtained. Then, a linear interpolation method is used to proportionally map the real-time difference value between these two historical values ​​to calculate the corresponding real-time single-time concurrent processing capacity for data recovery. Specifically, the real-time adjustment coefficient equals the lower historical adjustment coefficient plus the ratio of the real-time difference value to the historical low and high values ​​multiplied by the difference between the two historical adjustment coefficients. If the real-time difference value exceeds the historical range, if it is lower than the historical minimum, the minimum historical adjustment coefficient is directly used; if it is higher than the historical maximum, the maximum historical adjustment coefficient is used, thus obtaining the single-time concurrent processing capacity for data recovery.

[0058] By sorting the data contamination load values ​​of all derived data in ascending order, data with lower levels of contamination can be recovered first, ensuring minimal data loss and downstream propagation risks, and improving the overall efficiency of data recovery. If the data contamination load value of a certain derived data is too high, it indicates that the data references too much erroneous data, thus its data security and authenticity are poor, and its recovery is more difficult and resource-intensive. Therefore, isolating it can prevent other data from referencing this data again, which would reduce the authenticity of subsequent data, and also avoid the waste caused by data recovery of this derived data.

[0059] By determining the single concurrent processing volume based on the current resource scheduling margin of the equipment, computing and storage resources can be reasonably allocated while ensuring database access security. This avoids database performance degradation or further data attacks due to excessive resource consumption. At the same time, by detecting the number of derivative levels of the entire domain's derived data and comparing it with a preset threshold, data with excessively high derivative levels can be isolated and processed, cutting off the spread of abnormal data in multi-level derivative chains from the source and preventing further spread of pollution.

[0060] Data recovery is performed according to the sorting order and concurrent processing volume, which can maximize the recovery throughput under limited resources, achieve rapid processing and optimize data asset management. By combining resource scheduling margin with data recovery strategy, it can ensure that the database remains in a safe state during the recovery process and ensure that downstream data recovery is completed in priority, thereby improving the consistency and reliability of data assets.

[0061] Furthermore, the execution status of the full-domain derived data recovery process is adjusted as follows: The current resource scheduling reserve of the device during the full-domain derived data recovery process is obtained, and the difference between this reserve and a preset resource scheduling reserve threshold is analyzed to obtain a resource reserve difference value; based on the resource reserve difference value, it is matched with the database to obtain a concurrency adjustment coefficient; based on the concurrency adjustment coefficient and the single-time concurrency processing volume of data recovery, a multiplicative coupling process is performed to obtain the concurrency adjustment amount; based on the preset resource scheduling reserve threshold in the database, it is compared with the resource scheduling reserve. If the resource scheduling reserve is greater than the resource scheduling reserve threshold, the execution status of the full-domain derived data recovery process is adjusted to execute concurrency increase processing; if the resource scheduling reserve is less than the threshold, the execution status is adjusted to execute concurrency increase processing. If the resource scheduling margin is equal to the threshold, the global derived data recovery processing status will not be adjusted. If the resource scheduling margin is less than the threshold, the execution status of the global derived data recovery processing will be adjusted to execute concurrency reduction processing. If the resource scheduling margin is still less than the threshold after executing concurrency reduction processing, concurrency reduction processing will continue until the global derived data recovery processing is stopped. Concurrency increase processing will be executed by coupling the single concurrency processing volume of data recovery with the concurrency adjustment volume to obtain the concurrency increase adjustment execution volume. Concurrency decrease processing will be executed by reducing the single concurrency processing volume of data recovery based on the concurrency adjustment volume to obtain the concurrency decrease adjustment execution volume.

[0062] In this embodiment, the resource balance difference value is obtained by performing absolute difference processing (the absolute value of the difference) on the current resource scheduling balance of the device and the preset resource scheduling balance threshold to obtain the absolute difference value, and then dividing the absolute difference value by the preset resource scheduling balance threshold to obtain the resource balance difference value.

[0063] The concurrency adjustment coefficient is obtained by matching the resource balance difference value with the database. Specifically, the method involves: first, retrieving historical data from the database, including the resource balance difference value and its corresponding concurrency adjustment coefficient at each historical time point, forming historical data pairs. For the real-time measured resource balance difference value, the two closest historical difference values ​​are found in the historical data, designated as the lower and higher values, respectively, and their corresponding historical concurrency adjustment coefficients are obtained. Then, using linear interpolation, the real-time difference value is proportionally mapped between these two historical values ​​to calculate the corresponding real-time concurrency adjustment coefficient. Specifically, the real-time adjustment coefficient equals the lower historical adjustment coefficient plus the ratio of the real-time difference value to the historical low and high values ​​multiplied by the difference between the two historical adjustment coefficients. If the real-time difference value exceeds the historical range, if it is lower than the historical minimum, the minimum historical adjustment coefficient is used; if it is higher than the historical maximum, the maximum historical adjustment coefficient is used, thus obtaining the concurrency adjustment coefficient.

[0064] By acquiring the current resource scheduling margin of the device and performing differential analysis with a threshold, the availability of system resources can be determined in real time. This allows data recovery processing to be fully utilized when resources are sufficient, avoiding resource waste or bottlenecks. When the resource scheduling margin exceeds the threshold, concurrent processing is executed to increase the number of parallel processes in data recovery, thereby accelerating the recovery speed of all derived data, improving the overall recovery throughput, and shortening the time that abnormal data contaminates downstream derived data.

[0065] By employing a concurrency adjustment mechanism, the concurrent processing volume is reduced when resources are strained, preventing database performance degradation or potential data attacks due to excessive consumption of computing or storage resources, thus ensuring the database's safety and reliability during recovery. The dynamic increase or decrease of concurrency can be adjusted multiple times based on changes in resource availability, achieving a balance between resource utilization and recovery efficiency, ensuring continuous, efficient, and secure data recovery operations under varying load conditions.

[0066] Figure 3This application provides a flowchart for the data recovery process of a data asset management system based on multi-source heterogeneous data fusion. The flowchart describes the process of obtaining the data pollution load value of all derived data and sorting it in ascending order to determine the data recovery processing sequence. This ensures that less polluted data is recovered first, improving processing efficiency. The flowchart also describes obtaining the current resource scheduling margin of the device and calculating the single concurrent processing volume based on this margin to control the occupation of device resources during recovery. During recovery, the flowchart describes obtaining the derivation level of all derived data and comparing it with a preset derivation level threshold. If the derivation level of some data exceeds the threshold, it is isolated to prevent further pollution; otherwise, recovery processing continues. The flowchart describes calculating the resource margin difference value and obtaining a concurrent processing adjustment coefficient to calculate the final concurrent processing adjustment amount. If the current device resource margin is greater than the threshold, concurrent processing is performed to increase the concurrency and improve recovery efficiency; if the resource margin is equal to the threshold, the current execution state is maintained; if the resource margin is less than the threshold, concurrent processing is performed, continuing to decrease the concurrency as needed until safe stopping to prevent system resource overload.

[0067] Furthermore, a second adjustment is made to the execution interval of the data asset security assessment. The specific method is as follows: the average data pollution load value based on the data pollution load value of the whole domain derived data is processed to obtain the average data pollution load value; the preset average data pollution load threshold in the database is obtained and compared with the average data pollution load value to obtain the data asset assessment qualification level. If the average data pollution load value is above the average data pollution load threshold, the data asset assessment qualification level is that the assessment interval is unqualified, and the execution interval of the data asset security assessment is shortened. If the average data pollution load value is less than the average data pollution load threshold, the data asset assessment qualification level is that the assessment interval is qualified, and the second adjustment of the data asset security assessment execution interval is not performed.

[0068] In this embodiment, the average data pollution load value based on the global derived data is processed and compared with a threshold. When the average data pollution load value exceeds the threshold, the execution interval of data asset security assessment can be automatically shortened, enabling more frequent security checks and timely detection of potential attacks or abnormal data in the database. The higher the data pollution load value of the global derived data, the higher the overall degree of pollution of the database and its downstream derived data, which may indicate long-term malicious tampering or attacks. Adjusting the execution interval a second time ensures that the system performs multiple checks, preventing abnormal data from continuously spreading downstream and reducing the risk of continuous data asset pollution.

[0069] Compared to the initial adjustment, the secondary adjustment can dynamically update the assessment interval based on the latest data contamination load, making the security strategy more flexible and intelligent. This ensures that security assessments are not delayed when the database access robustness index declines or abnormal data increases, thereby improving the overall protection effectiveness. By shortening the execution interval, multiple rounds of security assessments can quickly identify abnormal data and take recovery or isolation measures, reducing the possibility of downstream derived data being contaminated, thus maintaining the authenticity, reliability, and reusability of research databases and derived data.

[0070] Furthermore, the execution interval of the data asset security assessment is shortened. Specifically, the following method is used: The data pollution load value of the entire domain's derived data is retrieved and averaged to obtain the average pollution load value; the average pollution load value is matched with the database to obtain the first interval adjustment coefficient; the data derivation levels of the entire domain's derived data are retrieved and averaged to obtain the average data derivation level; the average data derivation level is matched with the database to obtain the second interval adjustment coefficient; the execution interval of the data asset security assessment is multiplicatively coupled with the first and second interval adjustment coefficients to obtain a shortened execution time for the data asset security assessment, thereby shortening the execution interval of the data asset security assessment.

[0071] In this embodiment, the first interval adjustment coefficient is obtained by matching the average pollution load with the database. Specifically, the method involves first retrieving historical data from the database, including the average pollution load at each historical time point and the corresponding first interval adjustment coefficient, forming historical data pairs. For the real-time measured average pollution load, the two closest historical difference values ​​are first found in the historical data, denoted as the lower and higher values, and the corresponding historical first interval adjustment coefficients are obtained. Then, using linear interpolation, the real-time difference value is proportionally mapped between these two historical values ​​to calculate the corresponding real-time first interval adjustment coefficient. Specifically, the real-time adjustment coefficient equals the lower historical adjustment coefficient plus the ratio of the real-time difference value to the historical low and high values ​​multiplied by the difference between the two historical adjustment coefficients. If the real-time difference value exceeds the historical range, if it is lower than the historical minimum, the minimum historical adjustment coefficient is directly used; if it is higher than the historical maximum, the maximum historical adjustment coefficient is used, thus obtaining the first interval adjustment coefficient.

[0072] The second interval adjustment coefficient is obtained by matching the data-derived average level with the database. Specifically, the method involves first retrieving historical data from the database, including the data-derived average level and the corresponding second interval adjustment coefficient for each historical time point, forming historical data pairs. For the real-time measured data-derived average level, the two closest historical difference values ​​are found in the historical data, designated as the lower and higher values, respectively, and the corresponding historical second interval adjustment coefficients are obtained. Then, using linear interpolation, the real-time difference value is proportionally mapped between these two historical values ​​to calculate the corresponding real-time second interval adjustment coefficient. Specifically, the real-time adjustment coefficient equals the lower historical adjustment coefficient plus the ratio of the real-time difference value to the historical low and high values ​​multiplied by the difference between the two historical adjustment coefficients. If the real-time difference value exceeds the historical range, if it is lower than the historical minimum, the minimum historical adjustment coefficient is used; if it is higher than the historical maximum, the maximum historical adjustment coefficient is used, thus obtaining the second interval adjustment coefficient.

[0073] To shorten the execution time of data asset security assessments, the specific methods are as follows: In the formula, TH represents the shortened execution time of the data asset security assessment, HC represents the initial execution interval of the data asset security assessment, ra represents the first coefficient for interval adjustment, and rb represents the second coefficient for interval adjustment.

[0074] By retrieving and averaging the data contamination load values ​​of all derived data, the overall contamination level of the current data can be quantified. This allows for the determination of the first adjustment coefficient based on the actual contamination load, dynamically shortening the security assessment execution interval and ensuring more frequent checks when data risk is high. Simultaneously, retrieving and averaging the data derivation levels of all derived data yields a second adjustment coefficient, reflecting the diffusion and potential contamination impact of data in multi-level derivation processes. This further shortens the security assessment interval for cases with high data derivation complexity, effectively controlling the spread of potential risks. Multiplicatively coupling the average contamination load and the average number of data derivation levels with the original execution interval results in a shortened execution time, achieving comprehensive multi-factor driven security assessment adjustment. This enables the system to respond more accurately to database security status and data contamination risks. Frequent security assessments and dynamic interval adjustments allow for the timely detection of abnormal data and potential attack behaviors, enabling rapid recovery or isolation measures to prevent continuous contamination of downstream derived data and ensuring the integrity and reliability of the research database and its derived data. While shortening the execution interval, the coupled computing method can make reasonable use of system resources, balance the frequency of security monitoring and the computing load, and ensure that the database can still run efficiently in high-risk environments without affecting the overall data processing efficiency.

[0075] In summary, this embodiment performs periodic security assessments on multi-source heterogeneous data within the database and analyzes database access characteristic parameters in real time to calculate a database access security index. This allows for dynamic judgment of the database access status. If potential risks are detected in access behavior, the hash values ​​of each multi-source heterogeneous data are verified one by one. This enables rapid identification of tampered data at the data source level, thereby blocking the continuous transmission of maliciously tampered data from the source. This effectively solves the problem of malicious attacks on scientific research databases causing malicious tampering of the original data, which is then continuously transmitted during data sharing and reuse, leading to continuous contamination of derived data from downstream sources.

[0076] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0077] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0078] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0079] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0080] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0081] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A data asset management system based on multi-source heterogeneous data fusion, characterized in that, include: The database security assessment module is used to periodically conduct data asset security assessments on multi-source heterogeneous data in the database, obtain database access characteristic parameters, analyze and obtain the database access robustness index, and thus obtain the database security status. The data asset protection and control module is used to obtain the data asset protection level based on the database security status analysis. If the data asset protection level is secure, the execution interval of the data asset security assessment will not be shortened. Otherwise, abnormal data is obtained, and the execution interval of the data asset security assessment will be shortened based on the database access robustness index. The data recovery module is used to obtain the data pollution load parameters of downstream data and, in conjunction with the database access robustness index, analyze the comprehensive duration of permission lockout, thereby performing abnormal data recovery and downstream data update processing. The downstream data refers to the dataset generated by directly reusing or processing abnormal data. The full-domain data source tracing module is used to obtain the data pollution load value of the full-domain derived data, thereby generating a recovery judgment instruction. If the recovery judgment instruction is in a recovery state, the full-domain derived data is restored; otherwise, the full-domain derived data is isolated. The full-domain derived data includes each downstream derived data and each downstream derived data's multi-level derived data. The downstream derived data represents the dataset generated by directly reusing or processing the downstream data. The computing resource management module is used to obtain the current resource scheduling reserve of the device during the global derived data recovery process, and thereby adjust the execution status of the global derived data recovery process. The data asset assessment secondary adjustment module is used to adjust the execution interval of data asset security assessment based on the data pollution load value of the whole domain derived data. The specific method for obtaining the database access robustness index is as follows: The database access characteristic parameters include the number of concurrent IP connections, the proportion of full table scans, the number of authentication failures, the number of concurrent queries, and the connection request frequency; Obtain the preset allowed access set in the database and perform standardized normalization processing with the database access feature parameters to obtain the standardized normalization processing value. Based on the standardized normalization processing value, introduce the corresponding weighting factor for coupling processing to obtain the database access robustness index. The database access robustness index is used to reflect the stability of the database access within the current running cycle. The allowed access set includes the allowed value for concurrent IP connections, the allowed value for the proportion of full table scans, the allowed number of authentication failures, the allowed number of concurrent queries, and the allowed value for connection request frequency; To obtain the total lock duration, the specific method is as follows: The data pollution load parameters include the amount of referenced data, the reference ratio, and the reference frequency; Obtain the preset data pollution impact benchmark set in the database, and perform standardized normalization processing with the data pollution load parameters to obtain the standardized normalization processing value. Based on the standardized normalization processing value, introduce the corresponding weighting factor for coupling processing to obtain the data pollution load value. The data pollution load value is used to reflect the degree of pollution risk carried by the data due to the reuse of abnormal data. The initial duration of access control is obtained by matching the data pollution load value with the database. The database is matched based on the database access robustness index to obtain a permission lock influence coefficient; The permission lock comprehensive duration is obtained by multiplicative coupling processing based on the permission lock initial duration and the permission lock influence coefficient; The data pollution influence benchmark set includes a reference data volume benchmark value, a reference proportion benchmark value, and a reference frequency benchmark value.

2. The data asset management system based on multi-source heterogeneous data fusion according to claim 1, characterized in that: The database security state is obtained by comparing the database access robustness index with a preset database access security threshold value in the database. The database security state is obtained by comparing the database access robustness index with a preset database access security threshold value in the database.

3. The data asset management system based on multi-source heterogeneous data fusion according to claim 1, characterized in that: The abnormal data is obtained by comparing the hash value of each multi-source heterogeneous data in the database with a preset hash value. The hash value of each multi-source heterogeneous data in the database is compared with a preset hash value to obtain a data asset security determination result. The abnormal data is obtained by comparing the hash value of each multi-source heterogeneous data in the database with a preset hash value. The recovery determination instruction is obtained by comparing the data pollution load value with a preset data pollution influence threshold value in the database.

4. The data asset management system based on multi-source heterogeneous data fusion of claim 1, wherein: The data recovery sorting order is obtained by ascendingly sorting the data pollution load values of the global derivative data. The data recovery single-concurrent processing amount is obtained by matching the resource scheduling margin of the current device with the database.

5. The data asset management system based on multi-source heterogeneous data fusion of claim 1, wherein: The derivative series of the global derivative data is compared with a preset derivative series threshold value in the database. The execution state of the global derivative data recovery processing is adjusted by comparing the resource scheduling margin of the current device in the global derivative data recovery process with a preset resource scheduling margin threshold value. The resource scheduling margin of the current device in the global derivative data recovery process is compared with a preset resource scheduling margin threshold value to obtain a resource margin difference value. The concurrent processing adjustment coefficient is obtained by matching the resource margin difference value with the database.

6. The data asset management system based on multi-source heterogeneous data fusion of claim 1, wherein: The concurrent processing adjustment amount is obtained by multiplicative coupling processing based on the concurrent processing adjustment coefficient and the data recovery single-concurrent processing amount. The execution state of the global derivative data recovery processing is adjusted by comparing the resource scheduling margin of the current device in the global derivative data recovery process with a preset resource scheduling margin threshold value. ​ ​ ​ If the resource scheduling margin is equal to the resource scheduling margin threshold, the global derivative data recovery processing state is not adjusted. If the resource scheduling margin is less than the resource scheduling margin threshold, the execution state of the global derivative data recovery processing is adjusted to perform the concurrency reduction processing, and if the resource scheduling margin is still less than the resource scheduling margin threshold after performing the concurrency reduction processing, the concurrency reduction processing is continuously performed until the recovery processing of the global derivative data is stopped. The execution concurrency growth processing is specifically as follows: The single concurrency processing amount of data recovery is coupled with the concurrency processing adjustment amount to obtain a concurrency growth adjustment execution amount. The execution concurrency reduction processing is specifically as follows: The single concurrency processing amount of data recovery is reduced based on the concurrency processing adjustment amount to obtain a concurrency reduction adjustment execution amount.

7. The data asset management system based on multi-source heterogeneous data fusion of claim 1, wherein: The secondary adjustment of the data asset security assessment execution interval is specifically as follows: The data pollution load values of the global derivative data are processed by mean value processing to obtain a data pollution average load value. A preset data pollution average load threshold in the database is obtained, and is compared with the data pollution average load value to obtain a data asset assessment qualified judgment level. If the data pollution average load value is above the data pollution average load threshold, the data asset assessment qualified judgment level is an assessment interval unqualified, and the data asset security assessment execution interval is shortened. If the data pollution average load value is less than the data pollution average load threshold, the data asset assessment qualified judgment level is an assessment interval qualified, and the secondary adjustment of the data asset security assessment execution interval is not performed.

8. The data asset management system based on multi-source heterogeneous data fusion according to claim 7, characterized in that: The shortening processing of the data asset security assessment execution interval is specifically as follows: The data pollution load values of the global derivative data are called and are processed by mean value processing to obtain a pollution load average value. The pollution load average value is matched with the database to obtain a first interval adjustment coefficient. The data derivative series of the global derivative data are called and are processed by mean value processing to obtain a data derivative average series. The data derivative average series is matched with the database to obtain a second interval adjustment coefficient. The data asset security assessment execution interval is multiplicatively coupled with the first interval adjustment coefficient and the second interval adjustment coefficient to obtain a data asset security assessment shortened execution time length, so that the shortening processing of the data asset security assessment execution interval is performed.

Citation Information

Patent Citations

  • Power planning multi-source data integration method

    CN117350448A

  • A method and system for evaluating power grid assets

    CN118761594B

  • Multi-source data intelligent evaluation system based on security situation

    CN114884735A

  • Security risk analysis method and system based on multi-source data acquisition

    CN116127522A