A backup management method of a data file and related device

By identifying sensitive information in data files and generating sensitive tags, and adjusting encryption strategies in conjunction with the status data of the encryption management system, the problems of resource waste and insufficient security in existing technologies are solved, and efficient, secure and reliable differentiated management of data backup is achieved.

CN121478557BActive Publication Date: 2026-05-29GUANGZHOU LANBO NETWORK TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU LANBO NETWORK TECHNOLOGY CO LTD
Filing Date
2025-11-11
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing data file backup management methods generally lack sufficient intelligence, resulting in the application of the same security protection measures to data files of different values ​​and sensitivities, leading to resource waste or insufficient security.

Method used

By identifying sensitive information based on data timeliness and contextual information, generating sensitive tags, and embedding them into data files using non-visible methods, and adjusting encryption strategies in conjunction with the status data of the encryption management system, differentiated backup management can be achieved.

Benefits of technology

It improves the efficiency, security, and reliability of data backup, optimizes system resource utilization, and ensures differentiated protection for different data files.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121478557B_ABST
    Figure CN121478557B_ABST
Patent Text Reader

Abstract

The application discloses a kind of backup management method of data file and related device, it is related to data processing technical field, the method includes: based on data timeliness feature and context information, the sensitive information of original data stream is identified, determine its sensitive label;Based on non-visible method, the embedding of sensitive label is generated to data file;Read the sensitive label analysis information of file to determine its security level, determine initial encryption strategy based on security level;Based on the current state data and performance baseline of system, adjust initial encryption strategy to encrypt the data file;Integrity check is carried out to the data file to determine several copy files, store file to corresponding backup path based on security level;The synchronization state and health state of each copy file after storage are monitored to carry out integrity verification, copy repair and sensitive label re-embedding.The application intelligently adopts differentiated security backup strategy, improves the security and reliability of data backup.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and apparatus for backup management of data files. Background Technology

[0002] In today's era of widespread information technology application, various business systems and services generate and process massive amounts of data files every day. These data files are the core assets for organizational operations and making important decisions, and their security, integrity, and ability to be retrieved when needed are crucial. Therefore, establishing an efficient and reliable data backup management system has become an important technical means to ensure uninterrupted business operations and respond to system failures or security threats. Traditional backup methods typically involve saving data in its entirety or periodically according to a fixed schedule. In modern application scenarios where data volumes are increasing rapidly and types are becoming more complex, this is a necessary foundation for achieving systematic management and protection of data assets.

[0003] However, existing data file backup management methods generally suffer from insufficient intelligence. Specifically, most systems tend to adopt a one-size-fits-all strategy during backup, applying the same encryption strength and storage location to data files of varying importance and sensitivity. This crude management approach leads to two main problems: firstly, applying excessively high security protection measures to non-critical or publicly available data wastes computing resources and storage space; secondly, core sensitive data may lack sufficient security protection or be not allocated to the required priority storage location, thereby increasing the risk of data leakage or corruption. Therefore, how to enable backup systems to intelligently identify the value of data and automatically adopt differentiated security backup strategies based on that value has become an urgent technical challenge to be solved. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the prior art. This invention provides a data file backup management method and related apparatus, which can intelligently adopt differentiated security backup strategies according to the actual value and sensitivity of the data, significantly improving the efficiency, security and reliability of data backup.

[0005] To address the aforementioned technical problems, this invention provides a data file backup management method, the method comprising:

[0006] Sensitive information in the original data stream is identified based on the timeliness characteristics and contextual information, and sensitive tags for the sensitive information are determined.

[0007] A data file is generated based on the original data stream, and sensitive tags are embedded in the data file using a non-visible method to obtain the processed data file.

[0008] Read the sensitive tag parsing information of the processed data file, determine the security level of the processed data file based on the sensitive tag parsing information, and determine the initial encryption strategy based on the security level;

[0009] The current status data and performance baseline of the encryption management system are determined, and the initial encryption strategy is adjusted based on the performance baseline and the current status data to obtain the target encryption strategy. The processed data file is then encrypted based on the target encryption strategy to obtain the encrypted data file.

[0010] The encrypted data file is subjected to integrity verification. Based on the integrity verification result, several copy files corresponding to the encrypted data file are determined. Based on the security level, the encrypted data file and several copy files are stored in the corresponding backup path.

[0011] Monitor the synchronization and health status of each stored copy file, and perform integrity verification, copy repair, and sensitive tag re-embedding on each copy file based on the synchronization and health status.

[0012] Optionally, the step of identifying sensitive information in the original data stream based on data timeliness characteristics and contextual information, and determining the sensitive tags of the sensitive information, includes:

[0013] Obtain the data timeliness characteristics and business context information of the original data stream, and determine the weight and sensitivity threshold of the sensitive information pattern matching rule based on the data timeliness characteristics and business context information;

[0014] Sensitive information in the original data stream is identified based on the weights and sensitivity thresholds of the sensitive information pattern matching rules.

[0015] Based on the timeliness characteristics of the data and the business context information, the sensitive information is analyzed for type and security level recommendations to obtain the corresponding sensitive information type and security level recommendation information, and sensitive tags are generated based on the sensitive information type and security level recommendation information.

[0016] Optionally, the process of embedding sensitive tags into the data file using a non-visible method to obtain the processed data file includes:

[0017] The sensitive tags are encrypted to obtain encrypted sensitive tags, and an integrity check code is generated based on the encrypted sensitive tags and the content information of the data file.

[0018] Determine the type of data file and the expected storage environment, and determine the tag embedding mechanism based on the type of data file and the expected storage environment;

[0019] The data file is processed by embedding an integrity check code and an encrypted sensitive tag using the tag embedding mechanism based on the non-visible method, thereby obtaining the processed data file.

[0020] Optionally, determining the current state data and performance baseline of the encryption management system, and adjusting the initial encryption strategy based on the performance baseline and current state data to obtain the target encryption strategy, includes:

[0021] The current status data of the encryption management system is obtained based on the system performance probe, and the encryption management system is simulated for file encryption based on a preset time period to obtain encryption simulation data. The performance baseline is determined based on the encryption simulation data.

[0022] The current state data is compared with the performance baseline to obtain the comparison result. Based on the comparison result, the initial encryption strategy is adjusted in terms of concurrency, data block size, hardware security module interaction, and process priority to obtain the target encryption strategy.

[0023] Optionally, the monitoring of the synchronization status and health status of each stored copy file, and the performance of integrity verification, copy repair, and sensitive tag re-embedding on each copy file based on the synchronization status and health status, includes:

[0024] Based on metadata services, monitor the synchronization status and health status of each copy file after storage;

[0025] The verification strategy for the copy file is determined based on the synchronization status and health status.

[0026] Based on the verification strategy, integrity verification is performed on each copy file to obtain the corresponding integrity verification result;

[0027] The overall integrity verification result is determined based on the majority consensus principle and the integrity verification results of each copy file.

[0028] If the overall integrity verification result is that the overall integrity verification fails, then abnormal copy files are identified and repaired based on the distributed repair process;

[0029] After the abnormal copy file is repaired, the sensitive tags are re-embedded in the repaired abnormal copy file.

[0030] Optionally, the monitoring of the synchronization status and health status of each stored copy file based on metadata services includes:

[0031] Obtain the geographical distribution of the stored copies, the data update frequency, and the network bandwidth of each copy file after storage;

[0032] The initial monitoring frequency of metadata services is adjusted based on the storage geographic distribution, data update frequency, and network bandwidth to obtain the target monitoring frequency.

[0033] The metadata service monitors the synchronization status and health status of each copy file based on the target monitoring frequency.

[0034] Optionally, adjusting the initial monitoring frequency of the metadata service based on the storage geographic distribution, data update frequency, and network bandwidth to obtain the target monitoring frequency includes:

[0035] The historical storage geographic distribution, historical data update frequency, and historical network bandwidth are obtained, and parameter fluctuation analysis is performed on the historical storage geographic distribution, historical data update frequency, and historical network bandwidth to obtain the corresponding parameter fluctuation data.

[0036] Based on the storage geographic distribution, data update frequency, and network bandwidth, combined with the corresponding parameter fluctuation data, a significant deviation analysis is performed to obtain significant deviation information. Based on the significant deviation information, the initial monitoring frequency is incrementally adjusted to obtain the incrementally adjusted initial monitoring frequency.

[0037] Obtain the data monitoring effect of the initial monitoring frequency after incremental adjustment, and determine the adaptive adjustment factor based on the data monitoring effect;

[0038] The initial monitoring frequency after incremental adjustment is readjusted based on the adaptive adjustment factor to obtain the target monitoring frequency.

[0039] In addition, the present invention also provides a data file backup management device, the device comprising:

[0040] Tag determination module: used to identify sensitive information in the original data stream based on data timeliness characteristics and context information, and to determine the sensitive tags of the sensitive information;

[0041] Tag embedding module: used to generate a data file based on the original data stream, and to embed sensitive tags into the data file using a non-visible method to obtain the processed data file;

[0042] Initial strategy analysis module: used to read the sensitive tag parsing information of the processed data file, determine the security level of the processed data file based on the sensitive tag parsing information, and determine the initial encryption strategy based on the security level;

[0043] File encryption module: used to determine the current status data and performance baseline of the encryption management system, and adjust the initial encryption strategy based on the performance baseline and current status data to obtain the target encryption strategy, and encrypt the processed data file based on the target encryption strategy to obtain the encrypted data file;

[0044] File backup module: used to perform integrity verification on the encrypted data file, determine several copy files corresponding to the encrypted data file based on the integrity verification result, and store the encrypted data file and several copy files to the corresponding backup path based on the security level;

[0045] The copy file monitoring module is used to monitor the synchronization status and health status of each stored copy file, and to perform integrity verification, copy repair, and sensitive tag re-embedding on each copy file based on the synchronization status and health status.

[0046] In addition, the present invention also provides an electronic device, which includes a processor and a memory. The memory is used to store instructions, and the processor is used to call the instructions in the memory to cause the electronic device to execute the above-described data file backup management method.

[0047] In addition, the present invention also provides a computer-readable storage medium that stores computer instructions, which, when executed on an electronic device, cause the electronic device to perform the above-described data file backup management method.

[0048] In this embodiment of the invention, sensitive information in the original data stream is identified based on data timeliness characteristics and contextual information, and sensitive tags for the sensitive information are determined, resulting in more refined sensitive tags. Data files are generated based on the original data stream, and sensitive tags are embedded into the data files using a non-visible method, tightly integrating the sensitive tags with the data files and improving their concealment and security. The sensitive tag parsing information of the processed data files is read to determine the security level of the processed data files, and an initial encryption strategy is determined based on the security level. The initial encryption strategy is adjusted based on the current state data and performance baseline of the encryption management system to encrypt the processed data files, ensuring that security requirements are met while optimizing system resource utilization and improving encryption efficiency. Integrity verification is performed on the encrypted data files to determine several corresponding copy files. Based on the security level, the encrypted data files and several copy files are stored in the corresponding backup paths. The synchronization and health status of each stored copy file are monitored to perform integrity verification, copy repair, and sensitive tag re-embedding on each copy file. This allows for intelligent adoption of differentiated security backup strategies based on the actual value and sensitivity of the data, significantly improving the efficiency, security, and reliability of data backup. Attached Figure Description

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0050] Figure 1 This is a flowchart illustrating the data file backup management method in an embodiment of the present invention;

[0051] Figure 2 This is a flowchart illustrating a data file backup management method according to another embodiment of the present invention;

[0052] Figure 3 This is a schematic diagram of the structural composition of the data file backup management device in an embodiment of the present invention;

[0053] Figure 4 This is a schematic diagram of the structural composition of the electronic device in an embodiment of the present invention. Detailed Implementation

[0054] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0055] Example 1

[0056] Please see Figure 1 , Figure 1 This is a flowchart illustrating a data file backup management method according to an embodiment of the present invention. The method includes:

[0057] S11: Identify sensitive information in the original data stream based on data timeliness characteristics and contextual information, and determine the sensitive tags of the sensitive information;

[0058] In the specific implementation of this invention, the timeliness characteristics and business context information of the original data stream are acquired, and the weights and sensitivity thresholds of the sensitive information pattern matching rules are determined based on these characteristics. Sensitive information in the original data stream is identified based on these weights and thresholds. Type analysis and security level recommendation analysis are performed on the sensitive information based on the timeliness characteristics and business context information to obtain corresponding sensitive information types and security level recommendation information. Sensitive tags are then generated based on these types and recommendation information, assigning more refined sensitive tags to the sensitive information. This allows for differentiated processing of data with different levels of sensitivity, avoiding the resource waste caused by applying the same protection measures to all data in traditional methods.

[0059] S12: Generate a data file based on the original data stream, and embed sensitive tags into the data file using a non-visible method to obtain the processed data file;

[0060] In the specific implementation of this invention, a data file is generated based on the original data stream, sensitive tags are encrypted to obtain encrypted sensitive tags, and an integrity check code is generated based on the encrypted sensitive tags and the content information of the data file; the type of the data file and the expected storage environment are determined, and a tag embedding mechanism is determined based on the type of the data file and the expected storage environment; the integrity check code and the encrypted sensitive tags are embedded in the data file using a non-visible method and the tag embedding mechanism, so that the sensitive tags are tightly combined with the data file and are not easily detected or tampered with by conventional means, thereby improving the concealment and security of the sensitive tags.

[0061] S13: Read the sensitive tag parsing information of the processed data file, determine the security level of the processed data file based on the sensitive tag parsing information, and determine the initial encryption strategy based on the security level;

[0062] In the specific implementation of this invention, the sensitive tag parsing information of the processed data file is read by the parser, the security level of the processed data file is determined based on the sensitive tag parsing information, and the initial encryption strategy is determined based on the security level. This can improve the efficiency of initial encryption strategy matching and avoid consuming too many resources.

[0063] S14: Determine the current state data and performance baseline of the encryption management system, and adjust the initial encryption strategy based on the performance baseline and current state data to obtain the target encryption strategy. Then, encrypt the processed data file based on the target encryption strategy to obtain the encrypted data file.

[0064] In the specific implementation of this invention, the current status data of the encryption management system is obtained based on the system performance probe, and file encryption simulation is performed on the encryption management system based on a preset time period to obtain encryption simulation data. The performance baseline is then determined based on the encryption simulation data. The current status data is compared with the performance baseline to obtain the comparison result. Based on the comparison result, the initial encryption strategy is adjusted in terms of concurrency, data block size, hardware security module interaction, and process priority to obtain the target encryption strategy. The processed data file is then encrypted based on the target encryption strategy, which ensures that while meeting security requirements, system resource utilization is optimized and encryption efficiency is improved.

[0065] S15: Perform integrity verification on the encrypted data file, determine several copy files corresponding to the encrypted data file based on the integrity verification result, and store the encrypted data file and several copy files to the corresponding backup path based on the security level;

[0066] In the specific implementation of this invention, the encrypted data file is subjected to integrity verification. Based on the integrity verification result, several copy files corresponding to the encrypted data file are determined. Based on the security level, the encrypted data file and several copy files are stored in the corresponding backup path. Data files with different security levels are stored in different backup areas, which significantly improves the efficiency, security and reliability of data backup and provides strong technical support for the systematic management and protection of data assets.

[0067] S16: Monitor the synchronization status and health status of each stored copy file, and perform integrity verification, copy repair, and sensitive tag re-embedding on each copy file based on the synchronization status and health status.

[0068] In the specific implementation of this invention, the synchronization status and health status of each copy file after storage are monitored based on metadata services; the verification strategy for the copy files is determined based on the synchronization status and health status; the integrity of each copy file is verified based on the verification strategy to obtain the corresponding integrity verification result; the overall integrity verification result is determined based on the majority consensus principle combined with the integrity verification results of each copy file; if the overall integrity verification result is that the overall integrity verification fails, abnormal copy files are identified and repaired based on a distributed repair process; after the abnormal copy files are repaired, sensitive tags are re-embedded in the repaired abnormal copy files, introducing a continuous health monitoring and automatic repair mechanism for copy files, which greatly enhances the reliability of data backup and disaster recovery capabilities.

[0069] In this embodiment of the invention, sensitive information in the original data stream is identified based on data timeliness characteristics and contextual information, and sensitive tags for the sensitive information are determined, resulting in more refined sensitive tags. Data files are generated based on the original data stream, and sensitive tags are embedded into the data files using a non-visible method, tightly integrating the sensitive tags with the data files and improving their concealment and security. The sensitive tag parsing information of the processed data files is read to determine the security level of the processed data files, and an initial encryption strategy is determined based on the security level. The initial encryption strategy is adjusted based on the current state data and performance baseline of the encryption management system to encrypt the processed data files, ensuring that security requirements are met while optimizing system resource utilization and improving encryption efficiency. Integrity verification is performed on the encrypted data files to determine several corresponding copy files. Based on the security level, the encrypted data files and several copy files are stored in the corresponding backup paths. The synchronization and health status of each stored copy file are monitored to perform integrity verification, copy repair, and sensitive tag re-embedding on each copy file. This allows for intelligent adoption of differentiated security backup strategies based on the actual value and sensitivity of the data, significantly improving the efficiency, security, and reliability of data backup.

[0070] Example 2

[0071] Please see Figure 2 , Figure 2 This is a flowchart illustrating a data file backup management method according to another embodiment of the present invention, the method comprising:

[0072] S201: Identify sensitive information in the original data stream based on data timeliness characteristics and contextual information, and determine the sensitive tags of the sensitive information;

[0073] In a specific implementation of this invention, the step of identifying sensitive information in the original data stream based on data timeliness characteristics and context information, and determining the sensitive labels of the sensitive information, includes: acquiring data timeliness characteristics and business context information of the original data stream, and determining the weight and sensitivity threshold of the sensitive information pattern matching rule based on the data timeliness characteristics and business context information; identifying sensitive information in the original data stream based on the weight and sensitivity threshold of the sensitive information pattern matching rule; performing type analysis and security level recommendation analysis on the sensitive information based on the data timeliness characteristics and business context information to obtain the corresponding sensitive information type and security level recommendation information, and generating sensitive labels based on the sensitive information type and security level recommendation information.

[0074] Specifically, raw data streams refer to unprocessed sequences of raw data generated, transmitted, or received in real time by a system or application, such as network traffic, sensor data, and user input. Obtaining the timeliness characteristics and business context information of raw data streams involves real-time or near-real-time analysis of the raw data streams entering the system to extract their time-related attributes and business-related background information. Data timeliness characteristics can include the data's generation time, last modification time, expected lifespan, and the trend of its importance or sensitivity at different points in time. For example, financial transaction data may have extremely high timeliness and sensitivity within minutes of a transaction, but its sensitivity may gradually decrease over time. Business context information covers the business process in which the data is situated, the department to which it belongs, the data's purpose, access permissions, compliance requirements, and its correlation with other data. For example, a customer's personal identification information may have a different business context when used for identity verification than when used for market analysis. Based on the timeliness characteristics of the data and the business context information, the weights and sensitivity thresholds of the sensitive information pattern matching rules are determined. The weight of the sensitive information pattern matching rule refers to the proportion of importance of different matching rules when identifying sensitive information. Matching rules include regular expressions, keyword lists, machine learning models, etc., and this weight will be dynamically adjusted according to the timeliness of the data and the business context. The sensitivity threshold is the critical value used to determine whether information constitutes sensitive information, and it will also be adaptively adjusted according to the current timeliness characteristics of the data and the business context information.

[0075] Identifying sensitive information in the original data stream based on the weights and sensitivity thresholds of sensitive information pattern matching rules refers to scanning and analyzing the original data stream using dynamically adjusted weights and thresholds. For example, pattern matching of text in the data stream can be performed using a predefined regular expression library to identify structured sensitive information such as ID card numbers, bank card numbers, and phone numbers.

[0076] Based on data timeliness characteristics and business context information, the sensitive information undergoes type analysis and security level recommendation analysis to obtain corresponding sensitive information types and security level recommendations. Type analysis aims to categorize sensitive information into predefined sensitive information types, such as personal identification information, financial information, and health information, which helps to take differentiated protection measures for different types of sensitive information. Security level recommendation analysis assesses the potential risks of the sensitive information based on its type, current data timeliness characteristics, business context information, and preset security policies and compliance requirements, and recommends an appropriate security level, such as top secret, confidential, internal use, or public. Sensitive tags are generated based on the sensitive information type and security level recommendation information. This sensitive tag can be a composite tag containing metadata such as type identifier, security level, timeliness attribute, and business context identifier, used for subsequent data file encryption, storage, and management. Because the sensitive tag is dynamically generated based on data timeliness and business context, it can better adapt to the constantly changing data environment and business needs, and ensure that the generated sensitive tag accurately reflects the current sensitivity and risk status of the data, avoiding misjudgments or omissions that may be caused by traditional static rule identification.

[0077] S202: Generate a data file based on the original data stream, and embed sensitive tags into the data file using a non-visible method to obtain the processed data file;

[0078] In a specific implementation of this invention, the process of embedding sensitive tags into a data file using a non-visible method to obtain a processed data file includes: encrypting the sensitive tags to obtain encrypted sensitive tags, and generating an integrity check code based on the encrypted sensitive tags and the content information of the data file; determining the type of the data file and the expected storage environment, and determining a tag embedding mechanism based on the type of the data file and the expected storage environment; and using the tag embedding mechanism to embed the integrity check code and the encrypted sensitive tags into the data file using a non-visible method to obtain the processed data file.

[0079] Specifically, a data file is generated based on the original data stream, and the sensitive tags are encrypted to obtain encrypted sensitive tags. Encryption of the sensitive tags aims to enhance their security and prevent unauthorized access or tampering. An integrity check code is generated based on the encrypted sensitive tags and the content information of the data file. This integrity check code is used to subsequently verify the integrity and authenticity of the data file and its sensitive tags.

[0080] Determine the type of data file and the expected storage environment. The data file type can include, but is not limited to, text files, image files, video files, audio files, etc., while the expected storage environment may involve various scenarios such as local storage, cloud storage, and distributed storage. Based on the type of data file and the expected storage environment, determine the tag embedding mechanism. For example, for image files, steganography can be used to embed sensitive tags into the low-order pixels of the image; for text files, zero-width character embedding or metadata field embedding can be used.

[0081] The data file is processed by embedding an integrity check code and an encrypted sensitive tag using the tag embedding mechanism based on the non-visible method. The non-visible method refers to embedding the encrypted sensitive tag and integrity check code into the data file without affecting the normal use and perception of the data file. This step-by-step processing and dynamic adaptation strategy ensures the concealment, security and effective protection of the integrity of the data file by the embedded sensitive tag, making the embedding of sensitive tags more efficient and secure, and without affecting the normal use of the data file.

[0082] It should be noted that the step of generating an integrity verification code based on the encrypted sensitive tags and the content information of the data file includes: dividing the content information of the data file into blocks to obtain several data blocks; calculating the hash value of each data block to obtain the target hash value corresponding to each data block; generating an aggregated verification code based on the target hash value and the encrypted sensitive tags, and using the aggregated verification code as the integrity verification code.

[0083] Specifically, segmenting the content information of the data file means dividing the entire data file into several data blocks according to preset rules or sizes. For example, the data file can be divided into blocks of a fixed size or according to the logical structure of the file content. Segmentation is to improve the granularity of integrity verification, so that even if there are minor changes in the local content of the data file, they can be accurately located and detected.

[0084] Hash values ​​are calculated for each data block to obtain the target hash value corresponding to each data block. Hash value calculation refers to applying a hash function independently to each data block to generate a hash value of fixed length, which is the target hash value corresponding to each data block. The hash function has one-way and collision resistance, and can provide a unique digital fingerprint for each data block. Any tampering with the content of the data block will cause its hash value to change, thereby effectively detecting the integrity of the data block.

[0085] The aggregation checksum generated based on the target hash value and the encrypted sensitive tag combines the target hash value of all data blocks with the encrypted sensitive tag to form a comprehensive checksum. For example, the target hash values ​​of all data blocks can be concatenated according to their order in the data file, and then the concatenated hash value can be hashed again with the encrypted sensitive tag. Alternatively, a hash-based message authentication code or other mechanism can be used, with the encrypted sensitive tag as the key, to authenticate all target hash values. Thus, the generated aggregation checksum not only contains the integrity information of the data file content but also incorporates the encrypted information of the sensitive tag. Using the aggregation checksum as an integrity checksum ensures that any unauthorized modification of the content and embedded sensitive tags of the data file can be effectively detected during storage and transmission, thereby improving the overall security and trustworthiness of the data file.

[0086] S203: Read the sensitive tag parsing information of the processed data file, determine the security level of the processed data file based on the sensitive tag parsing information, and determine the initial encryption strategy based on the security level;

[0087] In the specific implementation of this invention, the sensitive tag parsing information of the processed data file is read. That is, the embedded sensitive tag information is parsed from the data file using a preset parser or application programming interface. For example, if the sensitive tag is embedded using steganography, corresponding decryption and extraction algorithms are needed to recover the sensitive tag. The parsed sensitive tag typically contains information such as the sensitive information type and security level recommendations. Based on this parsing information and preset security policy rules, the system determines the security level of the processed data file. For example, if a sensitive tag indicates that the data file contains highly sensitive financial data, the system may determine its security level as top secret. An initial encryption strategy is determined based on the security level. For example, for top secret files, the initial encryption strategy may include using the AES-256 encryption algorithm and a high-strength key. For confidential files, the AES-128 encryption algorithm may be used.

[0088] S204: Determine the current state data and performance baseline of the encryption management system, and adjust the initial encryption strategy based on the performance baseline and current state data to obtain a target encryption strategy. Then, encrypt the processed data file based on the target encryption strategy to obtain an encrypted data file.

[0089] In the specific implementation of this invention, determining the current state data and performance baseline of the encryption management system, and adjusting the initial encryption strategy based on the performance baseline and the current state data to obtain the target encryption strategy, includes: obtaining the current state data of the encryption management system based on a system performance probe, performing file encryption simulation on the encryption management system based on a preset time period to obtain encryption simulation data, and determining the performance baseline based on the encryption simulation data; comparing the current state data with the performance baseline to obtain a comparison result, and adjusting the concurrency, data block size, hardware security module interaction, and process priority of the initial encryption strategy based on the comparison result to obtain the target encryption strategy.

[0090] Specifically, the current status data of the encryption management system is obtained based on system performance probes. System performance probes are monitoring components deployed within the encryption management system, designed to collect various operational status data of the system in real-time or periodically, such as central controller utilization, latency, memory usage, network bandwidth, and encryption queue length. File encryption simulations are performed on the encryption management system within a preset time period to obtain encryption simulation data. The preset time period is a time window set during the file encryption simulation, which can be several minutes, several hours, or longer. Within this preset time period, the encryption management system is simulated to execute a series of file encryption tasks to generate encryption simulation data. This data can include indicators such as encryption throughput, latency, and resource consumption under different loads. A performance baseline is determined based on this encryption simulation data. The performance baseline can be understood as a performance reference standard for the encryption management system under normal or ideal operating conditions, used for subsequent comparison with the current status data.

[0091] The current state data is compared with a performance baseline, i.e., the deviation between the current state data and the performance baseline is analyzed. For example, if the current CPU utilization is much higher than the baseline, or the encryption throughput is much lower than the baseline, it indicates that the system may have a performance bottleneck, thus obtaining the comparison result. Based on the comparison result, the initial encryption strategy is adjusted in terms of concurrency, data block size, hardware security module interaction, and process priority to obtain the target encryption strategy. Concurrency adjustment refers to adjusting the number of encryption tasks running simultaneously to optimize the utilization of CPU and I / O device resources; data block size adjustment refers to changing the size of the data blocks processed in encryption operations; hardware security module interaction adjustment refers to optimizing the communication frequency and data transmission method between the encryption management system and the hardware security module to reduce latency and improve security; process priority adjustment refers to adjusting the priority of encryption-related processes in the operating system to ensure that critical encryption tasks can obtain sufficient computing resources. Through these fine-tuning adjustments, a target encryption strategy that is more in line with the current system operation can be obtained.

[0092] The processed data file is encrypted based on the target encryption strategy to ensure that the encryption process can maximize the use of system resources, while avoiding performance bottlenecks caused by insufficient adjustment in a single dimension. This dynamic, multi-dimensional adjustment mechanism enables the encryption strategy to better adapt to the complex environment of system load and resource changes, effectively improve the efficiency of data encryption, reduce resource consumption, and enhance the stability and responsiveness of the system under different loads.

[0093] S205: Perform integrity verification on the encrypted data file, determine several copy files corresponding to the encrypted data file based on the integrity verification result, and store the encrypted data file and several copy files to the corresponding backup path based on the security level;

[0094] In the specific implementation of this invention, the encrypted data file undergoes integrity verification. After the data file is encrypted, integrity verification is required to ensure its integrity during storage and transmission. One implementation method is to calculate a hash value for the encrypted data file and store this hash value along with the file. In subsequent operations, the integrity of the file can be verified by recalculating the hash value and comparing it with the stored hash value. Based on the integrity verification result, several copy files corresponding to the encrypted data file are determined. If the verification result indicates that the file integrity is good, the system may generate several copy files according to a preset redundancy strategy (e.g., a three-copy strategy). Based on the security level, the encrypted data file and several copy files are stored to the corresponding backup path. For example, for files of the top-secret level, they can be stored in physically isolated, high-security storage media or different geographical locations in the data center; for files of the confidential level, they can be stored in different areas of cloud storage.

[0095] S206: Monitor the synchronization status and health status of each stored copy file based on the metadata service, and determine the verification strategy for the copy file based on the synchronization status and health status;

[0096] In the specific implementation of this invention, the monitoring of the synchronization status and health status of each copy file after storage based on the metadata service includes: obtaining the storage geographical distribution, data update frequency and network bandwidth of each copy file after storage; adjusting the initial monitoring frequency of the metadata service based on the storage geographical distribution, data update frequency and network bandwidth to obtain a target monitoring frequency; and monitoring the synchronization status and health status of each copy file based on the target monitoring frequency.

[0097] Specifically, metadata service can be understood as a system component used to manage and maintain information related to data files and their copies. This service continuously collects and analyzes the metadata of each copy file, such as its storage location, size, creation time, modification time, version information, and association with the original data file, to keep track of the synchronization status and health status of each copy file in real time.

[0098] The system obtains the geographical distribution of the stored copies, the data update frequency, and the network bandwidth. Geographical distribution refers to the physical distribution of the copies across different locations or data centers. For example, copies may be distributed across different servers in the same region or across multiple data centers in different regions. Data update frequency refers to the rate at which the content of the copies changes. For example, some log files may be updated multiple times per second, while some archive files may only be updated once every few months or even years. Network bandwidth refers to the data transmission capability between the storage location of the copies and the metadata service, which directly affects the efficiency of monitoring data transmission.

[0099] The initial monitoring frequency of the metadata service is adjusted based on the storage geographic distribution, data update frequency, and network bandwidth to obtain the target monitoring frequency. The initial monitoring frequency can be understood as the default or preset monitoring cycle of the metadata service before any adjustments are made. By comprehensively analyzing the three key parameters of storage geographic distribution, data update frequency, and network bandwidth, the initial monitoring frequency can be intelligently adjusted. For example, when the replica files are widely distributed, the data update frequency is high, and the network bandwidth is low, the target monitoring frequency will be adjusted to a higher frequency to ensure timely detection and handling of potential problems. Conversely, when the replica files are concentrated in one region, the data update frequency is low, and the network bandwidth is sufficient, the target monitoring frequency will be appropriately reduced to save system resources.

[0100] The metadata service monitors the synchronization and health status of each copy file based on the target monitoring frequency. The synchronization status refers to the degree of data consistency between each copy file and its original data file or other copy files, such as whether there are data differences or update delays. The health status refers to the physical or logical integrity of the copy file itself, such as whether there are problems such as damage, loss, or inaccessibility. The initial monitoring frequency is adjusted to a more reasonable target monitoring frequency, thereby avoiding the inefficiency or response delay problems caused by fixed frequency monitoring.

[0101] The verification strategy for the replica files is determined based on their synchronization and health status. If a replica file's synchronization status shows significant differences from the master file, or its health status indicates a potential risk of corruption, a stricter or more frequent verification strategy may be assigned. The verification strategy may include verification frequency, verification depth, and verification priority. This allows for the optimization of verification resource allocation based on the actual condition of the replica files, ensuring that critical replica files are verified promptly and effectively.

[0102] Furthermore, adjusting the initial monitoring frequency of the metadata service based on the storage geographic distribution, data update frequency, and network bandwidth to obtain the target monitoring frequency includes: acquiring historical storage geographic distribution, historical data update frequency, and historical network bandwidth, and performing parameter fluctuation analysis on the historical storage geographic distribution, historical data update frequency, and historical network bandwidth to obtain corresponding parameter fluctuation data; performing significant deviation analysis based on the storage geographic distribution, data update frequency, and network bandwidth combined with the corresponding parameter fluctuation data to obtain significant deviation information, and incrementally adjusting the initial monitoring frequency based on the significant deviation information to obtain the incrementally adjusted initial monitoring frequency; acquiring the data monitoring effect of the incrementally adjusted initial monitoring frequency, and determining an adaptive adjustment factor based on the data monitoring effect; and readjusting the incrementally adjusted initial monitoring frequency based on the adaptive adjustment factor to obtain the target monitoring frequency.

[0103] Specifically, acquiring historical storage geographic distribution, historical data update frequency, and historical network bandwidth refers to the system continuously recording and storing the values ​​of these key parameters over a period of time, for example, by sampling and storing on an hourly, daily, or weekly basis. Performing parameter fluctuation analysis on the historical storage geographic distribution, historical data update frequency, and historical network bandwidth involves using statistical methods, such as calculating standard deviation, variance, moving average, or conducting time series analysis, to quantify the range and trend of these parameters' changes in historical data, thereby obtaining corresponding parameter fluctuation data. The purpose is to understand the inherent dynamic characteristics and expected range of change of the parameters.

[0104] Significant deviation analysis based on the storage geographic distribution, data update frequency, and network bandwidth, combined with corresponding parameter fluctuation data, involves comparing the current storage geographic distribution, data update frequency, and network bandwidth with historical parameter fluctuation data. For example, anomaly detection algorithms or statistical testing methods are applied to determine whether the current parameter value significantly deviates from the historical average level or expected fluctuation range, thereby obtaining significant deviation information. Incremental adjustment of the initial monitoring frequency based on this significant deviation information refers to making small, gradual adjustments to the initial monitoring frequency of the metadata service according to the degree and direction of the detected significant deviation, to obtain an incrementally adjusted initial monitoring frequency. The purpose is to ensure that the monitoring frequency can respond promptly to abnormal changes in the current environment.

[0105] Obtaining the data monitoring effect of the initial monitoring frequency after incremental adjustment refers to the system continuously monitoring its operational performance after incrementally adjusting the initial monitoring frequency. This includes evaluating indicators such as the detection latency of replica file synchronization status, the efficiency of detecting abnormal health status, resource consumption, and false alarm rate, thereby quantifying the actual effect of the adjustment. Determining an adaptive adjustment factor based on the data monitoring effect means calculating an adaptive adjustment factor for further optimizing the monitoring frequency using a preset algorithm or machine learning model based on the monitored effect data. Its purpose is to provide empirical feedback for subsequent frequency adjustments.

[0106] The process of readjusting the initial monitoring frequency after incremental adjustment based on the adaptive adjustment factor refers to applying the adaptive adjustment factor to the initial monitoring frequency after incremental adjustment for final fine-tuning to obtain the target monitoring frequency. The purpose is to ensure that the monitoring frequency not only responds to current changes but also learns from historical adjustment experience and continuously optimizes, significantly improving the accuracy and real-time performance of the synchronization status and health status monitoring of copy files. This reduces potential data risks or resource waste caused by inappropriate monitoring frequencies. Furthermore, through adaptive learning, this scheme enables the system to continuously optimize its monitoring strategy over time, thereby improving the robustness and efficiency of the entire data file backup management system.

[0107] It should be noted that the step of determining the verification strategy for the replica files based on the synchronization status and health status includes: obtaining the security level and verification real-time requirements of each replica file, and determining the replica consistency waiting time based on the security level and verification real-time requirements; determining the verification priority based on the synchronization status and health status; and determining the verification strategy for the replica files based on the replica consistency waiting time and verification priority.

[0108] Specifically, the system obtains the security level and verification real-time requirements for each copy file. Based on pre-set rules or from sensitive tag parsing information, the system acquires the security level corresponding to each copy file. The system also obtains the required verification real-time requirements for each copy file, which may depend on business needs, compliance requirements, or the importance of the data. For example, some critical business data may require extremely high verification real-time performance, while archived data may allow for lower real-time performance. Based on the security level and verification real-time requirements, the system determines the copy consistency wait time. The copy consistency wait time refers to the maximum time allowed for inconsistencies between copy files before integrity verification. For example, for files with high security levels and high real-time requirements, the copy consistency wait time is set to be shorter to ensure rapid detection and handling of inconsistencies; while for files with low security levels or low real-time requirements, this wait time can be appropriately extended to reduce verification frequency and resource consumption.

[0109] The system determines the verification priority based on the synchronization status and health status. According to the synchronization status and health status, the system will dynamically calculate and determine the verification priority of each copy file. For example, copy files with abnormal synchronization status or poor health status will be given higher verification priority so that integrity verification and repair can be performed first.

[0110] Determining a verification strategy for replica files based on the replica consistency wait time and verification priority means combining the determined replica consistency wait time with the verification priority to form a complete verification strategy. This strategy can include the verification frequency, verification depth (e.g., whether to perform metadata verification or full file content verification), verification timing, and the verification algorithm used. For example, for high-priority files with short replica consistency wait times, the verification strategy might be set to high-frequency, full-content verification, triggered in real time; while for low-priority files with long replica consistency wait times, the verification strategy might be set to low-frequency, metadata verification, performed when the system load is low, ensuring that potentially risky replica files are prioritized. This multi-dimensional, adaptive strategy determination mechanism allows verification resources to be allocated and utilized more rationally, thereby improving the efficiency and security of the entire backup management system.

[0111] S207: Perform integrity verification on each copy file based on the verification strategy to obtain the corresponding integrity verification result, and determine the overall integrity verification result based on the majority consensus principle and the integrity verification results of each copy file.

[0112] In the specific implementation of this invention, integrity verification refers to checking the data integrity of each copy file according to the determined verification strategy. This can be achieved by calculating the hash value of the copy file and comparing it with the original data file or a known correct hash value, or by using methods such as checksum and cyclic redundancy check. For example, for high-priority files with short copy consistency waiting time, the verification strategy may be set to high-frequency, full-content verification and triggered in real time; while for low-priority files with long copy consistency waiting time, the verification strategy may be set to low-frequency, metadata verification and performed when the system load is low. The verification process generates an integrity verification result, which indicates whether the copy file is consistent with its expected state, such as passing or failing, or provides a specific difference report.

[0113] The overall integrity verification result is determined based on the majority consensus principle combined with the integrity verification results of each copy file. Since there are multiple copy files, the majority consensus principle is adopted to improve the reliability and fault tolerance of the verification. This means that when there are multiple copy files for the same data, even if the integrity verification result of individual copy files shows that the integrity verification result fails, as long as the verification result of the majority of copy files passes, the overall data can be considered complete. For example, if there are three copy files, two of which pass the verification and one fails the verification, the overall integrity verification result will be judged as passing. This mechanism can effectively avoid the accidental error of a single copy file causing the entire data backup system to be misjudged as incomplete.

[0114] S208: If the overall integrity verification result is that the overall integrity verification fails, then identify the abnormal copy file and repair the abnormal copy file based on the distributed repair process;

[0115] In the specific implementation of this invention, if the overall integrity verification result is that the overall integrity verification fails, it indicates that there is a serious problem with the data backup system. The system will identify the abnormal copy file that caused the overall verification to fail and repair the abnormal copy file based on the distributed repair process. The distributed repair process refers to using other healthy and complete copy files in the system as data sources, and using network transmission and data synchronization technology to recover and rebuild the abnormal copy file. This process usually involves the cooperation between multiple storage nodes to ensure the efficiency and reliability of the repair process.

[0116] S209: After the abnormal copy file is repaired, the sensitive tags are re-embedded in the repaired abnormal copy file.

[0117] In the specific implementation of this invention, after the abnormal copy file is repaired, sensitive tags are re-embedded into the repaired abnormal copy file. The re-embedding of sensitive tags means that after the abnormal copy file is repaired and its data integrity is restored, in order to ensure that its security attributes are consistent with the original data file, the sensitive tag embedding process needs to be performed again. This includes steps such as re-identifying sensitive information, generating sensitive tags, and embedding tags based on non-visible methods in the repaired copy file. Its purpose is to ensure that even after the data is repaired, the copy file still carries the correct sensitive tags, so that it can continue to be subject to the corresponding security policies and management, which greatly improves the reliability, security and maintainability of data backup and effectively reduces the risk of data loss or leakage.

[0118] In this embodiment of the invention, sensitive information in the original data stream is identified based on data timeliness characteristics and contextual information, and sensitive tags for the sensitive information are determined, resulting in more refined sensitive tags. Data files are generated based on the original data stream, and sensitive tags are embedded into the data files using a non-visible method, tightly integrating the sensitive tags with the data files and improving their concealment and security. The sensitive tag parsing information of the processed data files is read to determine the security level of the processed data files, and an initial encryption strategy is determined based on the security level. The initial encryption strategy is adjusted based on the current state data and performance baseline of the encryption management system to encrypt the processed data files, ensuring that security requirements are met while optimizing system resource utilization and improving encryption efficiency. Integrity verification is performed on the encrypted data files to determine several corresponding copy files. Based on the security level, the encrypted data files and several copy files are stored in the corresponding backup paths. The synchronization and health status of each stored copy file are monitored to perform integrity verification, copy repair, and sensitive tag re-embedding on each copy file. This allows for intelligent adoption of differentiated security backup strategies based on the actual value and sensitivity of the data, significantly improving the efficiency, security, and reliability of data backup.

[0119] Example 3

[0120] Please see Figure 3 , Figure 3 This is a schematic diagram of the structural composition of a data file backup management device according to an embodiment of the present invention. The device includes:

[0121] Tag determination module 31: used to identify sensitive information in the original data stream based on data timeliness characteristics and context information, and to determine the sensitive tags of the sensitive information;

[0122] Tag embedding module 32: used to generate a data file based on the original data stream, and to perform sensitive tag embedding processing on the data file based on a non-visible method to obtain the processed data file;

[0123] Initial strategy analysis module 33: used to read the sensitive tag parsing information of the processed data file, determine the security level of the processed data file based on the sensitive tag parsing information, and determine the initial encryption strategy based on the security level;

[0124] File encryption module 34: used to determine the current status data and performance baseline of the encryption management system, and adjust the initial encryption strategy based on the performance baseline and current status data to obtain a target encryption strategy, and encrypt the processed data file based on the target encryption strategy to obtain an encrypted data file;

[0125] File backup module 35: used to perform integrity verification on the encrypted data file, determine several copy files corresponding to the encrypted data file based on the integrity verification result, and store the encrypted data file and several copy files to the corresponding backup path based on the security level;

[0126] Copy file monitoring module 36: used to monitor the synchronization status and health status of each copy file after storage, and to perform integrity verification, copy repair and sensitive tag re-embedding on each copy file based on the synchronization status and health status.

[0127] In the specific implementation of this invention, the specific implementation of the device item can be referred to the implementation of the method item above, and will not be repeated here.

[0128] In this embodiment of the invention, sensitive information in the original data stream is identified based on data timeliness characteristics and contextual information, and sensitive tags for the sensitive information are determined, resulting in more refined sensitive tags. Data files are generated based on the original data stream, and sensitive tags are embedded into the data files using a non-visible method, tightly integrating the sensitive tags with the data files and improving their concealment and security. The sensitive tag parsing information of the processed data files is read to determine the security level of the processed data files, and an initial encryption strategy is determined based on the security level. The initial encryption strategy is adjusted based on the current state data and performance baseline of the encryption management system to encrypt the processed data files, ensuring that security requirements are met while optimizing system resource utilization and improving encryption efficiency. Integrity verification is performed on the encrypted data files to determine several corresponding copy files. Based on the security level, the encrypted data files and several copy files are stored in the corresponding backup paths. The synchronization and health status of each stored copy file are monitored to perform integrity verification, copy repair, and sensitive tag re-embedding on each copy file. This allows for intelligent adoption of differentiated security backup strategies based on the actual value and sensitivity of the data, significantly improving the efficiency, security, and reliability of data backup.

[0129] This invention provides a computer-readable storage medium storing a computer program. When executed by a processor, this program implements the data file backup management method of any of the above embodiments. The computer-readable storage medium includes, but is not limited to, any type of disk (including floppy disk, hard disk, optical disk, CD-ROM, and magneto-optical disk), ROM (Read-Only Memory), RAM (Random Access Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory, magnetic cards, or optical cards. In other words, the storage device includes any medium that stores or transmits information in a readable form by a device (e.g., a computer, a mobile phone), and can be a read-only memory, a disk, or an optical disk, etc.

[0130] Example 4

[0131] Please see Figure 4 , Figure 4 This is a schematic diagram of the structural composition of the electronic device in an embodiment of the present invention.

[0132] This invention also provides an electronic device, such as... Figure 4 As shown, the electronic device includes a memory 41, a processor 43, and a computer program 42 stored in the memory 41 and executable on the processor 43. Those skilled in the art will understand that... Figure 4 The illustrated electronic device does not constitute a limitation on all devices and may include more or fewer components than illustrated, or combine certain components. Memory 41 can be used to store computer program 42 and various functional modules. Processor 43 runs the computer program 42 stored in memory 41, thereby performing various functional applications and data processing of the device. Memory can be internal memory or external memory, or both. Internal memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, or random access memory. External memory may include hard disks, floppy disks, ZIP disks, USB flash drives, magnetic tapes, etc. Processor 43 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, a single-chip microcomputer, or a processor 43, or any conventional processor, etc. The processors and memories disclosed in this invention include, but are not limited to, these types of processors and memories. The processors and memories disclosed in this invention are merely examples and not intended to be limiting.

[0133] As one embodiment, the electronic device includes: one or more processors 43, a memory 41, and one or more computer programs 42, wherein the one or more computer programs 42 are stored in the memory 41 and configured to be executed by the one or more processors 43, and the one or more computer programs 42 are configured to perform the data file backup management method in any of the above embodiments. For specific implementation processes, please refer to the above embodiments, which will not be repeated here.

[0134] Furthermore, the above provides a detailed description of a data file backup management method and related apparatus provided by the embodiments of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for backing up and managing data files, characterized in that, The method includes: Sensitive information in the original data stream is identified based on the timeliness characteristics and contextual information, and sensitive tags for the sensitive information are determined. A data file is generated based on the original data stream, and sensitive tags are embedded in the data file using a non-visible method to obtain the processed data file. Read the sensitive tag parsing information of the processed data file, determine the security level of the processed data file based on the sensitive tag parsing information, and determine the initial encryption strategy based on the security level; The current status data and performance baseline of the encryption management system are determined, and the initial encryption strategy is adjusted based on the performance baseline and the current status data to obtain the target encryption strategy. The processed data file is then encrypted based on the target encryption strategy to obtain the encrypted data file. The encrypted data file is subjected to integrity verification. Based on the integrity verification result, several copy files corresponding to the encrypted data file are determined. Based on the security level, the encrypted data file and several copy files are stored in the corresponding backup path. Monitor the synchronization and health status of each stored copy file, and perform integrity verification, copy repair, and sensitive tag re-embedding on each copy file based on the synchronization and health status.

2. The data file backup management method according to claim 1, characterized in that, The process of identifying sensitive information in the original data stream based on data timeliness characteristics and contextual information, and determining the sensitive tags of the sensitive information, includes: Obtain the data timeliness characteristics and business context information of the original data stream, and determine the weight and sensitivity threshold of the sensitive information pattern matching rule based on the data timeliness characteristics and business context information; Sensitive information in the original data stream is identified based on the weights and sensitivity thresholds of the sensitive information pattern matching rules. Based on the timeliness characteristics of the data and the business context information, the sensitive information is analyzed for type and security level recommendations to obtain the corresponding sensitive information type and security level recommendation information, and sensitive tags are generated based on the sensitive information type and security level recommendation information.

3. The data file backup management method according to claim 1, characterized in that, The method of embedding sensitive tags into the data file using a non-visible method to obtain the processed data file includes: The sensitive tags are encrypted to obtain encrypted sensitive tags, and an integrity check code is generated based on the encrypted sensitive tags and the content information of the data file. Determine the type of data file and the expected storage environment, and determine the tag embedding mechanism based on the type of data file and the expected storage environment; The data file is processed by embedding an integrity check code and an encrypted sensitive tag using the tag embedding mechanism based on the non-visible method, thereby obtaining the processed data file.

4. The data file backup management method according to claim 1, characterized in that, The process of determining the current state data and performance baseline of the encryption management system, and adjusting the initial encryption strategy based on the performance baseline and current state data to obtain the target encryption strategy includes: The current status data of the encryption management system is obtained based on the system performance probe, and the encryption management system is simulated for file encryption based on a preset time period to obtain encryption simulation data. The performance baseline is then determined based on the encryption simulation data. The current state data is compared with the performance baseline to obtain the comparison result. Based on the comparison result, the initial encryption strategy is adjusted in terms of concurrency, data block size, hardware security module interaction, and process priority to obtain the target encryption strategy.

5. The data file backup management method according to claim 1, characterized in that, The process of monitoring the synchronization and health status of each stored copy file, and performing integrity verification, copy repair, and sensitive tag re-embedding on each copy file based on the synchronization and health status, includes: Based on metadata services, monitor the synchronization status and health status of each copy file after storage; The verification strategy for the copy file is determined based on the synchronization status and health status. Based on the verification strategy, integrity verification is performed on each copy file to obtain the corresponding integrity verification result; The overall integrity verification result is determined based on the majority consensus principle and the integrity verification results of each copy file. If the overall integrity verification result is that the overall integrity verification fails, then the abnormal copy file is identified and repaired based on the distributed repair process; After the abnormal copy file is repaired, the sensitive tags are re-embedded in the repaired abnormal copy file.

6. The data file backup management method according to claim 5, characterized in that, The monitoring of the synchronization status and health status of each copy file after storage based on metadata services includes: Obtain the geographical distribution of the stored copies, the data update frequency, and the network bandwidth of each copy file after storage; The initial monitoring frequency of metadata services is adjusted based on the storage geographic distribution, data update frequency, and network bandwidth to obtain the target monitoring frequency. The metadata service monitors the synchronization status and health status of each copy file based on the target monitoring frequency.

7. The data file backup management method according to claim 6, characterized in that, The adjustment of the initial monitoring frequency of metadata services based on the storage geographic distribution, data update frequency, and network bandwidth to obtain the target monitoring frequency includes: The historical storage geographic distribution, historical data update frequency, and historical network bandwidth are obtained, and parameter fluctuation analysis is performed on the historical storage geographic distribution, historical data update frequency, and historical network bandwidth to obtain the corresponding parameter fluctuation data. Based on the storage geographic distribution, data update frequency, and network bandwidth, combined with the corresponding parameter fluctuation data, a significant deviation analysis is performed to obtain significant deviation information. Based on the significant deviation information, the initial monitoring frequency is incrementally adjusted to obtain the incrementally adjusted initial monitoring frequency. Obtain the data monitoring effect of the initial monitoring frequency after incremental adjustment, and determine the adaptive adjustment factor based on the data monitoring effect; The initial monitoring frequency after incremental adjustment is readjusted based on the adaptive adjustment factor to obtain the target monitoring frequency.

8. A data file backup management device, characterized in that, The device includes: Tag determination module: used to identify sensitive information in the original data stream based on data timeliness characteristics and context information, and to determine the sensitive tags of the sensitive information; Tag embedding module: used to generate a data file based on the original data stream, and to embed sensitive tags into the data file using a non-visible method to obtain the processed data file; Initial strategy analysis module: used to read the sensitive tag parsing information of the processed data file, determine the security level of the processed data file based on the sensitive tag parsing information, and determine the initial encryption strategy based on the security level; File encryption module: used to determine the current status data and performance baseline of the encryption management system, and adjust the initial encryption strategy based on the performance baseline and current status data to obtain the target encryption strategy, and encrypt the processed data file based on the target encryption strategy to obtain the encrypted data file; File backup module: used to perform integrity verification on the encrypted data file, determine several copy files corresponding to the encrypted data file based on the integrity verification result, and store the encrypted data file and several copy files to the corresponding backup path based on the security level; The copy file monitoring module is used to monitor the synchronization status and health status of each stored copy file, and to perform integrity verification, copy repair, and sensitive tag re-embedding on each copy file based on the synchronization status and health status.

9. An electronic device, the electronic device comprising a processor and a memory, characterized in that, The memory is used to store instructions, and the processor is used to invoke the instructions in the memory to cause the electronic device to execute the data file backup management method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed on an electronic device, cause the electronic device to perform a data file backup management method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Data security system based on dynamic data splitting

    CN119089482A

  • Data cross-border security guarantee method

    CN120378135A