Backup disaster recovery method and system combined with AI technology

By calculating the relevant redundancy, sensitive characteristics and access log analysis of the data nodes, a backup disaster recovery strategy is generated, which solves the problems of wasted data backup resources and low efficiency in the existing technology, and realizes efficient data backup and disaster recovery management.

CN120371609AActive Publication Date: 2025-07-25SHENZHEN FEICHUANGYUN INFORMATION TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510517121.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-25
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

Existing backup and disaster recovery systems cannot effectively process the diversity and complexity of enterprise data, resulting in excessive computing resources consumption, making it difficult to determine the urgency of data backup and generate efficient disaster recovery strategies.

Method used

By obtaining the text data of each data node in the enterprise database, calculating the relevant redundancy, filtering the data nodes to be backed up, extracting sensitive characteristics and access logs, determining privacy priorities and access abnormal factors, generating backup urgency, and formulating differentiated backup disaster recovery strategies.

Benefits of technology

Significantly reduce redundant backups, optimize storage resources, identify key data and formulate targeted strategies, improve data backup and recovery efficiency, and reduce costs and risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371609A_ABST
    Figure CN120371609A_ABST
Patent Text Reader

Abstract

The invention provides a backup disaster recovery method and system combined with an AI technology, and the method comprises the steps: obtaining the text data of each data node in an enterprise database, determining the related redundancy among all data nodes according to all text data, and screening out all to-be-backed-up data nodes in the enterprise database according to all related redundancy; respectively extracting sensitive features of each to-be-backed-up data node, and determining privacy priority of each to-be-backed-up data node according to the corresponding sensitive features; obtaining an access log of each to-be-backed-up data node, and determining an access exception factor of each to-be-backed-up data node based on the corresponding access log; and determining the backup urgency degree of each to-be-backed-up data node through the corresponding privacy priority and the corresponding access exception factor, and generating a backup disaster tolerance strategy of each to-be-backed-up data node according to the corresponding backup urgency degree. According to the scheme, the backup emergency degree of the data can be determined, and the disaster recovery strategy is generated, so that the efficiency of the data backup process is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of backup and disaster recovery, and more specifically, to a backup and disaster recovery method and system combining AI technology. Background Art

[0002] With the advancement of digital transformation, the importance and complexity of enterprise data have been continuously increasing. Traditional backup and disaster recovery systems can no longer meet the high requirements for data security, recovery efficiency, and business continuity. Backup and disaster recovery solutions combining AI technology have gradually become an important means for enterprises to cope with data disasters and improve the level of data protection. Through AI technology, backup and disaster recovery systems can automate processes such as data classification, priority judgment, and backup strategy optimization, thereby achieving more efficient and intelligent disaster recovery.

[0003] However, in practical applications, the diversity and complexity of data pose great challenges to the training and inference of AI algorithms. Enterprise data usually includes structured and unstructured data, and the formats, sizes, and update frequencies of the data are also different. AI models need to process these complex data sources, extract valuable information from them, and perform effective classification. Moreover, the introduction of AI technology requires the backup system to have powerful computing capabilities and storage capacities. The training and inference processes of AI algorithms usually require a large amount of computing resources. Especially when dealing with large-scale data sets, the training of AI models may consume a large amount of time and computing power. In addition, the requirements for real-time data backup and disaster recovery require the backup system to have high availability and low latency, and the computing process of the AI system may become a performance bottleneck. Therefore, how to determine the backup urgency of data and generate a disaster recovery strategy to improve the efficiency of the data backup process is a difficult problem faced by the industry. Summary of the Invention

[0004] The present application provides a backup and disaster recovery method and system combining AI technology, which can determine the backup urgency of data and generate a disaster recovery strategy to improve the efficiency of the data backup process.

[0005] In a first aspect, the present application provides a backup and disaster recovery method combining AI technology, and the disaster recovery method includes the following steps:

[0006] Obtain the text data of each data node in the enterprise database, determine the relevant redundancy between each data node in the enterprise database according to all the text data, and screen out all the data nodes to be backed up in the enterprise database based on each relevant redundancy;

[0007] Extract the sensitive features of each data node to be backed up respectively, and then determine the privacy priority of each data node to be backed up based on the corresponding sensitive features;

[0008] Obtain the access logs of each data node to be backed up, and determine the access anomaly factors of each data node to be backed up based on the corresponding access logs;

[0009] Determine the backup urgency of each data node to be backed up in the enterprise database through the corresponding privacy priority and the corresponding access anomaly factor, and generate a backup disaster recovery policy for each data node to be backed up according to the corresponding backup urgency.

[0010] Preferably, determining the relevant redundancy between each data node in the enterprise database according to all text data specifically includes:

[0011] Generate a hash signature set corresponding to each data node in the enterprise database based on the corresponding text data;

[0012] Select a data node among all data nodes in the enterprise database, and obtain the hash signature set corresponding to the selected data node;

[0013] Determine the data redundancy between the selected data node and other data nodes in the enterprise database based on the hash signature set corresponding to the selected data node;

[0014] Determine the relevant redundancy between the selected data node and other data nodes in the enterprise database through the corresponding data redundancy, and then obtain the relevant redundancy between each data node in the enterprise database.

[0015] Preferably, screening out all data nodes to be backed up in the enterprise database according to each relevant redundancy specifically includes:

[0016] Obtain a preset redundancy threshold, and then screen out all data node pairs with a relevant redundancy higher than the redundancy threshold;

[0017] Remove one data node from each of the screened data node pairs respectively, and then use the remaining data nodes and the data nodes in each data node pair with a relevant redundancy not higher than the redundancy threshold as the data nodes to be backed up in the enterprise database, and then obtain all data nodes to be backed up in the enterprise database.

[0018] Preferably, use natural language processing technology and a classification model based on machine learning to extract the sensitive features of each data node to be backed up respectively.

[0019] Preferably, determining the privacy priority of each data node to be backed up according to the corresponding sensitive features is to input the sensitive features of each data node to be backed up into the evaluation model for evaluation respectively, and then obtain the privacy priority of each data node to be backed up.

[0020] Preferably, determining the access anomaly factor of each data node to be backed up based on the corresponding access log specifically includes:

[0021] For each data node to be backed up, extract multiple access characteristics of the data node to be backed up from the access log of the data node to be backed up;

[0022] Determine the access anomaly factor of the data node to be backed up through all the access characteristics, and then obtain the access anomaly factors of each data node to be backed up.

[0023] Preferably, generating the backup and disaster recovery strategy for each data node to be backed up according to the corresponding backup urgency specifically includes:

[0024] Determine the backup priority group of each data node to be backed up according to the corresponding backup urgency;

[0025] Determine the backup and disaster recovery strategy of each data node to be backed up through the corresponding backup priority group.

[0026] In a second aspect, the present application provides a backup and disaster recovery system combined with AI technology for executing a backup and disaster recovery method combined with AI technology. The backup and disaster recovery system combined with AI technology includes a disaster recovery strategy generation unit, and the disaster recovery strategy generation unit includes:

[0027] A data acquisition module, configured to acquire the text data of each data node in the enterprise database, determine the relevant redundancy between each data node in the enterprise database according to all the text data, and screen out all the data nodes to be backed up in the enterprise database according to each relevant redundancy;

[0028] A priority determination module, configured to respectively extract the sensitive characteristics of each data node to be backed up, and then determine the privacy priority of each data node to be backed up according to the corresponding sensitive characteristics;

[0029] An abnormal access determination module, configured to acquire the access log of each data node to be backed up, and determine the access anomaly factor of each data node to be backed up based on the corresponding access log;

[0030] A strategy generation module, configured to determine the backup urgency of each data node to be backed up in the enterprise database through the corresponding privacy priority and the corresponding access anomaly factor, and generate the backup and disaster recovery strategy for each data node to be backed up according to the corresponding backup urgency.

[0031] In a third aspect, the present application provides a computer device, which includes a memory and a processor. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the computer device executes the above-mentioned backup and disaster recovery method combined with AI technology.

[0032] Fourthly, the present application provides a computer-readable storage medium storing instructions or codes, which, when running on a computer, enable the computer to implement the above backup and disaster recovery method combined with AI technology.

[0033] The technical solutions provided by the disclosed embodiments of the present application have the following beneficial effects:

[0034] In a backup and disaster recovery method and system combined with AI technology provided by the present application, text data of each data node in an enterprise database is obtained, the relevant redundancy between each data node in the enterprise database is determined according to all the text data, and all data nodes to be backed up in the enterprise database are screened out according to each relevant redundancy; sensitive features of each data node to be backed up are extracted respectively, and then the privacy priority of each data node to be backed up is determined according to the corresponding sensitive features, access logs of each data node to be backed up are obtained, and access anomaly factors of each data node to be backed up are determined based on the corresponding access logs; the backup urgency of each data node to be backed up in the enterprise database is determined through the corresponding privacy priority and the corresponding access anomaly factor, and a backup and disaster recovery strategy for each data node to be backed up is generated according to the corresponding backup urgency.

[0035] Thus, in the present application, firstly, by obtaining the text data of each data node in the enterprise database and calculating the relevant redundancy between each data node, redundant backups can be significantly reduced, storage resources can be optimized, and costs can be reduced. Screening data nodes to be backed up according to the relevant redundancy can avoid unnecessary repeated backups, so that resources can be allocated more efficiently, and faster data recovery and disaster recovery response can be achieved; secondly, by extracting the sensitive features of each data node to be backed up respectively and determining the privacy priority of the data node according to these sensitive features, it can be effectively identified which data needs to be protected preferentially, so as to formulate differentiated backup and disaster recovery strategies; thirdly, obtaining the access logs of each data node to be backed up and determining the access anomaly factors based on the access logs can effectively identify potential security risks and abnormal access behaviors, and help evaluate which data nodes to be backed up may face higher security threats, so that the data nodes to be backed up with higher risks can be preferentially backed up; finally, by dividing the backup priority groups of data nodes according to the backup urgency and formulating corresponding backup and disaster recovery strategies for each priority group, the efficiency of the data backup process can be improved.

[0036] In summary, the technical solution adopted by the present application can determine the backup urgency of data and generate a disaster recovery strategy to improve the efficiency of the data backup process. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 is an exemplary flowchart of a backup and disaster recovery method combined with AI technology according to some embodiments of the present application;

[0039] Figure 2 is an exemplary flowchart of determining the relevant redundancy between each data node in an enterprise database according to some embodiments of the present application;

[0040] Figure 3 is an exemplary flowchart of determining the access anomaly factor of each data node to be backed up according to some embodiments of the present application;

[0041] Figure 4 is a schematic diagram of the exemplary hardware and / or software of a disaster recovery policy generation unit according to some embodiments of the present application;

[0042] Figure 5 is a schematic diagram of the structure of a computer device for implementing a backup and disaster recovery method combined with AI technology according to some embodiments of the present application. Detailed implementation manners

[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.

[0044] The embodiments of the present application provide a backup and disaster recovery method and system combined with AI technology. The core is to obtain the text data of each data node in the enterprise database, determine the relevant redundancy between each data node in the enterprise database according to all the text data, and screen out all the data nodes to be backed up in the enterprise database according to each relevant redundancy; extract the sensitive features of each data node to be backed up respectively, and then determine the privacy priority of each data node to be backed up according to the corresponding sensitive features; obtain the access logs of each data node to be backed up, and determine the access anomaly factor of each data node to be backed up based on the corresponding access logs; determine the backup urgency of each data node to be backed up in the enterprise database through the corresponding privacy priority and the corresponding access anomaly factor, and generate a backup and disaster recovery strategy for each data node to be backed up according to the corresponding backup urgency. By adopting the above scheme, the backup urgency of the data can be determined and a disaster recovery strategy can be generated to improve the efficiency of the data backup process.

[0045] To better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners. Refer to Figure 1 , which is an exemplary flowchart of a backup and disaster recovery method combined with AI technology according to some embodiments of the present application. The disaster recovery method 100 mainly includes the following steps:

[0046] In step 101, obtain the text data of each data node in the enterprise database, determine the relevant redundancy between each data node in the enterprise database according to all the text data, and screen out all the data nodes to be backed up in the enterprise database according to each relevant redundancy.

[0047] Specifically, the text data of each data node in the enterprise database can be obtained by means of data query language and traversal; in actual implementation, all fields containing text data can be queried. For each data node, the text data it contains is extracted, and the structured data (such as table data) is converted into a form suitable for text processing, so that the text data can be cleaned and standardized, and noises such as irrelevant symbols and repeated words can be removed.

[0048] Preferably, in some embodiments, refer to Figure 2 shown, which is an exemplary flowchart for determining the relevant redundancy between each data node in the enterprise database in some embodiments of the present application. In this embodiment, the relevant redundancy between each data node in the enterprise database can be determined according to all the text data by the following steps:

[0049] In step 1011, generate a hash signature set corresponding to each data node in the enterprise database according to the corresponding text data;

[0050] In step 1012, select a data node among all data nodes of the enterprise database, and obtain the hash signature set corresponding to the selected data node;

[0051] In step 1013, determine the data redundancy between the selected data node and other data nodes in the enterprise database based on the hash signature set corresponding to the selected data node;

[0052] In step 1014, determine the relevant redundancy between the selected data node and other data nodes in the enterprise database through the corresponding data redundancy, and then obtain the relevant redundancy between each data node in the enterprise database.

[0053] Specifically, first, the hash signature set corresponding to each data node in the enterprise database can be generated according to the corresponding text data, that is, the SimHash method is used to perform hash processing on the text data of the data node, so as to obtain the hash signature set corresponding to the data node. Through the above method, the hash signature set corresponding to each data node in the enterprise database can be obtained; then, a data node can be selected among all data nodes of the enterprise database, and the hash signature set corresponding to the selected data node can be obtained; furthermore, the data redundancy between the selected data node and other data nodes in the enterprise database can be determined based on the hash signature set corresponding to the selected data node. Among them, the data redundancy represents the degree of similarity and overlap of the content between data nodes. In actual implementation, the data redundancy can be determined by the following method:

[0054]

[0055] Among them, λ(D i , D j ) represents the data redundancy between data node D i and data node D j . D i represents the i-th data node in the enterprise database, D j represents the j-th data node in the enterprise database, h(D i ) represents the hash signature set corresponding to data node D i , h(D j ) represents the hash signature set corresponding to data node D j . Through the above method, the data redundancy between the selected data node and other data nodes in the enterprise database can be obtained.

[0056] In addition, specifically, the relevant redundancy between the selected data node and other data nodes in the enterprise database can be determined through the corresponding data redundancy. Among them, the relevant redundancy represents the degree of mutual correlation of the content between data nodes in the enterprise database. In actual implementation, the relevant redundancy can be determined by the following method:

[0057]

[0058] Among them, γ(D i , D j ) represents the correlation redundancy between data node D i and data node D j . D i represents the i-th data node in the enterprise database, and D j represents the j-th data node in the enterprise database. λ(D i , D j ) represents the data redundancy between data node D i and data node D j . min(D i , D j ) and max(D i , D j ) respectively represent selecting the text data with a smaller and a larger data volume from data node D i and data node D j . By the above method, the correlation redundancy between each data node in the enterprise database can be obtained.

[0059] In some embodiments, to screen out all the data nodes to be backed up in the enterprise database according to each correlation redundancy, the following method can be specifically adopted, that is:

[0060] Obtain a preset redundancy threshold, and then screen out all data node pairs with a correlation redundancy higher than the redundancy threshold;

[0061] Respectively remove one data node from each of the screened data node pairs, and then use the remaining data nodes and the data nodes in each data node pair with a correlation redundancy not higher than the redundancy threshold as the data nodes to be backed up in the enterprise database, and then obtain all the data nodes to be backed up in the enterprise database.

[0062] Specifically, first, the redundancy threshold can be set according to historical experiments and data analysis, which will not be elaborated here; then, all the correlation redundancies can be compared with the redundancy threshold respectively to screen out all data node pairs with a correlation redundancy higher than the redundancy threshold; finally, one data node can be removed from each of the screened data node pairs respectively, and the removal can be performed by random removal or other methods, which is not limited here, so that the remaining data nodes and the data nodes in each data node pair with a correlation redundancy not higher than the redundancy threshold can be used as the data nodes to be backed up in the enterprise database. By the above method, all the data nodes to be backed up in the enterprise database can be obtained.

[0063] It should be noted that by obtaining the text data of each data node in the enterprise database and calculating the relevant redundancy between each data node, redundant backups can be significantly reduced, storage resources can be optimized, and costs can be reduced. Screening the data nodes to be backed up based on the relevant redundancy can avoid unnecessary repeated backups, so that resources can be allocated more efficiently, and faster data recovery and disaster tolerance response can be achieved.

[0064] In step 102, the sensitive features of each data node to be backed up are extracted respectively, and then the privacy priorities of each data node to be backed up are determined according to the corresponding sensitive features.

[0065] In some embodiments, natural language processing technology and a classification model based on machine learning are used to extract the sensitive features of each data node to be backed up respectively. Specifically, first, natural language processing (NLP) technology can be used to automatically extract sensitive features from the text data of each data node to be backed up, that is, sensitive features. Then, a feature extraction model of machine learning can be used to further classify the text data of the data node to determine which data in the text data belong to sensitive features. The sensitive features of each data node to be backed up can be extracted in the above way.

[0066] In some embodiments, determining the privacy priority of each data node to be backed up according to the corresponding sensitive feature is to input the sensitive features of each data node to be backed up into an evaluation model for evaluation, and then obtain the privacy priority of each data node to be backed up.

[0067] It should be noted that in this application, the privacy priority represents the importance degree of the sensitive features of the corresponding data node to be backed up in terms of privacy protection. Specifically, the sensitive features of each data node to be backed up can be input into an evaluation model for evaluation respectively. The evaluation model can evaluate the sensitivity degree of the sensitive features of the data node. Different types of sensitive features will have different priorities. For example, the privacy priorities of personal identity information and financial information are relatively high, while the privacy priorities of business logs or technical documents are relatively low. The privacy priorities of each data node to be backed up can be obtained in the above way.

[0068] It should be noted that by extracting the sensitive features of each data node to be backed up respectively and determining the privacy priority of the data node according to these sensitive features, it is possible to effectively identify which data need to be protected preferentially, so as to formulate differentiated backup and disaster tolerance strategies. Sensitive data usually means higher risks and compliance requirements. Therefore, more stringent protection measures are given to the data nodes with high privacy priority, such as more frequent backups, strong encryption, and distributed storage. While the data with low privacy priority can adopt conventional backup strategies, thus saving resources and costs.

[0069] In step 103, obtain the access logs of each data node to be backed up, and determine the access anomaly factors of each data node to be backed up based on the corresponding access logs.

[0070] When specifically implemented, the access logs of each data node to be backed up can be obtained from the enterprise database;

[0071] Preferably, in some embodiments, refer to Figure 3 As shown, this figure is an exemplary flowchart for determining the access anomaly factors of each data node to be backed up in some embodiments of the present application. In this embodiment, determining the access anomaly factors of each data node to be backed up based on the corresponding access logs can be implemented by the following steps:

[0072] In step 1031, for each data node to be backed up, extract multiple access characteristics of the data node to be backed up from the access log of the data node to be backed up;

[0073] In step 1032, determine the access anomaly factor of the data node to be backed up through all the access characteristics, and thus obtain the access anomaly factors of each data node to be backed up.

[0074] When specifically implemented, first, for each data node to be backed up, multiple access characteristics of the data node to be backed up can be extracted from the access log of the data node to be backed up by means of statistics. The access characteristics in the present application include: the access frequency of text data, the access interval, the file modification frequency, and the file size change rate. Then, the access anomaly factor of the data node to be backed up can be determined through all the access characteristics. Among them, the access anomaly factor represents the degree of access anomaly to the text data of the data node to be backed up. In actual implementation, for each access characteristic, the data mean and data standard deviation of the access characteristic can be calculated, and the current value of the access characteristic can be obtained. Divide the difference between the current value and the data mean of the access characteristic by the data standard deviation, and use the result as the anomaly factor of the access characteristic. This anomaly factor represents the anomaly degree of the corresponding access characteristic. Through the above method, the anomaly factors of each access characteristic can be obtained. Weighted sum the anomaly factors of each access characteristic. Among them, the corresponding weights can be set according to historical experience, and the final result can be used as the access anomaly factor of the corresponding data node to be backed up. Through the above method, the access anomaly factors of each data node to be backed up can be obtained.

[0075] It should be noted that obtaining the access logs of each data node to be backed up and determining the access anomaly factor based on the access logs can effectively identify potential security risks and abnormal access behaviors, help evaluate which data nodes to be backed up may face higher security threats, so as to be able to prioritize the backup of the data nodes to be backed up with higher risks, thereby ensuring that critical data can be quickly restored in the event of an attack or other disaster events and reducing the risk of data loss.

[0076] In step 104, the backup urgency of each data node to be backed up in the enterprise database is determined through the corresponding privacy priority and the corresponding access anomaly factor, and a backup disaster recovery policy for each data node to be backed up is generated according to the corresponding backup urgency.

[0077] In some embodiments, the backup urgency of each data node to be backed up in the enterprise database is determined through the corresponding privacy priority and the corresponding access anomaly factor; it should be noted that in this application, the backup urgency represents the urgency degree of backing up the corresponding data node to be backed up; specifically, for each data node to be backed up in the enterprise database, the privacy priority of the data node to be backed up and the access anomaly factor can be weighted and summed, where the corresponding weight can be set through data analysis and historical experiments, so as to use the calculation result as the backup urgency of the data node to be backed up. Through the above method, the backup urgency of each data node to be backed up in the enterprise database can be obtained.

[0078] In some embodiments, generating a backup disaster recovery policy for each data node to be backed up according to the corresponding backup urgency can specifically adopt the following method, that is:

[0079] Determine the backup priority group of each data node to be backed up according to the corresponding backup urgency;

[0080] Determine the backup disaster recovery policy of each data node to be backed up through the corresponding backup priority group.

[0081] Specifically, first, the backup priority group of each data node to be backed up can be determined according to the corresponding backup urgency, where the backup priority group represents the priority interval where the backup urgency of the corresponding data node to be backed up is located, and the backup urgency of the data node to be backed up can be mapped to a pre-set priority table. In this application, the priority table includes an emergency priority group, a normal priority group, and a low priority group; then, the backup disaster recovery policy of each data node to be backed up can be determined through the corresponding backup priority group, that is, once the data node to be backed up is divided into different backup priority groups according to the backup urgency, the backup disaster recovery policy of each data node to be backed up can be determined according to the characteristics of each backup priority group; different priority groups will adopt different backup frequencies, storage methods, encryption levels and other policies. For example:

[0082] The backup and disaster recovery strategy for the emergency priority group is as follows: perform real-time backup or hourly incremental backup to ensure data timeliness and avoid data loss; adopt off-site backup (such as cloud data centers in different geographical locations) and multi-cloud backup (distribute backup data across multiple cloud platforms) to ensure data recoverability even in the event of a disaster; perform high-intensity encryption on the backup data (such as AES-256 encryption) to ensure data security; set up real-time or near-real-time recovery mechanisms to quickly recover important data in the event of a disaster and avoid long-term business interruption; regularly verify the integrity of the backup data to ensure the accuracy and recoverability of the backup data.

[0083] The backup and disaster recovery strategy for the normal priority group is as follows: adopt regular backup (such as daily incremental backup) to ensure data integrity and reliability; combine data backup to the local data center and cloud storage to balance cost and security; perform conventional encryption on the data (such as AES-128 encryption) to ensure that the data backup is not leaked; set a longer recovery time so that data can be recovered in the event of a disaster without affecting normal operations; regularly perform backup integrity checks, but not as frequently as the high-priority group.

[0084] The backup and disaster recovery strategy for the low priority group is as follows: adopt low-frequency backup (such as full backup once a week or on-demand backup); store the backup data in the local data center or cloud storage, and off-site backup is not required; perform conventional encryption on the backup data, but there is no need for excessive encryption; in the event of a disaster, the time to recover these low-priority data can be appropriately extended, but it is still necessary to ensure that the data can be recovered; occasionally verify the backup data, but do not need to execute it frequently.

[0085] It should be noted that by dividing the backup priority groups of data nodes according to the backup urgency and formulating corresponding backup and disaster recovery strategies for each priority group, enterprises can achieve more efficient and targeted backup management, thereby ensuring that the most critical and sensitive data is protected first, while unimportant data is reasonably backed up under limited resources. Through flexible disaster recovery strategies and backup priority division, enterprises can not only optimize the use of backup resources, but also quickly recover critical data in the event of a disaster, reduce the risk of business interruption, and improve the efficiency of data backup and recovery.

[0086] It can be seen that in this application, first, by obtaining the text data of each data node in the enterprise database and calculating the relevant redundancy between each data node, redundant backups can be significantly reduced, storage resources can be optimized, and costs can be reduced. Screening the data nodes to be backed up based on the relevant redundancy can avoid unnecessary repeated backups, so that resources can be allocated more efficiently, and faster data recovery and disaster tolerance response can be achieved. Then, by separately extracting the sensitive features of each data node to be backed up and determining the privacy priority of the data node based on these sensitive features, it is possible to effectively identify which data needs to be protected first, so as to formulate differentiated backup and disaster tolerance strategies. Secondly, obtaining the access logs of each data node to be backed up and determining the access anomaly factor based on the access logs can effectively identify potential security risks and abnormal access behaviors, and help evaluate which data nodes to be backed up may face higher security threats, so that the data nodes to be backed up with higher risks can be preferentially backed up. Finally, by dividing the backup priority groups of the data nodes according to the backup urgency and formulating corresponding backup and disaster tolerance strategies for each priority group, the efficiency of the data backup process can be improved.

[0087] In summary, the technical solution adopted in this application can determine the backup urgency of the data and generate a disaster tolerance strategy to improve the efficiency of the data backup process.

[0088] In addition, on the other hand of this application, in some embodiments, this application provides a backup and disaster tolerance system combined with AI technology. The backup and disaster tolerance system combined with AI technology includes a disaster tolerance strategy generation unit. Refer to Figure 4 , which is a schematic diagram of the exemplary hardware and / or software of the disaster tolerance strategy generation unit shown in some embodiments of this application. The disaster tolerance strategy generation unit 400 includes: a data acquisition module 401, a priority determination module 402, an abnormal access determination module 403, and a strategy generation module 404, which are described as follows:

[0089] The data acquisition module 401. In this application, the data acquisition module 401 is mainly used to obtain the text data of each data node in the enterprise database, determine the relevant redundancy between each data node in the enterprise database according to all the text data, and screen out all the data nodes to be backed up in the enterprise database based on each relevant redundancy;

[0090] The priority determination module 402. In this application, the priority determination module 402 is mainly used to separately extract the sensitive features of each data node to be backed up, and then determine the privacy priority of each data node to be backed up based on the corresponding sensitive features;

[0091] An abnormal access determination module 403. In this application, the abnormal access determination module 403 is mainly used to obtain the access logs of each data node to be backed up, and determine the access abnormal factors of each data node to be backed up based on the corresponding access logs;

[0092] A policy generation module 404. In this application, the policy generation module 404 is mainly used to determine the backup urgency of each data node to be backed up in the enterprise database through the corresponding privacy priority and the corresponding access abnormal factor, and generate a backup disaster recovery policy for each data node to be backed up according to the corresponding backup urgency.

[0093] The above text has introduced in detail an example of a backup disaster recovery method and system combining AI technology provided by the embodiments of this application. It can be understood that, in order to implement the above functions, the corresponding device includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraint conditions of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0094] In some embodiments, this application also provides a computer device, the computer device includes a memory and a processor, the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the computer device executes the above-mentioned backup disaster recovery method combining AI technology.

[0095] In some embodiments, refer to Figure 5 , the dotted line in this figure indicates that this unit or this module is optional. This figure is a schematic structural diagram of a computer device for a backup disaster recovery method combining AI technology provided by the embodiments of this application. The above-mentioned backup disaster recovery method combining AI technology in the above embodiments can be implemented by Figure 5 the computer device shown. The computer device 500 includes at least one processor 501, a memory 502, and at least one communication unit 505. The computer device 500 can be a terminal device, a server, or a chip.

[0096] The processor 501 can be a general-purpose processor or a special-purpose processor. For example, the processor 501 can be a central processing unit (CPU), and the CPU can be used to control the computer device 500, execute software programs, and process the data of the software programs. The computer device 500 can also include a communication unit 505 for implementing signal input (reception) and output (transmission).

[0097] For example, the computer device 500 can be a chip, and the communication unit 505 can be the input and / or output circuit of the chip, or the communication unit 505 can be the communication interface of the chip. The chip can be a component of a terminal device, a network device, or other devices.

[0098] For another example, the computer device 500 can be a terminal device or a server, and the communication unit 505 can be the transceiver of the terminal device or the server, or the communication unit 505 can be the transceiver circuit of the terminal device or the server.

[0099] The computer device 500 can include one or more memories 502 on which a program 504 is stored. The program 504 can be run by the processor 501 to generate instructions 503, enabling the processor 501 to execute the methods described in the above method embodiments according to the instructions 503. Optionally, data (such as a target audit model) can also be stored in the memory 502. Optionally, the processor 501 can also read the data stored in the memory 502. This data can be stored at the same storage address as the program 504, or it can be stored at a different storage address from the program 504.

[0100] The processor 501 and the memory 502 can be set separately or integrated together. For example, they can be integrated on a system on chip (SOC) of a terminal device.

[0101] It should be understood that the steps of the above method embodiments can be completed by the logic circuit in hardware form or the instructions in software form in the processor 501. The processor 501 can be a central processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices. For example, discrete gate, transistor logic devices, or discrete hardware components.

[0102] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0103] For example, in some embodiments, the present application further provides a computer-readable storage medium storing instructions or code, which, when run on a computer, cause the computer to implement the above-mentioned backup and disaster recovery method combining AI technology.

[0104] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments as well as all changes and modifications falling within the scope of the present application.

[0105] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. A backup and disaster recovery method combined with AI technology, characterized in that, The disaster recovery method includes the following steps: Obtain the text data of each data node in the enterprise database, determine the relevant redundancy between each data node in the enterprise database according to all the text data, and screen out all the data nodes to be backed up in the enterprise database based on each relevant redundancy; Extract the sensitive features of each data node to be backed up respectively, and then determine the privacy priority of each data node to be backed up according to the corresponding sensitive features; Obtain the access logs of each data node to be backed up, and determine the access anomaly factor of each data node to be backed up based on the corresponding access logs; Determine the backup urgency of each data node to be backed up in the enterprise database through the corresponding privacy priority and the corresponding access anomaly factor, and generate the backup disaster recovery strategy for each data node to be backed up according to the corresponding backup urgency.

2. The backup and disaster recovery method combining AI technology according to claim 1, wherein, Determining the relevant redundancy between each data node in the enterprise database according to all the text data specifically includes: Generate the hash signature set corresponding to each data node in the enterprise database according to the corresponding text data; Select a data node from all the data nodes in the enterprise database, and obtain the hash signature set corresponding to the selected data node; Determine the data redundancy between the selected data node and other data nodes in the enterprise database based on the hash signature set corresponding to the selected data node; Determine the relevant redundancy between the selected data node and other data nodes in the enterprise database through the corresponding data redundancy, and then obtain the relevant redundancy between each data node in the enterprise database.

3. A backup and disaster recovery method combined with AI technology according to claim 1, characterized in that, Screening out all the data nodes to be backed up in the enterprise database based on each relevant redundancy specifically includes: Obtain the preset redundancy threshold, and then screen out all the data node pairs with a relevant redundancy higher than the redundancy threshold; Exclude one data node from each of the screened data node pairs respectively, and then use the remaining data nodes and the data nodes in each data node pair with a relevant redundancy not higher than the redundancy threshold as the data nodes to be backed up in the enterprise database, and then obtain all the data nodes to be backed up in the enterprise database.

4. A backup and disaster recovery method combining AI technology according to claim 1, characterized in that, Use natural language processing technology and a classification model based on machine learning to extract the sensitive features of each data node to be backed up respectively.

5. A backup and disaster recovery method combining AI technology according to claim 1, characterized in that, Determining the privacy priority of each data node to be backed up according to the corresponding sensitive features is to input the sensitive features of each data node to be backed up into the evaluation model for evaluation respectively, and then obtain the privacy priority of each data node to be backed up.

6. A backup and disaster recovery method combining AI technology according to claim 1, characterized in that, Determining the access anomaly factor of each data node to be backed up based on the corresponding access logs specifically includes: For each data node to be backed up, extract multiple access features of the data node to be backed up from the access log of the data node to be backed up; Determine the access anomaly factor of the data node to be backed up through all the access features, and then obtain the access anomaly factor of each data node to be backed up.

7. A backup and disaster recovery method combining AI technology according to claim 1, characterized in that Generating the backup disaster recovery strategy for each data node to be backed up according to the corresponding backup urgency specifically includes: Determine the backup priority group of each data node to be backed up according to the corresponding backup urgency; Determine the backup disaster recovery strategy of each data node to be backed up through the corresponding backup priority group.

8. A backup and disaster recovery system incorporating AI technology, which is used to execute a backup and disaster recovery method incorporating AI technology as described in any one of claims 1 to 7. The backup and disaster recovery system incorporating AI technology includes a disaster recovery policy generation unit, and is characterized in that, The disaster recovery policy generation unit includes: A data acquisition module, configured to acquire the text data of each data node in the enterprise database, determine the relevant redundancy between each data node in the enterprise database according to all the text data, and filter out all the data nodes to be backed up in the enterprise database based on each relevant redundancy; A priority determination module, configured to respectively extract the sensitive features of each data node to be backed up, and then determine the privacy priority of each data node to be backed up according to the corresponding sensitive features; An abnormal access determination module, configured to acquire the access logs of each data node to be backed up, and determine the access abnormal factor of each data node to be backed up based on the corresponding access log; A policy generation module, configured to determine the backup urgency of each data node to be backed up in the enterprise database through the corresponding privacy priority and the corresponding access abnormal factor, and generate a backup disaster recovery policy for each data node to be backed up according to the corresponding backup urgency.

9. A computer device, characterized in that, The computer device includes a memory and a processor. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the computer device executes the backup disaster recovery method combining AI technology according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Instructions or codes are stored in the computer-readable storage medium. When the instructions or codes run on a computer, the computer is caused to execute the backup disaster recovery method combining AI technology according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Data backup method and device and computer equipment

    CN118860738A

  • Data security and privacy protection method and system

    CN119646838A

  • Hard disk data protection method and device based on artificial intelligence

    CN119646906A

  • Terminal data automatic backup and recovery method and system

    CN119668939A

  • Protecting data based on a sensitivity level for the data

    US20200320215A1