Data security protection method based on big data

CN120104403AInactive Publication Date: 2025-06-06ANQING XINGLU TECH CONSULTING CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510128047.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-04
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure CN120104403A_ABST
    Figure CN120104403A_ABST
Patent Text Reader

Abstract

The invention discloses a data security protection method based on big data, which relates to the technical field of data security protection, and comprises the following steps: calculating an initial backup coefficient of each piece of sub-data, if a server does not have backup data of a user, marking the user as a new backup user, and if not, marking the user as a new backup user; the sub-data are sequentially uploaded to a server for backup according to the initial backup coefficients from large to small; if the server has the backup information of the user, the predicted backup completion time is calculated, if the predicted backup completion time is smaller than the preset backup time, a real-time backup strategy is adopted for the current backup, and the difference data of each piece of sub-data are sequentially uploaded to the server according to the descending order of the real-time update coefficient; if it is predicted that the backup completion time is longer than the preset backup time, a non-real-time backup strategy is adopted for this backup, the non-real-time update coefficients are sequentially uploaded to the server from large to small, and the user backup data are updated. In this way, the system can more intelligently determine which data need to be backed up preferentially, and the overall backup efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data security protection, and specifically to a data security protection method based on big data. Background Art

[0002] With the rapid development of information technology, big data has become an indispensable and important resource for modern enterprises and organizations. With its massive, diversified, complex and rapid characteristics, big data provides strong support for decision-making analysis, business optimization and innovative development in all walks of life. However, the widespread application of big data has also brought unprecedented data security challenges. Frequent security incidents such as data leakage, illegal access, and malicious tampering seriously threaten the interests of enterprises and users. At present, there are many data security backup technologies on the market, such as snapshots, data mirroring, RAID, and off-site backup. However, these technologies still have many limitations when dealing with big data backup needs: low backup efficiency. In the face of massive data, traditional backup methods often take too long and are difficult to meet the needs of rapid recovery; large resource consumption. The backup process requires a large amount of storage space and computing resources, which increases the operating costs of enterprises; insufficient flexibility. It is impossible to perform intelligent backup according to the real-time changes and importance of data, resulting in redundant backup data or omission of key data; security needs to be improved. There is still a risk of theft or tampering during the storage and transmission of backup data.

[0003] In the Chinese invention application with application publication number CN118277164A, a data backup and recovery method and system based on cloud computing are disclosed, including step S1: obtaining allocation parameters of files to be backed up; dividing files to be backed up into write hot data, read hot data and cold data; step S2: obtaining local parameters; backing up write hot data or read hot data locally according to the file size of the files to be backed up; step S3: configuring a non-local server computing platform according to local parameters to back up cold data; creating a snapshot for the files to be backed up; judging whether the files to be backed up are complete; if not, restoring the files to be backed up according to the snapshot to obtain a comparison file; if the data of the files to be backed up are complete, no processing is performed; step S4: summarizing the comparison files and feeding them back to the user;

[0004] In the above invention application, data is divided into important data (read hot data and write hot data) and cold data according to the frequency of data use by users. The important data (read hot data and write hot data) is backed up to the local server, and the cold data is backed up to the non-local server. According to the data processing rate of the local server, the data sending rate of the non-local server is adjusted to improve the user experience. However, in the actual backup process, when the amount of changed data is large and the network speed is slow, some write hot data may undergo multiple changes before being uploaded to the server. At this time, if such data is uploaded first, it will lead to frequent data uploads, causing the data to occupy too many resources during peak hours, while the remaining data has been waiting to be uploaded to the server for backup, resulting in low overall backup efficiency, and the occupation of system resources (such as CPU, memory, disk I / O, etc.) will also increase, resulting in a decline in the overall performance of the system and affecting the normal operation of other businesses.

[0005] To this end, the present invention provides a data security protection method based on big data. Summary of the invention

[0006] 1. Technical issues to be solved

[0007] In view of the deficiencies in the prior art, the present invention provides a data security protection method based on big data. The present invention analyzes the real-time backup status of users, marks it as new backup, real-time backup and non-real-time backup, and adopts different backup strategies. It dynamically adjusts according to changes in user data to adapt to different backup needs. It can more effectively utilize the storage and bandwidth resources of the server, which enables the system to more intelligently decide which data needs to be backed up first and which data can be processed later, thereby improving the overall efficiency of the backup and avoiding occupying too many resources during peak hours, thereby solving the technical problems recorded in the background technology.

[0008] (II) Technical solution

[0009] To achieve the above objectives, the present invention is implemented through the following technical solutions: A data security protection method based on big data, comprising the following steps:

[0010] Obtain user data from the user end and divide it into sub-data, and obtain the data volume Ls of each sub-data i , historical access time Fs ia and historical change time Gs ib , calculate the initial backup coefficient Cq of each sub-data i If the server does not have the backup data of the user, the user is marked as a new backup user, and each sub-data is backed up according to the initial backup coefficient Cq i Upload the server backups in order from largest to smallest;

[0011] If the server already has the backup information of the user, compare the server backup data with the user data to obtain the difference data volume Cy of each sub-data i , calculate the predicted backup completion time Tt, if the predicted backup completion time Tt is less than the preset backup time, then the real-time backup strategy is adopted for this backup, based on the difference data volume Cy of each sub-data i and the initial backup coefficient Cq i Calculate the real-time update coefficient Sg for each sub-data i , the difference data of each sub-data is updated according to the real-time update coefficient Sg i Upload the data to the server in order from largest to smallest, and update the user backup data;

[0012] If the predicted backup completion time Tt is greater than the preset backup time, the non-real-time backup strategy is adopted for this backup, based on the real-time change time Bs of each sub-data. ic and real-time change difference data volume Bc ic , calculate the real-time change coefficient Bx of each sub-data i , and calculate the non-real-time update coefficient Fg for each sub-data i , based on the non-real-time update coefficient Fg i Sort the data from largest to smallest and upload them to the server to update the user backup data.

[0013] Furthermore, the user data is obtained from the user end, and the user data is divided into sub-data according to the data storage directory, and the data volume Ls of each sub-data is obtained through a database management tool (such as MySQL Workbench, Navicat, etc.) i , and use system monitoring tools (such as Nagios, Zabbix, Prometheus, etc.) to obtain the historical access time Fs of each sub-data ia and historical change time Gs ib .

[0014] Database management tools (such as MySQL Workbench, Navicat, etc.) are important tools for database administrators and developers when performing database management and maintenance. These tools usually provide a graphical interface to facilitate users to create, modify, delete, backup, restore, optimize and monitor databases.

[0015] System monitoring tools (such as Nagios, Zabbix, Prometheus, etc.) System monitoring tools are an indispensable and important part of IT operation and maintenance. They can monitor the performance, health and security of the system in real time, and provide system administrators with necessary data support to ensure the stable operation of the system.

[0016] Further, the data volume Ls of each sub-data is obtained i , historical access time Fs ia and historical change time Gs ib , after dimensionless processing, calculate the initial backup coefficient Cq of each sub-data i :

[0017]

[0018] Among them, i represents the number of each sub-data, i=1, 2, ..., k, a represents the time sequence number of each access to the same sub-data, a=1, 2, ..., m, b represents the time sequence number of each historical change of the same sub-data, b=1, 2, ..., n.

[0019] Furthermore, the user's server account is used to log in to the corresponding cloud service website or application to check whether there is a device backup list for the user. If the server does not have the user's backup data, the user is marked as a new backup user, and each sub-data is backed up according to the initial backup coefficient Cq i Upload the server backups in order from largest to smallest.

[0020] Furthermore, if the server already has the backup information of the user, use a database comparison tool (such as DBeaver, Toad Data Point, etc.) or data comparison software (such as Beyond Compare, WinMerge, etc.) to compare the server backup data with the user-side data to obtain the difference data volume Cy of each sub-data. i . And use network performance testing tools (such as iperf, PsPing, etc.) to detect the maximum network communication volume Tx on the user side.

[0021] Network performance testing tools (such as iperf, PsPing, etc.) are important tools for evaluating and optimizing network performance. They can help users understand key indicators such as network transmission speed, latency, packet loss rate, etc., so as to discover potential network problems and optimize them.

[0022] Further, obtain the difference data volume Cy of each sub-data i and the maximum network traffic Tx, calculate the predicted backup completion time Tt:

[0023]

[0024] Among them, the difference data volume Cy i The unit of maximum network traffic is kilobyte (KB), and the unit of maximum network traffic Tx is kilobyte per second (KB / S).

[0025] Furthermore, if the predicted backup completion time Tt is less than the preset backup time, the real-time backup strategy is adopted for this backup to obtain the differential data volume Cy of each sub-data. i and the initial backup coefficient Cq i , after linear normalization, calculate the real-time update coefficient Sg of each sub-data i :

[0026] S i =Cy i *C i

[0027] The preset backup time may be any one of 1, 2, ..., 30 minutes, which is preset by the user. If the user does not preset it, the default is 10 minutes.

[0028] Furthermore, when the real-time backup strategy is adopted, the difference data of each sub-data is updated according to the real-time update coefficient Sg i Sort the data from largest to smallest and upload them to the server to update the user backup data.

[0029] Furthermore, if the predicted backup completion time Tt is greater than the preset backup time, the non-real-time backup strategy is adopted for this backup, and the system monitoring tool is used to obtain the real-time change time Bs of each sub-data after the user starts the backup. ic and real-time change difference data volume Bc ic , calculate the real-time change coefficient Bx of each sub-data i :

[0030]

[0031] Here, c represents the time sequence number of each real-time change of the same sub-data, c=1, 2, ..., w.

[0032] Furthermore, obtain the real-time change coefficient Bx of each sub-data i and the initial backup coefficient Cq i After linear normalization, the non-real-time update coefficient Fg of each sub-data is calculated i :

[0033] F i =Bx i *C i

[0034] When adopting the non-real-time backup strategy, the difference data of each sub-data is updated according to the non-real-time update coefficient Fg i Sort the data from largest to smallest and upload them to the server to update the user backup data.

[0035] (III) Beneficial effects

[0036] The present invention provides a data security protection method based on big data, which has the following beneficial effects:

[0037] 1. Obtain user data from the user end and divide it into sub-data, and obtain the data volume Ls of each sub-data i , historical access time Fs ia and historical change time Gs ib , calculate the initial backup coefficient Cq of each sub-data i If the server does not have the backup data of the user, the user is marked as a new backup user, and each sub-data is backed up according to the initial backup coefficient Cq i Uploading server backups in order from largest to smallest can identify which data is more critical or active for users. This helps prioritize the backup of important or frequently changing data, ensuring the timeliness and integrity of the data. Even if the initial data backup is interrupted, the user's critical data can be safely protected.

[0038] 2. If the server already has the backup information of the user, compare the server backup data with the user data to obtain the difference data volume Cy of each sub-data i , calculate the predicted backup completion time Tt, if the predicted backup completion time Tt is less than the preset backup time, then the real-time backup strategy is adopted for this backup, based on the difference data volume Cy of each sub-data i and the initial backup coefficient Cq i Calculate the real-time update coefficient Sg for each sub-data i , the difference data of each sub-data is updated according to the real-time update coefficient Sg i Uploading the data to the server in order from largest to smallest and updating the user backup data can make more efficient use of the server's storage and bandwidth resources. Important data is uploaded first, while less important data can be backed up when resources are more abundant, avoiding excessive resource usage during peak hours. Dynamic adjustments can be made based on changes in user data to meet different backup needs.

[0039] 3. If the predicted backup completion time Tt is greater than the preset backup time, the non-real-time backup strategy is adopted for this backup, based on the real-time change time Bs of each sub-data. ic and real-time change difference data volume Bc ic , calculate the real-time change coefficient Bx of each sub-data i , and calculate the non-real-time update coefficient Fg for each sub-data i , based on the non-real-time update coefficient Fg i The data is uploaded to the server in order from largest to smallest, and the user's backup data is updated. This enables the system to more intelligently decide which data needs to be backed up first and which data can be processed later, thereby improving the overall efficiency of the backup. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 The present invention is a flowchart of a data security protection method based on big data. DETAILED DESCRIPTION

[0041] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0042] See also Figure 1 The present invention provides a data security protection method based on big data, comprising the following steps:

[0043] Step 1: Obtain user data from the user end and divide it into sub-data, and obtain the data volume Ls of each sub-data i , historical access time Fs ia and historical change time Gs ib , calculate the initial backup coefficient Cq of each sub-data i If the server does not have the backup data of the user, the user is marked as a new backup user, and each sub-data is backed up according to the initial backup coefficient Cq i Sort the backup files from largest to smallest and upload them to the server.

[0044] The step 1 includes the following contents:

[0045] Step 101: Obtain user data from the user end, and divide the user data into sub-data according to the data storage directory, and obtain the data volume Ls of each sub-data through a database management tool (such as MySQL Workbench, Navicat, etc.) i , and use system monitoring tools (such as Nagios, Zabbix, Prometheus, etc.) to obtain the historical access time Fs of each sub-data ia and historical change time Gs ib .

[0046] Database management tools (such as MySQL Workbench, Navicat, etc.) are important tools for database administrators and developers when performing database management and maintenance. These tools usually provide a graphical interface to facilitate users to create, modify, delete, backup, restore, optimize and monitor databases.

[0047] System monitoring tools (such as Nagios, Zabbix, Prometheus, etc.) System monitoring tools are an indispensable and important part of IT operation and maintenance. They can monitor the performance, health and security of the system in real time, and provide system administrators with necessary data support to ensure the stable operation of the system.

[0048] Step 102: Obtain the data volume Ls of each sub-data i , historical access time Fs ia and historical change time Gs ib , after dimensionless processing, calculate the initial backup coefficient Cq of each sub-data i :

[0049]

[0050] Among them, i represents the number of each sub-data, i=1, 2, ..., k, a represents the time sequence number of each access to the same sub-data, a=1, 2, ..., m, b represents the time sequence number of each historical change of the same sub-data, b=1, 2, ..., n.

[0051] Step 103: Use the user's server account to log in to the corresponding cloud service website or application to check whether there is a device backup list for the user. If the server does not have the user's backup data, the user is marked as a new backup user, and each sub-data is backed up according to the initial backup coefficient Cq. i Sort the backup files from largest to smallest and upload them to the server.

[0052] When using, combine the contents in steps 101 to 103:

[0053] Obtain user data from the user end and divide it into sub-data, and obtain the data volume Ls of each sub-data i , historical access time Fs ia and historical change time Gs ib , calculate the initial backup coefficient Cq of each sub-data i If the server does not have the backup data of the user, the user is marked as a new backup user, and each sub-data is backed up according to the initial backup coefficient Cq i Uploading server backups in order from largest to smallest can identify which data is more critical or active for users. This helps prioritize the backup of important or frequently changing data, ensuring the timeliness and integrity of the data. Even if the initial data backup is interrupted, the user's critical data can be safely protected.

[0054] Step 2: If the server already has the backup information of the user, compare the server backup data with the user data to obtain the difference data volume Cy of each sub-data i, calculate the predicted backup completion time Tt, if the predicted backup completion time Tt is less than the preset backup time, then the real-time backup strategy is adopted for this backup, based on the difference data volume Cy of each sub-data i and the initial backup coefficient Cq i Calculate the real-time update coefficient Sg for each sub-data i , the difference data of each sub-data is updated according to the real-time update coefficient Sg i Sort the data from largest to smallest and upload them to the server to update the user backup data.

[0055] The step 2 includes the following contents:

[0056] Step 201: If the server already has the backup information of the user, use a database comparison tool (such as DBeaver, Toad Data Point, etc.) or data comparison software (such as Beyond Compare, WinMerge, etc.) to compare the server backup data with the user-side data to obtain the difference data volume Cy of each sub-data. i . And use network performance testing tools (such as iperf, PsPing, etc.) to detect the maximum network communication volume Tx on the user side.

[0057] Network performance testing tools (such as iperf, PsPing, etc.) are important tools for evaluating and optimizing network performance. They can help users understand key indicators such as network transmission speed, latency, packet loss rate, etc., so as to discover potential network problems and optimize them.

[0058] Step 202: Obtain the difference data volume Cy of each sub-data i and the maximum network traffic Tx, calculate the predicted backup completion time Tt:

[0059]

[0060] Among them, the difference data volume Cy i The unit of maximum network traffic is kilobyte (KB), and the unit of maximum network traffic Tx is kilobyte per second (KB / S).

[0061] Step 203: If the predicted backup completion time Tt is less than the preset backup time, the current backup adopts a real-time backup strategy to obtain the differential data volume Cy of each sub-data. i and the initial backup coefficient Cq i , after linear normalization, calculate the real-time update coefficient Sg of each sub-data i :

[0062] S i =Cy i *C i

[0063] The preset backup time may be any one of 1, 2, ..., 30 minutes, which is preset by the user. If the user does not preset it, the default is 10 minutes.

[0064] When adopting the real-time backup strategy, the difference data of each sub-data is updated according to the real-time update coefficient Sg i Sort the data from largest to smallest and upload them to the server to update the user backup data.

[0065] Step 204: Enable monitoring and logging functions in the backup task, record the difference data volume Sz and the actual backup data volume Bz in real time during the backup process, and calculate the user's redundant backup data ratio Pg:

[0066]

[0067] The total differential data volume Sz refers to the differential data volume between the data stored on the server before the backup starts and the data stored on the server after the backup is completed, and the actual backup data volume Bz refers to the total data volume transmitted from the user end to the server during the backup process.

[0068] Step 205: If the user's redundant backup data ratio Pg exceeds 1, and the predicted backup completion time Tt is greater than 0.8 of the preset backup time, adjust the preset backup time so that the current predicted backup completion time Tt is greater than the preset backup time, meet the non-real-time backup policy conditions, and start the non-real-time backup policy.

[0069] When using, combine the contents in steps 201 to 205:

[0070] If the server already has the backup information of the user, compare the server backup data with the user data to obtain the difference data volume Cy of each sub-data i , calculate the predicted backup completion time Tt, if the predicted backup completion time Tt is less than the preset backup time, then the real-time backup strategy is adopted for this backup, based on the difference data volume Cy of each sub-data i and the initial backup coefficient Cq i Calculate the real-time update coefficient Sg for each sub-data i , the difference data of each sub-data is updated according to the real-time update coefficient Sg i Uploading the data to the server in order from largest to smallest and updating the user backup data can make more efficient use of the server's storage and bandwidth resources. Important data is uploaded first, while less important data can be backed up when resources are more abundant, avoiding excessive resource usage during peak hours. Dynamic adjustments can be made based on changes in user data to meet different backup needs.

[0071] Step 3: If the predicted backup completion time Tt is greater than the preset backup time, the non-real-time backup strategy is adopted for this backup, based on the real-time change time Bs of each sub-data. ic and real-time change difference data volume Bc ic , calculate the real-time change coefficient Bx of each sub-data i , and calculate the non-real-time update coefficient Fg for each sub-data i , based on the non-real-time update coefficient Fg i Sort the data from largest to smallest and upload them to the server to update the user backup data.

[0072] The step three includes the following contents:

[0073] Step 301: If the predicted backup completion time Tt is greater than the preset backup time, the current backup adopts a non-real-time backup strategy and uses a system monitoring tool to obtain the real-time change time Bs of each sub-data after the user starts the backup. ic and real-time change difference data volume Bc ic , calculate the real-time change coefficient Bx of each sub-data i :

[0074]

[0075] Here, c represents the time sequence number of each real-time change of the same sub-data, c=1, 2, ..., w.

[0076] Step 302: Obtain the real-time change coefficient Bx of each sub-data i and the initial backup coefficient Cq i After linear normalization, the non-real-time update coefficient Fg of each sub-data is calculated i :

[0077] F i =Bx i *C i

[0078] When adopting the non-real-time backup strategy, the difference data of each sub-data is updated according to the non-real-time update coefficient Fg i Sort the data from largest to smallest and upload them to the server to update the user backup data.

[0079] When used, combine the contents in steps 301 and 302:

[0080] If the predicted backup completion time Tt is greater than the preset backup time, the non-real-time backup strategy is adopted for this backup, based on the real-time change time Bs of each sub-data. ic and real-time change difference data volume Bc ic , calculate the real-time change coefficient Bx of each sub-data i, and calculate the non-real-time update coefficient Fg for each sub-data i , based on the non-real-time update coefficient Fg i The data is uploaded to the server in order from largest to smallest, and the user's backup data is updated. This enables the system to more intelligently decide which data needs to be backed up first and which data can be processed later, thereby improving the overall efficiency of the backup.

[0081] The above embodiments may be implemented in whole or in part by software, hardware, firmware or any other combination thereof. When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product. A person of ordinary skill in the art may appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein may be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution.

[0082] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0083] The above description is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.

Claims

1. A data security protection method based on big data, characterized by: The steps include: Obtain user data from the user end and divide it into sub-data, and obtain the data volume Ls of each sub-data i , historical access time Fs ia and historical change time Gs ib , calculate the initial backup coefficient Cq of each sub-data i If the server does not have the backup data of the user, the user is marked as a new backup user, and each sub-data is backed up according to the initial backup coefficient Cq i Upload the server backups in order from largest to smallest; If the server already has the backup information of the user, compare the server backup data with the user data to obtain the difference data volume Cy of each sub-data i , calculate the predicted backup completion time Tt, if the predicted backup completion time Tt is less than the preset backup time, then the real-time backup strategy is adopted for this backup, based on the difference data volume Cy of each sub-data i and the initial backup coefficient Cq i Calculate the real-time update coefficient Sg for each sub-data i , the difference data of each sub-data is updated according to the real-time update coefficient Sg i Upload the data to the server in order from largest to smallest, and update the user backup data; If the predicted backup completion time Tt is greater than the preset backup time, the non-real-time backup strategy is adopted for this backup, based on the real-time change time Bs of each sub-data. ic and real-time change difference data volume Bc ic , calculate the real-time change coefficient Bx of each sub-data i , and calculate the non-real-time update coefficient Fg for each sub-data i , based on the non-real-time update coefficient Fg i Sort the data from largest to smallest and upload them to the server to update the user's backup data.

2. According to a data security protection method based on big data according to claim 1, it is characterized by: Obtain user data from the user end, and divide the user data into sub-data according to the data storage directory, and obtain the data volume Ls of each sub-data through the database management tool i , and use the system monitoring tool to obtain the historical access time Fs of each sub-data ia and historical change time Gs ib .

3. The data security protection method based on big data according to claim 2 is characterized in that: Get the data size Ls of each sub-data i , historical access time Fs ia and historical change time Gs ib , after dimensionless processing, calculate the initial backup coefficient Cq of each sub-data i : Among them, i represents the number of each sub-data, i=1, 2, ..., k, a represents the time sequence number of each access to the same sub-data, a=1, 2, ..., m, b represents the time sequence number of each historical change of the same sub-data, b=1, 2, ..., n.

4. The data security protection method based on big data according to claim 3 is characterized in that: Use the user's server account to log in to the corresponding cloud service website or application to check whether there is a device backup list for the user. If the server does not have the user's backup data, the user is marked as a new backup user, and each sub-data is backed up according to the initial backup coefficient Cq i Sort the backup files from largest to smallest and upload them to the server.

5. The data security protection method based on big data according to claim 4 is characterized in that: If the server already has the backup information of the user, use the database comparison tool to compare the server backup data with the user-side data to obtain the difference data volume Cy of each sub-data i And use the network performance test tool to detect the maximum network communication volume Tx on the user side.

6. The data security protection method based on big data according to claim 5 is characterized in that: Get the difference data volume Cy of each sub-data i and the maximum network traffic Tx, calculate the predicted backup completion time Tt: Among them, the difference data volume Cy i The unit of maximum network traffic is kilobyte (KB), and the unit of maximum network traffic Tx is kilobyte per second (KB / S).

7. The data security protection method based on big data according to claim 6 is characterized in that: If the predicted backup completion time Tt is less than the preset backup time, the real-time backup strategy is adopted for this backup to obtain the differential data volume Cy of each sub-data i and the initial backup coefficient Cq i , after linear normalization, calculate the real-time update coefficient Sg of each sub-data i : Sg i =Cy i *Cq i The preset backup time can be any one of 1, 2, ..., 30 minutes, which is preset by the user. If the user does not preset it, the default is 10 minutes; When adopting the real-time backup strategy, the difference data of each sub-data is updated according to the real-time update coefficient Sg i Sort the data from largest to smallest and upload them to the server to update the user's backup data.

8. The data security protection method based on big data according to claim 7 is characterized in that: Enable the monitoring and logging functions in the backup task, record the difference in data volume Sz and the actual backup data volume Bz in real time during the backup process, and calculate the user's redundant backup data ratio Pg: The difference data volume Sz refers to the difference between the data stored on the server before the backup starts and the data stored on the server after the backup is completed. The actual backup data volume Bz refers to the total amount of data transmitted from the user end to the server during the backup process. If the user's excess backup data ratio Pg exceeds 1, and the predicted backup completion time Tt is greater than 0.8 of the preset backup time, the preset backup time is adjusted so that the current predicted backup completion time Tt is greater than the preset backup time, the non-real-time backup policy conditions are met, and the non-real-time backup policy is enabled.

9. The data security protection method based on big data according to claim 6 is characterized by: If the predicted backup completion time Tt is greater than the preset backup time, the non-real-time backup strategy is adopted for this backup, and the system monitoring tool is used to obtain the real-time change time Bs of each sub-data after the user starts the backup. ic and real-time change difference data volume Bc ic , calculate the real-time change coefficient Bx of each sub-data i : Here, c represents the time sequence number of each real-time change of the same sub-data, c=1, 2, ..., w.

10. A data security protection method based on big data according to claim 9, characterized in that: Get the real-time change coefficient Bx of each sub-data i and the initial backup coefficient Cq i After linear normalization, the non-real-time update coefficient Fg of each sub-data is calculated i : Fg i =Bx i *Cq i When adopting the non-real-time backup strategy, the difference data of each sub-data is updated according to the non-real-time update coefficient Fg i Sort the data from largest to smallest and upload them to the server to update the user's backup data.

Citation Information

Patent Citations

  • Data backup and recovery method and system based on cloud computing

    CN118277164A