Data backup method and system based on artificial intelligence
By constructing an AI-based data importance assessment model and an anomaly data detection model, and combining user behavior data and historical backup data, the backup frequency is dynamically adjusted, solving the problem that existing technologies fail to consider user preferences and data changes, and achieving more efficient and reliable data backup.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAINAN UNITED UNIVERSITY
- Filing Date
- 2025-12-09
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies fail to effectively consider user preferences and access volume in data backup methods, and fail to adaptively adjust backup frequency according to data changes, resulting in insufficient reliability and accuracy of backup methods.
An AI-based data backup method is adopted. By constructing a data importance assessment model and an anomaly data detection model, combined with user behavior data and historical backup data, the importance and reliability of the data are evaluated, a backup assessment value is generated, and the backup frequency is adjusted according to the backup assessment value.
It improves the reliability and accuracy of data backup methods, and ensures that important data is backed up in a timely manner by dynamically adjusting the backup frequency, reducing redundant backups and improving storage efficiency.
Smart Images

Figure CN122064531A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, specifically to a data backup method and system based on artificial intelligence. Background Technology
[0002] Data backup is primarily used to prevent data loss due to operational errors or system failures by copying all or part of the data to other storage media. This backup process helps improve data security, ensuring that no data is lost or inaccurate when restoring or retrieving data. Traditional data backup methods use built-in or external tape drives for data processing; however, with the continuous growth of data volumes, traditional data backup methods can no longer meet the efficiency and security requirements of individuals and businesses for data backup.
[0003] Currently, existing technologies for data backup still have shortcomings. On the one hand, existing technologies only assess the priority and importance of data based on content and type, without considering user preferences for data usage, access volume, or the degree of data anomalies and redundancy. On the other hand, existing technologies do not adaptively adjust the data backup frequency according to data changes, which significantly reduces the reliability and accuracy of data backup methods.
[0004] To address this, an artificial intelligence-based data backup method and system are proposed. Summary of the Invention
[0005] The purpose of this invention is to provide a data backup method and system based on artificial intelligence. First, historical backup data, input data, and user behavior data are acquired. Then, the input data is processed using a user behavior-based data importance assessment model to obtain a first data importance assessment value. Next, the input data is processed using a pre-trained abnormal data detection model to obtain the amount of abnormal data. Based on the amount of abnormal data and redundant data, a second data reliability assessment value is obtained. The first data importance assessment value and the second data reliability assessment value are evaluated to obtain a backup assessment value. Finally, the backup assessment value is compared with a backup assessment threshold. If it is lower than the backup assessment threshold, no data backup is required; otherwise, real-time input data is fed into a backup frequency prediction model, and backup frequency assessment information is generated based on the data results.
[0006] To achieve the above objectives, the present invention provides the following technical solution: An artificial intelligence-based data backup method includes...
[0007] Acquire historical backup data, entered data, and user behavior data; A data importance assessment model based on user behavior is constructed. The historical backup data, the entered data, and the user behavior data are input into the data importance assessment model to obtain a first data importance assessment value. An abnormal data detection model is constructed, and the historical backup data is input into the abnormal data detection model for training to obtain a pre-trained abnormal data detection model. Then, the pre-trained abnormal data detection model is used to process the input data to obtain the abnormal data volume. Based on the abnormal data volume and the redundant data volume of the input data, a second data reliability assessment value is calculated. The first data importance assessment value and the second data reliability assessment value are weighted and evaluated to obtain the backup assessment value of the entered data; The backup evaluation value is compared with the backup evaluation threshold. If it is lower than the backup evaluation threshold, it is determined that no data backup is required; otherwise, it is determined that data backup is to be performed, and the real-time recorded data is input into the backup frequency prediction model to generate backup frequency evaluation information based on the data results.
[0008] Furthermore, the specific implementation process of constructing a data importance assessment model based on user behavior, and inputting the historical backup data, the entered data, and the user behavior data into the data importance assessment model to obtain a first data importance assessment value includes: Acquire the types, quantities, and content of historical backup data and entered data; Based on user behavior data, the volume of high-access data and the volume of low-access data in the entered data are obtained; The number of categories, the content data, the amount of high-access data, and the amount of low-access data are input into the data importance assessment model to obtain the first data importance assessment value. The specific process of the data importance assessment model includes: A threshold evaluation is performed on the high access data volume and the low access data volume to obtain an access evaluation value; The number of identical types and identical content data in the historical backup data and the input data are evaluated to obtain the type evaluation value and the content evaluation value. The access evaluation value, the category evaluation value, and the content evaluation value are weighted to obtain the first data importance evaluation value.
[0009] Further, an abnormal data detection model is constructed, and the historical backup data is input into the abnormal data detection model for training to obtain a pre-trained abnormal data detection model; then, the pre-trained abnormal data detection model is used to process the input data to obtain the abnormal data volume; the specific implementation process of calculating the second data reliability assessment value based on the abnormal data volume and the redundant data volume of the input data includes: Retrieve historical backup data and input data; The historical backup data includes: normal historical backup data and abnormal historical backup data; Furthermore, the historical backup data is preprocessed to obtain preprocessed historical backup data; the preprocessed historical backup data is then divided into a training set and a test set in an 8:2 ratio. Furthermore, an anomaly detection model is constructed, and the training set is input into the anomaly detection model for training to obtain the initial trained anomaly detection model; Furthermore, the test set is input into the initially trained anomaly detection model, the model parameters are optimized, and a pre-trained anomaly detection model is obtained. Furthermore, the entered data is input into the pre-trained abnormal data detection model for processing to obtain the amount of abnormal data; Furthermore, the amount of abnormal data and the amount of redundant data in the entered data are evaluated to obtain a second data reliability assessment value.
[0010] Furthermore, the weighted evaluation of the first data importance assessment value and the second data reliability assessment value yields the following formula for calculating the backup assessment value of the entered data: ; Wherein, BFPG represents the backup evaluation value; The weight is represented by the data importance assessment weight; ZYPG represents the first data importance assessment value. is represented as the data reliability assessment weight; KKPG is represented as the second data reliability assessment value.
[0011] Furthermore, the specific implementation process of inputting real-time recorded data into the backup frequency prediction model and generating backup frequency assessment information based on the data results includes: Acquire historical backup data and real-time input data in time series; Furthermore, the historical backup data is input into the backup frequency prediction model for training to obtain the backup frequency influence coefficient; Further, the backup frequency impact coefficient and the real-time input data are input into the backup frequency prediction model to update the model parameters and output the final backup frequency impact coefficient; wherein, the final backup frequency impact coefficient includes: the final data importance impact coefficient and the final data reliability impact coefficient; Furthermore, the backup frequency assessment information is obtained by weighting and summing the first data importance assessment value and the second data reliability assessment value of the real-time entered data using the final backup frequency impact coefficient.
[0012] An artificial intelligence-based data backup system includes: a system control module, a data acquisition module, a data backup evaluation module, a backup frequency evaluation module, and an output module; The system control module is used to control the system's start, pause, and stop. The data acquisition module is used to acquire historical backup data, input data, and user behavior data. The data backup evaluation module is used to comprehensively evaluate the importance and reliability of the data to obtain a backup evaluation value; wherein, the data backup evaluation module includes: a data importance evaluation unit and a data reliability evaluation unit; The backup frequency evaluation module is used to predict and evaluate the data backup frequency to obtain backup frequency evaluation information; the output module is used to output the evaluation judgment result of data backup, including: a judgment unit and a prompt unit.
[0013] Furthermore, the specific implementation process of the data importance assessment unit inputting the historical backup data, the entered data, and the user behavior data into the data importance assessment model to obtain the first data importance assessment value includes: Acquire the types, quantities, and content of historical backup data and entered data; Based on user behavior data, the volume of high-access data and the volume of low-access data in the entered data are obtained; The number of categories, the content data, the amount of high-access data, and the amount of low-access data are input into the data importance assessment model to obtain the first data importance assessment value. The specific process of the data importance assessment model includes: A threshold evaluation is performed on the high access data volume and the low access data volume to obtain an access evaluation value; The number of identical types and identical content data in the historical backup data and the input data are evaluated to obtain the type evaluation value and the content evaluation value. The access evaluation value, the category evaluation value, and the content evaluation value are weighted to obtain the first data importance evaluation value.
[0014] Further, the data reliability assessment unit inputs the historical backup data into the abnormal data detection model for training to obtain a pre-trained abnormal data detection model; then, it uses the pre-trained abnormal data detection model to process the entered data to obtain the amount of abnormal data; the specific implementation process of calculating the second data reliability assessment value based on the amount of abnormal data and the amount of redundant data in the entered data includes: Retrieve historical backup data and input data; The historical backup data includes: normal historical backup data and abnormal historical backup data; Furthermore, the historical backup data is preprocessed to obtain preprocessed historical backup data; the preprocessed historical backup data is then divided into a training set and a test set in an 8:2 ratio. Furthermore, an anomaly detection model is constructed, and the training set is input into the anomaly detection model for training to obtain the initial trained anomaly detection model; Furthermore, the test set is input into the initially trained anomaly detection model, the model parameters are optimized, and a pre-trained anomaly detection model is obtained. Furthermore, the entered data is input into the pre-trained abnormal data detection model for processing to obtain the amount of abnormal data; Furthermore, the amount of abnormal data and the amount of redundant data in the entered data are evaluated to obtain a second data reliability assessment value.
[0015] Furthermore, the data backup evaluation module is used to comprehensively evaluate the importance and reliability of the data, and the formula for calculating the backup evaluation value is as follows: ; Wherein, BFPG represents the backup evaluation value; The weights represent the data importance assessment weights; ZYPG represents the first data importance assessment value. It represents the data reliability assessment weight; KKPG represents the second data reliability assessment value.
[0016] Furthermore, the backup frequency evaluation module is used to predict and evaluate the data backup frequency. The specific implementation process for obtaining backup frequency evaluation information includes: Acquire historical backup data and real-time input data in time series; Furthermore, the historical backup data is input into the backup frequency prediction model for training to obtain the backup frequency influence coefficient; Further, the backup frequency impact coefficient and the real-time input data are input into the backup frequency prediction model to update the model parameters and output the final backup frequency impact coefficient; wherein, the final backup frequency impact coefficient includes: the final data importance impact coefficient and the final data reliability impact coefficient; Furthermore, the backup frequency assessment information is obtained by weighting and summing the first data importance assessment value and the second data reliability assessment value of the real-time entered data using the final backup frequency impact coefficient.
[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention proposes a data importance assessment function to determine whether to perform data backup based on data importance. This function uses a user behavior-based data importance assessment model to output a first data importance assessment value, and judges the data backup operation based on the first data importance assessment value. The first data importance assessment value is calculated by comprehensively evaluating the types, quantities, and content of user behavior data, historical backup data, and input data. This function combines user behavior data and historical backup data, which can effectively improve the reliability and accuracy of data backup methods.
[0018] 2. This invention proposes a data reliability assessment function to determine whether to perform data backup processing from the perspective of data reliability. This function assesses the amount of abnormal data output by the abnormal data detection model and the amount of redundant data in the input data, and uses the calculated second data reliability assessment value to judge the data backup operation. The second data reliability assessment value is obtained by measuring the proportion of the abnormal data amount and the redundant data amount. This function can effectively improve the reliability and accuracy of the data backup method.
[0019] 3. This invention proposes a data backup frequency evaluation function to adjust the data backup frequency based on changes in data status. This function trains a backup frequency prediction model using historical backup data to obtain a backup frequency influence coefficient. Then, it inputs the backup frequency influence coefficient and real-time data into the backup frequency prediction model to obtain a final backup frequency influence coefficient. This function evaluates the evaluation value of the real-time data with the final backup frequency influence coefficient to obtain backup frequency evaluation information. Using this backup frequency evaluation information to adjust the data backup frequency effectively improves the reliability and accuracy of the data backup method. Attached Figure Description
[0020] Figure 1 This is a flowchart illustrating an artificial intelligence-based data backup method according to the present invention. Figure 2This is a schematic diagram of the abnormal data detection model of the present invention; Figure 3 This is a schematic diagram of the backup frequency prediction model of the present invention; Figure 4 This is a schematic diagram of the structure of an artificial intelligence-based data backup system according to the present invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Data backup is primarily used to prevent data loss due to operational errors or system failures by copying all or part of the data to other storage media. This backup process helps improve data security, ensuring that no data is lost or inaccurate when restoring or retrieving data. Traditional data backup methods use built-in or external tape drives for data processing; however, with the continuous growth of data volumes, traditional data backup methods can no longer meet the efficiency and security requirements of individuals and businesses for data backup.
[0023] Currently, existing technologies for data backup still have shortcomings. On the one hand, existing technologies only assess the priority and importance of data based on content and type, without considering user preferences for data usage, access volume, or the degree of data anomalies and redundancy. On the other hand, existing technologies do not adaptively adjust the data backup frequency according to data changes, which significantly reduces the reliability and accuracy of data backup methods.
[0024] Example 1 The specific implementation process in this application embodiment will be implemented by an artificial intelligence-based data backup method of the present invention, see below. Figure 1 The present invention provides a flowchart of the method proposed in this invention; the artificial intelligence-based data backup method includes: S10. Obtain historical backup data, input data, and user behavior data; S20. The historical backup data, the entered data, and the user behavior data are evaluated using a user behavior-based data importance assessment model to obtain a first data importance assessment value; S30. Calculate the second data reliability assessment value based on the amount of abnormal data output by the pre-trained abnormal data detection model and the amount of redundant data in the input data; S40. Perform a weighted evaluation on the first data importance evaluation value and the second data reliability evaluation value to obtain a backup evaluation value; S50. Compare the backup evaluation value with the backup evaluation threshold. If it is lower than the backup evaluation threshold, no data backup is required. Otherwise, determine that data backup should be performed, and input the real-time recorded data into the backup frequency prediction model to generate backup frequency evaluation information based on the data results.
[0025] Furthermore, the specific implementation process of an artificial intelligence-based data backup method is as follows: Acquire historical backup data, entered data, and user behavior data; A data importance assessment model based on user behavior is constructed. The historical backup data, the entered data, and the user behavior data are input into the data importance assessment model to obtain a first data importance assessment value. For example, the number of times a file has been read and modified in the past 30 days can be obtained through the file system log of the operating system.
[0026] An abnormal data detection model is constructed, and the historical backup data is input into the abnormal data detection model for training to obtain a pre-trained abnormal data detection model. Then, the pre-trained abnormal data detection model is used to process the input data to obtain the abnormal data volume. Based on the abnormal data volume and the redundant data volume of the input data, a second data reliability assessment value is calculated. The first data importance assessment value and the second data reliability assessment value are weighted and evaluated to obtain the backup assessment value of the entered data; The backup assessment value is compared with the backup assessment threshold. If it is lower than the backup assessment threshold, it is determined that no data backup is needed; otherwise, it is determined that a data backup should be performed, and the real-time input data is input into the backup frequency prediction model. Backup frequency assessment information is generated based on the data results. Specifically, after the real-time input data undergoes data importance / reliability assessment, it enters each layer of the backup frequency prediction model sequentially, outputs frequency assessment information, and synchronizes it with the system control module. The control submodule then determines whether to trigger a backup action.
[0027] This embodiment proposes an artificial intelligence-based data backup method. First, historical backup data, input data, and user behavior data are acquired. Then, the input data is processed using a user behavior-based data importance assessment model to obtain a first data importance assessment value. Next, the input data is processed using a pre-trained abnormal data detection model to obtain the amount of abnormal data. Based on the amount of abnormal data and redundant data, a second data reliability assessment value is obtained. The first data importance assessment value and the second data reliability assessment value are evaluated to obtain a backup assessment value. Finally, the backup assessment value is compared with a backup assessment threshold. If it is lower than the backup assessment threshold, no data backup is required; otherwise, real-time input data is fed into a backup frequency prediction model, and backup frequency assessment information is generated based on the data results. This method can effectively improve the reliability and accuracy of data backup methods.
[0028] For specific illustration, the following examples are provided: Acquire historical backup data, entered data, and user behavior data; In this embodiment, historical backup data and user behavior data are collected from the corresponding device; the device can be a mobile phone, computer or other hardware storage device.
[0029] Furthermore, the specific implementation process of constructing a data importance assessment model based on user behavior, and inputting the historical backup data, the entered data, and the user behavior data into the data importance assessment model to obtain a first data importance assessment value includes: Acquire the types, quantities, and content of historical backup data and entered data; In this embodiment, the data types include documents, images, videos, and other types; the content data includes text content, pixel content, video content, and other data content.
[0030] Furthermore, the amount of high-access data and low-access data in the entered data are obtained based on user behavior data; The high access data volume and the low access data volume refer to the number of data entries entered by the user with high frequency of access and the number of data entries entered by the user with low frequency of access, respectively. They are determined based on the access data frequency and a comparison between a preset high access threshold and a preset low access threshold. Specifically, the high access data volume refers to the number of files accessed more than a set threshold (e.g., 10 times / day) within the statistical period (day / week / month), and the low access data volume refers to the number of files accessed less than or equal to the threshold.
[0031] Furthermore, the number of categories, the content data, the amount of high-access data, and the amount of low-access data are input into the data importance assessment model to obtain a first data importance assessment value; The specific process of the data importance assessment model includes: A threshold evaluation is performed on the high access data volume and the low access data volume to obtain an access evaluation value; The formula for calculating the access evaluation value is: ; ; Here, FWPG represents the access evaluation value; the higher the value, the more important the user is to the data. Represented as a weighting factor for high-access data; Represented as the number of data points; This indicates high-access data within the user behavior data; This is expressed as the frequency of the accessed data; This is represented by the preset high access threshold; Represented as the intersection of data; The data to be entered is shown below; This is represented as a weighting factor for low-access data; This indicates low-access data within the user behavior data; This is represented by the preset low access threshold.
[0032] In this embodiment, the weighting factors for high-access data and low-access data are set to 0.6 and 0.4, respectively.
[0033] In this embodiment, the preset high access threshold and preset low access threshold are not unique, and those skilled in the art can flexibly adjust them according to the actual data size and application scenario.
[0034] Furthermore, the number of identical types and identical content data in the historical backup data and the entered data are evaluated to obtain type evaluation values and content evaluation values; wherein, the evaluation of identical content data is specifically achieved by calculating the MD5 hash value of the data file, and data with the same hash value are evaluated as identical content data.
[0035] The formulas for calculating the category assessment value and the content assessment value are as follows: ; ; Wherein, ZLPG represents the evaluation value of the species; This represents the number of data types. This refers to the type of historical backup data; This is represented as an intersection relationship; The data type is represented by NRPG; the content evaluation value is represented by NRPG. This is represented as the quantity of content data; This refers to the content data of the historical backup data; This refers to the content data of the entered data.
[0036] Further, the access evaluation value, the category evaluation value, and the content evaluation value are weighted to obtain the first data importance evaluation value. The larger the first data importance evaluation value, the higher the importance of the data, and the more necessary it is to be backed up. The first data importance evaluation value can be expressed by the following formula: ; ; Wherein, ZYPG represents the first data importance assessment value. The larger the value, the higher the importance of the data and the higher the likelihood of data backup. FWPG, ZLPG, and NRPG are the access assessment value, the category assessment value, and the content assessment value, respectively. , and These represent the access evaluation weight, category evaluation weight, and content evaluation weight, respectively.
[0037] In this embodiment , and The values are set to 0.35, 0.3, and 0.35 respectively. Of course, these weight values can be flexibly adjusted according to the actual situation and are not unique.
[0038] To facilitate the explanation of the user behavior-based data importance assessment function proposed in this invention, this embodiment selects three sets of input data samples for testing, denoted as Sample 1, Sample 2, and Sample 3, each with a size of 10G; the input data samples come from the same type of device but different users; simultaneously, user access data for each sample is acquired, and preset high access thresholds and preset low access thresholds are set to 10 times / day and 3 times / day, respectively; then, combined with the specific process of the data importance assessment model, the access assessment value, category assessment value, content assessment value, and first data importance assessment value of each sample are obtained. The data importance test results are shown in Table 1: Table 1. Results of Data Importance Test Test samples Access Evaluation Value Species Assessment Value Content Evaluation Value First data importance assessment value Sample 1 0.91 0.47 0.72 0.71 Sample 2 0.84 0.59 0.63 0.69 Sample 3 0.93 0.62 0.76 0.78 This embodiment proposes a data importance assessment function to determine whether to perform data backup based on data importance. This function uses a user behavior-based data importance assessment model to output a first data importance assessment value, and then judges the data backup operation based on this first data importance assessment value. The first data importance assessment value is calculated by comprehensively evaluating the types, quantities, and content of user behavior data, historical backup data, and entered data. This function combines user behavior data and historical backup data, which can effectively improve the reliability and accuracy of the data backup method.
[0039] Further, an abnormal data detection model is constructed, and the historical backup data is input into the abnormal data detection model for training to obtain a pre-trained abnormal data detection model; then, the pre-trained abnormal data detection model is used to process the input data to obtain the abnormal data volume; the specific implementation process of calculating the second data reliability assessment value based on the abnormal data volume and the redundant data volume of the input data includes: Retrieve historical backup data and input data; The historical backup data includes: normal historical backup data and abnormal historical backup data; Furthermore, the historical backup data is preprocessed to obtain preprocessed historical backup data; the preprocessed historical backup data is then divided into a training set and a test set in an 8:2 ratio. Furthermore, an anomaly detection model is constructed, and the training set is input into the anomaly detection model for training to obtain the initial trained anomaly detection model; Furthermore, the test set is input into the initially trained anomaly detection model, the model parameters are optimized, and a pre-trained anomaly detection model is obtained. Furthermore, the entered data is input into the pre-trained abnormal data detection model for processing to obtain the amount of abnormal data; The structure of the abnormal data detection model is as follows: Figure 2 As shown, it includes: an input layer, a multi-scale feature extraction layer, an anomaly data feature labeling layer, an anomaly data feature recognition layer, and an output layer. Its main implementation process includes: The entered data is fed into the input layer to obtain data features; The data features are processed using a multi-scale feature extraction layer to obtain multi-scale data features; Before building the anomaly detection model, the entered data needs to undergo visual mapping preprocessing. The specific process is as follows: Binary stream truncation and reshaping: The binary data stream of the file to be backed up is segmented into fixed byte lengths (e.g., 1024 bytes).
[0040] Grayscale mapping: Each segment of binary data is mapped to a 32*32 pixel grayscale matrix, where the value of each byte (0-255) corresponds to the grayscale value of a pixel. Insufficient bit segments are padded with zeros.
[0041] Multi-channel synthesis: Three consecutive grayscale matrices are synthesized into a 32*32*3 pseudo-color feature map.
[0042] After the above processing, the abstract input data is transformed into pseudo-image feature data.
[0043] The multi-scale feature extraction layer obtains features at different scales by continuously employing four convolutional kernels of size 3×3 with a stride of 2 and channel numbers of 16, 32, 64, and 128 respectively. An anomaly feature labeling layer (specifically a Region Proposal Network) labels the multi-scale data features to obtain anchor boxes; the number of anchor boxes is set to 5. An anomaly feature recognition layer (specifically a fully connected layer consisting of 256 neurons and a Softmax classifier) recognizes the labeled features in the anchor boxes to obtain anomaly recognition features. These anomaly recognition features are then input to the output layer to obtain the amount of anomalous data. In this embodiment, the anomaly detection model is trained using the cross-entropy loss function and optimized using the Adam optimizer. The multi-scale feature extraction layer obtains features at different scales by continuously employing convolutional kernels of size 3×3 with a stride of 2. The multi-scale data features are labeled using an anomaly data feature labeling layer to obtain anchor boxes; the number of anchor boxes is set to 5. The abnormal data feature recognition layer is used to identify the marker features in the anchor box to obtain abnormal data recognition features; The abnormal data identification features are input into the output layer to obtain the amount of abnormal data.
[0044] Furthermore, the amount of abnormal data and the amount of redundant data in the entered data are evaluated to obtain a second data reliability assessment value; the formula for calculating the second data reliability assessment value is: ; ; Wherein, KKPG represents the second data reliability assessment value. The larger the value, the higher the reliability of the data and the higher the possibility of data backup. Represented as the number of data points; This refers to the entered data; and These are respectively represented as abnormal data and redundant data; and These represent the weights for abnormal assessment and redundant assessment, respectively.
[0045] In this embodiment and The values are set to 0.6 and 0.4 respectively. Of course, these weight values can be flexibly adjusted according to the actual situation and are not unique.
[0046] This embodiment selects three sets of input data samples for testing to illustrate the data reliability assessment function proposed in this invention. These are designated as Input Sample 1, Input Sample 2, and Input Sample 3, each with a size of 5GB. The input data samples come from the same type of device but from different users. Each sample data is input into a trained abnormal data detection model to obtain the amount of abnormal data for each sample. Simultaneously, algorithms (compression algorithms, hash algorithms, etc.) are used to obtain the amount of redundant data for each sample. Then, combined with the calculation formula for the second data reliability assessment value, the second data reliability assessment value for each test sample is obtained. The data reliability test results are shown in Table 2. Table 2. Data Reliability Test Results Test samples Redundant data volume ratio percentage of abnormal data Second data reliability assessment value Input Sample 1 0.18 0.04 0.91 Input Sample 2 0.31 0.07 0.83 Input Sample 3 0.24 0.05 0.88 This embodiment proposes a data reliability assessment function to determine whether to perform data backup processing from the perspective of data reliability. This function evaluates the amount of abnormal data output by the abnormal data detection model and the amount of redundant data in the input data, and uses the calculated second data reliability assessment value to judge the data backup operation. The second data reliability assessment value is obtained by measuring the ratio of the amount of abnormal data and the amount of redundant data. This function can effectively improve the reliability and accuracy of the data backup method.
[0047] Furthermore, the weighted evaluation of the first data importance assessment value and the second data reliability assessment value yields the following formula for calculating the backup assessment value of the entered data: ; Wherein, BFPG represents the backup evaluation value. The higher the value, the higher the probability that the data needs to be backed up. The weight is represented by the data importance assessment weight; ZYPG represents the first data importance assessment value. The data reliability assessment weight is represented by KKPG; KKPG represents the second data reliability assessment value. and Set all to 0.5; This embodiment proposes a backup evaluation function to determine whether data needs to be backed up. The function first evaluates the data from the perspectives of importance and reliability, and then weights the first data importance evaluation value and the second data reliability evaluation value to obtain a backup evaluation value. The backup evaluation value can accurately reflect the necessity of data backup. By using the backup evaluation value for judgment, this function can effectively improve the reliability and accuracy of data backup methods.
[0048] Furthermore, the backup evaluation value is compared with the backup evaluation threshold. If it is lower than the backup evaluation threshold, it is determined that no data backup needs to be performed; otherwise, it is determined that data backup can be performed, and the real-time recorded data is input into the backup frequency prediction model to generate backup frequency evaluation information based on the data results.
[0049] Furthermore, the specific implementation process of inputting real-time recorded data into the backup frequency prediction model and generating backup frequency assessment information based on the data results includes: Acquire historical backup data and real-time input data in time series; Furthermore, the historical backup data is input into the backup frequency prediction model for training to obtain the backup frequency influence coefficient; In this embodiment, the backup frequency prediction model uses a hybrid network combining GRU and ConvLSTM. The structure of this model is as follows: Figure 3 As shown, the hybrid network adopts a dual-branch structure, including: an input layer, a feature extraction layer, three GRU layers, three ConvLSTM layers, a feature fusion layer, and an output layer.
[0050] The specific implementation process of the backup frequency prediction model in this embodiment includes: first, inputting historical backup data sequentially into the input layer and the feature extraction layer for processing to obtain historical backup data features; then, inputting the historical backup data features into two sub-networks, the GRU branch and the ConvLSTM branch, respectively for feature processing; finally, fusing the output features of different networks through the feature fusion layer to obtain fused features; and finally, using the output layer to transform the fused features and output the backup frequency influence coefficient.
[0051] Further, the backup frequency impact coefficient and the real-time input data are input into the backup frequency prediction model to update the model parameters and output the final backup frequency impact coefficient; wherein, the final backup frequency impact coefficient includes: the final data importance impact coefficient and the final data reliability impact coefficient; Furthermore, the backup frequency assessment information is obtained by weighting and summing the first data importance assessment value and the second data reliability assessment value of the real-time entered data using the final backup frequency impact coefficient; wherein, the calculation formula for the backup frequency assessment information is: ; ; Wherein, PLPG represents the backup frequency evaluation information used to adjust the frequency of data backup in subsequent time periods; The final data importance impact coefficient is represented by ZYPG; ZYPG represents the first data importance assessment value. The final data reliability impact coefficient is represented by KKPG; the second data reliability assessment value is represented by KKPG.
[0052] This embodiment selects three sets of real-time input data samples for testing to illustrate the backup frequency prediction function proposed in this invention. The sizes are 5G, 10G, and 15G, respectively. The real-time input data samples come from the same type of device but different users. Each sample data is input into the trained backup frequency prediction model to obtain the final backup frequency influence coefficients for each sample, which are (0.63, 0.37), (0.51, 0.49), and (0.55, 0.45), respectively. Then, combined with the calculation formula of backup frequency evaluation information, the backup frequency evaluation information of each test sample is obtained, and the backup frequency evaluation thresholds are set to 0.7 and 0.5. The backup frequency evaluation thresholds indicate that data above 0.7 should increase the backup frequency, data below 0.5 should decrease the backup frequency, and the rest remain unchanged. The backup frequency evaluation test results are shown in Table 3. Table 3. Backup Frequency Evaluation Test Results Test samples First data importance assessment value Second data reliability assessment value Backup frequency assessment information Backup frequency adjustment Real-time sample 1 0.42 0.59 0.48 reduce Real-time Sample 2 0.63 0.69 0.66 constant Real-time Sample 3 0.72 0.74 0.73 improve This embodiment proposes a data backup frequency evaluation function to adjust the data backup frequency based on changes in data status. This function trains a backup frequency prediction model using historical backup data to obtain a backup frequency influence coefficient. Then, it inputs the backup frequency influence coefficient and real-time data into the backup frequency prediction model to obtain a final backup frequency influence coefficient. This function evaluates the evaluation value of the real-time data with the final backup frequency influence coefficient to obtain backup frequency evaluation information. Using this backup frequency evaluation information to adjust the data backup frequency can effectively improve the reliability and accuracy of the data backup method.
[0053] Example 2 As one embodiment of the present invention, refer to Figure 4An artificial intelligence-based data backup system includes: a system control module, a data acquisition module, a data backup evaluation module, a backup frequency evaluation module, and an output module; The system control module is used to control the system's start, pause, and stop. The data acquisition module is used to acquire historical backup data, input data, and user behavior data. The data backup evaluation module is used to comprehensively evaluate the importance and reliability of the data to obtain a backup evaluation value; wherein, the data backup evaluation module includes: a data importance evaluation unit and a data reliability evaluation unit; Furthermore, the specific implementation process of the data importance assessment unit inputting the historical backup data, the entered data, and the user behavior data into the data importance assessment model to obtain the first data importance assessment value includes: Acquire the types, quantities, and content of historical backup data and entered data; Based on user behavior data, the volume of high-access data and the volume of low-access data in the entered data are obtained; The number of categories, the content data, the amount of high-access data, and the amount of low-access data are input into the data importance assessment model to obtain the first data importance assessment value. The specific process of the data importance assessment model includes: A threshold evaluation is performed on the high access data volume and the low access data volume to obtain an access evaluation value; The number of identical types and identical content data in the historical backup data and the input data are evaluated to obtain the type evaluation value and the content evaluation value. The access evaluation value, the category evaluation value, and the content evaluation value are weighted to obtain the first data importance evaluation value.
[0054] Further, the data reliability assessment unit inputs the historical backup data into the abnormal data detection model for training to obtain a pre-trained abnormal data detection model; then, it uses the pre-trained abnormal data detection model to process the entered data to obtain the amount of abnormal data; the specific implementation process of calculating the second data reliability assessment value based on the amount of abnormal data and the amount of redundant data in the entered data includes: Retrieve historical backup data and input data; The historical backup data includes: normal historical backup data and abnormal historical backup data; Furthermore, the historical backup data is preprocessed to obtain preprocessed historical backup data; the preprocessed historical backup data is then divided into a training set and a test set in an 8:2 ratio. To verify the rationality of the dataset partitioning ratio selected in this invention, this embodiment selects 5G of historical backup data and preprocesses it to obtain test data. The test data contains anomalous data, and the total amount of anomalous data is recorded beforehand. Then, three ratios (7:3, 8:2, and 9:1) are used to partition the test data, denoted as Test 1, Test 2, and Test 3. Combining the training and processing of the anomalous data detection model, the ratio of the detected anomalous data amount to the total amount of anomalous data in each test is obtained. The test results of the data partitioning ratio are shown in Table 4. Table 4. Test Results of Data Division Ratio test Proportion The ratio of the amount of detected outliers to the total amount of outliers. Test 1 7:3 0.89 Test 2 8:2 0.93 Test 3 9:1 0.91 As shown in Table 4, Test 2, which uses an 8:2 ratio, has the best performance in detecting abnormal data, indicating that the 8:2 ratio for the dataset partitioning in this invention is more reasonable.
[0055] Furthermore, an anomaly detection model is constructed, and the training set is input into the anomaly detection model for training to obtain the initial trained anomaly detection model; Furthermore, the test set is input into the initially trained anomaly detection model, the model parameters are optimized, and a pre-trained anomaly detection model is obtained. Furthermore, the entered data is input into the pre-trained abnormal data detection model for processing to obtain the amount of abnormal data; Furthermore, the amount of abnormal data and the amount of redundant data in the entered data are evaluated to obtain a second data reliability assessment value.
[0056] Furthermore, the data backup evaluation module is used to comprehensively evaluate the importance and reliability of the data, and the formula for calculating the backup evaluation value is as follows: ; Wherein, BFPG represents the backup evaluation value; The weights represent the data importance assessment weights; ZYPG represents the first data importance assessment value. It represents the data reliability assessment weight; KKPG represents the second data reliability assessment value.
[0057] To verify the actual effect of the data backup evaluation scheme proposed in this invention, and in conjunction with the data backup method process S10-S40 based on artificial intelligence described in Embodiment 1, this embodiment designed multiple sets of comparative experiments; this invention obtains 100 sets of sample data from device A, including: input data, historical backup data and user behavior data for the past 3 months; wherein, the input data comes from different personnel to ensure the uniqueness of the collected data.
[0058] Scheme 1 applies a data backup evaluation scheme proposed in this invention, which includes a data importance evaluation model based on user behavior and a data reliability evaluation model based on anomaly detection. Option 2 uses a simple data importance assessment model and an anomaly detection-based data reliability assessment model for evaluation, omitting user behavior data; Option 3 uses a data importance assessment model and a simple data reliability assessment model, omitting outlier detection; Option 4 uses only a simple data importance assessment model and a simple data reliability assessment model.
[0059] Based on the backup evaluation values obtained from the four schemes, a data backup evaluation was performed on 100 sets of the input data to obtain the percentage of erroneous data backup evaluation for each scheme. Table 5 shows a comparison of the effects of different backup evaluation schemes: Table 5. Comparison of the effects of different backup evaluation schemes plan Model Percentage of data backup assessments with errors (%) Option 1 Data importance assessment model and data reliability assessment model 3 Option 2 Simple data importance assessment model and data reliability assessment model 10 Option 3 Data importance assessment model and simple data reliability assessment model 7 Option 4 Simple data importance assessment model and simple data reliability assessment model 14 As shown in Table 5, Scheme 1, which combines a data importance assessment model based on user behavior and a data reliability assessment model based on anomaly detection, performs best, indicating that the data backup scheme based on artificial intelligence proposed in this invention is the most accurate in terms of data backup assessment.
[0060] The backup frequency evaluation module is used to predict and evaluate the data backup frequency to obtain backup frequency evaluation information. Furthermore, the backup frequency evaluation module is used to predict and evaluate the data backup frequency. The specific implementation process for obtaining backup frequency evaluation information includes: Acquire historical backup data and real-time input data in time series; Furthermore, historical backup data is input into the backup frequency prediction model for training to obtain the backup frequency impact coefficient; Further, the backup frequency impact coefficient and the real-time input data are input into the backup frequency prediction model to update the model parameters and output the final backup frequency impact coefficient; wherein, the final backup frequency impact coefficient includes: the final data importance impact coefficient and the final data reliability impact coefficient; Furthermore, the backup frequency assessment information is obtained by weighting and summing the first data importance assessment value and the second data reliability assessment value of the real-time entered data using the final backup frequency impact coefficient.
[0061] The output module is used to output the evaluation and judgment results of data backup, including: a judgment unit and a prompting unit; The judgment unit is used to compare the backup evaluation value with the backup evaluation threshold, and transmit the judgment result to the prompting unit, which outputs the result in the form of text or image.
[0062] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A data backup method based on artificial intelligence, characterized in that, include: Acquire historical backup data, entered data, and user behavior data; A data importance assessment model based on user behavior is constructed. The historical backup data, the entered data, and the user behavior data are input into the data importance assessment model to obtain a first data importance assessment value. An abnormal data detection model is constructed, and the historical backup data is input into the abnormal data detection model for training to obtain a pre-trained abnormal data detection model. Then, the pre-trained abnormal data detection model is used to process the input data to obtain the abnormal data volume. Based on the abnormal data volume and the redundant data volume of the input data, a second data reliability assessment value is calculated. The first data importance assessment value and the second data reliability assessment value are weighted and evaluated to obtain the backup assessment value of the entered data; The backup evaluation value is compared with the backup evaluation threshold. If it is lower than the backup evaluation threshold, it is determined that no data backup needs to be performed. Otherwise, it is determined that a data backup will be performed, and the real-time data will be input into the backup frequency prediction model to generate backup frequency evaluation information based on the data results.
2. The data backup method based on artificial intelligence according to claim 1, characterized in that, The specific implementation process of constructing a data importance assessment model based on user behavior, and inputting the historical backup data, the entered data, and the user behavior data into the data importance assessment model to obtain a first data importance assessment value, includes: Acquire the types, quantities, and content of historical backup data and entered data; Based on user behavior data, the volume of high-access data and the volume of low-access data in the entered data are obtained; The number of categories, the content data, the amount of high-access data, and the amount of low-access data are input into the data importance assessment model to obtain the first data importance assessment value. The specific process of the data importance assessment model includes: A threshold evaluation is performed on the high access data volume and the low access data volume to obtain an access evaluation value; The number of identical types and identical content data in the historical backup data and the input data are evaluated to obtain the type evaluation value and the content evaluation value. The access evaluation value, the category evaluation value, and the content evaluation value are weighted to obtain the first data importance evaluation value.
3. The data backup method based on artificial intelligence according to claim 1, characterized in that, An abnormal data detection model is constructed by inputting the historical backup data into the abnormal data detection model for training, thereby obtaining a pre-trained abnormal data detection model. Then, the pre-trained abnormal data detection model is used to process the input data to obtain the abnormal data volume. The specific implementation process for calculating the second data reliability assessment value based on the amount of abnormal data and the amount of redundant data in the entered data includes: Retrieve historical backup data and input data; The historical backup data includes: normal historical backup data and abnormal historical backup data; The historical backup data is preprocessed to obtain preprocessed historical backup data; the preprocessed historical backup data is then divided into a training set and a test set in an 8:2 ratio. An abnormal data detection model is constructed by inputting the training set into the abnormal data detection model for training, thereby obtaining the initial trained abnormal data detection model. The test set is input into the initially trained anomaly detection model, and the model parameters are optimized to obtain the pre-trained anomaly detection model. The entered data is input into the pre-trained abnormal data detection model for processing to obtain the amount of abnormal data. The amount of abnormal data and the amount of redundant data in the entered data are evaluated to obtain a second data reliability assessment value.
4. The data backup method based on artificial intelligence according to claim 1, characterized in that, The specific implementation process of inputting real-time data into the backup frequency prediction model and generating backup frequency assessment information based on the data results includes: Acquire historical backup data and real-time input data in time series; The historical backup data is input into the backup frequency prediction model for training to obtain the backup frequency impact coefficient. The backup frequency impact coefficient and the real-time input data are input into the backup frequency prediction model to update the model parameters and output the final backup frequency impact coefficient; wherein, the final backup frequency impact coefficient includes: the final data importance impact coefficient and the final data reliability impact coefficient; The backup frequency assessment information is obtained by weighting and summing the first data importance assessment value and the second data reliability assessment value of the real-time entered data using the final backup frequency impact coefficient.
5. A data backup system based on artificial intelligence, characterized in that, include: The system comprises a control module, a data acquisition module, a data backup evaluation module, a backup frequency evaluation module, and an output module. The system control module controls the system's startup, pause, and stop. The data acquisition module acquires historical backup data, input data, and user behavior data. The data backup evaluation module comprehensively evaluates the importance and reliability of the data to obtain a backup evaluation value. This data backup evaluation module includes a data importance evaluation unit and a data reliability evaluation unit. The backup frequency evaluation module predicts and evaluates the data backup frequency to obtain backup frequency evaluation information. The output module outputs the data backup evaluation results and includes a judgment unit and a prompt unit.
6. A data backup system based on artificial intelligence according to claim 5, characterized in that, The specific implementation process of the data importance assessment unit inputting the historical backup data, the entered data, and the user behavior data into the data importance assessment model to obtain the first data importance assessment value includes: Acquire the types, quantities, and content of historical backup data and entered data; Based on user behavior data, the volume of high-access data and the volume of low-access data in the entered data are obtained; The number of categories, the content data, the amount of high-access data, and the amount of low-access data are input into the data importance assessment model to obtain the first data importance assessment value. The specific process of the data importance assessment model includes: A threshold evaluation is performed on the high access data volume and the low access data volume to obtain an access evaluation value; The number of identical types and identical content data in the historical backup data and the input data are evaluated to obtain the type evaluation value and the content evaluation value. The access evaluation value, the category evaluation value, and the content evaluation value are weighted to obtain the first data importance evaluation value.
7. A data backup system based on artificial intelligence according to claim 5, characterized in that, The data reliability assessment unit inputs the historical backup data into the abnormal data detection model for training to obtain a pre-trained abnormal data detection model; then, the pre-trained abnormal data detection model is used to process the input data to obtain the amount of abnormal data. The specific implementation process for calculating the second data reliability assessment value based on the amount of abnormal data and the amount of redundant data in the entered data includes: Retrieve historical backup data and input data; The historical backup data includes: normal historical backup data and abnormal historical backup data; The historical backup data is preprocessed to obtain preprocessed historical backup data; the preprocessed historical backup data is then divided into a training set and a test set in an 8:2 ratio. An abnormal data detection model is constructed by inputting the training set into the abnormal data detection model for training, thereby obtaining the initial trained abnormal data detection model. The test set is input into the initially trained anomaly detection model, and the model parameters are optimized to obtain the pre-trained anomaly detection model. The entered data is input into the pre-trained abnormal data detection model for processing to obtain the amount of abnormal data. The amount of abnormal data and the amount of redundant data in the entered data are evaluated to obtain a second data reliability assessment value.
8. A data backup system based on artificial intelligence according to claim 5, characterized in that, The backup frequency evaluation module is used to predict and evaluate the data backup frequency. The specific implementation process for obtaining backup frequency evaluation information includes: Acquire historical backup data and real-time input data in time series; The historical backup data is input into the backup frequency prediction model for training to obtain the backup frequency impact coefficient. The backup frequency impact coefficient and the real-time input data are input into the backup frequency prediction model to update the model parameters and output the final backup frequency impact coefficient; wherein, the final backup frequency impact coefficient includes: the final data importance impact coefficient and the final data reliability impact coefficient; The backup frequency assessment information is obtained by weighting and summing the first data importance assessment value and the second data reliability assessment value of the real-time entered data using the final backup frequency impact coefficient.