Incremental data backup platform and method based on Bayesian model

By introducing an incremental data backup platform based on Bayesian model in the data backup system, using event notification, data compression and feature extraction technologies, the problem of inaccurate data classification in traditional backup methods is solved, and efficient, orderly and secure data backup effects are achieved.

CN120179463AInactive Publication Date: 2025-06-20GUANGZHOU SHANGZHIJIE NETWORK SAFETY TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510287961.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-20
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional data backup methods have shortcomings in data transmission, classification management and backup decisions, especially inaccurate classification of data files of unknown categories, resulting in inaccurate classification of data during the backup process, reducing backup efficiency and reliability of data recovery.

Method used

The incremental data backup platform based on Bayesian model is adopted to detect data changes in real time through the client's event notification mechanism, and the data is processed using data compression algorithms and feature extraction technology. The Bayesian classification model is trained based on data samples of known categories, and outputs the category with the greatest posterior probability as the classification result of unknown category data files to achieve accurate classification and orderly management.

Benefits of technology

It realizes a low bandwidth, high efficiency, orderly management, safe and reliable backup effect, improves backup efficiency and data recovery reliability, and solves the shortcomings of traditional backup technology in data transmission, classification management and backup decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179463A_ABST
    Figure CN120179463A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer data, and particularly relates to an incremental data backup platform and method based on a Bayesian model, and the platform comprises a client which is used for storing user data; detecting a data change event in real time through a built-in event notification mechanism, and pushing the data change event to the data processing module; the data processing module is used for retrieving and acquiring the changed data file from the client according to the metadata contained in the event notification, and compressing the changed data file; the feature extraction module is also used for performing feature extraction on the compressed data file; and the Bayesian classification module is also used for constructing and training a Bayesian classification model, and outputting the category with the maximum posterior probability as a classification result of the data file of the unknown category, namely a file path. The Bayesian model is adopted for intelligent classification, so that the data files of unknown categories can be accurately classified into different data categories when being backed up, and the incremental data backup efficiency and security with ordered management are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer data, and specifically, relates to an incremental data backup platform and method based on a Bayesian model. Background Art

[0002] With the rapid development of information technology and the sharp increase in data volume, various business systems have an increasing dependence on data and the data volume shows an explosive growth. Data backup and disaster recovery have become important means for various enterprises and organizations to ensure information security and maintain business continuity. Traditional data backup methods are mainly divided into full backup and incremental backup. Full backup is to copy and store the entire data set. As the data scale continues to expand, the time and storage costs of traditional full backup methods are constantly increasing. At the same time, data storage shows the characteristics of multi-level non-standardization, resulting in quality differences during the data storage process. This redundant data transmission will cause serious occupation of network bandwidth, and at the same time, it also prolongs the backup time and reduces the backup efficiency.

[0003] In addition, although the current incremental backup technology can reduce the data transmission volume to a certain extent, in terms of data classification, simple rules or threshold judgments are mostly used, and it is impossible to achieve accurate classification of data files of unknown categories. As a result, the backup data is often mixed together, the management is not orderly enough, and it is difficult to retrieve when using it later. Especially for data files of unknown categories, there is a lack of in-depth analysis and statistical learning of the data content, and it is impossible to make full use of the statistical characteristics of historical data to guide the backup. This leads to inaccurate data classification during the backup process, reducing the backup efficiency and the reliability of data recovery. Summary of the Invention

[0004] To solve the technical problems of inaccurate classification and disorderly management of data files of unknown categories, the present invention provides an incremental data backup platform and method based on a Bayesian model. It solves the deficiencies of traditional backup technologies in data transmission, classification management, and backup decision-making, and realizes the backup effects of low bandwidth, high efficiency, orderly management, and security and reliability.

[0005] The object of the present invention can be achieved by the following technical solutions:

[0006] An incremental data backup platform based on a Bayesian model, including a client, a data processing module, a backup server, and a network communication module;

[0007] The client includes a plurality of user data memories for storing user data; and it detects data change events in real time through a built-in event notification mechanism and pushes them to the data processing module;

[0008] The data processing module is used to retrieve and obtain the changed data files from the client according to the metadata contained in the event notification after receiving the event notification, and compress the changed data files;

[0009] It is also used to extract features from the compressed data files and represent them using an m-dimensional feature vector Y;

[0010] It is also used to construct and train a Bayesian classification model according to the data samples of known categories in the backup server, and output the category with the highest posterior probability as the classification result of the data files of unknown categories, that is, the file path;

[0011] The backup server includes a cloud server or a local server, which is used to receive and store the backup data uploaded by the client; it is also used to construct the classification criteria of the data category tree and back up data files of different categories to the corresponding file paths;

[0012] The network communication module is used for data transmission between the client, the backup server and the data processing module.

[0013] Preferably, the event notification mechanism in the client is that the client generates a packet containing metadata for the data change event and pushes it to the data processing module; the data change event includes data modification or data deletion operations; the metadata includes the file path, change timestamp and user ID of the data.

[0014] Preferably, the method for compressing the changed data files in the data processing module is to use the LZW data compression algorithm to generate an adaptive string of the data file in the form of data encoding, and completely record non-repeating characters; the specific steps include:

[0015] a. Initialize a string table containing all single characters, including single characters represented by 0 to 255

[0016] characters;

[0017] b. Set the S string to be empty and mark it with code 0;

[0018] c. Select characters of the string class of the character sequence according to the data arrangement of the data category tree;

[0019] d. Determine whether there is a concatenation result of s + string in the middle string; if it exists, set s + string = s and re-select characters;

[0020] e. Substitute the s code in the middle string into the code stream to complete the data encoding of the string;

[0021] f. Loop steps b to e in this way to complete the compression of the data file.

[0022] Preferably, the m - dimensional feature vector Y in the data processing module is expressed as: Y = {y1, y2, …, y m};

[0023] In the formula, y1 represents the normalized value of the string feature vector after data file compression, y2 represents a certain - dimensional feature vector after the text feature of the data file compression is processed by word vectors, and so on. y m represents a certain - dimensional feature vector obtained by feature extraction after data file compression; ensuring that the feature vectors of each data sample have the same m dimensions.

[0024] Preferably, the specific implementation method for constructing and training a Bayesian classification model in the data processing module includes:

[0025] Define the prior probability distribution: Count the occurrence frequencies of each category in the training data set, that is, the prior probabilities of each category; Let the number of samples of data category D j be X, and the total number of samples be Z. Then the expression of the prior probability is:

[0026]

[0027] In the formula, P(D j ) is the prior probability, which reflects the initial distribution of samples belonging to each data category before observing the feature vector Y;

[0028] Construct the likelihood function: Count the occurrence frequencies of each feature vector given the category D j to construct the likelihood function; its expression is:

[0029]

[0030] In the formula, P(Y∣D j ) represents the likelihood function of observing the entire feature vector set Y given the data category D j , that is, the probability that the data generates the entire feature vector set; P(y i ∣D j ) represents the probability of observing a single feature vector y j given the data category D i ;

[0031] Posterior probability calculation: According to Bayes' theorem, calculate the posterior probability that the feature vector Y belongs to the data category D j ; its expression is:

[0032]

[0033] In the formula, P(D jP(D∣Y) represents the posterior probability that the data belongs to class D after observing the feature vector Y; P(D j ) represents the prior probability; P(Y∣D j ) represents the likelihood function; P(Y) represents the marginal probability of observing the feature vector Y, which is usually a normalization constant. j

[0034] Preferably, the backup server is further configured to construct a classification criterion for the data category tree, including:

[0035] An node, as the starting point of the file path, contains n first-level data categories; Bn node, under each first-level data category, contains n second-level data attribute categories; Cn node, under each second-level data category, contains n third-level data categories.

[0036] The present invention also provides an incremental data backup method based on a Bayesian model, including the following steps:

[0037] S1. Construct a data category tree with a tree structure on the backup server according to a preset data classification criterion;

[0038] S2. Identify the changes of data files in the client through an event notification mechanism at the client side, and only back up the changed data files;

[0039] S3. Perform data compression processing on the changed data files by using a data compression algorithm;

[0040] S4. For data files with known categories, directly transfer the compressed data to the corresponding data categories of the backup server after data compression, that is, the file path;

[0041] S5. For data files with unknown categories, it is necessary to extract features from the compressed data files and represent them with an m-dimensional feature vector;

[0042] S6. Represent each data sample in the backup server with an m-dimensional feature vector Y = {y1, y2, …, y m}; and label each data sample with a known data category Di to form a training data set;

[0043] S7. Use the training data set to construct and train a Bayesian model, and initialize the prior probability and the likelihood function;

[0044] S8. Input the m-dimensional feature vector represented by the compressed changed data file into the trained Bayesian model, and output the posterior probability of belonging to a certain data category; select the data category with the largest posterior probability as the classification result.

[0045] The beneficial effects of the present invention:

[0046] ​1. By utilizing the event notification mechanism built into the client, this method can capture change events such as data modification and deletion in real time, and only back up the data files that have changed. This not only significantly reduces unnecessary data transmission, but also significantly reduces the network bandwidth occupancy and backup time during the backup process, improving the overall efficiency of the system.

[0047] 2. A tree-like data category structure including first-level, second-level, and third-level data attribute categories is constructed in the backup server, and intelligent classification is combined with the Bayesian model, enabling data files of unknown categories to be accurately classified into different data categories (i.e., file paths) during backup, achieving efficient and secure incremental data backup with orderly management.

[0048] 3. Use the data samples of known categories in the backup server to train the Bayesian model, and use prior and likelihood information to make classification decisions on the data, ensuring sufficient classification basis for unknown data and making the classification backup process more intelligent and accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for describing the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0050] Figure 1 It is a structural framework diagram of an incremental data backup platform based on the Bayesian model of the present invention.

[0051] Figure 2 It is a step flow chart for constructing a Bayesian classification model in an incremental data backup platform based on the Bayesian model of the present invention.

[0052] Figure 3 It is a step flow chart of an incremental data backup method based on the Bayesian model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0054] Please refer to Figures 1 - 3 As shown, an incremental data backup platform based on the Bayesian model includes a client, a data processing module, a backup server, and a network communication module;

[0055] The client includes multiple user data memories for storing user data; it uses a built-in event notification mechanism to detect data change events in real time and push them to the data processing module;

[0056] The data processing module is used to retrieve and obtain the changed data files from the client according to the metadata included in the event notification after receiving the event notification, and compress the changed data files;

[0057] It is also used to extract features from the compressed data files and represent them using an m-dimensional feature vector Y;

[0058] It is also used to construct and train a Bayesian classification model based on the data samples of known categories in the backup server, and output the category with the highest posterior probability as the classification result of the data files of unknown categories, that is, the file path;

[0059] The backup server includes a cloud server or a local server, which is used to receive and store the backup data uploaded by the client; it is also used to construct the classification criteria of the data category tree and back up different categories of data files to the corresponding file paths;

[0060] The network communication module is used for data transmission between the client, the backup server and the data processing module.

[0061] Specifically, the client monitors file changes in real time, generates a packet containing metadata through the event notification mechanism, and starts data compression when data changes. The compressed data files and their related information are uploaded through the network communication module.

[0062] After receiving the notification, the data processing module obtains the changed data from the client, compresses the data (using the LZW algorithm), extracts features (constructs an m-dimensional feature vector), and merges these vectors with the existing training data set in the backup server to construct and train a Bayesian model. Subsequently, the feature vector of the new data is input into the model to obtain the classification result with the highest posterior probability, that is, the corresponding file path. In the specific implementation process, the data processing module is the UEFI firmware driver manager, which runs a specific function program written by developers in the UEFI environment, and can be an open-source tool using EDK2 to implement functions such as data acquisition, compression, feature extraction, and Bayesian model construction and training.

[0063] The backup server receives the classified backup data and stores different categories of data files in their respective corresponding file paths according to the preset data category tree, realizing orderly management and quick recovery.

[0064] The network communication module runs through the entire process, ensuring the secure, stable, and efficient transmission of data between modules and achieving seamless docking between modules within the system. During the specific implementation process, the network communication module uses the network protocol stack provided by UEFI (such as EFI_TCP4_PROTOCOL) for communication.

[0065] Specifically, the backup platform of the present invention constitutes a backup system with clear logic and perfect functions through event notification and data compression of the client, feature extraction and Bayesian classification of the data processing module, tree-like classification storage of the backup server, and efficient data transmission of the network communication module. The modules cooperate closely with each other, which can not only accurately identify and classify backup data, but also reduce bandwidth occupancy and improve the efficiency of data management and recovery, thus meeting the high requirements for data backup in modern large-scale data environments.

[0066] Furthermore, the event notification mechanism in the client is that the client generates a packet containing metadata for the data change event and pushes it to the data processing module; the data change event includes data modification or data deletion operations; the metadata includes the file path, change timestamp, and user ID of the data.

[0067] Specifically, the client captures file changes through the file monitoring API or log listening mechanism provided by the operating system. When a data change event (including modification or deletion) occurs, the system generates a notification packet containing metadata, and the metadata includes information such as the file path, change timestamp, and user ID of the data. Through the network communication module, the event packet is sent to the data processing module, providing a trigger signal and data location basis for subsequent processing. In this way, an efficient and reliable event-driven data file synchronization solution is realized, which is applicable to data backup in local or cloud service environments.

[0068] Furthermore, the method for compressing the changed data file in the data processing module is to use the LZW data compression algorithm to generate an adaptive string of the data file in the form of data encoding, and completely record non-repeating characters; the specific steps include:

[0069] a. Initialize a string table containing all single characters, including single characters represented by 0 to 255;

[0070] b. Set the S string to be empty and mark it as the 0 code;

[0071] c. Select characters of the string class of the character sequence according to the data arrangement of the root data category tree;

[0072] d. Determine whether the concatenation result of s + string appears in the intermediate string; if it exists, set s + string = s and re-select characters;

[0073] e. Substitute the s - code in the intermediate string into the code stream to complete the data encoding of the string;

[0074] f. Loop through steps b - e like this to complete the compression of the data file.

[0075] Specifically, the LZW algorithm first needs to initialize a string table that contains all possible single characters. For the data in the database, it is usually stored in bytes. Therefore, all ASCII characters from 0 to 255 need to be added to the string table in advance. This initialization process ensures that each possible character has an initial encoding, laying a foundation for subsequent string matching and compression. Among them, the calculation formula for the encoding efficiency of the LZW algorithm is as follows:

[0076] E=(Bits - per - sym)·(Cnt - of - str) / indexBits·g(ψ)

[0077] Where: indexBits represents the bit modulation result of the position information output within the LZW string. Bits - per - sym represents the number of bits of a single input encoding within the string. Cnt - of - str represents the number of characters of the input encoding within the string.

[0078] After receiving the event notification, the data processing module locates and obtains the changed data file from the client according to the metadata, and compresses the data using data compression algorithms such as LZW. This compression process ensures that the information is not lost when the compressed data file is restored by initializing the character string table, gradually matching strings, and encoding to generate adaptive strings. At the same time, it reduces the volume of the transmitted data and the bandwidth occupancy.

[0079] Furthermore, the m - dimensional feature vector Y in the data processing module is expressed as: Y = {y1, y2, …, y m};

[0080] y1 represents the normalized value of the string feature vector after the data file is compressed, y2 represents a certain - dimensional feature vector after the text feature of the compressed data file is processed by the word vector, and so on. y m represents a certain - dimensional feature vector obtained through feature extraction after the data file is compressed; ensuring that the feature vectors of each data sample have the same m dimensions.

[0081] Specifically, an appropriate feature extraction method is selected according to the data type. Comparing with the compressed structured data in the present invention, the compressed string is directly used as the feature vector, or the standardized value is used as the feature. In addition, for text data, the bag-of-words model or word vectors (such as Word2Vec, BERT, etc.) are used to convert the text into a vector representation. And so on; each single feature that has been preprocessed and selected is combined into a feature vector dataset Y in a certain order. And it is saved in the form of an array set or matrix for the input of the subsequent model. This feature extraction processing method can provide solid basic information for subsequent classification, probability calculation, and incremental backup, and realize the fine quantization description and efficient processing of data samples.

[0082] Further, the specific implementation method for constructing and training the Bayesian classification model in the data processing module includes:

[0083] Define the prior probability distribution: count the occurrence frequencies of each category in the training dataset, that is, the prior probabilities of each category; assume that the number of samples of data category D j is X, and the total number of samples is Z, then the expression of the prior probability is:

[0084]

[0085] In the formula, P(D j ) is the prior probability, which reflects the initial distribution of the samples belonging to each data category before observing the feature vector Y;

[0086] Construct the likelihood function: count the occurrence frequencies of each feature vector under the given category D j to construct the likelihood function; its expression is:

[0087]

[0088] In the formula, P(Y∣D j ) represents the likelihood function of observing the entire feature vector set Y under the data category D j , that is, the probability that the data generates the entire feature vector set; P(y i ∣D j ) represents the probability of observing a single feature vector y j under the data category D i ;

[0089] Posterior probability calculation: According to Bayes' theorem, calculate the posterior probability that the feature vector Y belongs to the data category D j ; its expression is:

[0090]

[0091] where P(D j ∣Y) represents the posterior probability that the data belongs to class D after observing the feature vector Y j ; P(D j ) represents the prior probability; P(Y∣D j ) represents the likelihood function; P(Y) represents the marginal probability of observing the feature vector Y, which is usually a normalization constant.

[0092] Specifically, the Bayesian theory is a statistical method based on probabilistic inference. Its core idea is to use prior knowledge (prior probability) and conditional probability to combine newly obtained data evidence to calculate the posterior probability. It is an intuitive and mathematically rigorous method of probabilistic inference. When training a Bayesian model, we need to estimate the prior probability and conditional probability through labeled data, and then use methods such as smoothing techniques and maximum likelihood estimation to ensure the accuracy of the model parameters. The trained Bayesian model can output the posterior probability of each class based on the input multi-dimensional feature vector, thereby achieving accurate classification of the data. This method shows the advantages of high efficiency, accuracy, and interpretability in the fields of data classification, incremental backup, etc.

[0093] In the specific implementation process, use the data samples of known classes in the backup server, label each sample with its known data class Di to form a training data set. Based on the training data set, construct a Bayesian classification model, and initialize the prior probability P(Di) and conditional probability (likelihood function) P(Y∣Di); use maximum likelihood estimation and necessary smoothing techniques to train the model so that it can accurately capture the data class distribution. Use the training data, through statistical analysis (frequency statistics or maximum likelihood estimation), to estimate the prior probability P(D) under all classes D and the conditional probability P(yi∣D) of each feature. Then input the m-dimensional feature vector obtained by feature extraction from the compressed changed data file into the trained Bayesian model to calculate the posterior probability P(Di∣Y) of each class. The system selects the class with the maximum posterior probability as the classification result of the data according to the Bayesian formula, and feeds back the classification result to the backup server, which decides to store the data under the corresponding file path (i.e., the corresponding node in the data class tree).

[0094] Furthermore, the backup server is also used to construct the classification criteria of the data class tree, including:

[0095] The An node, as the starting point of the file path, contains n first-level data classes; the Bn node, under each first-level data class, contains n second-level data attribute classes; the Cn node, under each second-level data class, contains n third-level data classes.

[0096] Specifically, while receiving client data, the backup server classifies and stores the data according to the data category tree, and data files of different categories are stored in corresponding file paths. The representation form of the file path is: An / Bn / Cn. This not only realizes the hierarchical management of data but also provides a clear indexing basis for the rapid retrieval and recovery of data.

[0097] The present invention also provides an incremental data backup method based on a Bayesian model, including the following steps:

[0098] S1. Build a tree-structured data category tree on the backup server according to a preset data classification standard;

[0099] S2. On the client side, identify the changes of data files in the client through an event notification mechanism, and only back up the changed data files; reduce bandwidth consumption and improve backup efficiency.

[0100] S3. Use a data compression algorithm to perform data compression processing on the changed data files;

[0101] S4. For data files of known categories, directly transfer them to the corresponding data categories in the backup server, that is, the file paths, after data compression; reduce the data volume and lower the bandwidth occupancy during the backup process.

[0102] S5. For data files of unknown categories, it is necessary to extract the features of the compressed data files and represent them using an m-dimensional feature vector; ensure that each sample has a unified and quantified description to provide sufficient information for subsequent classification.

[0103] S6. Represent each data sample in the backup server using an m-dimensional feature vector Y = {y1, y2, …, y m}; and label each data sample with a known data category Di to form a training data set; lay the foundation for the training of the Bayesian model.

[0104] S7. Use the training data set to build and train a Bayesian model, and initialize the prior probability and likelihood function; thus train a classifier that can output the posterior probability.

[0105] S8. Input the m-dimensional feature vector represented by the compressed changed data file into the trained Bayesian model, and output the posterior probability of belonging to a certain data category; select the data category with the largest posterior probability as the classification result; that is, determine the file path to which the data file of the unknown category should be assigned. Ensure that the data is stored hierarchically by category for subsequent management and recovery.

[0106] The incremental data backup method based on the Bayesian model of the present invention identifies data changes through an event notification mechanism, and only backs up the changed parts, greatly reducing unnecessary data transmission and reducing backup time and network bandwidth occupancy. A tree-like data category structure including first-level, second-level, and third-level data attribute categories is constructed on the backup server side, and intelligent classification is performed in combination with the Bayesian model, enabling data files of unknown categories to be accurately classified into different categories during backup, facilitating orderly storage and rapid recovery. At the same time, the present invention uses the data samples of known categories in the backup server to construct a training data set through fixed-dimensional feature vectors formed after feature extraction. Subsequently, the Bayesian model is trained using these data, and by statistically calculating prior probabilities and conditional probabilities, sufficient basis is ensured when classifying unknown data. The model can output the category with the highest posterior probability as the classification result, making the entire classification backup process more intelligent and accurate.

[0107] In summary, the incremental data backup method based on the Bayesian model of the present invention solves the deficiencies of traditional backup technologies in data transmission, classification management, and backup decision-making through accurate identification of data changes, intelligent compression and feature extraction, and scientific probability model classification, achieving a backup effect of low bandwidth, high efficiency, orderly management, and security and reliability, providing innovative technical support for data protection in modern large-scale data environments.

[0108] The above content is only an example and explanation of the structure of the present invention. Those skilled in the art of the present technology make various modifications or supplements or use similar methods to replace the described specific embodiments, as long as they do not deviate from the structure of the invention or exceed the scope defined by this claim book, they should fall within the protection scope of the present invention. In the specific embodiments provided in this application, it should be understood that the disclosed methods and processes can be implemented in other ways. For example, the module embodiments described above are only illustrative. If the calculation methods or steps are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. And the aforementioned local data storage includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical discs that can store program codes.

Claims

1. An incremental data backup platform based on a Bayesian model, characterized by: It includes client, data processing module, backup server and network communication module; The client includes a plurality of user data storage devices for storing user data; The built-in event notification mechanism detects data change events in real time and pushes them to the data processing module; The data processing module is used to retrieve and obtain the changed data files from the client according to the metadata contained in the event notification after receiving the event notification, and compress the changed data files; it is also used to extract features from the compressed data files and represent them using m-dimensional feature vectors; it is also used to build and train a Bayesian classification model based on data samples of known categories in the backup server, and output the category with the largest posterior probability as the classification result of the data file of unknown category, that is, the file path; The backup server includes a cloud server or a local server, which is used to receive the backup data uploaded by the client for storage; It is also used to construct the classification standard of the data category tree and back up the data files of different categories to the corresponding file paths; The network communication module is used for data transmission between the client, backup server and data processing module.

2. The incremental data backup platform based on the Bayesian model according to claim 1, characterized in that: The data change event includes a data modification or data deletion operation.

3. The incremental data backup platform based on the Bayesian model according to claim 1, characterized in that: The event notification mechanism in the client generates a package containing metadata for the data change event and pushes it to the data processing module.

4. The incremental data backup platform based on the Bayesian model according to claim 3, characterized in that: The metadata includes the file path, change timestamp and user ID of the data.

5. The incremental data backup platform based on the Bayesian model according to claim 1, characterized in that: The method for compressing the changed data file in the data processing module is to use the LZW data compression algorithm to generate an adaptive character string of the data file in a data encoding manner, and completely record non-repeated characters; The specific steps include: a. Initialize a string table containing all single characters, including single characters represented by 0 to 255; b. Set the S string to empty and mark it as 0 code; c. Select characters of the character sequence string class according to the data arrangement of the data category tree; d. Determine whether the concatenation result of s+string appears in the middle string; if so, set s+string=s and reselect characters; e. Substitute the s code in the middle string into the code stream to complete the data encoding of the string; f. Repeat steps b to e to complete the compression of the data file.

6. The incremental data backup platform based on the Bayesian model according to claim 1, characterized in that: The m-dimensional feature vector in the data processing module is represented as: Y = {y1, y2, ..., y m }; In the formula, y1 represents the standardized value of the string feature vector after the data file is compressed, y2 represents a certain dimension feature vector after the text feature of the data file is compressed and processed by the word vector, and so on. m Represents a feature vector of a certain dimension obtained by feature extraction after data file compression; ensure that the feature vector of each data sample has the same m dimension.

7. The incremental data backup platform based on the Bayesian model according to claim 1, characterized in that: The specific implementation method of constructing and training the Bayesian classification model in the data processing module includes: Define the prior probability distribution: Count the frequency of occurrence of each category in the training data set, that is, the prior probability of each category; let the data category D j The number of samples is X, and the total number of samples is Z, then the expression of prior probability is: In the formula, P(D j ) is the prior probability, which reflects the initial distribution of samples belonging to each data category before the feature vector Y is observed; Constructing the likelihood function: Statistics in a given category D j The frequency of occurrence of each eigenvector is used to construct the likelihood function; its expression is: In the formula, P(Y|D j ) indicates that in data category D j The likelihood function of the entire feature vector set Y observed under , that is, the probability that the data generates the entire feature vector set; P(y i ∣D j ) indicates that in data category D j A single eigenvector y is observed under i probability; Posterior probability calculation: According to Bayes' theorem, calculate whether the feature vector Y belongs to data category D j The posterior probability of .

8. The incremental data backup platform based on the Bayesian model according to claim 7, characterized in that: The calculation expression of the posterior probability in Bayes' theorem is: In the formula, P(D j |Y) means that after observing the feature vector Y, the data belongs to category D j The posterior probability of j ) represents the prior probability; P(Y|D j ) represents the likelihood function; P(Y) represents the marginal probability of observing the feature vector Y, which is usually a normalization constant.

9. The incremental data backup platform based on the Bayesian model according to claim 1, characterized in that: The backup server is also used to construct a classification standard for the data category tree, including: The An node, as the starting point of the file path, contains n first-level data categories; the Bn node, under each first-level data category, contains n second-level data attribute categories; the Cn node, under each second-level data category, contains n third-level data categories.

10. The incremental data backup method based on a Bayesian model according to claim 1 is applied to the incremental data backup platform based on a Bayesian model according to any one of claims 1 to 9, characterized in that: The following steps are involved: S1. Construct a data category tree with a tree structure according to a preset data classification standard on the backup server; S2. The client uses an event notification mechanism to identify changes in data files in the client and only backs up the changed data files; S3. Using a data compression algorithm to compress the changed data files; S4. After data files of known categories are compressed, they are directly transferred to the corresponding data categories of the backup server, that is, the file paths; S5. For data files of unknown categories, it is necessary to extract features from the compressed data files and represent them using m-dimensional feature vectors; S6. Each data sample in the backup server is represented by an m-dimensional feature vector Y = {y1, y2, ..., y m }; and label each data sample with a known data category Di to form a training data set; S7. Use the training data set to build and train the Bayesian model and initialize the prior probability and likelihood function; S8. Input the m-dimensional feature vector represented by the compressed change data file into the trained Bayesian model, and output the posterior probability of belonging to a certain data category; select the data category with the largest posterior probability as the classification result.

Citation Information

Patent Citations

  • XML file classification method and system

    CN104281573A

  • Data classification method and apparatus

    CN104951791A

  • A data backup method and device

    CN109725895A

  • File importing and filing method, electronic equipment and storage medium

    CN111611211A