Triggering the backup of an object based on the relevance of the changed category of the object
By categorizing and weighting data changes to prioritize critical backups, the method addresses inefficiencies in existing backup systems, enhancing efficiency and reducing data loss risks in IT systems.
Patent Information
- Application Number
- CN202111118645.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-28
- Filing Date
- 2021-09-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-09-24
AI Technical Summary
In the process of data backup, the prior art cannot effectively distinguish the nature of data changes, resulting in the failure to effectively reduce the risk of backup resources and data loss.
By detecting the change category and its correlation weight of the data object, calculating the trigger indicator, selecting the key object for instant backup, and postponing the backup of non-critical objects, semantic control of the backup is realized.
Improves the efficiency and effectiveness of backups, saves storage space and time, and reduces the risk of loss of high-value data.
Smart Images

Figure CN114328006B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of information technology. More specifically, the present disclosure relates to the backup of information technology systems. Background Art
[0002] Information technology systems typically store large amounts of data, including quite critical data (e.g., because it is related to business / production activities, services provided, etc.). Due to many reasons, the data stored in information technology systems may be (permanently) lost. For example, this data loss may be due to failures (such as hardware crashes, data corruption, software vulnerabilities, supplier dissolution), disasters (such as fires, natural events), accidental actions (such as deletions, errors), or criminal attacks (such as theft, malicious code). When the data cannot be recovered (e.g., passwords cannot be recovered in the case of transactions), the data loss may lead to serious consequences. In any case, data loss results in relatively high costs (e.g., due to the temporary unavailability of services based on them and the time spent on recreating the data).
[0003] Therefore, backups of the data are performed regularly to avoid or at least limit the risk of its loss. Generally, a backup is a copy of data that can be used to restore its original version in case its original version is lost; for security reasons, the backup is preferably stored in a location different from one of the original data (so as to be still available even in case of its disruption).
[0004] Generally, backups are performed regularly (such as nightly) or in response to a manual request (such as after a large update of the data). Backups can be of different types. For example, a full backup involves a complete copy of the data; full backups are easy to restore, but they are time-consuming and occupy large storage spaces. An incremental backup only involves the changes relative to the previous backup (usually, between successive full backups at a lower frequency); incremental backups are faster than other types of backups and occupy less storage space; however, the corresponding restoration is time-consuming because they need to start from the last full backup and then apply all subsequent incremental backups. A differential backup only involves the changes relative to the previous full backup; differential backups are faster than full backups and occupy less storage space, and their restoration is faster than that of incremental backups; however, the restoration of differential backups is slower than that of full backups, and differential backups occupy more storage space than incremental backups. Summary of the Invention
[0005] A simplified overview of the present disclosure is presented herein to provide a basic understanding thereof; however, the sole purpose of this overview is to introduce some concepts of the present disclosure in a simplified form as a prelude to its more detailed description below, and this overview should not be construed as an identification of its key elements nor as a delineation of its scope.
[0006] Generally speaking, the present disclosure is based on the idea of triggering backups according to the relevance of changes.
[0007] Specifically, an exemplary embodiment provides a method for controlling backups of objects in an information technology system. The method includes operations such as determining a change category of changes of an object that has been changed since its previous backup. Additionally, in this method, based on the change category of the change of the object and its relevance weight, a backup of the changed object is triggered according to a corresponding trigger indicator.
[0008] According to one aspect of the present invention, there is a method, computer program product, and / or system that performs the following operations (not necessarily in the following order): detecting, by a computing system, one or more changed objects among objects that have been changed since their previous backup; determining, by the computing system, corresponding changes of the changed objects; determining, by the computing system, corresponding change categories of the changes; calculating, by the computing system, a corresponding trigger indicator for each changed object according to the change category of the change of the changed object and the relevance weight of the change category, the relevance weight indicating the relevance of the corresponding change category; selecting, by the computing system, one or more critical objects among the changed objects according to its trigger indicator; and triggering, by the computing system, a corresponding backup of the critical objects. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The solution of the present disclosure, as well as its further features and advantages, will be best understood by referring to its detailed description given below only by way of non - limiting indication and read in conjunction with the drawings (wherein, for simplicity, corresponding elements are denoted by the same or similar reference numerals and their explanations are not repeated, and the name of each entity is generally used to represent its type and its attributes, such as value, content, and representation).
[0010] Figure 1A 、 Figure 1B 、 Figure 1C and Figure 1D illustrate the general principle of the solution according to an embodiment of the present invention;
[0011] Figure 2 illustrates a schematic block diagram of an information technology system according to an embodiment of the present invention;
[0012] Figure 3 illustrates the main software components that can be used to implement the solution according to an embodiment of the present invention; and
[0013] Figure 4A and Figure 4B illustrates an activity diagram depicting the activity flow according to an embodiment of the present invention. DETAILED DESCRIPTION
[0014] Specifically refer toFigures 1A - 1D , which shows the general principle of a solution according to an embodiment of the present disclosure.
[0015] Starting from Figure 1A , the information technology system 100A includes a number of objects to be backed up to limit the risk of their loss; the objects are (software) artifacts that store data (such as files, databases, programs, etc.). Detect the objects that have been changed (since their previous backup); for each of these changed objects, determine one or more (object) changes. Figure 1A It includes the following: object 102A, object 104A, object 106A, object 108A, and change list 110A.
[0016] Moving to Figure 1B , determine the (change) category of each change (as shown by the change list 110B in the environment 100B); each change category groups together changes with similar natures. For example, change categories can involve content (such as the text of a document, the values / formulas of a spreadsheet, the fields of a database), comments (such as sentences / cells), operation parameters (such as applications, systems, networks), and so on.
[0017] Moving to Figure 1C , calculate a trigger indicator for each changed object according to the change category of its change (as shown by the change list 110C in the environment 100C) and according to the relevance weight of the change category (such as according to its weighted sum). The relevance weight of each change category indicates the relevance of the change category (such as for the information technology system); for example, the relevance weight is high for critical operation parameters and low for comments.
[0018] Moving to Figure 1D , select possible critical objects among the changed objects according to the corresponding trigger indicator. For example, a critical object is an object whose change belongs to a change category with high relevance (such as critical operation parameters). Then, trigger the backup of the critical objects. On the contrary, for other changed objects, such as objects whose changes belong to change categories with low relevance (such as comments), no action is taken. Figure 1D It includes the following: object 102D, object 104D, object 106D, object 108D, and change list 110D.
[0019] The above solution adds semantic functions to the control of backups. In fact, backups are now triggered according to the nature of the changes of the objects (not just according to the amount of the changes and / or the type of the objects, even though they can be further considered for this purpose).
[0020] In this way, backups can be performed when they are actually useful. This makes the control of backups more efficient. For example, it is now possible to distinguish between actually relevant changes (and then backup them as soon as possible) and other changes that are essentially irrelevant (and then can be backed up later). For example, when an object is affected by a change in even a single critical operation parameter, its backup is immediately triggered; in contrast, when an object (even of a relevant type) is affected by (even a large number of) comment changes, its backup can be postponed.
[0021] Therefore, time and storage space for relatively low-importance data can be saved, and at the same time, the risk of loss of high-value data can be reduced.
[0022] Now refer to Figure 2 , which shows a schematic block diagram of an information technology system 200 in which the solutions according to the embodiments of the present disclosure can be practiced.
[0023] The information technology system 200 includes a number of server machines, or simply servers 205d, 205c, that communicate with each other via a (telecommunication) network 210 (e.g., Internet-based). Servers 205d and 205c include one or more data servers 205d and a control server 205c (or more). Each data server 205d stores one or more objects to be backed up, and the objects are used to implement the services provided by the data server 205d (e.g., WAS, DBRM, SaaS, etc.). The control server 205c controls the backup of the objects stored in the data servers 205d.
[0024] Each of the above-mentioned servers 205d and 205c includes a number of units connected to each other via a bus structure 215 having one or more levels. Specifically, one or more microprocessors (μP) 220 provide the logical capabilities of the servers 205d, 205c; the non-volatile memory (ROM) 225 stores the basic code for the boot programs of the servers 205d, 205c, and the volatile memory (RAM) 230 is used by the microprocessor 220 as a working memory. The servers 205d, 205c are provided with a mass storage 235 (e.g., a storage device implementing the data center of the servers 205d, 205c) for storing programs and data. In addition, the servers 205d, 205c include a number of controllers for peripheral devices, or input / output (I / O) units 240; for example, the peripheral devices 240 include a network adapter that is used to insert the servers 205d, 205c into the data center and then connect them to the console of the data center (e.g., a personal computer, which is also provided with a drive for reading / writing removable storage units such as USB keys) for its control and the switch / router subsystem of the data center for its communication with the network 210.
[0025] Now referring to Figure 3 , the main software components can be used to implement the solutions according to the embodiments of the present disclosure.
[0026] Specifically, all software components (programs and data) are collectively denoted by reference numeral 300. The software components 300 are typically stored in a mass storage and (at least partially) loaded into the working memories of the servers 250d, 205c during program execution. For example, the programs are installed into the mass storage by reading from a removable storage unit and / or downloading from a network. In this regard, each program can be a module, segment, or portion of code including one or more executable instructions for implementing a specified logical function.
[0027] Starting from each data server 205d (only one is shown in the figure), it includes the following components. The data server 205d stores the corresponding objects to be backed up, denoted by reference numeral 305 (e.g., the configuration files of WAS, the databases of DBRM, the documents of SaaS, etc.). The configuration repository 310 defines the configuration information for the backup of the objects 305. Specifically, the configuration repository 310 lists the objects 305 to be backed up by corresponding identifiers (e.g., the absolute path of its file). The configuration repository 310 stores the control period for controlling the backup according to the trigger indicator and the forced period for performing the backup indiscriminately (e.g., to comply with audit requirements); the forced period is strictly higher than the control period (e.g., equal to 10 - 50 times of it). In this way, the control of the backup can be performed at a high frequency (such as every 1 - 24 hours) to backup critical objects as soon as possible (fast processing), and in any case, backup all objects at a low frequency (such as every 1 - 7 days) (slow processing). The configuration repository 310 stores the check instructions for the objects 305; for each (object) type of the objects 305 (indicated by the corresponding identifier, e.g., its file extension for configuration files, databases, documents, spreadsheets, etc.), the check instructions specify how to determine that the (object) change belongs to one or more change categories (e.g., parameters, permissions and comments of configuration files, fields, tables, permissions and comments of databases, text, permissions and comments of documents, cells, permissions and comments of spreadsheets, etc.).
[0028] The configuration repository 310 stores the corresponding correlation weights of the (trigger) parameters for calculating the trigger indicator of the objects 305. Specifically, the trigger parameters include all possible change categories of the objects 305; and, the trigger parameters may further include the (elapsed) time since the previous backup of each of the objects 305, one or more topics defined by the user of the data server 205d (indicated by the corresponding identifier, such as its user ID) that can change the objects 305, all possible object types of the objects 305, and the (change) frequency of the change of each object 305.
[0029] The backup agent 315 runs the backup of the objects 305. The backup agent 315 accesses (in read mode) the objects 305 and it copies them to a backup repository that is stored in another server (not shown in the figure) at a remote location. The change detector 320 detects the objects 305 that have changed since its previous backup. The change detector 320 accesses (in read mode) the objects 305 and the configuration repository 310. The scheduler 325 activates the change detector 320 (according to a control cycle) and the backup agent 315 (according to a forced cycle). The scheduler 325 accesses (in read mode) the configuration repository 310. The change checker 330 determines the changes of the objects 305. The change checker 330 accesses (in read mode) the configuration repository 310 and it utilizes one or more tracking services 335. The tracking services 335 track the changes of one or more change categories. The tracking services 335 can be local (as shown in the figure) or remote, such as running on a control server 205c (not shown in the figure); for example, the tracking services 335 include: a version control service (VCS), such as RCS, for determining the changes of the content (such as parameters, comments, text, and cells) of a document / spreadsheet; a pattern matching engine for determining the changes of binary type objects via system / application logs or audit files (such as permissions, fields, and tables).
[0030] Moving to the control server 205c, it includes the following components. The change repository 340 indicates the changes that have been determined in the data server 205c. For example, the change repository includes an entry for each changed object; the entry stores the identifier of the changed object, its object type, the change category of the corresponding change, the identifier of the data server 205d (e.g., its network address) where the changed object is stored, the identifier of the (change) agent that changed it, the identifiers of the possible (affected) objects affected by the change (data server 205d and objects), the (backup) result (i.e., success or failure), and the (backup) time of the possible backup of the changed object. The change repository 340 is accessed (in write mode) by the change checker 330 of each data server 205c.
[0031] Like the IBM Tivoli Application Dependency Discovery Manager (TADDM) of IBM Corporation (its trademark), the dependency detector 345 detects the affected objects of each changed object of an information technology system in the same or different data servers. The dependency detector 345 accesses (in read / write mode) the system topology repository 350, which stores the system topology of the information technology system defined by the correlations between its objects (e.g., defined by clusters, dependencies, etc.). The backup manager 355 manages the control of the backup of the data server 205d. For example, the backup manager 355 is implemented by a neural network based on a linear regression model.
[0032] A neural network is a data processing component that operates similarly to the human brain. A neural network includes: basic processing elements (neurons) that perform operations defined by an activation function based on corresponding weights; the neurons are connected via one-way channels (synapses) that transmit information between them. The neurons are organized in layers that perform different operations, always including an input layer for receiving input information and an output layer for providing output information. In a linear regression model, the neural network has only two layers (i.e., the input layer and the output layer, with no hidden layers between them), and the activation function is linear. The backup manager 355 accesses (in read mode) the change repository 340, and it controls the backup agent 315 of each data server 205d. The training module 360 uses supervised learning techniques to train the neural network of the backup manager 355. The training of the neural network is a process of finding the (optimal) values of its weights, which optimizes the performance of the neural network. In supervised learning techniques, the training is based on a training set of input information and corresponding output information that has been manually determined to be correct. The training module 360 accesses (in read mode) the configuration repository 310 of each data server 205d, and it receives the corresponding feedback indicators of the backups input by the operator; each feedback indicator provides an indication of the utility of the corresponding backup. The anomaly detector 365 detects possible anomalies in the changes of the changed objects of the system topology of the information technology system. The anomaly detector 365 accesses (in read mode) the change repository 340 and the system topology repository 350.
[0033] Now refer to Figures 4A - 4B , which shows an activity diagram that describes the activity flow related to the implementation of the solution according to an embodiment of the present disclosure.
[0034] Specifically, this activity diagram represents an exemplary process that can be used to control backups in an information technology system using the method 400. In this regard, each block may correspond to one or more executable instructions for implementing a specified logical function on each workstation.
[0035] Starting from the lane of the general data server, whenever the scheduler detects that the control cycle has expired (since the previous control of the backup), the process passes from block 402 to block 404. In response to this, a change detector that takes into account the (current) objects listed in the configuration repository enters a loop (starting from the first in any arbitrary order). The change detector verifies in block 406 whether an object has changed since its last backup. For this purpose, for example, the change detector is registered with the corresponding API of the operating system of the data server in order to receive notifications whenever each object is modified; in response to this, the change detector asserts the corresponding change flag (which is de-asserted whenever an object is backed up). Depending on the result of this verification, the active process branches at block 408. If the object has been changed (the change flag is asserted), the change checker determines the object type of the changed object (based on its file extension) at block 410. The change checker determines any changes to the changed object (via the check instructions corresponding to its object type) at block 412. Thus, the change checker determines the change category of the change at block 414 (in response to their detection by the corresponding tracking service). The change checker determines the body of the changed object that has been changed at block 416 (e.g., based on its metadata). The change checker uploads the information thus determined for the changed object (indicated by its identifier), i.e., the object type, change category, and body, together with the identifier of the data server, to the control server at block 418. Then, the process continues to block 420; if the object has not been changed (the change flag is de-asserted), it also reaches the same point directly from block 408. At this time, the change detector verifies whether the last object has been processed. If not, the process returns to block 404 to repeat the same operation for the next object. On the contrary, once all objects have been processed, the loop is exited by returning to block 402 and waiting for the expiration of the next control cycle.
[0036] Move to the lane of the control server, (in this case at block 418), information related to any changed objects being uploaded by the corresponding data server is added to the change repository (added to its corresponding (new) entry) at block 422. At block 424, the backup manager determines the time elapsed since the previous backup of the same changed object (equal to the difference between the current time and the backup time of the most recent entry in the change repository having the same identifier for the changed object). The backup manager determines the change frequency of the changed object at block 426 (equal to the average of the differences between the backup times of each pair of consecutive entries in the change repository having the identifier of the changed object). At block 428, the backup manager determines the input information for its neural network based on the trigger parameters of the changed object. Specifically, the trigger parameters include the change category, the elapsed time, the subject, the object type, and the change frequency of the changed object. The input information alternatively includes the corresponding (category) flag for the change category (asserted for the change category of the changed object, otherwise de-asserted), the elapsed time, the corresponding (subject) flag for the subject (asserted for the subject of the changed object, otherwise de-asserted), the corresponding (type) flag for the object type (asserted for the object type of the changed object, otherwise de-asserted), and the change frequency. Then, the change manager applies the input information to the neural network, which provides corresponding output information including a trigger indicator for the changed object (e.g., in the range from 0 to 1 in ascending order of the utility of the backup of the changed object).
[0037] According to the trigger indicator, the active process branches at block 430. If the trigger indicator is (possibly strictly) higher than the trigger threshold (e.g., 0.6 - 0.8), this means that the backup of the changed object may be very useful. Therefore, the dependency detector determines at block 432 the possible (affected) objects that may be (directly or indirectly) affected by the change of the changed object, and it updates the change repository accordingly; for this purpose, the dependency detector determines (based on the system topology in the corresponding repository) the affected objects related to the changed object, the (additional) affected objects related to these affected objects, and so on, until no additional affected objects are determined. At block 434, the backup manager triggers the backup of the changed object and its affected objects on the corresponding data server, which may be the same or different from each other (as indicated in the change repository). In this way, a consistent configuration of the information technology system (given by the changed object and its affected objects) can be backed up; for example, in the case of a change to a WAS application, the WAS application is backed up together with the corresponding database. Then, the process returns from block 434 or directly from block 430 to block 422 (if the trigger indicator is (possibly strictly) lower than the trigger threshold, meaning that the backup of the changed object is most likely quite useless), waiting for the upload of information related to the next changed object.
[0038] Move to the lane of each data server where the backup has been triggered by the control server at block 434 (the same one as in the figure above), and activate the backup agent accordingly at block 436 for the changed / affected objects stored in the data server. Whenever the scheduler detects that the forced period has expired since the previous (full) backup of the data server, it also reaches the same point from block 438; in this case, the backup agent is activated for all the objects of the data servers listed in the configuration repository. In both cases, the backup agent copies the objects to be backed up to the backup repository at block 440. The backup agent returns the backup (backup) result to the control server (success or failure of the corresponding object) at block 442. Then, the process returns to block 436, waiting for the next activation of the backup agent.
[0039] Referring back to the control server again, the backup manager receives the backup result from any data server at block 444 (from block 442 in the example being discussed). In response to this, the control manager updates the corresponding entry in the change repository accordingly at block 446 (by adding the backup result and its backup time set to the current time). Then, the process returns to block 444, waiting for the next backup result.
[0040] In a completely independent manner, whenever a (training) event occurs that triggers the training of the neural network of the backup manager by the training module; for example, this occurs in response to adding a (new) data server to the information technology system (manually notified by the operator) or in response to the running of a (previous) backup (notified by the backup manager). The active process then branches according to the training event. Specifically, in the case of a new data server, blocks 452 - 454 are executed, while in the case of a previous backup, blocks 456 - 460 are executed; in both cases, the process then returns to block 448, waiting for the next training event. Specifically referring to block 452 (new data server), the training module retrieves the relevance weights from the configuration repository of the new data server (where the relevance weights have been set manually by the operator or have been set automatically to random values). The training module initializes the neural network accordingly at block 454. In the particular example being discussed (a neural network based on a linear regression model), the neural network has an input layer including corresponding neurons for all the trigger parameters (change category, elapsed time, subject, object type, and change frequency) of all the data servers and an output layer with a single neuron for the trigger indicator. The activation function of each neuron in the input layer multiplies an input message corresponding to its trigger parameter by its relevance weight, and the activation function of the neuron in the output layer sums the results provided by all the neurons in the input layer. Thus, the trigger indicator is calculated as the (weighted) sum corresponding to the trigger parameters:
[0041]
[0042] where TI is a trigger indicator (e.g., further normalized to a range from 0 to 1), Ncc is the number of changed categories, Wcc i is the relevance weight of the i-th changed category, Fcc i is the category flag of the i-th changed category (1 when asserted, 0 when de-asserted), Wet is the relevance weight of the elapsed time, te is the elapsed time, Nu is the number of subjects, Wu i is the relevance weight of the i-th subject, Fu i is the subject flag of the i-th subject (1 when asserted, 0 when de-asserted), Not is the number of object types, Wot i is the relevance weight of the i-th object type, Fot i is the type flag of the i-th object type (1 when asserted, 0 when de-asserted), Wcf is the relevance weight of the change frequency, and fc is the change frequency. Now referring to block 456 (previous backup), the training module prompts the operator to input a corresponding feedback indicator (based on its utility); e.g., the feedback indicator ranges from 0 (the previous backup is considered completely useless) to 1 (the previous backup is considered completely useful). The training module retrieves, at block 458, the trigger parameters that have been used to calculate the trigger indicator of the previous backup (from the change repository). At block 460, the training module trains the neural network (for the backup object) based on the trigger parameters and the feedback indicator. For example, the training is an iterative process based on the Stochastic Gradient Descent (SGD) algorithm. In this case, the training module determines the change in the relevance weight that should reduce the difference between the trigger indicator (provided by the neural network in response to the input information corresponding to the trigger parameters) and the feedback indicator (considered as its correct value); the direction and amount of the change are given by the gradient of the error function with respect to the relevance weight, which is approximated using the backpropagation algorithm. The same operation is repeated until an acceptable difference (between the trigger indicator and the feedback indicator) is obtained or the change in the relevance weight does not provide any significant improvement (meaning that the minimum, at least local, or flat region of the error function has been found). The relevance weight is changed with the addition of random noise to find different (and possibly better) local minima and to distinguish flat regions of the error function. In this way, the relevance weight adapts over time, thereby automatically improving the accuracy of the neural network.
[0043] In a completely independent manner, whenever a (verification) event occurs that triggers a change verification, the process passes from block 462 to block 464. In response to this, the anomaly detector retrieves information about the changes detected after its previous verification (from the corresponding repository), such as for each of them, the change category and the identifier of the data server. At block 466, the anomaly detector verifies these changes against the system topology (from the corresponding repository) according to one or more verification rules, for example. Depending on the result of this verification, the active flow branches at block 468. If one or more (non-conforming) changes do not conform to the system topology, the anomaly detector detects the corresponding anomaly at block 470. This can occur, for example, when a critical operating parameter is changed only in a part of a set of data servers that contribute to providing a particular service; a typical scenario is when the data servers are nodes of a cluster (such as a WAS type) that act as a single logical entity and provide services in response to corresponding requests submitted to them, which are automatically distributed among all the data servers in the cluster that perform the same task. In this case, the anomaly detector outputs an appropriate error message (such as by displaying it on the monitor of the control server); for example, the error message indicates each non-conforming change and the explanation for its non-conformance to the system topology. In this way, the operator can intervene promptly to correct the anomaly. The process then returns from block 470 or directly from block 468 (if no anomaly is detected) to block 462, waiting for the next verification event.
[0044] Naturally, to meet local and specific requirements, those skilled in the art can apply many logical and / or physical modifications and changes to the present disclosure. More specifically, although the present disclosure has been described with a certain degree of particularity with reference to one or more embodiments of the present disclosure, it should be understood that various omissions, substitutions, and changes in form and details, as well as in other embodiments, are possible. Specifically, different embodiments of the present disclosure can even be practiced without the specific details (such as numerical values) set forth in the foregoing description to provide a more comprehensive understanding thereof; conversely, well-known features may have been omitted or simplified so as not to obscure the description with unnecessary details. In addition, it is explicitly contemplated that the specific elements and / or method steps described in connection with any embodiment of the present disclosure can be incorporated in any other embodiment as a matter of general design choice. Moreover, items presented in the same group and in different embodiments, examples, or alternatives should not be construed as being in fact equivalent to each other (but they are separate and autonomous entities). In any case, each numerical value should be read with due allowance for the applicable tolerances; in particular, unless otherwise specified, the terms "substantially", "about", "approximately", etc. should be understood to be within 10%, preferably 5% and still more preferably 1%. In addition, each range of numerical values should be intended to clearly specify any possible number along the continuum (including its endpoints) within that range. Conventional or other qualifiers are only used as labels to distinguish elements having the same name, and they do not in themselves imply any precedence, priority, or order. The terms including, comprising, having, containing, involving, etc. should be intended to have an open-ended, non-exhaustive meaning (i.e., not limited to the items listed), depending on, according to, function, etc. should be intended to be a non-exclusive relationship (i.e., involving possible other variables), the term a / an should be intended to mean one or more items (unless otherwise clearly indicated), and the term means (or any means-plus-function formulation) should be intended to mean any structure adapted or configured to perform the relevant function.
[0045] For example, an embodiment provides a method for controlling the backup of an object in an information technology system. However, the object can be of any number and any type (e.g., parts of the above-mentioned object, different or additional objects), and they can be provided in any information computing system (e.g., having a distributed architecture, having an independent architecture, stored in any number and type of physical / virtual computing machines, etc.).
[0046] In one embodiment, the method includes the following steps under the control of a computing system. However, the computing system can be of any type (see below).
[0047] In one embodiment, the method includes (by a computing system) detecting one or more changed objects among the objects that have been changed since their previous backup. However, the changed objects can be detected in any way (e.g., by registering corresponding notifications, by polling the objects, etc.); in any case, the changed objects can be detected centrally (by receiving an indication thereof from the computing machine storing the objects) or locally (directly on each of them).
[0048] In one embodiment, the method includes (by a computing system) determining the corresponding changes of the changed objects. However, the changes can be of any type (e.g., related to content, metadata, permissions, etc.), and they can be determined in any way (e.g., using a partial, different, or additional local tracking service relative to the above-mentioned part of the local tracking service, using a centralized tracking service such as GIT, etc.); in any case, the changes can be determined centrally (by receiving an indication thereof from the computing machine storing the objects) or locally (directly on each of them).
[0049] In one embodiment, the method includes (by a computing system) determining the corresponding change categories of the changes. However, the change categories can be of any number and any type (e.g., partial, different, or additional change categories relative to the above-mentioned change categories), and they can be determined in any way (e.g., directly defined by the tracking service that has determined the corresponding changes, by differentiating different change categories for the same tracking service according to corresponding rules (such as global and local parameters of a configuration file, values and schemas of a database, etc.)); in any case, the change categories can be determined centrally (by receiving an indication thereof from the computing machine storing the objects) or locally (directly on each of them).
[0050] In one embodiment, the method includes (by a computing system) calculating a corresponding trigger indicator for the changed objects. However, the trigger indicator can be of any type (e.g., an index, a logical value, etc. that can take any range of discrete / continuous values) and they can be calculated in any way (e.g., using machine learning techniques, algorithms, etc.).
[0051] In one embodiment, each trigger indicator is calculated based on the change category of the change of the changed object and the correlation weight of the change category. However, the trigger indicator can be calculated in any way according to the change category and the correlation weight (e.g., based on any linear / nonlinear function, such as a weighted sum, a logarithmic / exponential combination, based only on the change category or further based on partial, different, or additional other trigger parameters regarding the above trigger parameters, etc.).
[0052] In one embodiment, the relevance weight indicates the relevance of the corresponding change category. However, the relevance weight may indicate the relevance of the change category for any purpose (e.g., for an information technology system, its users, its organization, etc.) and in any manner (e.g., according to any linear / nonlinear law, based on switches, increasing / decreasing as its value increases / decreases, etc.).
[0053] In one embodiment, the method includes (by a computing system) selecting one or more critical objects in the change object according to its trigger indicator. However, the critical objects can be selected in any way (e.g., by comparing the trigger indicator with any static / dynamic threshold, directly by the logical value of the trigger indicator, according to the trend of the value of each trigger indicator over time, etc.).
[0054] In one embodiment, the method includes (by a computing system) triggering the corresponding backup of the critical objects. However, the backup of the critical objects can be triggered in any way (e.g., by causing them to run on other computing machines storing the critical objects or by causing them to run directly on the same computing machine, for each critical object or a group thereof stored on the same computing machine, etc.).
[0055] Additional embodiments provide additional advantageous features, which, however, can be completely omitted in the basic implementation.
[0056] Specifically, in one embodiment, the method includes (by a computing system) receiving the corresponding feedback indicator of the backup, which indicates its utility. However, the feedback indicators can be of any type (e.g., the same or different relative to the trigger indicator) and they can be received in any way (e.g., local input, transmitted from the computing machine on which the backup has been run, etc.).
[0057] In one embodiment, the method includes (by a computing system) updating the relevance weight according to the feedback indicator. However, the relevance weight can be updated in any way (e.g., using machine learning techniques, using adaptive algorithms, etc.) and at any time (e.g., after each backup, after a predefined number of backups, periodically, etc.).
[0058] In one embodiment, the method includes (by a computing system) using machine learning techniques to update the relevance weight according to the feedback indicator. However, the machine learning techniques can be of any type (e.g., based on any neural network, such as linear regression, convolution, and similar types, trained in any way, such as based on stochastic gradient descent, real-time regression learning, and similar algorithms, etc.).
[0059] In one embodiment, the method includes (by a computing system) determining one or more possible affected objects among the objects affected by a change to each of the critical objects. However, the affected objects can be any number (down to none) and any type (e.g., directly affected only by the changed object, indirectly affected at any level or up to a maximum level, etc.). The affected objects can be determined in any way; for example, the objects affected by each critical object can be determined, the objects affected by each changed object can be determined, and then their trigger indicators can be calculated for selecting the critical objects among all these objects.
[0060] In one embodiment, the affected objects are determined based on the object correlation between the objects. However, the object correlation can be of any type (e.g., clustering, dependency, prerequisite, interaction, etc.).
[0061] In one embodiment, the method includes (by a computing system) triggering a corresponding backup of the affected objects of each of the critical objects. However, the backup of the affected objects can be triggered in any way (the same or different with respect to the critical objects).
[0062] In one embodiment, an information technology system includes a plurality of computer machines, each computer machine storing one or more objects. However, these computer machines can be any number and any type, each storing any number of objects (e.g., physical / virtual machines communicating over any local, wide area, global, cellular or satellite network using any type of wired and / or wireless connection, etc.).
[0063] In one embodiment, the method includes (by a computing system) causing the critical objects and the corresponding backups of the affected objects to be on the corresponding computer machines. However, this result can be achieved in any way (e.g., via remote commands, messages, etc.).
[0064] In one embodiment, the method includes (by a computing system) detecting one or more anomalies based on a comparison of the change and the object correlation. However, the anomalies can be detected in any way (e.g., using compliance rules, artificial intelligence techniques, etc.).
[0065] In one embodiment, the method includes (by a computing system) outputting an indication of the anomaly. However, the indication of the anomaly can be output in any way (e.g., by display, printing, remote notification, etc.).
[0066] In one embodiment, the method includes (by the computing system) further calculating a trigger indicator for each of the changed objects based on the time elapsed since a previous backup in the backup of the changed objects. However, the trigger indicator can be calculated in any manner based on the elapsed time (e.g., according to any corresponding relevance weights (the same or different from those above), according to its value, or according to its predefined range, etc.) until there is none.
[0067] In one embodiment, the method includes (by the computing system) further calculating a trigger indicator for each of the changed objects based on the corresponding agent that has performed the change of the changed object. However, the trigger indicator can be calculated in any manner based on the agent (e.g., according to any corresponding relevance weights (the same or different from those above), for individual agents or for their roles, etc.) until there is none.
[0068] In one embodiment, the method includes (by the computing system) further calculating a trigger indicator for each of the changed objects based on the object type of the changed objects. However, the trigger indicator can be calculated in any manner based on the object type (e.g., according to any corresponding relevance weights (the same or different from those above), for object types defined in any way, such as by file extension, logical function, etc.) until there is none.
[0069] In one embodiment, the method includes (by the computing system) determining the corresponding change frequency of the change of the changed objects. However, the change frequency can be determined in any manner (e.g., over any number of previous backups, where the backups are triggered and / or run locally or remotely notified, etc.).
[0070] In one embodiment, the method includes (by the computing system) further calculating a trigger indicator for each of the changed objects based on the change frequency of the changed objects. However, the trigger indicator can be calculated in any manner based on the change frequency (e.g., according to any corresponding relevance weights (the same or different from those above), according to its value, or according to its predefined range, etc.) until there is none.
[0071] In one embodiment, the method includes (by the computing system) triggering the backup of an object in response to a maximum latency from a previous backup in the backup. However, the maximum latency can have any value (individually or globally), and it can be used to trigger such an indiscriminate backup of the object in any manner (e.g., each object relative to its previous backup, all objects of each computing machine relative to the previous backup of any one of them, all objects of an information technology system relative to the previous backup of any one of them, etc.).
[0072] In one embodiment, the method includes (by a computing system) running each backup in response to a trigger for backup. However, the backup can be run in any way (e.g., by copying the corresponding objects to a local / remote disk, tape, etc., by a computing machine that is the same as or different from the computing machine that triggers the backup, etc.).
[0073] Generally, if the same scenario is implemented by an equivalent method (by using similar steps with the same function that have more steps or parts thereof, removing some non-essential steps or adding additional optional steps), then similar considerations apply; furthermore, these steps can be (at least partially) executed in a different order, simultaneously, or in an interleaved manner.
[0074] Embodiments provide a computer program configured to cause a computing system to perform the above method. Embodiments provide a computer program product that includes a computer-readable storage medium having program instructions implemented therewith; the program instructions can be executed by a computing system to cause the computing system to perform the same method. However, the computer program can be implemented as a stand-alone module, a plug-in to a pre-existing software program (e.g., a backup application), or directly therein. Furthermore, the computer program can be executed on any computing system (as described below). In any case, the solution according to the embodiments of the present disclosure is adapted to be implemented even by using a hardware structure (e.g., by electronic circuits integrated in one or more chips of a semiconductor material) or by using a combination of software and hardware that is appropriately programmed or otherwise configured.
[0075] Embodiments provide a system that includes means configured to perform the steps of the above method. Embodiments provide a system that includes a circuit (i.e., any hardware appropriately configured by software, for example) for performing each step of the above method. However, the system can be of any type (e.g., a control computing machine for controlling one or more data computing machines that store objects, each data computing machine, a combination of a control computing machine and data computing machines, etc.).
[0076] Generally, if the system has a different structure or includes equivalent components or it has other operating characteristics, then similar considerations apply. In any case, each of its components can be separated into more elements, or two or more components can be combined together into a single element; furthermore, each component can be replicated to support parallel execution of the corresponding operations. Additionally, unless otherwise stated, any interaction between different components generally does not need to be continuous, and it can be direct or indirect through one or more intermediate media.
[0077] The present invention can be a system, method, and / or computer program product at any possible level of integration technology detail. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to perform aspects of the present invention. The computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device.
[0078] The computer-readable storage medium can be, by way of example and not limitation, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in a groove having instructions recorded thereon), and any suitable combination of the foregoing.
[0079] As used herein, a computer-readable storage medium should not be construed to be a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire. The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.
[0080] The computer-readable program instructions for carrying out operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0081] In some embodiments, an electronic circuit, including for example a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit so as to perform aspects of the present invention. Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions. These computer-readable program instructions may be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus to produce a machine, which, when executed by the processor of the computer or other programmable data processing apparatus, creates a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0082] These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium having instructions stored therein comprises an article of manufacture that includes instructions for implementing aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram. The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices for a series of operational steps to be executed on a computer, other programmable apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0083] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function. In some alternative embodiments, the functions noted in the blocks may occur out of the order noted in the figures. For example, depending on the functionality involved, two blocks shown in succession may actually be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order. It will also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a system based on dedicated hardware that performs the specified functions or acts or executes a combination of dedicated hardware and computer instructions.
Claims
1. A method for triggering backup of objects, comprising: detecting, by a computing system, one or more changed objects in a set of objects, the one or more changed objects having been changed since a previous backup of the set of objects; determining, by the computing system, corresponding changes of the changed objects; determining, by the computing system, corresponding change categories of the changes; calculating, by the computing system, a trigger indicator for each of the one or more changed objects, the trigger indicator being calculated based on the change categories of the changes of the one or more changed objects and a correlation weight of the change categories, the correlation weight indicating the relevance of the corresponding change category to the computing system; selecting, by the computing system, one or more critical objects from the changed objects according to their trigger indicators; and triggering, by the computing system, corresponding backups of the critical objects.
2. The method according to claim 1, further comprising: receiving, by the computing system, a corresponding feedback indicator indicating the utility of the corresponding backup; and updating, by the computing system, the correlation weight according to the feedback indicator.
3. The method according to claim 2, further comprising: updating, by the computing system, the correlation weight according to the feedback indicator by using machine learning techniques.
4. The method according to claim 1, further comprising: determining, by the computing system, one or more possible affected objects in the set of objects that are affected by the changes of each of the critical objects according to the object correlation between the set of objects; and triggering, by the computing system, corresponding backups of the affected objects of each of the critical objects.
5. The method according to claim 4, further comprising: causing, by the computing system, the critical objects and the corresponding affected objects to be backed up on corresponding computer machines in computer machines included in the computing system.
6. The method according to claim 4, further comprising: detecting, by the computing system, one or more anomalies according to a comparison of the corresponding changes with the object correlation; and outputting, by the computing system, an indication of the anomalies.
7. The method according to claim 1, further comprising: further calculating, by the computing system, a trigger indicator for each of the changed objects according to the time elapsed since a previous backup among the backups of the changed objects.
8. The method according to claim 1, further comprising: further calculating, by the computing system, a trigger indicator for each of the changed objects according to the corresponding entity that has performed the changes of the changed objects.
9. The method according to claim 1, further comprising: further calculating, by the computing system, a trigger indicator for each of the changed objects according to the object type of the changed objects.
10. The method according to claim 1, further comprising: determining, by the computing system, the corresponding change frequency of the changes of the changed objects; and further calculating, by the computing system, a trigger indicator for each of the changed objects according to the change frequency of the changed objects.
11. The method according to claim 1 further comprises: triggering the corresponding backup by the computing system in response to the maximum latency from the previous backup.
12. The method according to claim 1 further comprises: running each of the backups by the computing system in response to the triggering of the backup.
13. A computer program product comprising: computer code including instructions and data for causing one or more processors to start performing operations of the method according to any one of claims 1 to 12.
14. A system for controlling the backup of objects in an information technology system, wherein the system comprises circuitry for performing the steps of the method according to any one of claims 1 to 12, respectively.
Citation Information
Patent Citations
Data backup method and device
CN106569911A
Architecture for back up and / or recovery of electronic data
US20080126442A1