Method and system for backup and recovery of credential database

By using a partitioning method that abstracts features and evaluates spatial heat of the trusted computing database, the problem of low backup and recovery efficiency of the trusted computing database in the existing technology is solved, and efficient and accurate data backup and recovery is achieved.

CN120723540AActive Publication Date: 2025-09-30JIANGSU MOBILE INFORMATION SYST INTEGRATION CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511186601.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-09-30
Estimated Expiration
2045-08-25

AI Technical Summary

Technical Problem

When faced with large-scale massive attribute data, the existing trusted database backup and recovery algorithm is unable to fully reflect the comprehensive differences between samples, resulting in insufficient partition rationality and affecting the efficiency, accuracy and reliability of backup and recovery.

Method used

By abstracting the features of each row of data in the Xinchuang database, we obtain the data vector, determine the termination evaluation sequence based on the spatial heat and neighbor relationship of the data points, perform scientific partitioning, and export the area data as a target file for backup and recovery.

Benefits of technology

It achieves rapid data location and accurate restoration, improves the efficiency, accuracy and reliability of backup and recovery, reduces the risk of data confusion, and ensures the integrity of the database structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723540A_ABST
    Figure CN120723540A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a credential database backup recovery method and system.The credential database backup recovery method comprises the steps that feature abstraction is conducted on each row of data of a credential database to be backed up, a data vector of each row of data is obtained, and each row of data comprises at least one information attribute; determining the spatial popularity of each data point according to the distance between adjacent data points of each data point and each data point and a data vector formed by the adjacent data points and each data point; based on the spatial popularity and the data vector of each data point, determining termination evaluation of each data point to obtain a termination evaluation sequence, and partitioning data in the credential database based on the termination evaluation sequence to obtain a plurality of districts; and exporting the data of each area as a target file so as to back up the data of the credential database, and recovering the data of each area in the target file to the corresponding partition of the credential database. According to the method, the efficiency, accuracy and reliability of backup and recovery of the credential database are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a method and system for backing up and restoring a trusted computing database. Background Art

[0002] With the continuous advancement of the information technology application innovation industry, the Xinchuang database, as an independent and controllable complete database ecosystem, has gradually become the core support for data storage and management in various industries. With its flexible use of network hardware at different levels, it can fully mobilize various resources from software to hardware during the database backup and recovery process, greatly improving the overall performance utilization rate, and providing a solid foundation for data security and business continuity. In the actual operation of database backup and recovery, the current mainstream solution maximizes the retention of the database's branch structure by splitting the collected data into indexes and placing the indexed data in the same partition. This design allows the recovery process to proceed step by step according to different branches. After the recovery of a single partition is completed, the database can operate on the partition, effectively reducing the downtime during backup and recovery, and achieving earlier resumption of normal use while restricting some functions, greatly reducing the losses caused by business interruptions. However, existing algorithms still have obvious limitations when faced with massive attribute data in large-scale databases. In massive attribute databases, the more descriptive attributes a single object has, the more comprehensive the record of that object is. However, existing backup algorithms can only sort by a single attribute and achieve object splitting by counting the distribution of data on that attribute. This single-dimensional processing method makes it difficult to fully reflect the comprehensive differences between samples, resulting in insufficient partitioning rationality, which in turn affects the efficiency, accuracy, and reliability of the backup and recovery of the trusted database. Summary of the Invention

[0003] In order to solve the technical problems of low efficiency, accuracy and reliability of backup and recovery of trusted databases, the purpose of the present invention is to provide a method and system for backup and recovery of trusted databases.

[0004] In order to solve the above technical problems, the technical solutions adopted are as follows: In a first aspect, an embodiment of the present invention provides a method for backing up and restoring a trusted database, comprising: performing feature abstraction on each row of data in the trusted database to be backed up to obtain a data vector for each row of data, wherein each row of data includes at least one information attribute; determining the spatial heat of each data point based on the distance between the neighboring data points of each data point and each data point and the data vector formed by the neighboring data points and each data point; determining the termination evaluation of each data point based on the spatial heat and data vector of each data point to obtain a termination evaluation sequence, and partitioning the data in the trusted database based on the termination evaluation sequence to obtain multiple partitions; exporting the data of each partition as a target file to back up the data of the trusted database, and restoring the data of each partition in the target file to the corresponding partition of the trusted database.

[0005] Optionally, determining the spatial heat of each data point based on the distance between each data point's neighboring data points and each data point and the data vector formed by the neighboring data points and each data point includes: taking the current data point as the starting point and any neighboring data point of the current data point as the ending point to obtain the data vector of the current data point and any neighboring data point; taking the current data point as the starting point and the sample center origin as the ending point to obtain the test vector of the current data point; calculating the angle between the data vectors of each neighboring data point of the current data point and the test vector; determining the spatial heat of the current data point based on the distance between each neighboring data point of the current data point and the current data point and the angle between the data vectors of each neighboring data point of the current data point and the test vector.

[0006] Optionally, determining the spatial heat of the current data point based on the distances between each neighboring data point of the current data point and the current data point, and the angle between the data vectors of each neighboring data point of the current data point and the test vector includes: selecting the minimum distance and the maximum distance from each distance, and determining the representative angle between the data vector of the neighboring data point corresponding to the minimum distance and the test vector; calculating a first ratio between the minimum distance and the maximum distance, and a first difference between the angle between the data vector of each neighboring data point and the test vector and the representative angle, and a second difference between the angles between the data vectors of each neighboring data point and the test vector; selecting the maximum difference from the second differences, and calculating the second ratio between the first difference and the maximum difference; and determining the spatial heat of the current data point based on the first ratio and the second ratio.

[0007] Optionally, the window size of the neighboring data points is calculated based on the average and standard deviation of the distance between each data point and its nearest data point.

[0008] Optionally, based on the spatial heat and data vector of each data point, the termination evaluation of each data point is determined, and the termination evaluation sequence obtained includes: sorting the spatial heat of each data point in descending order, and selecting the data point corresponding to the highest spatial heat as the starting point for area division; taking all the neighboring data points of the starting point as the initial area, and accumulating the data vectors of all data points in the initial area to obtain the first sum data vector of the initial area, the first sum data vector indicating the convergence characteristics between the neighboring data points of the starting point; selecting the data point to be divided that is closest to the end position of the area data vector, and determining the data vector between the data point to be divided and its neighboring data points; superimposing the data vectors between the data point to be divided and its neighboring data points to obtain an updated second sum data vector; determining the termination evaluation of the data point based on the first sum data vector, the second sum data vector, the spatial heat of the data point to be divided and the spatial heat of all data points in the initial area, and sorting the termination evaluation of each of the data points according to the acquisition order of the data points to obtain the termination evaluation sequence.

[0009] Optionally, determining the termination evaluation of the data point based on the first sum data vector, the second sum data vector, the spatial heat of the data point to be divided, and the spatial heat of all data points in the initial area includes: calculating the angle between the first sum data vector and the second sum data vector, the absolute value of the third difference between the first sum data vector and the second sum data vector, and the fourth difference between the absolute value of the first sum data vector and the absolute value of the third difference; calculating the heat mean and standard deviation of the spatial heat of all data points in the initial area; determining the termination evaluation based on the angle between the first sum data vector and the second sum data vector, the fourth difference, the heat mean and the standard deviation.

[0010] Optionally, partitioning the data in the ICT database based on the termination evaluation sequence to obtain multiple partitions includes: taking the first-order difference of the termination evaluation of each data point in the termination evaluation sequence to obtain a differential numerical sequence; partitioning the data in the ICT database based on the signs of the differential values ​​in the differential numerical sequence to obtain multiple partitions.

[0011] Optionally, the data in the ICT database is partitioned based on the signs of the differential values ​​in the differential numerical sequence to obtain multiple partitions, including: if the signs of the previous and next differential values ​​of the current differential value in the differential numerical sequence are different from the sign of the current differential value, then the sign of the current differential value is updated to the signs of the previous and next differential values ​​to obtain an updated differential numerical sequence; starting from the second differential value of the updated differential numerical sequence, determine whether the signs of the subsequent differential values ​​of the second differential value in the updated differential numerical sequence are the same as the sign of the second differential value; when the sign of the target differential value appears, which is different from the sign of the first and second differential values, the data point of the termination evaluation corresponding to the target differential value at the subtrahend position is not acquired; the remaining data points except the data point of the termination evaluation corresponding to the target differential value at the subtrahend position are acquired to obtain the partitioned partitions.

[0012] Optionally, after obtaining the remaining data points in the initial area except for the data point at the terminating evaluation position corresponding to the target differential value to obtain the partitioned area, the method also includes: sorting all the split areas in descending order according to the number of data points in the area to obtain an area sequence, and obtaining two target areas with the largest difference in the number of data points between adjacent sorted areas in the area sequence; taking the one with fewer data points in the two target areas as the starting area, and taking the area with the least number of data points in the area sequence as the ending area; taking all the areas between the starting area and the ending area in the area sequence as fragment areas, and merging the fragment areas, the starting area and the ending area to obtain a synthetic area.

[0013] In the second aspect, an embodiment of the present invention provides a system for backing up and recovering a trusted database, comprising: a processor and a memory; wherein the memory is used to store computer programs that can run on the processor; and the processor is used to execute the programs stored in the memory to implement the steps of the trusted database backup and recovery method mentioned in the first aspect.

[0014] The present invention has the following beneficial effects: By abstracting the features of each row of data to obtain a data vector, the multi-dimensional information attributes of the data can be converted into a quantifiable vector form, breaking through the limitations of traditional single-attribute sorting and comprehensively capturing the inherent connections between data. Spatial heat is determined based on the distance between a data point's neighboring data points and the data vector, and a termination evaluation is derived by combining the spatial heat and data vector, making the evaluation of the data point more consistent with its actual distribution characteristics in the overall data space. Data is partitioned based on the termination evaluation sequence, taking into account the spatial heat and multi-dimensional feature differences of the data points, allowing data with similar characteristics and high correlation to be clustered in the same area. The correlation and consistency of data within each area are ensured, making the partitioning more rational and paving the way for subsequent file backup and recovery. Furthermore, by exporting the data of each area as a target file for backup, targeted backup can be achieved based on the characteristics of the area, improving the utilization and efficiency of backup resources. During the recovery process, the data of each area in the target file is restored to the corresponding partition. Relying on the rationality of the partitioning, the data can be quickly located and accurately restored, reducing data matching time during the recovery process and improving the efficiency and accuracy of data recovery. Therefore, the ICT database is scientifically partitioned through the above-mentioned method of the embodiment of the present invention. The data in each partition is relatively independent and closely related internally. The risk of backup file damage caused by data mixing can be reduced during backup, and the data security is high. During recovery, the data in each partition is restored to the designated partition, which can effectively avoid data confusion, ensure the integrity of the database structure and the accuracy of the data after recovery, and improve the overall efficiency, accuracy and reliability of ICT database backup and recovery. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0016] Figure 1 A flowchart of a method for backing up and restoring a trusted database disclosed in one embodiment of the present invention; Figure 2 A structural diagram of a trusted database backup and recovery system disclosed in an embodiment of the present invention. DETAILED DESCRIPTION

[0017] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the following is a detailed description of a method and system for backing up and restoring a trusted database proposed by the present invention, in combination with the accompanying drawings and preferred embodiments, including its specific implementation, structure, features and effects. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments may be combined in any suitable form.

[0018] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0019] The specific scheme of the trusted database backup and recovery method disclosed in the present invention is described in detail below with reference to the accompanying drawings. Example

[0020] See also Figure 1 , which shows a flowchart of a method for backing up and restoring a trusted database provided by an embodiment of the present invention, including: Step S101, perform feature abstraction on each row of data in the credible innovation database to be backed up to obtain a data vector for each row of data, where each row of data includes at least one information attribute.

[0021] Specifically, before performing data backup, it is necessary to extract files from the current trusted database to be backed up. In actual application, the trusted database runs in real time and there are constant changes in files. Therefore, the embodiment of the present invention takes a snapshot of the trusted database at the start of the backup, extracts and analyzes the snapshot obtained at the backup moment, and incrementally updates the data of subsequent operations to reduce the impact of backup updates on the operation of the trusted database.

[0022] Furthermore, the embodiment of the present invention extracts data values ​​from each row in the snapshot of the credible database obtained above in the order of row labels, and further abstracts it through a word vector algorithm (Word to Vector, word2vec), thereby obtaining word vectors corresponding to each information attribute in each row, which represent the characteristic values ​​of each information attribute in each row in the table, and then obtains the corresponding data vector. Furthermore, the information attributes stored in different rows on a single column in the credible database are the same, thus completing the alignment of data attributes between different rows, and further placing the data points corresponding to all rows in the current table into the sample space and performing the following steps.

[0023] Step S102 : determining the spatial heat of each data point based on the distance between each data point and its neighboring data points and the data vector formed by the neighboring data points and each data point.

[0024] Specifically, if the trend between the stored data points and the rest of the local data points in the sample space is coordinated, it means that the fluctuation direction reflected by these data is close. Therefore, the superposition of the fluctuation trends of multiple data points can be used as effective data for partitioning and segmentation, so that the segmented data points can play a close synergistic role in the trusted computing database, thereby completing the dynamic adjustment of the partitions.

[0025] Furthermore, as an optional embodiment of the present invention, determining the spatial heat of each data point based on the distance between each data point's neighboring data points and each data point and the data vector formed by the neighboring data points and each data point includes: taking the current data point as the starting point and any neighboring data point of the current data point as the ending point to obtain the data vector of the current data point and any neighboring data point; taking the current data point as the starting point and the sample center origin as the ending point to obtain the test vector of the current data point; calculating the angle between the data vector of each neighboring data point of the current data point and the test vector; determining the spatial heat of the current data point based on the distance between each neighboring data point of the current data point and the current data point and the angle between the data vector of each neighboring data point of the current data point and the test vector. The window size of the neighboring data points is calculated based on the average value and standard deviation of the distance between each data point and its nearest data point.

[0026] Specifically, the data points to be screened by the embodiments of the present invention should be less isolated from the remaining data points. Since the Xinchuang database stores only text distributed by columns, the similarity measurement of the data is unclear. However, when the data points are extracted as data vectors, the similarity between the data points can be judged from the sample space. Therefore, the embodiments of the present invention construct data vectors between each data point, and then judge the similarity between the data points in the sample space based on the data vectors.

[0027] More specifically, the embodiment of the present invention first calculates the distance between each data point in the current sample space and its nearest data point The mean , and various Standard deviation The window size (number) of neighboring data points of each data point is , for each data point, w data points within its window are extracted as the neighboring data points of each data point.

[0028] More specifically, the embodiment of the present invention uses the current data point As a starting point, the current data point Any neighboring data point As the end point, we get the data vector . Traverse the current data point All w neighboring data points of , get w data vectors. Then take the current data point As the starting point, the sample center origin is used as the end point to get the current data point The test vector . And calculate the data points separately The data vectors of all w nearest neighboring data points and the test vector The angle between At this point, the present invention has The spatial distribution position of the data is analyzed. The intervals between isolated data points in the database are different from those of the other data points, and the distance away from the clustered position leads to a concentrated distribution direction, which means that the deviation between the current data point and other data points in the current database is higher, which means that the current data point is far away from the concentrated position of the data.

[0029] Further, as an optional embodiment of the present invention, determining the spatial heat of the current data point based on the distances between each neighboring data point of the current data point and the current data point, and the angle between the data vectors of each neighboring data point of the current data point and the test vector includes: selecting the minimum distance and the maximum distance from each distance, and determining the representative angle between the data vector of the neighboring data point corresponding to the minimum distance and the test vector; calculating a first ratio between the minimum distance and the maximum distance, and a first difference between the angle between the data vector of each neighboring data point and the test vector and the representative angle, and a second difference between the angles between the data vectors of each neighboring data point and the test vector; selecting the maximum difference from the second differences, and calculating the second ratio between the first difference and the maximum difference; and determining the spatial heat of the current data point based on the first ratio and the second ratio.

[0030] Specifically, the embodiment of the present invention uses the following formula to calculate the space heat: In the above formula, Indicates the current data point space heat. Indicates the current data point The minimum distance between the points and their neighboring data points within the window. Indicates the current data point The maximum distance between the points and their neighboring data points within the window. The smaller the value, the more the current data point The farther away from the area in the dataset, the closer the current data point is to the The area is the edge of data distribution. Indicates the current data point The number of neighboring data points in the window. Indicates the current data point The jth data vector in the window and the test vector The angle between them. Indicates the current data point The data vector and test vector of the nearest data point corresponding to the minimum distance in the window The representative angle between them. The data vector and test vector representing the pairwise nearest neighbor data points The second difference in the angle between them. () represents the maximum function, which is used to obtain the maximum value from each second difference. Indicates the current data point All neighboring data points in the window are traversed. The larger the value is, the closer the current data point is. The greater the difference in the distribution direction of the neighboring data points in the window, the greater the difference in the distribution direction of the neighboring data points in the window. The angle distribution range is used to measure the deviation between the angles to determine the current data point The distribution location is far away from the concentrated area of ​​data distribution. The larger the value, the higher the current data point The farther the distribution location is from the concentrated area of ​​data distribution.

[0031] In this way, all data points in the current sample space are traversed to obtain the spatial heat corresponding to each data point.

[0032] Step S103: Determine the termination evaluation of each data point based on the spatial heat and data vector of each data point, obtain a termination evaluation sequence, and partition the data in the ICT database based on the termination evaluation sequence to obtain multiple areas.

[0033] Specifically, after obtaining the spatial heat of each data point, the embodiment of the present invention extracts the local high-heat area according to the spatial heat of each data point, thereby obtaining the effective area corresponding to the outlier data, reducing the waste of backup space caused by the separate storage of outlier data and the increase in operation time of the independent storage area during the recovery process, which affects the recovery efficiency. In addition, the outlier data has strong specificity, so the recovery process ensures that the area of ​​the database maintains a certain function, so the outlier data is classified into the same area to reduce the forced inclusion of the remaining areas, which causes the data quality of the remaining areas to be damaged. Based on this, the embodiment of the present invention divides the area by radiating outward from the spatial heat of the neighboring data points adjacent to the data point.

[0034] Furthermore, as an optional embodiment of the present invention, based on the spatial heat and data vector of each data point, the termination evaluation of each data point is determined, and the termination evaluation sequence is obtained, including: sorting the spatial heat of each data point in descending order, and selecting the data point corresponding to the highest spatial heat as the starting point for area division; taking all the neighboring data points of the starting point as the initial area, and accumulating the data vectors of all data points in the initial area to obtain the first sum data vector of the initial area, the first sum data vector indicating the convergence characteristics between the neighboring data points of the starting point; selecting the data point to be divided that is closest to the end position of the area data vector, and determining the data vector between the data point to be divided and its neighboring data points; superimposing the data vectors between the data point to be divided and its neighboring data points to obtain an updated second sum data vector; determining the termination evaluation of the data point based on the first sum data vector, the second sum data vector, the spatial heat of the data point to be divided and the spatial heat of all data points in the initial area, and sorting the termination evaluation of each of the data points according to the acquisition order of the data points to obtain the termination evaluation sequence.

[0035] Specifically, the embodiment of the present invention sorts the spatial heat of all data points in descending order, obtains the data point corresponding to the highest spatial heat as the starting point for the partition, and then takes all the neighboring data points within the window of the starting point as the initial partition p, and accumulates the data vectors of all the neighboring data points within the window of the starting point to obtain the first sum data vector of the initial partition p. , It represents the convergence characteristics of all data points within the window from the starting point, thereby completing the screening of similar data and making it easier to divide the areas with consistent data functions. The other data point closest to the endpoint , the data points The data point to be divided is used as the initial area p. The data point to be divided is used as the end point and the data vector between the data point to be divided and its neighboring data points is determined according to the above embodiment. Then the data vector between the data point to be divided and all its neighboring data points Accumulate and get the updated second sum data vector 。 Thus, the embodiment of the present invention sorts the termination evaluations according to the acquisition order of each data point to obtain a termination evaluation sequence, that is, sorts the termination evaluations according to the time sequence of calculating the termination evaluation of each data point to obtain a termination evaluation sequence.

[0036] Further, as an optional embodiment of the present invention, determining the termination evaluation of the data point based on the first and data vectors, the second and data vectors, the spatial heat of the data points to be divided, and the spatial heat of all data points in the initial area includes: calculating the angle between the first and data vectors and the second and data vectors, and the absolute value of the third difference between the first and data vectors and the second and data vectors, and the fourth difference between the absolute value of the first and data vectors and the absolute value of the third difference; calculating the heat mean and standard deviation of the spatial heat of all data points in the initial area; determining the termination evaluation based on the angle between the first and data vectors and the second and data vectors, the fourth difference, the heat mean and the standard deviation.

[0037] Specifically, the embodiment of the present invention uses the following formula to calculate the termination evaluation: In the above formula, Represents the data points in the initial patch p Termination evaluation. Represents the first sum data vector of the initial patch p. Denotes the second sum data vector. Represents the first and data vector with the second and data vector The angle between them. Represents a data point space heat. Represents the mean heat value of the spatial heat of all data points in the initial patch p. Represents the standard deviation of the spatial heat of all data points in the initial patch p. The larger the value, the more data points After adding the current initial area p, it causes too much deviation from the vector, thus obtaining the data point This will cause changes in the function of the initial area. By data point The comparison between the theoretical change represented by the modulus of and the actual change represented by the modulus of the difference between the vectors, and The closer it is to the current data point The more obvious the effect of the change on the current initial area p, the further acquisition of subsequent data points should be terminated to avoid deviation of the data characteristics of the area. The larger the value, the more data points There is a big difference between the heat distribution in the initial area p and the data point The larger the termination evaluation value, the greater the difference is. The difference is increased by squaring and the value remains non-negative. The absolute value of a vector refers to the modulus of the vector. The calculation method of the absolute value of a vector can refer to known techniques and will not be described in detail in this embodiment of the present invention.

[0038] Thus, the embodiment of the present invention has achieved calculation of termination evaluation in the process of continuously acquiring data points.

[0039] Furthermore, after obtaining new data points, it is necessary to cut off according to the changes in the termination evaluation to reduce the inaccuracy of the area division. As an optional embodiment of the present invention, partitioning the data in the credible innovation database based on the termination evaluation sequence to obtain multiple areas includes: taking the first-order difference of the termination evaluation of each data point in the termination evaluation sequence to obtain a differential numerical sequence; partitioning the data in the credible innovation database based on the signs of the differential values ​​in the differential numerical sequence to obtain multiple areas.

[0040] Specifically, this embodiment of the present invention uses the final evaluation of the initial data point region p as the starting point of the final evaluation sequence. For each new data point, its final evaluation is calculated and placed after the starting point. The final evaluation sequence is then first-order differencing (i.e., subtracting the next final evaluation from the previous final evaluation in the sequence) to obtain a final evaluation difference sequence.

[0041] Further, as an optional embodiment of the present invention, the data in the xinchuang database is partitioned based on the signs of the differential values ​​in the differential numerical sequence to obtain multiple partitions including: if the signs of the previous and next differential values ​​of the current differential value in the differential numerical sequence are different from the sign of the current differential value, then the sign of the current differential value is updated to the signs of the previous and next differential values ​​to obtain an updated differential numerical sequence; starting from the second differential value of the updated differential numerical sequence, it is determined whether the signs of the subsequent differential values ​​of the second differential value in the updated differential numerical sequence are the same as the sign of the second differential value; when the sign of the target differential value appears, which is different from the sign of the first and second differential values, the data point of the termination evaluation corresponding to the target differential value at the subtrahend position is not acquired; the remaining data points except the data point of the termination evaluation corresponding to the target differential value at the subtrahend position are acquired to obtain the partitioned partitions.

[0042] Specifically, if the signs of the difference values ​​before and after the current difference value in the difference value sequence are different from the sign of the current difference value (the case where the difference value is 0 is not determined), the sign of the current difference value is updated to the sign of the difference values ​​before and after, and an updated difference value sequence is obtained. Starting from the second difference value in the updated difference value sequence, it is determined whether the signs of the subsequent difference values ​​remain the same. When the first different sign of the difference value appears, the data point corresponding to the termination evaluation of the subtrahend position of the current difference value is not acquired, and the remaining data points after the data point corresponding to the termination evaluation of the subtrahend position in the difference value sequence continue to be acquired. The above steps are repeated to complete the segmentation. Among them, the remaining data points before the data point corresponding to the termination evaluation of the subtrahend position in the difference value sequence can be used as a segment.

[0043] Furthermore, there are different distributions among the split areas. The areas split from the areas with high concentration are of complete size and have similar data characteristics, so that storage recovery can quickly complete the corresponding tasks. However, in the remaining areas, there are fragmented areas with very little data content divided by some special data. In order to reduce the phenomenon of setting up multiple storage heads for the fragmented areas during the database backup and recovery process, the fragmented areas need to be screened and merged, thereby completing the intelligent adjustment of the configuration strategy during the database backup and recovery process. As an optional embodiment of the present invention, after obtaining the remaining data points in the initial area except for the data point at the terminating evaluation position corresponding to the target differential value to obtain the partitioned area, the method also includes: sorting all the split areas in descending order according to the number of data points in the area to obtain an area sequence, and obtaining two target areas with the largest difference in the number of data points between adjacent sorted areas in the area sequence; taking the one with fewer data points in the two target areas as the starting area, and the area with the least number of data points in the area sequence as the ending area; taking all the areas between the starting area and the ending area in the area sequence as fragment areas, and merging the fragment areas, the starting area and the ending area to obtain a synthetic area.

[0044] Specifically, the embodiment of the present invention first sorts all the split slices in descending order according to the number of data points in the slices, and obtains the two slices with the largest difference in the number of data points between the adjacent sorted slices. The slice with the smaller number of data points is used as the starting point, and the slice with the least number of data points in the sequence is used as the end point. The slice between the starting point and the end point, as well as the slices corresponding to the starting point and the end point are used as fragment areas. The data in these fragment areas are treated as the same slice and merged, completing the merging of the fragment areas to obtain a merged slice.

[0045] At this point, the merged area and other areas are regarded as the divided areas.

[0046] Step S104: export the data of each area into a target file to back up the data of the trusted database, and restore the data of each area in the target file to the corresponding partition of the trusted database.

[0047] Specifically, after the sharding is completed, the embodiment of the present invention uses a database management system (DBMS) tool to export the shard data into a Structured Query Language (SQL) file, which is also a backup file. The SQL file stores the content of each shard.

[0048] Furthermore, after obtaining the backup file, the file is restored. The embodiment of the present invention first confirms the recovery scope and specifies the partition, table or time point that needs to be restored. Then, an isolated environment is built, and the recovery process is verified in a test environment. Finally, the embodiment of the present invention imports the SQL file to restore the data to the trusted database, and then applies the transaction log to push the data in the SQL file to the earliest time point of the stored data in the current partition, and then cancels the uncommitted transaction to ensure database consistency.

[0049] For example, when importing an SQL file, you can use the following command: a)MySQL: mysql -u root -p my_database <backup.sql; b) SQL Server: RESTORE DATABASE my_database FROM DISK = 'C:\backup\my_database.ba k'; By abstracting the features of each row of data to generate a data vector, the multi-dimensional information attributes of the data are converted into a quantifiable vector form. This breaks the limitations of traditional single-attribute sorting and comprehensively captures the inherent connections between data. Spatial popularity is determined based on the distance between a data point's neighbors and the data vector. This is combined with the spatial popularity and data vector to generate a termination evaluation, making the evaluation of data points more accurate to their actual distribution characteristics within the overall data space. Data is partitioned based on the termination evaluation sequence, taking into account the spatial popularity and multi-dimensional feature differences of data points. This allows data with similar characteristics and high correlation to be clustered in the same partition. This ensures the correlation and consistency of data within each partition, making the partitioning more rational and facilitating subsequent file backup and recovery. Furthermore, exporting each partition's data as a target file for backup enables targeted backup based on the characteristics of the partition, improving backup resource utilization and efficiency. During the recovery process, data from each partition in the target file is restored to the corresponding partition. This rationalization of the partitioning allows for rapid data location and accurate restoration, reducing data matching time during the recovery process and improving data recovery efficiency and accuracy. Therefore, the ICT database is scientifically partitioned through the above-mentioned method of the embodiment of the present invention. The data in each partition is relatively independent and closely related internally. The risk of backup file damage caused by data mixing can be reduced during backup, and the data security is high. During recovery, the data in each partition is restored to the designated partition, which can effectively avoid data confusion, ensure the integrity of the database structure and the accuracy of the data after recovery, and improve the overall efficiency, accuracy and reliability of ICT database backup and recovery. Example

[0050] Corresponding to the credential creation database backup and recovery method provided in the above embodiment, based on the same technical concept, the embodiment of the present invention also provides a credential creation database backup and recovery system, which is used to execute the credential creation database backup and recovery method. Figure 2 A structural diagram of another credential creation database backup and recovery system provided by an embodiment of the present invention is as follows Figure 2 The Xinchuang database backup and recovery system may have relatively large differences due to different configurations or performances, and may include one or more processors 401 and memory 402, the memory 402 is used to store computer programs that can be run on the processor 401, and the processor 401 is used to execute the programs stored in the memory 402 to achieve the above Figure 1 The various steps in the method embodiment. Memory 402 may be a temporary storage or a persistent storage. The application stored in memory 402 may include one or more modules (not shown), each of which may include a series of computer-executable instructions for the Xinchuang database backup and recovery system.

[0051] Furthermore, the processor 401 can be configured to communicate with the memory 402 to execute a series of computer-executable instructions in the memory 402 on the credentialed database backup and recovery system. The credentialed database backup and recovery system can also include one or more power supplies 403, one or more wired or wireless network interfaces 404, one or more input and output interfaces 405, and one or more keyboards 406.

[0052] Specifically in this embodiment, the Xinchuang database backup and recovery system includes a processor, a communication interface, a memory and a communication bus; wherein the processor, the communication interface and the memory communicate with each other through the bus; the memory is used to store computer programs; the processor is used to execute the programs stored in the memory to achieve the above Figure 1 The various steps in the method embodiment have the beneficial effects of the above method embodiments. To avoid repetition, the embodiments of the present invention will not be described again here.

[0053] It should be noted that the trusted database backup and recovery system provided by the embodiment of the present invention and the trusted database backup and recovery method provided by the embodiment of the present invention are based on the same application concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the aforementioned trusted database backup and recovery method, and has the same or similar beneficial effects, and the repeated parts will not be repeated.

[0054] It should be noted that the order in which the embodiments of the present invention are described above is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0055] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

[0056] The embodiment of the present invention further provides a computer-readable storage medium, which stores one or more programs. When the one or more programs are executed by an electronic device including multiple application programs, the electronic device executes Figure 1 The methods disclosed in the illustrated embodiments implement the functions and beneficial effects of the various methods in the preceding method embodiments, which will not be described in detail here.

[0057] Among them, the computer readable storage medium includes read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.

Claims

1. A method for backing up and restoring a trusted database, characterized in that: include: Perform feature abstraction on each row of data in the backed-up credential creation database to obtain a data vector for each row of data, where each row of data includes at least one information attribute; Determining the spatial heat of each data point based on the distance between each data point and its neighboring data points and a data vector formed by the neighboring data points and each data point; Based on the spatial heat of each data point and the data vector, determining the termination evaluation of each data point to obtain a termination evaluation sequence, and partitioning the data in the information creation database based on the termination evaluation sequence to obtain multiple partitions; Export the data of each of the areas as a target file to back up the data of the trusted database, and restore the data of each area in the target file to the corresponding partition of the trusted database.

2. A method for backing up and restoring a trusted database according to claim 1, characterized in that: Determining the spatial heat of each data point based on the distance between each data point and its neighboring data points and the data vector formed by the neighboring data points and each data point includes: Taking the current data point as the starting point and any neighboring data point of the current data point as the ending point, a data vector of the current data point and any neighboring data point is obtained; Taking the current data point as the starting point and the sample center origin as the ending point, a test vector of the current data point is obtained; Calculating the angle between the data vectors of each neighboring data point of the current data point and the test vector; The spatial heat of the current data point is determined based on the distances between the current data point and each neighboring data point, and the angles between the data vectors of each neighboring data point and the test vector.

3. A method for backing up and restoring a trusted database according to claim 2, characterized in that: Determining the spatial heat of the current data point based on the distances between each neighboring data point of the current data point and the current data point, and the angles between the data vectors of each neighboring data point of the current data point and the test vector includes: Selecting a minimum distance and a maximum distance from each of the distances, and determining a representative angle between a data vector of a neighboring data point corresponding to the minimum distance and the test vector; Calculating a first ratio of the minimum distance to the maximum distance, a first difference between the angle between the data vector of each of the neighboring data points and the test vector and the representative angle, and a second difference between the angles between the data vectors of each of the neighboring data points and the test vector; selecting a maximum difference from the second differences, and calculating a second ratio between the first difference and the maximum difference; The spatial heat of the current data point is determined based on the first ratio and the second ratio.

4. A method for backing up and restoring a trusted database according to claim 2, characterized in that: The window size of the neighboring data points is calculated based on the average value and standard deviation of the distance between each data point and its nearest data point.

5. A method for backing up and restoring a trusted database according to claim 1, characterized in that: Determining the termination evaluation of each data point based on the spatial heat of each data point and the data vector to obtain a termination evaluation sequence includes: Sort the spatial heat of each data point in descending order, and select the data point corresponding to the highest spatial heat as the starting point for area division; All neighboring data points of the starting point are used as an initial patch, and data vectors of all data points in the initial patch are accumulated to obtain a first sum data vector of the initial patch, where the first sum data vector indicates convergence characteristics between neighboring data points of the starting point; Selecting a data point to be divided that is closest to the end point of the data vector of the slice, and determining a data vector between the data point to be divided and its neighboring data points; Superimposing the data vectors between the data point to be divided and its neighboring data points to obtain an updated second sum data vector; Determining a termination evaluation of the data point based on the first sum data vector, the second sum data vector, the spatial heat of the data point to be divided, and the spatial heat of all data points in the initial area; The termination evaluations of the data points are sorted according to the order in which the data points are acquired to obtain the termination evaluation sequence.

6. A method for backing up and restoring a trusted database according to claim 5, characterized in that: The determining of the termination evaluation of the data point according to the first sum data vector, the second sum data vector, the spatial heat of the data point to be divided, and the spatial heat of all data points in the initial area includes: Calculating an angle between the first sum data vector and the second sum data vector, an absolute value of a third difference between the first sum data vector and the second sum data vector, and a fourth difference between the absolute value of the first sum data vector and the absolute value of the third difference; Calculate the mean and standard deviation of the spatial heat of all data points in the initial area; The termination evaluation is determined based on the angle between the first sum data vector and the second sum data vector, the fourth difference, the heat mean and the standard deviation.

7. A method for backing up and restoring a trusted database according to claim 5, characterized in that: Partitioning the data in the information creation database based on the termination evaluation sequence to obtain multiple partitions includes: Taking first-order differences of the termination evaluations of the data points in the termination evaluation sequence to obtain a differential value sequence; The data in the ICT database is partitioned based on the signs of the differential values ​​in the differential numerical sequence to obtain multiple partitions.

8. A method for backing up and restoring a trusted database according to claim 7, characterized in that: Partitioning the data in the information creation database based on the signs of the differential values ​​in the differential numerical sequence to obtain multiple partitions includes: If the signs of the difference values ​​before and after the current difference value in the difference value sequence are different from the sign of the current difference value, then the sign of the current difference value is updated to the sign of the difference values ​​before and after, to obtain an updated difference value sequence; Starting from the second differential value of the updated differential value sequence, determining whether the signs of subsequent differential values ​​of the second differential value in the updated differential value sequence are the same as the sign of the second differential value; When a sign of a first target difference value that is different from a sign of the second difference value appears, the data point corresponding to the target difference value and at the subtrahend position for terminating evaluation is not acquired; The remaining data points except the data point corresponding to the target difference value and at the subtrahend position for terminating the evaluation are acquired to obtain the partitioned area.

9. A method for backing up and restoring a trusted database according to claim 8, characterized in that: After acquiring the remaining data points except for the data point corresponding to the target difference value and at the subtrahend position for terminating evaluation to obtain the partitioned areas, the method further includes: Sort all the slices obtained by splitting in descending order according to the number of data points in the slices to obtain a slice sequence, and obtain two target slices in the slice sequence with the largest difference in the number of data points between adjacent sorted slices; The one of the two target slices with the smaller number of data points is used as the starting slice, and the slice with the smallest number of data points in the slice sequence is used as the ending slice; All the slices between the starting slice and the ending slice in the slice sequence are taken as fragments, and the fragments, the starting slice and the ending slice are merged to obtain a composite slice.

10. A system for backing up and restoring a trusted database, characterized in that: include: A processor and a memory; wherein the memory is used to store a computer program that can be run on the processor; A processor is used to execute the program stored in the memory to implement the steps of the method for backing up and restoring the trusted database as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Database backup and recovery method and device in ERP system

    CN101714107A

  • Data backup method, device, equipment, system, medium and program product

    CN115904800A

  • Storage partition updating method and device, electronic equipment and storage medium

    CN116909477A

  • Method and system for backing up and restoring a multi-user relational database management system

    US9424265B1