A data archiving method and device

The method addresses data archiving challenges by using cloud storage to archive data based on primary keys, ensuring efficient data retrieval without disrupting the original structure and reducing costs.

CN114116675BActive Publication Date: 2025-07-15BEIJING JINGDONG ZHENSHI INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111477886.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-06
Publication Date
2025-07-15
Estimated Expiration
2041-12-06

AI Technical Summary

Technical Problem

The prior art requires a large amount of disk space to be used during data archiving, and the data structure needs to be processed when data is pulled back, resulting in high cost and difficulty.

Method used

By uploading data from the first database to the cloud storage database based on primary key information, archived data is generated, and archived records are saved in the archive library, data in the original database is deleted, and data compression and deletion of redundant fields are avoided.

Benefits of technology

Reduces the disk space required for data archiving, avoids the damage to data structures, and reduces the cost and difficulty of data pullback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114116675B_ABST
    Figure CN114116675B_ABST
Patent Text Reader

Abstract

The present invention discloses a data archiving method and apparatus, relating to the field of computer technologies. A specific embodiment of the method includes: determining first data in a first database according to primary key information, where the first data is data to be archived; uploading the first data to a second database to generate second data, where the second data is the archived data of the first data in the second database; saving an archiving record of the first data in an archive library, and deleting the first data in the first database. This embodiment does not require a large amount of disk space, avoids data compression and deletion of redundant fields, thus does not damage the original data structure, and does not require processing of the data structure when pulling back data, reducing the cost and difficulty of pulling back data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a data archiving method and apparatus. Background Art

[0002] Currently, the data archiving solution is to archive through disk storage with low configuration and large space. At the same time, the archived data is compressed, and unnecessary indexes and redundant fields in the data are deleted to save the usage space of the database.

[0003] In the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art:

[0004] A large amount of disk space is required. Compressing the table structure and cleaning some unnecessary indexes make it difficult to pull back the data. Cleaning the redundant fields destroys the original data structure. When the data needs to be used again, the data structure needs to be processed, and the cost and difficulty of pulling back the data are relatively high. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a data archiving method and apparatus, which do not require a large amount of disk space, avoid data compression and deletion of redundant fields, thus not destroying the original data structure, and do not need to process the data structure when pulling back the data, reducing the cost and difficulty of pulling back the data.

[0006] To achieve the above object, according to one aspect of the embodiments of the present invention, a data archiving method is provided.

[0007] A data archiving method includes: determining first data in a first database according to primary key information, where the first data is data to be archived; uploading the first data to a second database to generate second data, where the second data is the archived data of the first data in the second database; saving an archiving record of the first data in an archiving library, and deleting the first data in the first database.

[0008] Optionally, the determining first data in the first database according to primary key information includes: determining target primary key information according to a time limit and an archiving type configured in an archiving task, where the target primary key information includes the creation time and business code of business data, and the creation time is within the range of the time limit, and the business code is determined according to the archiving type; determining the business data corresponding to the target primary key information in the first database as the first data.

[0009] Optionally, uploading the first data to a second database to generate second data includes: judging whether there is an available archiving address corresponding to the target primary key information currently according to each archiving record in the archiving library; if there is, uploading the first data to the second database according to the available archiving address corresponding to the target primary key information to generate the second data; if not, uploading the first data to the second database to generate the second data, and obtaining the archiving address of the first data according to the storage address of the second data.

[0010] Optionally, the archiving record includes primary key information, a hash code of the primary key information, an archiving address, and a status of the archiving address; judging whether there is an available archiving address corresponding to the target primary key information currently according to each archiving record in the archiving library includes: querying the hash code of the target primary key information in each archiving record of the archiving library, if the hash code of the target primary key information is queried, judging whether the status of the archiving address corresponding to the hash code of the target primary key information indicates that the archiving address is available, if so, there is an available archiving address corresponding to the target primary key information currently; if the hash code of the target primary key information is not queried, or the status of the archiving address corresponding to the queried hash code of the target primary key information indicates that the archiving address is not available, there is no available archiving address corresponding to the target primary key information currently.

[0011] Optionally, uploading the first data to the second database according to the available archiving address corresponding to the target primary key information includes: obtaining historically archived data from the available archiving address corresponding to the target primary key information, comparing the last business data of the historically archived data with the first data, and incrementally archiving the first data to the second database when the comparison result is different.

[0012] Optionally, when the size of the file stored at the archiving address exceeds a preset threshold, the status of the archiving address indicates that the archiving address is not available.

[0013] Optionally, the archiving record of the first data includes a unique identifier of the first data; before deleting the first data in the first database, it includes: obtaining the unique identifier of the first data from the archiving library; obtaining the corresponding second data from the second database according to the unique identifier of the first data; comparing the first data and the second data, and determining that the comparison results are the same.

[0014] Optionally, the second database is a cloud storage database.

[0015] Optionally, before uploading the first data to the second database to generate second data, it includes: converting the unique identifier of the first data and the first data into a JSON-formatted string through serialization.

[0016] According to another aspect of the embodiments of the present invention, a data archiving device is provided.

[0017] A data archiving device includes: a first data determination module, configured to determine first data in a first database according to primary key information, where the first data is data to be archived; a second data generation module, configured to upload the first data to a second database to generate second data, where the second data is the archived data of the first data in the second database; and an archiving record storage module, configured to store the archiving record of the first data in an archiving library and delete the first data in the first database.

[0018] Optionally, the first data determination module is further configured to: determine target primary key information according to the time limit and archiving type configured in the archiving task, where the target primary key information includes the creation time and business code of the business data, and the creation time is within the range of the time limit, and the business code is determined according to the archiving type; and determine the business data corresponding to the target primary key information in the first database as the first data.

[0019] Optionally, the second data generation module is further configured to: judge whether there is an available archiving address corresponding to the target primary key information currently according to each archiving record in the archiving library; if so, upload the first data to the second database according to the available archiving address corresponding to the target primary key information to generate the second data; if not, upload the first data to the second database to generate the second data, and obtain the archiving address of the first data according to the storage address of the second data.

[0020] Optionally, the archiving record includes primary key information, a hash code of the primary key information, an archiving address, and a status of the archiving address; the second data generation module is further configured to: query the hash code of the target primary key information in each archiving record of the archiving library, if the hash code of the target primary key information is queried, judge whether the status of the archiving address corresponding to the hash code of the target primary key information indicates that the archiving address is available, if so, there is an available archiving address corresponding to the target primary key information currently; if the hash code of the target primary key information is not queried, or the status of the archiving address corresponding to the queried hash code of the target primary key information indicates that the archiving address is unavailable, there is no available archiving address corresponding to the target primary key information currently.

[0021] Optionally, the second data generation module is further configured to: obtain historical archived data from an available archived address corresponding to the target primary key information, compare the last business data of the historical archived data with the first data, and incrementally archive the first data to the second database when the comparison result is different.

[0022] Optionally, when the size of the file stored at the archived address exceeds a preset threshold, the status of the archived address indicates that the archived address is unavailable.

[0023] Optionally, the archival record of the first data includes the unique identifier of the first data; it further includes a comparison module configured to: obtain the unique identifier of the first data from the archive library; obtain the corresponding second data from the second database according to the unique identifier of the first data; compare the first data and the second data, and determine that the comparison results are the same.

[0024] Optionally, the second database is a cloud storage database.

[0025] Optionally, it further includes a first data conversion module configured to: serialize and convert the unique identifier of the first data and the first data into a json-formatted string.

[0026] According to another aspect of the embodiments of the present invention, an electronic device is provided.

[0027] An electronic device includes: one or more processors; a memory for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the data archiving method provided by the embodiments of the present invention.

[0028] According to another aspect of the embodiments of the present invention, a computer-readable medium is provided.

[0029] A computer-readable medium has a computer program stored thereon, and when the computer program is executed by a processor, the data archiving method provided by the embodiments of the present invention is implemented.

[0030] One embodiment of the above invention has the following advantages or beneficial effects: Determine the first data in the first database according to the primary key information, where the first data is the data to be archived; Upload the first data to the second database to generate the second data, where the second data is the archived data of the first data in the second database; Save the archiving record of the first data in the archive library and delete the first data in the first database. It is possible to perform data archiving based on the cloud storage database, without the need for a large amount of disk space, avoiding data compression and deletion of redundant fields, thus not destroying the original data structure, and without the need to process the data structure when pulling back the data, reducing the cost and difficulty of data pulling back.

[0031] The further effects of the above non-conventional optional methods will be described below in conjunction with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The drawings are used to better understand the present invention and do not constitute an improper limitation of the present invention. Among them:

[0033] Figure 1 is a schematic diagram of the main steps of a data archiving method according to an embodiment of the present invention;

[0034] Figure 2 is one of the schematic flowcharts of data archiving according to an embodiment of the present invention;

[0035] Figure 3 is another schematic flowchart of data archiving according to an embodiment of the present invention;

[0036] Figure 4 is a schematic flowchart of the processing of archived failure data according to an embodiment of the present invention;

[0037] Figure 5 is a schematic diagram of the main modules of a data archiving device according to an embodiment of the present invention;

[0038] Figure 6 is an exemplary system architecture diagram to which the embodiments of the present invention can be applied;

[0039] Figure 7 is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The exemplary embodiments of the present invention will be described below with reference to the accompanying drawings. Various details of the embodiments of the present invention are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.

[0041] Figure 1 It is a schematic diagram of the main steps of a data archiving method according to an embodiment of the present invention.

[0042] As Figure 1 shown, the data archiving method according to an embodiment of the present invention mainly includes the following steps S101 to S103.

[0043] Step S101: Determine the first data in the first database according to the primary key information, and the first data is the data to be archived.

[0044] Determining the first data in the first database according to the primary key information may include: determining the target primary key information according to the time limit and archiving type configured in the archiving task, where the target primary key information includes the creation time and business code of the business data, and the creation time is within the time limit, and the business code is determined according to the archiving type; determining the business data corresponding to the target primary key information in the first database as the first data.

[0045] Step S102: Upload the first data to the second database to generate the second data, and the second data is the archived data of the first data in the second database.

[0046] Uploading the first data to the second database to generate the second data may include: judging whether there is an available archiving address corresponding to the target primary key information according to each archiving record in the archiving library; if so, uploading the first data to the second database according to the available archiving address corresponding to the target primary key information to generate the second data; if not, uploading the first data to the second database to generate the second data, and obtaining the archiving address of the first data according to the storage address of the second data.

[0047] Based on each archived record in the archive library, determine whether there is currently an available archive address corresponding to the target primary key information, which may include: querying the hash code of the target primary key information in each archived record of the archive library. If the hash code of the target primary key information is queried, then determine whether the status of the archive address corresponding to the hash code of the target primary key information indicates that the archive address is available. If so, there is currently an available archive address corresponding to the target primary key information; if the hash code of the primary key information is not queried, or the status of the archive address corresponding to the queried hash code of the primary key information indicates that the archive address is unavailable, then there is currently no available archive address corresponding to the target primary key information.

[0048] Based on the available archive address corresponding to the target primary key information, upload the first data to the second database, which may include: obtaining the historically archived data from the available archive address corresponding to the target primary key information, and comparing the last business data of the historically archived data with the first data. In the case where the comparison result is different, incrementally archive the first data to the second database. Incremental archiving means archiving in an incremental manner, with the aim of not archiving data that has already been archived and only archiving data that has not been archived.

[0049] When the size of the file stored in the archive address exceeds the preset threshold, the status of the archive address indicates that the archive address is unavailable.

[0050] Before uploading the first data to the second database to generate the second data, it may include: converting the unique identifier of the first data and the first data into a json-formatted string through serialization. The unique identifier of the data is the data ID.

[0051] Among them, the first database is the original database where the data to be archived is located, and the second database is the database used for data archiving. The second database can be a cloud storage database.

[0052] Step S103: Save the archived record of the first data in the archive library and delete the first data in the first database.

[0053] The archived record may include primary key information, the hash code of the primary key information, the archive address, the status of the archive address, and may also include the unique identifier of the corresponding data.

[0054] The archived record of the first data may include the target primary key information and the corresponding hash code, as well as the archive address and the corresponding status of the first data. The archived record of the first data may also include the unique identifier of the first data. Before deleting the first data in the first database, it may include: obtaining the unique identifier of the first data from the archive library; obtaining the corresponding second data from the second database according to the unique identifier of the first data; comparing the first data and the second data, and determining that the comparison results are the same.

[0055] Figure 2 One of the schematic flowcharts of data archiving according to an embodiment of the present invention Figure 3 Another schematic flowchart of data archiving according to an embodiment of the present invention

[0056] As Figure 2 and Figure 3 shown, the primary key information is determined according to the archiving type of the business data. The archiving types of the business data can be divided into pre - processed archiving data, standard single archiving data, and billing result archiving data. The business code corresponding to the pre - processed archiving data is the business line code, and the business codes corresponding to the standard single archiving data and the billing result archiving data are the merchant codes. Therefore, the primary key information of the pre - processed archiving data includes the business line code and the creation time of the business data (i.e., Figure 2 "business line + year - month" in Figure 2 ), and the primary key information of the standard single archiving data and the billing result archiving data includes the merchant code and the creation time of the business data (i.e.,

[0057] "merchant + year - month" in Figure 3 Figure 2 ). In the archiving task, the time limit and the archiving type are configured, and the scanning conditions are determined according to the time limit and the archiving type. Specifically, the creation time of the business data within the time limit and the business code determined according to the archiving type are used as the scanning conditions, that is, the business data in the first database is scanned according to the target primary key information. The target primary key information includes the creation time and the business code of the business data, and the creation time in the target primary key information is within the time limit, and the business code in the target primary key information is determined according to the archiving type. For example, in the above example, if the archiving type is standard single archiving data, the business code is the merchant code. Assuming that the time limit is within one year before May 1, 2020, the creation time of the business data needs to meet the date < '2020 - 05 - 01', that is, the scanning condition is merchant / business line + date < '2020 - 05 - 01'. By querying the business data corresponding to the primary key information whose primary key information is "merchant + year - month" and whose "year - month" meets the above scanning condition (i.e., the target primary key information), the first data is determined. After determining the first data, the unique identifier of the first data and the first data are serialized and converted into a json - formatted string, and the primary key information of the first data is converted into a hash code.

[0057] According to each archiving record in the archiving library, it is judged whether there is an available archiving address corresponding to the target primary key information, and the first data is uploaded to the second database (i.e., the cloud storage database, Figure 3In "JFS (a cloud storage service)", second data is generated. Query the hash code of the target primary key information in each archival record of the archival library. If the hash code of the target primary key information is queried, determine whether the status of the archival address corresponding to the hash code of the target primary key information indicates that the archival address is available. If so, there is currently an available archival address corresponding to the target primary key information. According to the available archival address corresponding to the target primary key information, upload the first data to the second database to generate the second data. If the hash code of the target primary key information is not queried, or the status of the archival address corresponding to the queried hash code of the target primary key information indicates that the archival address is unavailable, there is currently no available archival address corresponding to the target primary key information. Upload the first data to the second database to generate the second data, and obtain the archival address of the first data based on the storage address of the second data. Specifically, determine whether there is an available archival address corresponding to the target primary key information through the hash code of the target primary key information. Among them, each archival record in the archival library is shown in Table 1. Taking the archival record of the first data as an example, ID is the unique identifier of the first data, hashKey is the hash code of the primary key information of the first data, sourceKey is the primary key information of the first data, URL is the storage address (i.e., the archival address) of the second data corresponding to the first data in the second database, and status is the status of this archival address. If status is 1, the status of this archival address indicates availability. If status is 2, the status of this archival address indicates unavailability.

[0058] As Figure 3 shown, determine whether it is a new upload. If it is a new upload, it means there is no available archival address currently, and it is necessary to upload to a new storage address in the second database as the archival address. Then, add a new archival record and set the status (i.e., the status of the archival address) to 1, indicating that the archival address is available. If it is not a new upload, it means there is an available archival address currently, and the first data can be uploaded to this available archival address. Then, update the archival record, and the status is 1 (the archival address is available). Specifically, obtain a JSON-formatted string by serializing Map.put(id, data), upload the corresponding data to the database for data archiving, and add or update the archival record. The uploaded data is, for example, uploaded to JFS, and the archival address URL is obtained.

[0059] Table 1 Each archival record in the archival library

[0060] ID hashKey sourceKey URL status 1

[0061] In one embodiment, the unique identifier of business data is generated according to an incremental rule. When uploading the first data to the second database, the historical archived data is obtained from the available archive address corresponding to the target primary key information, and the last piece of business data in the historical archived data is compared with the first data. In the case where the comparison result is different, the first data is incrementally archived to the second database for incremental upload of the first data.

[0062] In one embodiment, a threshold is set for the size of the file stored at the archive address. When the size of the file stored at the archive address exceeds the preset threshold, the status of the archive address indicates that the archive address is unavailable. For example Figure 3 As shown, it is determined whether the uploaded file content (i.e., the archived data, such as the second data) is greater than 100M (the preset threshold). If so, the archive record is updated to completed and the status is 2, that is, the status indication of this URL (archive address) is unavailable; if not, it returns to the step of determining whether to add a new upload (this step has been introduced above and will not be elaborated here).

[0063] In one embodiment, the first data is uploaded to the second database through the set Map, where Map is a data storage set in a programming language.

[0064] In one embodiment, the unique identifier of the first data is obtained from the target primary key information recorded in the archive library; according to the unique identifier of the first data, the corresponding second data is obtained from the second database; the first data and the second data are compared. After determining that the comparison results are the same, the first data in the first database is deleted to release the storage space of the first database. This step corresponds to Figure 2 the post-archive pull-back data check. If the check passes ("yes"), the original data (i.e., the first data) is deleted. If the check fails (i.e., "no"), it is retried. If the retry fails, the original data is not deleted. And after deleting the original data, the disk fragments of the original database (i.e., the first database) can be cleaned regularly. Specifically, for the first data that has been completed archived in the archive library, it is pulled to obtain its unique identifier and compared with the corresponding second data in the second database. If the comparison results are the same, a deletion operation is performed on the first data in the original database (i.e., the first database). If the comparison results are different, the first data is archived again. By comparing the first data and the second data, the accuracy of data archiving can be verified, and the original data (i.e., the first data) is cleared based on the unique identifier to release space.

[0065] Figure 4 It is a schematic diagram of the processing flow of archived failed data according to an embodiment of the present invention.

[0066] As Figure 4As shown, in one embodiment, data is scanned according to the business line / merchant + year and month, that is, the data to be archived (the first data) in the first database is determined according to the primary key information, and this data is archived. When an archiving failure occurs during the data archiving process, the unique identifier of the data with archiving failure can be recorded, and the data with archiving failure is used as the new data to be archived, and the data with archiving failure during the archiving process is re-archived, and the data with successful archiving among the re-archived data is recorded in the archive library, and the data with archiving failure continues to be used as the new data to be archived for archiving retry. Archiving the new data to be archived corresponds to Figure 4 the "update / add" branch, that is, when archiving the new data to be archived, it is archived in the update or add manner. Archiving in the add manner means the case of new upload introduced above, indicating that there is no available archive address currently, and it needs to be uploaded to a new storage address in the second database as the archive address. Archiving in the update manner means the case of non-new upload introduced above, indicating that there is an available archive address currently, and the first data can be uploaded to this available archive address. In addition, the ID of the data with archiving failure can also be uploaded to JFS. Figure 4 the "success" branch corresponding to judging whether there is data with archiving failure in

[0067] Figure 5 is a schematic diagram of the main modules of a data archiving device according to an embodiment of the present invention.

[0068] As Figure 5 shown, a data archiving device 500 according to an embodiment of the present invention mainly includes: a first data determination module 501, a second data generation module 502, and an archiving record storage module 503.

[0069] The first data determination module 501 is used to determine the first data in the first database according to the primary key information, and the first data is the data to be archived.

[0070] The second data generation module 502 is used to upload the first data to the second database to generate the second data, and the second data is the archived data of the first data in the second database.

[0071] The archiving record storage module 503 is used to save the archiving record of the first data in the archive library and delete the first data in the first database.

[0072] In one embodiment, the first data determination module is specifically configured to: determine target primary key information according to the time limit and archiving type configured in the archiving task, where the target primary key information includes the creation time and business code of the business data, the creation time is within the time limit, and the business code corresponds to the archiving type; and determine the business data corresponding to the target primary key information in the first database as the first data.

[0073] In one embodiment, the second data generation module is specifically configured to: determine whether there is an available archiving address corresponding to the target primary key information according to each archiving record in the archiving library; if so, upload the first data to the second database according to the available archiving address corresponding to the target primary key information to generate the second data; if not, upload the first data to the second database to generate the second data, and obtain the archiving address of the first data according to the storage address of the second data.

[0074] In one embodiment, the archiving record may include primary key information, the hash code of the primary key information, the archiving address, and the status of the archiving address; the second data generation module is specifically configured to: query the hash code of the target primary key information in each archiving record of the archiving library, if the hash code of the target primary key information is queried, determine whether the status of the archiving address corresponding to the hash code of the target primary key information indicates that the archiving address is available, if so, there is an available archiving address corresponding to the target primary key information currently; if the hash code of the primary key information is not queried, or the status of the archiving address corresponding to the queried hash code of the primary key information indicates that the archiving address is not available, then there is no available archiving address corresponding to the target primary key information currently.

[0075] In one embodiment, the second data generation module is specifically configured to: obtain the historically archived data from the available archiving address corresponding to the target primary key information, compare the last business data of the historically archived data with the first data, and in the case where the comparison result is different, incrementally archive the first data to the second database.

[0076] In one embodiment, when the size of the file stored at the archiving address exceeds the preset threshold, the status of the archiving address indicates that the archiving address is not available.

[0077] In one embodiment, the archiving record of the first data may include the target primary key information and the corresponding hash code, as well as the archiving address and the corresponding status of the first data; it may also include a comparison module, which is configured to: obtain the unique identifier of the first data from the target primary key information recorded in the archiving library; obtain the corresponding second data from the second database according to the unique identifier of the first data; compare the first data and the second data, and determine that the comparison results are the same.

[0078] In one embodiment, the second database may be a cloud storage database.

[0079] In one embodiment, a first data conversion module may also be included, which is configured to: serialize and convert the unique identifier of the first data and the first data into a JSON-formatted string.

[0080] In addition, the specific implementation details of the data archiving device in the embodiments of the present invention have been described in detail in the above data archiving method, so the repeated content will not be described here.

[0081] Figure 6 An exemplary system architecture 600 to which the data archiving method or data archiving device of the embodiments of the present invention can be applied is shown.

[0082] As Figure 6 shown, the system architecture 600 may include terminal devices 601, 602, 603, a network 604, and a server 605. The network 604 is used as a medium to provide a communication link between the terminal devices 601, 602, 603 and the server 605. The network 604 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0083] Users may use the terminal devices 601, 602, 603 to interact with the server 605 through the network 604 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 601, 602, 603, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0084] The terminal devices 601, 602, 603 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0085] The server 605 may be a server that provides various services, such as a background management server that supports the shopping websites browsed by users using the terminal devices 601, 602, 603 (only as an example). The background management server may analyze and process data such as product information query requests received, and feedback the processing results (such as target push information, product information - only as examples) to the terminal devices.

[0086] It should be noted that the data archiving method provided by the embodiments of the present invention is generally executed by the server 605. Correspondingly, the data archiving device is generally set in the server 605.

[0087] It should be understood that Figure 6 the numbers of terminal devices, networks, and servers in

[0088] Refer to the following Figure 7 , which shows a schematic structural diagram of a computer system 700 of a terminal device or a server suitable for implementing the embodiments of the present invention. Figure 7 The shown terminal device or server is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.

[0089] As Figure 7 shown, the computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage section 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the system 700 are also stored. The CPU 701, ROM 702, and RAM 703 are connected to each other via a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.

[0090] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed so that a computer program read from it can be installed into the storage section 708 as needed.

[0091] Specifically, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 709, and / or installed from the removable medium 711. When the computer program is executed by the central processing unit (CPU) 701, the above functions defined in the system of the present invention are executed.

[0092] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0093] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0094] The modules involved in the embodiments of the present invention can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, it can be described as: a processor includes a first data determination module, a second data generation module, and an archived record storage module. Among them, the names of these modules do not constitute a limitation on the modules themselves in some cases. For example, the first data determination module can also be described as "a module for determining the first data in the first database according to the primary key information".

[0095] As another aspect, the present invention also provides a computer-readable medium, which can be included in the device described in the above embodiments; or can exist separately without being assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the device, the device includes: determining the first data in the first database according to the primary key information, where the first data is the data to be archived; uploading the first data to the second database to generate the second data, where the second data is the archived data of the first data in the second database; saving the archived record of the first data in the archive library and deleting the first data in the first database.

[0096] According to the technical solution of the embodiments of the present invention, determining the first data in the first database according to the primary key information, where the first data is the data to be archived; uploading the first data to the second database to generate the second data, where the second data is the archived data of the first data in the second database; saving the archived record of the first data in the archive library and deleting the first data in the first database. It can perform data archiving based on the cloud storage database, without the need for a large amount of disk space, avoiding data compression and deletion of redundant fields, and reducing the cost and difficulty of data retrieval.

[0097] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A data archiving method, characterized in that, Including: Determine the first data in the first database according to the primary key information, where the first data is the data to be archived; Upload the first data to the second database to generate second data, where the second data is the archived data of the first data in the second database; Save the archive record of the first data in the archive library and delete the first data in the first database; The archive record includes primary key information, the hash code of the primary key information, the archive address, and the status of the archive address; The uploading the first data to the second database to generate second data includes: Query the hash code of the target primary key information in each archive record of the archive library. If the hash code of the target primary key information is queried, then determine whether the status of the archive address corresponding to the hash code of the target primary key information indicates that the archive address is available. If so, there is currently an available archive address corresponding to the target primary key information. According to the available archive address corresponding to the target primary key information, upload the first data to the second database to generate the second data; If the hash code of the target primary key information is not queried, or the status of the archive address corresponding to the queried hash code of the target primary key information indicates that the archive address is not available, then there is currently no available archive address corresponding to the target primary key information. Upload the first data to the second database to generate the second data, and obtain the archive address of the first data according to the storage address of the second data.

2. The method according to claim 1, characterized in that The determining the first data in the first database according to the primary key information includes: Determine the target primary key information according to the time limit and archive type configured in the archive task. The target primary key information includes the creation time and business code of the business data, and the creation time is within the range of the time limit, and the business code is determined according to the archive type; Determine the business data corresponding to the target primary key information in the first database as the first data.

3. The method according to claim 1, wherein The uploading the first data to the second database according to the available archive address corresponding to the target primary key information includes: Obtain the historically archived data from the available archive address corresponding to the target primary key information, compare the last business data of the historically archived data with the first data, and in the case where the comparison result is different, incrementally archive the first data to the second database.

4. The method according to claim 1, wherein When the size of the file stored at the archive address exceeds a preset threshold, the status of the archive address indicates that the archive address is not available.

5. The method according to claim 1 or 2, characterized in that, The archive record of the first data includes the unique identifier of the first data; Before deleting the first data in the first database, it includes: Obtain the unique identifier of the first data from the archive library; According to the unique identifier of the first data, obtain the corresponding second data from the second database; Compare the first data and the second data and determine that the comparison results are the same.

6. The method according to claim 1, characterized in that The second database is a cloud storage database.

7. A data archiving device, characterized in that, Including: The first data determination module is configured to determine first data in the first database according to the primary key information, where the first data is data to be archived; The second data generation module is configured to upload the first data to the second database to generate second data, where the second data is the archived data of the first data in the second database; The archive record storage module is configured to store the archive record of the first data in the archive library and delete the first data in the first database; The archive record includes primary key information, a hash code of the primary key information, an archive address, and a status of the archive address; The uploading the first data to the second database to generate second data includes: Querying for the hash code of the target primary key information in each archive record of the archive library. If the hash code of the target primary key information is queried, determining whether the status of the archive address corresponding to the hash code of the target primary key information indicates that the archive address is available. If so, there is currently an available archive address corresponding to the target primary key information, and according to the available archive address corresponding to the target primary key information, uploading the first data to the second database to generate the second data; If the hash code of the target primary key information is not queried, or the status of the archive address corresponding to the queried hash code of the target primary key information indicates that the archive address is unavailable, then there is currently no available archive address corresponding to the target primary key information. Upload the first data to the second database to generate the second data, and obtain the archive address of the first data according to the storage address of the second data.

8. An electronic device, characterized in that, Comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method according to any one of claims 1-6.

9. A computer-readable medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Database filing method, device and system, equipment and readable storage medium

    CN109684270A