Real-time data archiving method and system based on hybrid cloud
The computer device provides remote function calls and network query, compresses and archives object system data, solves the problems of slow data archiving speed and insufficient storage space in the prior art, and realizes efficient remote near-line data archiving.
Patent Information
- Application Number
- CN202080084429.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-08-27
- Filing Date
- 2020-12-22
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2040-12-22
AI Technical Summary
The existing technology is difficult to provide remote near-line data archiving functions with high data compression rates and fast querying. Especially in the field of big data, existing investments are concentrated on infrastructure and lack efficient data archiving technology.
Remote function calls are provided through computer devices, the data stored in the object system is compressed and archived into the storage system, and archived by the network query archived data to realize compressed storage and management of data, including data filtering, compression, index generation and life cycle management.
It realizes efficient data archiving and query functions, improves data archiving speed and query efficiency, reduces storage space requirements, and supports a variety of memory types and network topology.
Smart Images

Figure CN114761939B_ABST
Abstract
Description
Technical Field
[0001] The following description relates to a real-time data archiving method and system based on a hybrid cloud. Background Art
[0002] Recently, with the strengthening of data regulations, the development of the medical industry, the increasing importance of patient data storage and management, and the growing focus on data management within businesses, the demand for data archiving has been increasing. For example, to protect consumer rights, laws require that data such as financial transaction data and medical information be retained for years or even decades. Consequently, various data regulations necessitate long-term data storage. Furthermore, within the medical industry, the increasing reliance on diagnostic imaging is leading to an increase in the volume of medical imaging data. Consequently, the demand for archiving systems capable of managing large amounts of data, including storage and backup, is also increasing. Furthermore, within enterprises, data management is becoming increasingly important, not only for storing large amounts of data sent and received by the enterprise on servers and for real-time recovery and backup of stored data, but also for security reasons, the ability to preserve and manage critical data is becoming increasingly important. Furthermore, within the manufacturing industry, advancements in robotics are accelerating process automation through the development of converged robotic factories that improve production efficiency and quality.
[0003] With the advent of the Fourth Industrial Revolution, the big data sector is attracting significant attention. However, current investments in South Korea's big data sector are primarily focused on infrastructure such as servers, storage, and networks. Therefore, the development of archiving technology is needed to diversify infrastructure investments and expand opportunities in the software and service sectors. To achieve this, the development of archiving technology that offers higher data compression rates and speeds than existing technologies, as well as the ability to quickly search data, is crucial.
[0004] Prior art literature
[0005] Korean Patent Publication No. 2014-0072929 (Invention Title: Method for Automating Filing Work, Publication Date: June 16, 2014) Summary of the Invention
[0006] Technical issues
[0007] The object of the present invention is to provide the following data archiving method and device, namely, receiving a remote function call from an object system that stores data, and in response to such remote function call, providing the above-mentioned object system with a first function for archiving at least a portion of the data stored in the object system in the storage system through a network, and providing the above-mentioned object system with a second function for querying the data archived in the above-mentioned storage system through the above-mentioned network, thereby providing a remote near-line data archiving function.
[0008] Solutions to the Problem
[0009] The data archiving method provided by the present invention is executed by a computer device including at least one processor, and includes: a step of receiving a remote function call from an object system storing data through the above-mentioned at least one processor; a first function providing step of providing a first function to the above-mentioned object system through the above-mentioned at least one processor and the above-mentioned network in response to the above-mentioned remote function call, and the above-mentioned first function is used to archive at least a part of the data stored in the above-mentioned object system in the storage system; and a second function providing step of providing a second function to the above-mentioned object system through the above-mentioned at least one processor and the above-mentioned network, and the above-mentioned second function is used to query the data archived in the above-mentioned storage system.
[0010] According to one embodiment, the present invention is characterized in that the first function providing step may provide a function of compressing at least a portion of data stored in a local database of the target system and archiving the data in a list of the local database.
[0011] According to another embodiment, the present invention is characterized in that the first function providing step may provide a function of compressing at least a portion of the data stored in the local database of the target system and archiving the data in a list of an external database of the target system.
[0012] According to another embodiment, the present invention is characterized in that the above-mentioned first function providing step can provide the following function: compressing at least a part of the data stored in the local database of the above-mentioned object system into a file and archiving it in a memory included in the external system of the above-mentioned object system.
[0013] According to another embodiment, the present invention is characterized in that the above-mentioned first function providing step can provide: a 1-1 function, which determines the data record-related partitions included in the list of the local database of the above-mentioned object system based on the filtering information of the data records by controlling the above-mentioned object system; a 1-2 function, which compresses the data records according to the above-mentioned partitions by controlling the above-mentioned object system, thereby generating compressed partitions; a 1-3 function, which controls the above-mentioned object system so that the above-mentioned compressed partitions and the storage keys that uniquely identify the above-mentioned compressed partitions are connected and stored in the compressed list; and a 1-4 function, which controls the above-mentioned object system so that the above-mentioned storage keys and the above-mentioned filtering information are connected and stored in the index list of the above-mentioned local database.
[0014] According to another embodiment, the present invention is characterized in that the above-mentioned filtering information includes any field value of the corresponding data record, and the above-mentioned 1-4 functions can control the above-mentioned object system to connect the above-mentioned storage key and the above-mentioned any field value and store them in the group index list of the above-mentioned local database.
[0015] According to another embodiment, the present invention is characterized in that the above-mentioned filtering information includes time-related information of the corresponding data record, and the above-mentioned 1-4 functions can control the above-mentioned object system to connect the above-mentioned storage key and the above-mentioned time-related information and store them in the time index list.
[0016] According to another embodiment, the present invention is characterized in that the above-mentioned first function providing step can also provide functions 1-5, by controlling the above-mentioned object system to respectively connect and store the data records contained in the above-mentioned list as the primary key, the key index information of the corresponding data record position in the compressed partition compressed by the corresponding data record, and the storage key corresponding to the compressed partition containing the corresponding data record compression in the key index list.
[0017] According to another embodiment, the present invention is characterized in that the above-mentioned 1-5 functions can control the above-mentioned object system to search for data records whose primary keys are the same as the data records contained in the above-mentioned list from the data records contained in the above-mentioned second compressed partition, relative to the second compressed partition generated by compressing data records in the connection list in which the above-mentioned primary key is connected to the above-mentioned list, and relative to the searched data records, the sub-index information of the position within the above-mentioned second compressed partition is further stored in the data records with the same primary key on the above-mentioned key index list.
[0018] According to another embodiment of the present invention, the first function providing step may further provide a 1-6 function of deleting the compressed data record from the list by controlling the target system.
[0019] According to another embodiment, the present invention is characterized in that the above-mentioned first function providing step can also provide functions 1-7, by controlling the above-mentioned object system to respond to the restoration request of the deleted data record to search for the storage key connected to the identification information containing the above-mentioned restoration request from the above-mentioned index list and searching for the compressed partition connected to the searched storage key from the above-mentioned compression list, restoring the deleted data record by releasing the compression of the searched compressed partition and recording the restored data record in the above-mentioned list based on the above-mentioned identification information.
[0020] According to another embodiment, the present invention is characterized in that the function 1-2 may generate the compressed partition by controlling the target system to compress the data records included in the determined partition into a binary object.
[0021] According to another embodiment, the present invention is characterized in that the above-mentioned second function providing step can provide: a 2-1 function, which receives a search condition containing filtering information of a data record by controlling the above-mentioned object system; a 2-2 function, which searches for a storage key connected to the filtering information contained in the above-mentioned search condition from an index list of storage keys that are connected and store the filtering information of the data record and the compressed partition that uniquely identifies the compressed partition containing the corresponding data record on a local database of the above-mentioned object system by controlling the above-mentioned object system; and a 2-3 function, which searches for a compressed partition connected to the searched storage key from a compressed list that is connected and stores the storage key and the compressed partition by controlling the above-mentioned object system.
[0022] According to another embodiment, the present invention is characterized in that the above-mentioned data archiving method may also include a third function providing step, providing a third function to the above-mentioned object system through the above-mentioned network, and the above-mentioned third function is used to manage the life cycle of data archived in the above-mentioned storage system.
[0023] According to another embodiment, the present invention is characterized in that the above-mentioned third function providing step can provide: function 3-1, when the data archived in the above-mentioned storage system exceeds the retention time, the data managed in the list of the database in a compressed state is archived into a file for storage by controlling the above-mentioned object system; and function 3-2, deleting the data archived into the above-mentioned file by controlling the above-mentioned object system.
[0024] The computer program provided by the present invention is combined with a computer device and is stored in a computer-readable recording medium in order to enable the computer device to execute the above method.
[0025] The computer-readable recording medium provided by the present invention stores a program for causing a computer device to execute the above method.
[0026] The computer device provided by the present invention includes at least one processor for executing computer-readable instructions, receiving a remote function call from an object system storing data through the at least one processor, providing a first function to the object system through the network in response to the remote function call, wherein the first function is used to archive at least a portion of the data stored in the object system in a storage system, and providing a second function to the object system through the network, wherein the second function is used to query the data archived in the storage system.
[0027] Effects of the Invention
[0028] The present invention has the following effects, namely, receiving a remote function call from an object system storing data, and in response to such a remote function call, providing the object system with a first function for archiving at least a portion of the data stored in the object system in the storage system via the above-mentioned network, and providing the object system with a second function for querying the data archived in the above-mentioned storage system via the above-mentioned network, thereby providing a remote near-line data archiving function. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 A diagram showing an example of a network environment according to an embodiment of the present invention.
[0030] Figure 2 This is a block diagram showing an example of a computer device according to an embodiment of the present invention.
[0031] Figure 3 FIG2 is a diagram showing an overview of a computer system for archiving according to an embodiment of the present invention.
[0032] Figure 4 FIG. 1 is a flowchart illustrating an example of a data archiving method according to an embodiment of the present invention.
[0033] Figure 5 FIG. 1 is a flowchart illustrating an example of a process of archiving data through a first function according to an embodiment of the present invention.
[0034] Figure 6 A diagram showing a first example of the structure of a compression list according to an embodiment of the present invention.
[0035] Figure 7 This is a diagram showing a second example of the structure of a compression list according to an embodiment of the present invention.
[0036] Figure 8 A diagram showing an example of the structure of a time index list according to an embodiment of the present invention.
[0037] Figure 9 A diagram showing an example of the structure of a group index list according to an embodiment of the present invention.
[0038] Figure 10 This is a diagram showing a second example of the structure of a compression list according to an embodiment of the present invention.
[0039] Figure 11 This figure shows an example of the structure of an index list in which a time index list and a group index list are combined according to an embodiment of the present invention.
[0040] Figure 12 FIG. 1 is a flowchart illustrating another example of a process of archiving data through the first function according to an embodiment of the present invention.
[0041] Figure 13 This is a diagram showing an example of the structure of a compression list and a key index list according to an embodiment of the present invention.
[0042] Figure 14 This is a diagram showing another example of the structure of the compression list and key index list according to one embodiment of the present invention.
[0043] Figure 15 FIG. 1 is a diagram illustrating an example of a process of searching archived data by using a second function according to an embodiment of the present invention.
[0044] Figure 16 and Figure 17 A diagram showing an example of searching archived data according to an embodiment of the present invention.
[0045] Figure 18 FIG. 1 is a diagram illustrating an example of a process for efficiently storing data according to an embodiment of the present invention.
[0046] Figure 19 A diagram illustrating an example of a data de-identification method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0047] Best Mode for Carrying Out the Invention
[0048] The present invention is susceptible to numerous variations and embodiments. Therefore, specific embodiments will be described in detail with reference to the accompanying drawings. However, the present invention is not limited to specific embodiments and should be understood to encompass all variations, equivalents, and alternatives within the spirit and technical scope of the present invention. Similar structural elements will be denoted by similar reference numerals throughout the various drawings.
[0049] Terms such as "first," "second," "A," and "B" are used only to describe various structural elements, and the aforementioned structural elements are not limited to the aforementioned terms. The aforementioned terms are used only to distinguish one structural element from other structural elements. For example, without departing from the scope of protection of the present invention, the first structural element may be named the second structural element, and similarly, the second structural element may be named the first structural element. The term "and / or" includes a combination of multiple related recorded items or any one of multiple related recorded items.
[0050] When a structural element is said to be "connected" or "coupled" to another structural element, while it may be directly connected to or in contact with the other structural element, it should also be understood that other structural elements may exist in between. Conversely, when a structural element is said to be "directly connected" or "directly coupled" to another structural element, it should be understood that no other structural elements exist in between.
[0051] The terms used in this specification are only used to illustrate specific embodiments and are not intended to limit the present invention. Unless the context clearly indicates otherwise, singular expressions include plural expressions. It should be understood that the terms "including" or "having" in this specification are only used to specify the presence of features, numbers, steps, actions, structural elements, components or combinations thereof described in this specification, and do not preclude the presence or additional possibility of one or more other features, numbers, steps, actions, structural elements, components or combinations thereof.
[0052] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meanings as those commonly understood by those skilled in the art to which this invention belongs. Terms defined in commonly used dictionaries should be interpreted as having the same meanings as those in the context of the relevant art and should not be interpreted in an idealized or overly formal sense unless explicitly defined in this specification.
[0053] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings.
[0054] The data archiving system according to an embodiment of the present invention can be implemented by at least one computer device, and the data archiving method according to an embodiment of the present invention can be executed by at least one computer device included in the data archiving system. A computer program according to an embodiment of the present invention can be installed and driven by the computer device. The computer device can execute the data archiving method according to an embodiment of the present invention under the control of the driven computer program. To be integrated with the computer device and execute the data archiving method on the computer, the computer program can be stored on a computer-readable recording medium.
[0055] Figure 1 FIG is a diagram showing an example of a network environment according to an embodiment of the present invention. Figure 1In the illustrated example, the network environment includes a plurality of electronic devices 110 , 120 , 130 , 140 , a plurality of servers 150 , 160 , and a network 170 . Figure 1 This is just an example for explaining the present invention. The number of electronic devices or servers is not limited to Figure 1 The quantity shown. And, Figure 1 The network environment is only used to illustrate an example of an environment that can be applied to this embodiment, and the environment that can be applied to this embodiment is not limited to Figure 1 The network environment shown.
[0056] The plurality of electronic devices 110, 120, 130, 140 may be fixed terminals or mobile terminals implemented as computer devices. For example, the plurality of electronic devices 110, 120, 130, 140 may be smart phones, mobile phones, navigation devices, computers, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), tablet computers, etc. As an example, in Figure 1 In the figure, the shape of a smartphone is shown as an example of the electronic device 110, but in an embodiment of the present invention, the electronic device 110 can be one of a variety of physical computer devices that actually utilizes wireless communication or wired communication and communicates with other electronic devices 120, 130, 140 and / or servers 150, 160 through a network 170.
[0057] The communication method is not limited to this, and may be a communication method utilizing a communication network that can be included in the scope of network 170 (e.g., a mobile communication network, a wired network, a wireless network, a broadcast network), and may also include short-range wireless communication between multiple devices. For example, network 170 may include any one or more of a personal area network (PAN), a local area network (LAN), a campus area network (CAN), a metropolitan area network (MAN), a wide area network (WAN), a broadband network (BBN), and the like. Furthermore, network 170 may include one or more network topologies such as a bus network, a star network, a ring network, a mesh network, a star bus network, a tree network, or a hierarchical network, but is not limited thereto.
[0058] The servers 150 and 160 may be implemented by a computer device or multiple computer devices that communicate with the multiple electronic devices 110, 120, 130, and 140 via the network 170 to provide instructions, codes, files, content, services, etc. For example, the server 150 may be a system that provides services (e.g., archiving services, file publishing services, content services, group call services (or voice conferencing services), messaging services, email services, social networking services, mapping services, translation services, financial services, settlement services, and search services) to the multiple electronic devices 110, 120, 130, and 140 connected via the network 170.
[0059] Figure 2 1 is a block diagram showing an example of a computer device according to an embodiment of the present invention. The plurality of electronic devices 110, 120, 130, 140 and the plurality of servers 150, 160 described above can be respectively Figure 2 The computer device 200 shown is implemented.
[0060] like Figure 2 As shown, the computer device 200 may include a memory 210, a processor 220, a communication interface 230, and an input / output interface 240. The memory 210, as a computer-readable recording medium, may include a random access memory (RAM), a read-only memory (ROM), and a non-volatile mass storage device such as a disk drive. Non-volatile mass storage devices such as a read-only memory and a disk drive may also be included in the computer device 200 as separate permanent storage devices distinct from the memory 210. Furthermore, the memory 210 may store an operating system and at least one program code. Multiple such software components may be loaded into the memory 210 from a separate computer-readable recording medium distinct from the memory 210. Such a separate computer-readable recording medium may include a floppy disk drive, a hard disk, a magnetic tape, a DVD / CD-ROM drive, a memory card, or other computer-readable recording medium. In another embodiment, multiple software components may also be loaded into the memory 210 via the communication interface 230 rather than via a computer-readable recording medium. For example, a plurality of software structural elements may be loaded into the memory 210 of the computer device 200 based on a computer program configured from a plurality of files received through the network 170 .
[0061] The processor 220 can process the instructions of the computer program by performing basic arithmetic, logical, and input / output operations. The instructions can be provided to the processor 220 through the memory 210 or the communication interface 230. For example, the processor 220 can execute the received instructions using program code stored in a storage device such as the memory 210.
[0062] The communication interface 230 can provide a function for enabling the computer device 200 to communicate with other devices (e.g., the multiple storage devices described above) via the network 170. As an example, the processor 220 of the computer device 200 can cause requests or instructions, data, files, etc. generated by program code stored in a storage device such as the memory 210 to be transmitted to other devices via the network 170 under the control of the communication interface 230. Conversely, the computer device 200 can receive signals or instructions, data, files, etc. sent from other devices via the network 170 using the communication interface 230 of the computer device 200. Thus, the signals or instructions, data, etc. received via the communication interface 230 can be transmitted to the processor 220 or the memory 210, and the files, etc. can be stored in a storage medium (the aforementioned permanent storage device) that the computer device 200 may also include.
[0063] The input / output interface 240 may be a unit for connecting to an input / output device 250. For example, an input device may include a microphone, keyboard, or mouse, and an output device may include a display, speakers, or other devices. As another example, the input / output interface 240 may be a unit for connecting to a device that integrates input and output functions, such as a touch screen. The input / output device 250 may also be integrated with the computer device 200.
[0064] Furthermore, in another embodiment, the computer device 200 may also include Figure 2 Fewer or more structural elements may be shown. However, most structural elements in the prior art do not need to be explicitly shown. For example, the computer device 200 may include at least a portion of the input / output device 250 described above, or may further include other structural elements such as a radio transceiver and a database.
[0065] Figure 3 FIG2 is a diagram showing an overview of a computer system for archiving according to an embodiment of the present invention.
[0066] like Figure 3 As shown, the data archiving system 310 can be Figure 2The computer device 200 shown may be implemented as a physical device or a combination of multiple physical devices, and may include a data compression module 311, a query module 312, a display and control module 313, and a near-line interface module 314. The data compression module 311, the query module 312, the display and control module 313, and the near-line interface module 314 may each represent functional expressions related to tasks performed by the processor 220 of the computer device 200 based on code in the archiving solution program implemented in the data archiving system 310. For example, the archiving solution program may include code for providing data compression functionality, and the processor 220 may provide the data compression functionality using such code. In this case, the term "data compression module 310" may be used to represent the functional expressions related to the tasks performed by the processor 220 for providing the data compression functionality.
[0067] In other words, data archiving system 310 can be implemented by installing and operating an archiving solution program on computer device 200. For example, the archiving solution program can be developed as a cloud Software as a Service (SaaS) product that can be registered on the cloud systems of various cloud providers and provide archiving functions through the object system 320 described below. As another example, data archiving system 310 can also be in the form of an application server that integrates remote and near-line data archiving technology (archiving solution program) and hardware. In the case of an application server, the product format allows for rapid delivery and simplified maintenance and repair, thereby providing stable product quality and price competitiveness.
[0068] like Figure 3 As shown, the object system 320 can also be Figure 2 The computer device 200 shown may be implemented as a physical device or a combination of multiple physical devices, and may include a database 321, a control module 322, and a near-line interface module 323. In this case, the control module 322 and the near-line interface module 323 may also be functional expressions of the tasks performed by the processor 220 of the computer device 200 that implements the target system 320.
[0069] The data archiving system 310 and the object system 320 can be connected to each other via a network (eg, Figure 1 and Figure 2The target system 320 and the data archiving system 310 communicate with each other over the network 170 shown in FIG. Under the control of the control module 322, the target system 320 can call functions provided by the data archiving system 310 through the near-line interface module 323. In this case, the data archiving system 310 can provide the target system 320 with the functions called by the target system 320. For example, the target system 320 can be a comprehensive information system for enterprise resource planning (ERP). As an example, the near-line interface module 323 can be based on the remote function call (RFC) used in SAP ERP.
[0070] Figure 4 The flowchart of an example of a data archiving method according to an embodiment of the present invention is shown. The data archiving method according to this embodiment can be executed by the computer device 200 that implements the data archiving system 310 described above. In this case, the processor 220 of the computer device 200 can execute the control instructions (instructions) of the operating system code or the code based on at least one computer program included in the memory 210. The processor 220 can control the computer device 200 to execute the control instructions provided by the code stored in the computer device 200. Figure 4 The method shown includes steps 410 to 430. Furthermore, the computer program may correspond to the archiving solution program described above.
[0071] In step 410, the computer device 200 may receive a remote function call from a target system storing data. Figure 3 In the object system 320 shown, a remote function call may be generated through a near-line interface module 323 of the object system 320 .
[0072] In step 420 , the computer device 200 may respond to the remote function call and provide a first function to the target system via the network. The first function is used to archive at least a portion of the data stored in the target system to the storage system.
[0073] For example, refer again to Figure 3 Based on the call of the object system 320, the data archiving system 310 can provide the object system 320 with a first function for archiving at least a portion of the data stored in the database 321 of the object system 320 to the storage system 330 through the network.
[0074] The storage system 330 may be a local database (eg, database 321 ) included in the object system 320 of the embodiment, or an external database of the object system 320 and / or a memory of an external system (eg, a file server or a cloud server) including the object system 320 .
[0075] For example, the data archiving system 310 can provide a first function of compressing at least a portion of the data stored in the database 321 of the target system 320 and archiving the data in a list in the database 321. In this case, the compressed data is not stored as a file in the list in the database 321 of the target system 320, thereby improving not only archiving speed but also data query speed.
[0076] As another example, the data archiving system 310 may provide a first function of compressing at least a portion of data stored in the database 321 of the target system 320 and archiving the data in a list of databases external to the target system 320. For example, at the level of the data archiving system 310, assuming that the target system 320 is a client, the data archiving system 310 may store the compressed data in a list of databases included in other clients.
[0077] As another example, the data archiving system 310 may provide a first function of compressing at least a portion of the data stored in the database 321 of the target system 320 into a file and archiving the file in a storage system external to the target system 320. For example, if the data archiving system 310 exists in a cloud system, the data archiving system 310 may store the file included in the compressed data in the storage system of the cloud system.
[0078] As a more specific example, the data archiving system 310 can provide a user interface to the object system 320 through the display and control module 313 for providing archiving services such as managing retention cycles, archiving configuration, archiving execution, monitoring, data query and data management functions.
[0079] In this case, if archiving is requested via the user interface provided by the display and control module 313, the data archiving system 310 can provide the target system 320 with a first function for archiving at least a portion of the data stored in the database 312 of the target system 320 to the storage system 330 using the archiving configuration set by the data compression module 311. In other words, the target system 320 can archive at least a portion of the data stored in the local database 321 to the storage system 330 using the first function provided by the data archiving system 310.
[0080] In step 430, the computer device 200 may provide the target system with a second function for querying data archived in the storage system via the network. This second function may also be provided via a remote function call from the target system.
[0081] For example, refer again to Figure 3Based on the call of the object system 320, the data archiving system 310 can provide the object system 320 with a second function for querying the data archived in the storage system 330 through the network.
[0082] If a data query is requested through the user interface provided by the display and control module 313, the data archiving system 310 can provide the target system 320 with a second function for querying the data archived in the storage system 330 through the query module 312. In other words, the target system 320 can query the data archived in the storage system 330 through the second function provided by the data archiving system 310.
[0083] As such, the object system 320 may utilize multiple functions provided by the data archiving system 310 to archive the data stored in the database 321 .
[0084] As described above, the first function provided by the data archiving system 310 may include storing compressed data in a database (database 321 of the target system 320 or an external database) or as a file. In this case, because archived data stored in a database list as compressed data also increases the database size, the data archiving system 310 can manage the data lifecycle. For example, the data archiving system 310 can manage the data lifecycle through a process of "database → data compression and archiving → file archiving → archive extinction." "Database" may mean managing data as it is stored in the database 321 of the target system 320. Furthermore, "data compression and archiving" may mean managing data in a compressed state in a database list (database 321 of the target system 320 or an external database) due to compression. Furthermore, "file archiving" may mean archiving data stored in a compressed state in a database list as a file due to the compressed data having expired. "Archiving extinction" may mean deleting data from the archived file that is no longer needed.
[0085] "File archiving" can be performed in the memory of the target system 320, but it can also be performed in the memory of a system external to the target system 320. As a more specific example, the data archiving system 310 can also receive archived data from the target system 320 after completing compression target extraction in order to transmit it to a cloud system external to the target system 320. In this case, the data archiving system 310 can call the target system via the nearline interface module 314. This call can be based on an API call. Since compressed data can be stored in a variety of storage types, it can connect to various storage types, such as databases, disks, files, in-memory, quantum memory devices, non-relational databases (NoSQL), graph databases (graph-DB), and blockchain databases. Furthermore, the data archiving system 310 can define transmission scenarios based on business types such as finance, cost, production, sales, materials, quality, and systems. According to embodiments, the data archiving system 310 can also create groups of transmission scenarios based on network bandwidth. Furthermore, in transmission scenarios, the data archiving system 310 can be equivalent to an object. When there are subgroups for each transmission scenario, the data archiving system 310 can be considered as the extracted object within each transmission scenario group. Furthermore, the data archiving system 310 can convert the extracted objects into binary objects and construct a transmission history status table based on the object capacity and number for each transmission scenario and / or subgroup. Furthermore, the data archiving system 310 can perform transmission simulations. In this case, the data archiving system 310 can select simulation targets for each transmission scenario and / or subgroup. After running the transmission simulation to confirm the transmission time for each object, the system can predict the optimal transmission time based on the object data ratio. After the transmission simulation, the data archiving system 310 can use the scenario, subgroup, and / or object information to perform the actual data transmission. In this case, based on the transmission simulation information, the data archiving system 310 can sort the subgroups and / or objects with the longest transmission times to those with the least transmission time, thereby optimizing the overall completion time. In this case, the data archiving system 310 can store data in different storage locations based on data attributes, and the number of transmissions and execution times can be confirmed in real time using a transmission status monitoring tool. Furthermore, the data archiving system 310 can update the extraction execution status in the transmission execution graph. When an error occurs, execution can be resumed from the completed sequence, maintaining speed and integrity. Data transmission can be performed in a streaming manner or by object unit. Furthermore, after confirming whether the scene data and / or group data transmitted from the object system 320 has been transmitted to the storage system 330, the data archiving system 310 can verify the transmission process of the archived data by comparing the object capacity and quantity status table categorized by transmission scene and / or group with the transmission data. In this case, data transmission can be performed in a 1:1 relationship or a 1:N relationship through different servers simultaneously. In this case, a transmission history status table can be created for each server.
[0086] Figure 5 The flowchart of an example of a process for archiving data using a first function according to an embodiment of the present invention is shown. The process of this embodiment utilizes the first function provided by the data archiving system 310 and can be executed by the computer device 200 that implements the target system 320. In this case, the processor 220 of the computer device 200 can execute control instructions (instructions) based on the code of the operating system or at least one computer program contained in the memory 210. The processor 220 can control the computer device 200 to execute the control instructions provided by the code stored in the computer device 200. Figure 5 The method shown includes steps 510 to 550. The code may include code for the data archiving system 310 to provide the first function.
[0087] In step 510, the computer device 200 may determine the data record related partitions included in the list of the database based on the filtering information of the data record. Figure 3 The database 321 of the target system 320 is shown. The filtering information may include time-related information of the data record and / or the value of any field in the data record. The computer device 200 may determine the partitions related to the data record based on the time-related information and / or field values. A list is a basic unit of data storage in a database. The list mentioned in step 510 may be a list used to implement archiving in a capacity-saving manner among multiple lists contained in the database.
[0088] For example, the computer device 200 may filter data records whose field values fall within a specified range into a partition. In this case, the field value may be the value of the most frequently searched field in the list. This is because when future searches of archived data are performed, index information generated based on the corresponding field values can be used to maximize search efficiency. As another example, the computer device may filter data records whose time-related information falls within a specified range into a partition.
[0089] Furthermore, partitions can be formed by selecting a set of data records from the entire list. Partitions can be generated for at least one or more data records, or, as needed, for only a subset of data records, rather than the entire list. For example, in a list, in addition to data records from 2015 onwards, a partition can be generated for archiving purposes, targeting only data records before 2015.
[0090] On the other hand, the number of data records contained in a partition can be determined by comprehensively analyzing and checking the number of overall records contained in the list, the computer performance of searching the database, and the most frequently searched conditions in the database.
[0091] In yet another embodiment, if an excess partition exists in the screened partitions where the number of data records exceeds a critical value, the excess partition can be split into multiple partitions with fewer records than the critical value. For example, the number of data records a partition can contain can be set to 100,000, i.e., the critical value can be set to 100,000. However, if the screened partitions contain data records exceeding the critical value, this can lead to computer overload and reduced efficiency, thus causing a loss of capacity. Therefore, if a partition contains more than 100,000 data records, it can be split into multiple partitions of 100,000 each to generate multiple partitions. For example, if a partition contains 250,000 data records, the computer device 200 can split the excess partition into three partitions: two partitions each containing 100,000 data records and one partition containing 50,000 data records.
[0092] On the other hand, since the multiple partitions separated in the above manner are classified based on the same classification criteria using the same field value, there may be no method for distinguishing the multiple partitions. Therefore, a sequence number (e.g., 1, 2, 3, 4, etc.) can be assigned to each of the multiple separated record groups, and further stored in the sequence number field of the index list. In this case, when searching archived data, the search can be performed by distinguishing each of the multiple separated partitions. The sequence number described above may correspond to the sequence described below.
[0093] In step 520, the computer device 200 may compress the data records according to the partition to generate a compressed partition. As an example, the computer device 200 may compress the data records included in the determined partition into a binary object to generate the compressed partition.
[0094] As an example, to generate a compressed partition, computer device 200 may first store the data records to be included in the compressed partition in a buffer. The size of the buffer used to store the data records can be determined based on the structure of the list (the number, type, and size of the fields) and the critical number of data records to be included in the compressed partition. For example, a list may contain three fields: DATE (8 characters), NAME (30 characters), and AGE (a 4-byte integer). If the critical number of data records included in the compressed partition is 100,000, then, assuming that one character is equal to two bytes, the buffer size may be at least 100,000 * (8 * 2 + 30 * 2 + 4) = 8 million bytes (approximately 8 megabytes). Computer device 200 may then sequentially read all the data records and their field values included in the compressed partition and store them in the buffer.
[0095] Subsequently, the computer device 200 may generate a compressed partition by compressing the data stored in the buffer. The compressed partition may be a binary object generated by compressing the data stored in the buffer. In this case, to prevent loss due to compression, a lossless compression algorithm may be used, such as ZIP, CTW, LZ77, LZW, gzip, bzip2, DEFLATE, etc.
[0096] In this case, the computer device 200 may generate a uniquely identified storage key according to the generated compressed partition.
[0097] In step 530, the computer device 200 may link the compressed partitions and the storage keys that uniquely identify the compressed partitions and store them in a compression list. In the above description, the compressed data may be stored in a list in the database 321 of the target system 320 or in a list in an external database. The compression list may include fields generated according to partition compression for storing compressed partitions and fields for storing storage keys that uniquely identify the corresponding compressed partitions. The storage key is a key that includes a unique value assigned to each compressed partition. The value of the shared storage key may be stored in the field of the compression list corresponding to the storage key for each compressed partition. Furthermore, the corresponding fields of the storage key may be more than one. By combining the values of multiple storage keys stored in more than one field, a unique storage key may be formed for each compressed partition.
[0098] In step 540, the computer device 200 may link the storage key and the filtering information and store them in an index list of the database. As an example, when the filtering information includes any field value of the corresponding data record, the computer device 200 may link the storage key and the arbitrary field value and store them in a group index list through step 540. The storage key and field value stored in the group index list can be used as an index for searching for compressed stored data records based on a search condition including any field value. As another example, when the filtering information includes time-related information of the data record, the computer device 200 may link the storage key and the time-related information and store them in a time index list. The storage key and time-related information stored in the time index list can be used as an index for searching for compressed stored data records based on a search condition including any time-related information. In other words, the index list including the group index list and / or the time index list can be used to obtain the field value included in the search condition and / or the storage key corresponding to the time-related information, and the storage key can be used to obtain the compressed partition corresponding to the storage key from the compressed list.
[0099] In step 550, computer device 200 may delete the compressed data records from the list. The purpose of compressing an archived database is to reduce database storage space. Therefore, computer device 200 may delete the archived data records from the list to reduce database storage space. However, according to an embodiment, the compressed data records may be deleted from the list after a specified period of time, rather than being deleted directly from the list.
[0100] Alternatively, deleted data records can be later restored to the corresponding list. For example, in response to a request to restore a deleted data record, computer device 200 searches the index list for a storage key associated with the identification information included in the restoration request, and then searches the compressed list for a compressed partition associated with the searched storage key. Computer device 200 can then restore the deleted data record by decompressing the searched compressed partition and store the restored data record in the list based on the identification information. In this case, information in the key index list, described below, can also be used to identify the specific data record requested to be restored among the multiple data records that include the compressed partition.
[0101] The above steps 510 to 550 may be implemented by the first function provided by the data archiving system 310. In other words, the data archiving system 310 may provide the first function including the following functions, ie, for controlling the target system 320 to execute steps 510 to 550.
[0102] Figure 6 A diagram showing a first example of the structure of a compression list according to an embodiment of the present invention. Figure 6 The list 610 may include a Doc.No. field 611, a Date field 612 for time, and a Col1 field 613 for a specific attribute. In this case, as filtering information, the computer device 200 may classify and compress a plurality of data records in the list 610 based on the field value of the Date field 612 or the field value of the Col1 field 613 in the time-related information list 610 to generate a compressed partition. In this case, the computer device 200 may generate the compressed list 600 by linking and storing a storage key for uniquely identifying the compressed partition and the corresponding compressed partition. For example, in Figure 6 In the embodiment of FIG. 5 , the compression list 600 may include: an OBJECT ID field 621 having a storage key as a field value; and a COMPRESSED DATA field 622 having a compressed partition as a field value.
[0103] Figure 7 FIG2 is a diagram showing a second example of the structure of a compression list according to an embodiment of the present invention. Figure 8 FIG. 1 is a diagram showing an example of the structure of a time index list according to an embodiment of the present invention. Figure 9A diagram showing an example of the structure of a group index list according to an embodiment of the present invention.
[0104] Figure 7 To illustrate the Figure 6 The list 610 shown in FIG. 6 is another embodiment of generating a compressed list 700. For example, as filtering information, the computer device 200 can classify and compress multiple data records in the list 610 based on the field value of the Date field 612 in the time-related information list 610 to generate compressed partitions. Furthermore, the computer device 200 can generate the compressed list 700 by linking and storing the filtering information and the corresponding compressed partitions. For example, in Figure 7 In the embodiment of FIG. 5 , the compression list 700 may include: a PERIOD field 710 , which takes time-related information as a field value; and a COMPRESSED DATA field 720 , which takes the compressed partition as a field value.
[0105] on the other hand, Figure 8 An example of a time index list 800 is shown, which can be generated and applied when the compression list 700 includes compressed partitions generated by classifying and compressing data records based on the field value of the Date field 612 (time-related information). In this case, the time index list 800 may include: a PERIOD field 810 with time-related information as a field value; and an OBJECT ID field 820 with a storage key as a field value. For example, when the computer device 200 receives a search condition containing time-related information (e.g., "2020.01") as filtering information, the computer device 200 can search the time index list 800 for the corresponding storage key using the time-related information included in the search condition (e.g., searching the time index list 800 for the storage key "00001" corresponding to the time-related information "2020.01"). As a result, the computer device 200 can search the compression list 620 for the compressed partition corresponding to the storage key using the searched storage key (e.g., searching the compression list 620 for the compressed partition "50000 Rows" corresponding to the storage key "00001").
[0106] and, Figure 9An example of a group index list 900 is shown, which can be generated and applied when a compression list 600 includes compressed partitions generated by classifying and compressing data records based on the field value of the Col1 field 613. In this case, the group index list 900 may include a PERIOD field 910 having the field value of the Col1 field 613 as its own field value, and an OBJECT ID field 920 having the storage key as its field value. For example, when the computer device 200 receives a search condition containing a field value of the Col1 field 613 (for example, "1000") as filtering information, the field value contained in the search condition can be used to search for the corresponding storage key from the group index list 900 (for example, searching for the storage key "O0001" corresponding to the field value "1000" from the group index list 900), and thus, the searched storage key can be used to search for the compressed partition corresponding to the storage key from the compression list 600 (for example, searching for the compressed partition of "50000Rows" corresponding to the storage key "O0001" from the compression list 600).
[0107] Figure 10 FIG. 2 is a diagram showing a second example of the structure of a compression list according to an embodiment of the present invention. Figure 11 This figure shows an example of the structure of an index list in which a time index list and a group index list are combined according to an embodiment of the present invention.
[0108] Figure 10 To illustrate the Figure 6 The illustrated list 610 generates another embodiment of the compressed list 1000. For example, the computer device 200 may classify and compress the data records of the list 610 based on two field values of time-related information, namely, the field value of the Date field 612 and the field value of the Col1 field 613, to generate compressed partitions.
[0109] As a more specific example, the computer device 200 may generate a first compressed partition by compressing multiple data records whose field value of the Data field 612 is "2002.01" and whose field value of the Col1 field 613 is "1000", generate a second compressed partition by compressing multiple data records whose field value of the Data field 612 is "2002.01" and whose field value of the Col1 field 613 is "2000", and generate a third compressed partition by compressing multiple data records whose field value of the Data field 612 is "2002.02" and whose field value of the Col1 field 613 is "1000". A third compressed partition is generated, a fourth compressed partition is generated by compressing multiple data records in which the field value of the Data field 612 is "2002.02" and the field value of the Col1 field 613 is "2000", a fifth compressed partition is generated by compressing multiple data records in which the field value of the Data field 612 is "2002.03" and the field value of the Col1 field 613 is "1000", and a sixth compressed partition is generated by compressing multiple data records in which the field value of the Data field 612 is "2002.03" and the field value of the Col1 field 613 is "2000".
[0110] In this case, the computer device 200 may generate the compression list 1000 by concatenating and storing a storage key for uniquely identifying a compression partition and the corresponding compression partition. Figure 10 In the embodiment of FIG. 1 , the compression list 1000 may include: an OBJECT ID field 1010 having a storage key as a field value; and a COMPRESSED DATA field 1020 having a compressed partition as a field value.
[0111] on the other hand, Figure 11 An example of an index list 1100 in which a time index list and a group index list are combined is shown. In this case, index list 1100 may include: a PERIOD field 1110 having time-related information as a field value; a Col1 field 1120 having the field value of Col1 field 613 as its own field value; and an OBJECT ID field 1130 having a storage key as a field value. For example, when computer device 200 receives a search condition including time-related information (e.g., "2020.02") and the field value of Col1 field 613 as filtering information, it can search index list 1100 for a storage key (e.g., storage key "00003" in index list 1100) that satisfies both the time-related information and the field value included in the search condition. Thus, the searched storage key can be used to search compressed list 1000 for a compressed partition corresponding to the storage key (e.g., searching compressed list 1000 for the compressed partition "30000Rows" corresponding to the storage key "00003").
[0112] Figure 12 FIG. 1 is a flowchart illustrating another example of a process of archiving data through the first function according to an embodiment of the present invention. Figure 5 After step 540 , the process of this embodiment may further include step 1210 .
[0113] In step 1210, for each data record included in the list, the computer device 200 may concatenate key index information, including the primary key, the corresponding data record position within the compressed partition in which the corresponding data record was compressed, and the storage key corresponding to the compressed partition in which the corresponding data record was compressed, and store the information in a key index list. Step 1210 may be implemented using a first function provided by the data archiving system 310. In other words, the data archiving system 310 may provide a first function that includes controlling the target system 320 to execute step 1210.
[0114] A primary key can refer to a corresponding value in a field that has a value that uniquely identifies each record in a database. This field can also be called a relational key, primary key, or unique key. Furthermore, a list can have more than one primary key. Furthermore, key index information is information used to search for the location within a compressed partition where a data record with a specific primary key value is stored. For example, information related to the storage order of a data record stored 1000th time among related information about 100,000 data records contained in a compressed partition can be stored as key index information.
[0115] On the other hand, the reason for storing the primary key in the key index list is that, in addition to being based on other field values and time-related information, the list that is the search target can be directly searched through the above-mentioned primary key. That is, in the case where the user inputs a specific primary key, if you want to search for data records with its primary key from the list, you can use the key index list. More specifically, the computer device 200 can search the key index information and storage key of the data record with the specific primary key from the key index list. In this case, the computer device 200 can obtain the compressed partition corresponding to the storage key from the compressed list based on the obtained storage key, and can enable the user to obtain the desired specific data record using the key index information. As described above, in the process of restoring data records with specific conditions from the list, the key index information of the above-mentioned key index list can also be used to identify data records with specific conditions from multiple data records containing compressed partitions.
[0116] Figure 13 This is a diagram showing an example of the structure of a compression list and a key index list according to an embodiment of the present invention.
[0117] The compressed list 1310 may include an OBJECT ID field 1311, which has the storage key as its value; a SEQ field 1312, which has the order (sequence) of the target list being processed as its value; and a COMPRESSED DATA field 1313, which has the compressed partition as its value. When the sequence consists of a main list and sublists, the main list is first extracted, and then the extracted data of the main list can be used to define the order of processing the sublists.
[0118] As described above, the key index list 1320 may include: a Doc.No. field 1321, which uses the primary key as the field value; an OBJECT ID field 1322, which uses the storage key as the field value; and a Key Location info. field 1323, which uses the key index information as the field value. For example, in the key index information "1@1001," the "1" before the "@" is a sequence corresponding to the field value of the SEQ field 1312, and the "1001" after the "@" may indicate the 1001th data record among the multiple data records contained in the corresponding compressed partition. As a more specific example, in the first record of the key index list 1320, the storage key of the data record with the primary key "1" is "O0001," which may indicate that the data record is stored as the 1001th data record among the multiple data records contained in the compressed partition with the sequence "1." Similarly, in the second record of the key index list 1320, the storage key of the data record with the primary key "2" is "O0001", which indicates that it is stored as the 2001th data record among multiple data records of the compressed partition with the sequence "2".
[0119] As such, the key index information may include information related to the location of a specific data record within the compressed partition, and thus, the key index list (e.g., Figure 13 The key index list 1320) is used to reduce the number of data records that the user needs to query according to the search conditions.
[0120] In another embodiment, with respect to a second compressed partition generated by compressing data records in a connected list whose primary key is connected to a first list (for example, the list described in step 410), the computer device 200 may search for data records whose primary key is the same as that of the data records contained in the first list from the data records contained in the second compressed partition, so that the sub-index information of the position within the second compressed partition is further stored in the data record with the same primary key on the key index list with respect to the searched data record. The connected list is a list connected to the first list by a primary key. That is, the primary key may exist in both the first list and the connected list. When a connected list connected to the first list by a primary key exists, the second compressed partition may be data generated by compressing the data records of the corresponding connected list. In this case, the second compressed partition may be generated by compressing the data records of the corresponding connected list. Figure 4 The compressed partition shown above is generated in the same manner and can be stored in the compressed list together with the unique storage key of the compressed partition. The sub-index information is information used to search for the location of the data record with a specific primary key stored in the second compressed partition. For example, information related to the storage order of the data record stored for the 1000th time among the 100,000 data records contained in the second compressed partition can be stored as sub-index information. For example, a connection list connected to the first list by a primary key exists in the database, and the user may need field value information of a field that does not exist in the first list but exists in the connection list. In this case, the computer device 200 can also store sub-index information for data records with the same primary key on the key index list so that its connection list can be searched later.
[0121] In another embodiment, when multiple connected lists exist relative to the first list, computer device 200 may aggregate and compress the sub-index information of each connected list and then store the sub-index information as new sub-index information in the key index list. For example, computer device 200 may aggregate all sub-index information for locations within two or more second compressed partitions for data records with the same primary key in the connected lists, and then compress the aggregated values and store them as new sub-index information in data records containing the same primary key value in the key index list.
[0122] Figure 14 This is a diagram showing another example of the structure of the compression list and key index list according to one embodiment of the present invention.
[0123] Compressed list 1410 may include an OBJECT ID field 1411 with the storage key as the field value; a TABLE field 1412 with the list identifier as the field value; a SEQ field 1413 with the sequence as the field value; and a COMPRESSED DATA field 1414 with the compressed partition as the field value. TABLE field 1412 may include the list identifier as the field value, thereby identifying the list from which the corresponding compressed partition contains data records extracted.
[0124] The key index list 1420 of this embodiment may include: a Doc.No. field 1421, which uses the primary key as the field value; an OBJECT ID field 1422, which uses the storage key as the field value; a Key Location info. field 1423, which uses the key index information as the field value; and a Sub Location info. field 1424, which uses the sub-index information as the field value.
[0125] For example, the first record in the key index list 1420 indicates the 10001th data record among multiple data records that include a data record with a primary key of "1" as a compressed partition with a storage key of "00001" and a sequence of "1." In this case, the field value "TAB1@1001-2 / TAB2@2001-3" in the Sub Location info. field 1424 indicates a location within the second compressed partition generated relative to the connected list of the data record with the primary key of "1." For example, in the field value "TAB1@1001-2 / TAB2@2001-3," "TAB1" and "TAB2" preceding the "@" may indicate multiple connected lists connected by the same primary key, and "1001-2" following the "@" indicates two data records starting from the 1001th data record (the 1001th data record (the first data record) and the 1002nd data record (the second data record)) among the multiple data records included in the second compressed partition relative to the connected list "TAB1." Furthermore, "2001-3" after the "@" sign indicates that, among the multiple data records included in the second compressed partition corresponding to the linked list "TAB2," the three data records starting with the "2001"th data record (the 2001th data record (the third data record), the 2002th data record (the fourth data record), and the 2003th data record (the fifth data record)) are all identified by the same primary key. In this case, the first through fifth data records can all be identified by the same primary key.
[0126] Figure 15 This diagram illustrates an example of a process for searching archived data using a second function according to an embodiment of the present invention. The process of this embodiment utilizes the second function provided by the data archiving system 310 and can be executed by the computer device 200 implementing the target system 320.
[0127] In step 1510, the computer device 200 may receive search criteria, including filter information for data records. This filter information may include any field values of the data records to be searched and / or time-related information of the corresponding data records. The field values and / or time-related information included in the filter information may also be included in the form of ranges.
[0128] In step 1520, the computer device 200 may search for a storage key associated with the filtering information included in the search criteria from an index list that links and stores filtering information for data records and storage keys used to uniquely identify compressed partitions containing corresponding data records in the database. As described above, the index list may include a group index list and / or a time index list. The group index list may be used to link and store specific field values and storage keys, while the time index list may be used to link and store time-related information and storage keys. Therefore, the computer device 200 may search for storage keys corresponding to field values and / or time-related information included in the filtering information from the group index list and / or the time index list. For example, when the filtering information includes any field value of a data record, the computer device 200 may search for storage keys associated with any field value included in the filtering information used as the search criteria from the group index list that links and stores storage keys and any field values. As another example, when the filtering information includes time-related information for a data record, the computer device 200 may search for storage keys associated with the time-related information included in the filtering information used as the search criteria from the time index list that links and stores storage keys and time-related information.
[0129] In step 1530, the computer device 200 may search for a compressed partition linked to the searched storage key from a compression list that links and stores storage keys and compressed partitions. As described above, the compression list links and stores compressed partitions and storage keys that uniquely identify them. Therefore, the computer device 200 can search for a corresponding compressed partition from the compression list using the storage key.
[0130] As described above, by further utilizing a key index list, the user can use the primary key for searching. As described above, for each data record included in any table on the database, the key index list can link and store key index information containing the primary key, the location of the corresponding data record within the compressed partition in which the corresponding data record was compressed, and the storage key corresponding to the compressed partition containing the corresponding data record. In this case, when the search criteria also include the primary key of the data record, the computer device 200 can search the key index list for key index information and storage keys linked to the primary key also included in the search criteria. Subsequently, based on the searched key index information and storage keys, the computer device 200 can search the compressed partitions searched in step 1530 for specific data records that meet the search criteria.
[0131] Furthermore, when a linked list exists that is connected to any list via a primary key, the key index list may also include sub-index information indicating the position of the data record in the second compressed partition, relative to a second compressed partition generated by compressing the data record in the linked list. Therefore, if the search criteria also include a primary key, the computer device 200 may further search the key index list for sub-index information linked to the primary key also included in the search criteria, and further search the second compressed partition for data records that meet the search criteria based on the second compressed partition and the sub-index information. This allows the computer device 200 to obtain not only the field values of the first list being searched, but also the field values of the linked list connected to the first list via the primary key, for a specific data record.
[0132] On the other hand, as described above, the compressed list may also include compressed lists of databases of other computer devices connected to computer device 200 via a network. In this case, computer device 200 can search, in step 1530, the compressed lists of the databases of the other computer devices over the network for compressed partitions connected to the storage key searched in step 1520.
[0133] The above steps 1510 to 1530 may be implemented by the second function provided by the data archiving system 310. In other words, the data archiving system 310 may provide the second function for controlling the target system 320 to execute steps 1510 to 1530.
[0134] Figure 16 and Figure 17 A diagram showing an example of searching archived data according to an embodiment of the present invention.
[0135] Figure 16 An example of searching archived data from a compressed list 1620 by query 1610 is shown. Figure 16 In an embodiment, compression list 1620 is combined with an index list and may include a PERIOD field 1621, a COL1 field 1622, a TABLE field 1623, an OBJECT ID field 1624, a SEQ field 1625, and a COMPRESSED DATA field 1626. Depending on the embodiment, PERIOD field 1621 and COL1 field 1622 may also exist in a separate index list. In this case, to connect compression list 1620 and the index list, the OBJECT ID field 1624 may exist in both lists. Depending on the embodiment, TABLE field 1623 and SEQ field 1625 may also exist in the index list.
[0136] In this case, query 1610 may indicate an instruction to search for data records in table "TAB1" whose PERIOD field 1621 value is "2002.01" and whose COL1 field 1622 value is "1000." At this point, computer device 200 may identify the data record corresponding to query 1610 in compressed table 1620 as a compressed partition stored in COMPRESSED DATA field 1626 of the first record in compressed table 1620. Therefore, computer device 200 may decompress the corresponding compressed partition and provide the multiple data records ("50,000 rows" data records) contained in the corresponding compressed partition as search results.
[0137] Figure 17 To search the columns of archived data from the compressed list 1620 by query 1710. Figure 17 In the embodiment of FIG, since the query 1710 applies the primary key as the search condition, a key index list 1720 may be applied. The key index list 1720 may include a Doc.No. field 1721, an OBJECT ID field 1722, a Key Location Info. field 1723, and a Sub Location Info. field 1724.
[0138] In this case, query 1710 may represent an instruction to search for data records whose Doc. No. field 1721, serving as the primary key, has a value of "1" from the lists "TAB1" and "TAB2." In this case, computer device 200 can identify the first record in key index list 1720 whose Doc. No. field 1721 has a value of "1" and then search for data records whose primary key is "1" from compressed list 1620 using the value of the SubLocation Info. field 1724 of the first record. For example, computer device 200 can retrieve data records whose primary key is "1" from compressed list 1620 using the sub-index information "TAB1@1001-2 / TAB2@2001-3" of the first record in key index list 1720 whose Doc. No. field 1721 has a value of "1." In this case, data records with a specific primary key value can be easily and quickly searched based on the location information contained in the sub-index information, without having to search all data records in the compressed partition.
[0139] Refer again Figure 3In an embodiment where another system (e.g., a cloud storage system) external to the object system 320 includes the storage system 330, the data archiving system 310 can optimize the data of the object system 320 and the storage system 330 by applying a data query catalog. For example, the data archiving system 310 can continuously optimize the data capacity and user access speed between the object system 320 and the storage system 330 by analyzing at least one of (1) a history list access log of a local (on-premise, stored and run by an enterprise through a local device rather than through a cloud environment) database, (2) access volume predicted based on the history list access log through machine learning, and (3) an access log after data is converted to the storage system 330.
[0140] Figure 18 FIG. 1 is a diagram illustrating an example of a process for efficiently storing data according to an embodiment of the present invention. Figure 18 The object system 320 and the cloud system 1810 are shown. Figure 18 In an embodiment, the storage system 330 and the data archiving system 310 may both be implemented on the cloud system 1810. To efficiently store data in the remote storage (the storage system 330 implemented on the cloud system 1810), the data archiving system 310 may manage storage levels based on data usage. For example, the data archiving system 310 may provide the target system 320 with a function for controlling the target system 320 to transmit data based on the data usage of the local database. In this case, the data archiving system 310 may analyze the data usage of the target system 310 using this function to separate the data into different levels. The data may then be separated by level before being transmitted to the cloud system 1810. In this case, the cloud system 1810 may also include hierarchical storage based on the level, storing data corresponding to the level of storage at a specific level.
[0141] Furthermore, the data archiving system 310 can monitor the data usage status of the data transmitted to the cloud system 1810 by business objects and time, and can separate and store the data usage status. For example, the data archiving system 310 can manage the storage of the cloud system 1810 based on the data usage rate in the storage.
[0142] On the other hand, the data archiving system 310 controls the data usage status of the object system 320 to be transmitted to the cloud system 1810 and applies machine learning to analyze the data usage rate, and then stores it in various hierarchical storages. For example, the data archiving system 310 can control the object system 310 to transfer the data usage status within the enterprise to the cloud system 1810 within a specified time, and can predict the data usage rate based on the relevant machine learning application of the transferred data usage status. In addition, the data archiving system 310 can process the data transfer between the object system 320 and the cloud system 1810 in a manner that optimizes the data based on the predicted data usage rate. For example, among the data stored in the memory of the cloud system 1810 (storage system 320), the data with a data usage rate above the first critical value can be transferred to the memory of the object system 320 (database 321), and among the data stored in the memory of the object system 320, the data with a data usage rate above the second critical value can be transferred to the memory of the cloud system 1810. Data transfer may require the above through Figures 3 to 17 The embodiments describe data compression or decompression.
[0143] As such, the data archiving system 310 can continuously perform storage optimization tasks based on the data usage status of the object system 320 (past), the data usage status of the cloud system (current), and the data usage rate predicted by machine learning (future).
[0144] As another example, the data archiving system 310 can provide functionality for optimizing the performance of the target system 320. For example, the target system 320 can be considered to be located in a cloud environment in real-time. In this scenario, after the target system 320 deletes data (or reduces storage usage through the ongoing storage optimization described above), the data archiving system 310 monitors the overall performance (central processing unit (CPU), memory usage, system response speed, etc.) of the target system 320 in the cloud environment in real-time based on the target system 320's database capacity. Based on this monitored performance, the target system 320's specifications can be modified to a server type that reduces costs, thereby reducing costs at the target system 320 level. For example, without optimizing data volume, the data archiving system 310 can provide real-time optimization functionality that also considers CPU and memory efficiency. To this end, the data archiving system 310 can examine the potential for optimizing additional resources resulting from the reduction in data volume. As a more specific example, data archiving system 310 can measure the time required for each process by analyzing the technical bill of materials (BOM) and internal structure of a program that has been used frequently within a specified period of time (e.g., one year). This can reduce the CPU and memory requirements by reducing the processing time of database-related logic. Furthermore, data archiving system 310 can change the real-time level used to implement target system 320 to a lower level of real-time from an economical perspective than the initially set level. In measuring the time required for each process, in addition to the program's technical bill of materials and internal structure, other factors can be considered, such as system response rate, CPU utilization, processing time, and database response time.
[0145] In another embodiment, data archiving system 310 may provide data de-identification capabilities. When collecting data for archiving, de-identification may be implemented based on business requirements and / or legal requirements. Alternatively, de-identification may be required to apply data archived in storage system 330 to systems other than target system 320. Figure 19 A diagram illustrating an example of a data de-identification method according to an embodiment of the present invention.
[0146] As described above, according to an embodiment of the present invention, the present invention can receive a remote function call from an object system that stores data, and in response to such a remote function call, provide the object system with a first function for archiving at least a portion of the data stored in the object system in the storage system via the above-mentioned network, and provide the object system with a second function for querying the data archived in the above-mentioned storage system via the above-mentioned network, thereby providing a remote near-line data archiving function.
[0147] The systems or devices described above can be implemented using hardware structural elements, or a combination of hardware structural elements and software structural elements. For example, the devices and structural elements described in the embodiments can be implemented using at least one general-purpose computer or special-purpose computer, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor (DSP), a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or other device capable of executing and responding to instructions. The processing device can execute an operating system (OS) and at least one software application executed on the operating system. Furthermore, the processing device can access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, although only one processing device is described, it should be understood by those skilled in the art that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, the processing device may include multiple processors or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.
[0148] Software may include a computer program, code, instructions, or a combination of more than one of these, capable of configuring a processing device in a manner that performs work as needed or independently or collectively issuing instructions to a processing device. Software and / or data may be embodied by any type of machinery, component, physical device, virtual equipment, computer storage medium, or device for interpretation by a processing device or for providing instructions or data to a processing device. Software may be distributed across computer systems connected via a network so that it can be stored or executed in a distributed manner. Software and data may be stored in at least one computer-readable recording medium.
[0149] The method of the embodiment can be implemented in the form of program instructions executed by various computer units and recorded on a computer-readable medium. The above-mentioned computer-readable medium may include program instructions, data files, data structures, etc., alone or in combination. The medium may also be temporarily stored for continuous storage, execution, or downloading of computer executable programs. In addition, the medium may be various recording devices or storage devices combined with a single or multiple hardware, and is not limited to a medium directly connected to a computer system, but may also be a medium distributed on a network. As an example, the medium may include magnetic media such as hard disks, floppy disks and magnetic disks, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floppy disks (magneto-optical medium), and devices for storing program instruction languages such as read-only memories, random access memories, and flash memories. In addition, as an example of other media, there may also be application stores for circulating application programs, websites that provide or circulate various other software, and recording media or storage media managed by servers, etc. As an example, program instructions include not only machine language codes generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc.
[0150] Modes for Carrying Out the Invention
[0151] Although the embodiments are described above by way of limiting examples and accompanying drawings, a person skilled in the art may make various modifications and variations based on the above. For example, even if the described techniques are performed in a different order than the described methods and / or the described systems, structures, devices, circuits, and other structural elements are combined or combined in different implementations than the described methods, or even if they are replaced or substituted by other structural elements or equivalent technical solutions, appropriate results may still be achieved.
[0152] Therefore, other implementation modes, other examples and contents equivalent to the scope of protection of the invention also fall within the scope of protection of the appended invention.
Claims
1. A computer program product comprising a computer program, wherein the computer program is stored in a computer-readable recording medium for causing a computer device included in a cloud system to execute a data archiving method, the computer device comprising at least one processor, the data archiving method comprising: receiving, by the at least one processor, a remote function call from a target system storing data; providing, by the at least one processor, a first function to the target system via a network in response to the remote function call, wherein the first function is used to archive at least a portion of the data stored in the target system to a storage system; as well as The at least one processor provides a second function to the target system via the network, wherein the second function is used to query data archived in the storage system. The object system directly processes the archiving and querying of the data archived in the storage system through the network based on the first function and the second function. wherein, under the control of the computer device, the first function and the second function are provided to the target system via the network using Software as a Service (SaaS) registered in the cloud system; Wherein, providing the first function includes providing a third function for compressing at least a portion of the data stored in the local database of the target system and archiving the compressed data in a first list of the local database of the target system, The data archiving method further includes: providing, by the at least one processor, a fourth function to the target system, for instructing the target system to archive and save the data filed in the first list in a file format when a retention period of the data filed in the first list has expired; and providing, by the at least one processor, a fifth function to the target system, for instructing the target system to delete data corresponding to the data filed in the form of the file from the first list, Wherein, providing the first function also includes: providing a 1-1 function for controlling the target system to determine the data record-related partitions included in the first list based on the filtering information of the data records; providing a 1-2 function for controlling the target system to generate compressed partitions by compressing data records for each partition; providing functions 1-3 for controlling the object system so that the compressed partition and a storage key uniquely identifying the compressed partition are linked and stored in a compression list; providing functions 1-4 for controlling the object system so that the storage key and the filter information are linked and stored in an index list of the local database; and providing functions 1-5 for controlling the target system to concatenate, for each of the data records included in the first list, a primary key, key index information indicating a position of the corresponding data record within the compressed partition containing the corresponding data record, and a storage key corresponding to the compressed partition containing the corresponding data record, and store the concatenated data in a key index list; wherein, with respect to a second compressed partition generated by compressing data records in a connection list connected to the index list by the primary key, Functions 1-6 control the object system: searching the data records included in the second compressed partition for data records whose primary keys are the same as the primary keys of the data records included in the index list, and For data records having the same primary key in the key index list, storing sub-index information as the position of the retrieved data record in the second compressed partition, The key index information includes: a sequence being a sequence number of the compressed partition containing a first data record corresponding to a first primary key; and an order of the first data records in the compressed partition corresponding to the sequence, and The sub-index information includes: Identifier of the join list connected by the second primary key; a first delimiter, separating identifiers of the connection lists connected by the second primary key; The range of data records identified by the second primary key for each identifier; and The second separator separates the identifier and range of the same connection list in the connection lists connected by the second primary key.
2. The computer program product according to claim 1, wherein The filtering information includes a given field value of the corresponding data record, and The first to fourth functions control the object system so that the storage key and the given field value are linked and stored in a group index list of the local database.
3. The computer program product according to claim 1, wherein The screening information includes time-related information of the corresponding data record, and The first to fourth functions control the target system so that the storage key and the time-related information are linked and stored in a time index list.
4. The computer program product of claim 1 , wherein: Providing the first function includes also providing functions 1-7 for controlling the object system: In response to a request to restore the deleted data record, searching the index list for a storage key linked to identification information included in the restore request, searching the compressed list for a compressed partition joined with the retrieved storage key, restoring the deleted data records by decompressing the retrieved compressed partitions, and The restored data record is recorded in the first list based on the identification information.
5. The computer program product of claim 1 , wherein: The first-2 function controls the target system to generate the compressed partition by compressing data records included in the determined partition into a binary object.
6. The computer program product of claim 1 , wherein: Providing the second function includes providing: The second-1 function is for controlling the target system to receive a search condition including filtering information of a data record; Function 2-2 is for controlling the target system to search for a storage key connected with the filtering information included in the search condition from the index list connected and storing filtering information of data records and a storage key uniquely identifying a compressed partition including corresponding data records on a local database of the target system; and The second-third function is to control the object system to search for the compressed partition linked to the retrieved storage key from the compressed list that links and stores storage keys and compressed partitions.
7. A data archiving method, performed by a computer device comprising at least one processor, wherein: include: receiving, by the at least one processor, a remote function call from a target system storing data; providing, by the at least one processor, a first function to the target system via a network in response to the remote function call, wherein the first function is used to archive at least a portion of the data stored in the target system to a storage system; as well as The at least one processor provides a second function to the target system via the network, wherein the second function is used to query data archived in the storage system. The object system directly processes the archiving and querying of the data archived in the storage system through the network based on the first function and the second function. wherein, under the control of the computer device, the first function and the second function are provided to the target system via the network using Software as a Service (SaaS) registered in a cloud system; Wherein, providing the first function includes providing a third function for compressing at least a portion of the data stored in the local database of the target system and archiving the compressed data in a first list of the local database of the target system, The data archiving method further includes: providing, by the at least one processor, a fourth function to the target system, for instructing the target system to archive and save the data filed in the first list in a file format when a retention period of the data filed in the first list has expired; and providing, by the at least one processor, a fifth function to the target system, for instructing the target system to delete data corresponding to the data filed in the file form from the first list, Wherein, providing the first function also includes: providing a 1-1 function for controlling the target system to determine the data record-related partitions included in the first list based on the filtering information of the data records; providing a 1-2 function for controlling the target system to generate compressed partitions by compressing data records for each partition; providing functions 1-3 for controlling the object system so that the compressed partition and a storage key uniquely identifying the compressed partition are linked and stored in a compression list; providing functions 1-4 for controlling the object system so that the storage key and the filter information are linked and stored in an index list of the local database; and providing functions 1-5 for controlling the target system to concatenate, for each of the data records included in the first list, a primary key, key index information indicating a position of the corresponding data record within the compressed partition containing the corresponding data record, and a storage key corresponding to the compressed partition containing the corresponding data record, and store the concatenated data in a key index list; wherein, with respect to a second compressed partition generated by compressing data records in a connection list connected to the index list by the primary key, Functions 1-6 control the object system: searching the data records included in the second compressed partition for data records whose primary keys are the same as the primary keys of the data records included in the index list, and For data records having the same primary key in the key index list, storing sub-index information as the position of the retrieved data record in the second compressed partition, The key index information includes: a sequence being a sequence number of the compressed partition containing a first data record corresponding to a first primary key; and an order of the first data records in the compressed partition corresponding to the sequence, and The sub-index information includes: Identifier of the join list connected by the second primary key; a first delimiter separating identifiers of the connection lists connected by the second primary key; The range of data records identified by the second primary key for each identifier; and The second separator separates the identifier and range of the same connection list in the connection lists connected by the second primary key.
8. The data archiving method according to claim 7, wherein: Providing the second function includes providing: The second function is to control the target system to receive a search condition including filtering information of a data record; Function 2-2 is for controlling the target system to search for a storage key connected with the filtering information included in the search condition from the index list connected and storing filtering information of data records and a storage key uniquely identifying a compressed partition including corresponding data records on a local database of the target system; and The second-third function is to control the object system to search for the compressed partition linked to the retrieved storage key from the compressed list that links and stores storage keys and compressed partitions.
9. A computer-readable recording medium, wherein: A computer program is stored in order to cause a computer device to execute the data archiving method according to claim 7 or 8 .
10. A computer device, wherein: comprising at least one processor for executing computer-readable instructions, configured to: receiving remote function calls from the object system storing the data; providing a first function to the target system via a network in response to the remote function call, wherein the first function is used to archive at least a portion of the data stored in the target system in a storage system; as well as providing a second function to the target system via the network, wherein the second function is used to query data archived in the storage system; The object system directly processes the archiving and querying of the data archived in the storage system through the network based on the first function and the second function. wherein, under the control of the computer device, the first function and the second function are provided to the target system via the network using Software as a Service (SaaS) registered in a cloud system; Wherein, providing the first function includes providing a third function for compressing at least a portion of the data stored in the local database of the target system and archiving the compressed data in a first list of the local database of the target system, The at least one processor is further configured to: providing the target system with a fourth function for instructing the target system to archive and preserve the data filed in the first list in a file form when a preservation period of the data filed in the first list has expired; and providing a fifth function to the target system for instructing the target system to delete data corresponding to the data filed in the file form from the first list, The at least one processor is configured to provide the first function as follows: providing a 1-1 function for controlling the target system to determine the data record-related partitions included in the first list based on the filtering information of the data records; providing a 1-2 function for controlling the target system to generate compressed partitions by compressing data records for each partition; providing functions 1-3 for controlling the object system so that the compressed partition and a storage key uniquely identifying the compressed partition are linked and stored in a compression list; providing functions 1-4 for controlling the object system so that the storage key and the filter information are linked and stored in an index list of the local database; and providing functions 1-5 for controlling the target system to concatenate, for each of the data records included in the first list, a primary key, key index information indicating a position of the corresponding data record within the compressed partition containing the corresponding data record, and a storage key corresponding to the compressed partition containing the corresponding data record, and store the concatenated data in a key index list; wherein, with respect to a second compressed partition generated by compressing data records in a connection list connected to the index list by the primary key, Functions 1-6 control the object system: searching the data records included in the second compressed partition for data records whose primary keys are the same as the primary keys of the data records included in the index list, and For data records having the same primary key in the key index list, storing sub-index information as the position of the retrieved data record in the second compressed partition, The key index information includes: a sequence being a sequence number of the compressed partition containing a first data record corresponding to a first primary key; and an order of the first data records in the compressed partition corresponding to the sequence, and The sub-index information includes: Identifier of the join list connected by the second primary key; a first delimiter, separating identifiers of the connection lists connected by the second primary key; The range of data records identified by the second primary key for each identifier; and The second separator separates the identifier and range of the same connection list in the connection lists connected by the second primary key.
Citation Information
Patent Citations
Database-archiving method and apparatus that generate index information, and method and apparatus for searching archived database comprising index information
CN108604249A
Network archiver system, and computer-readable recording medium recording program constructing the system
JP1998161913A
Printer data processing system and printer data processing method
JP2016091556A
Searchable archive
US20050187962A1