Data archiving method, system, electronic apparatus and storage medium
By using streaming processing and distributed storage, the problems of slow querying and high maintenance costs after data archiving are solved, enabling fast storage and querying of data archiving results and saving disk space.
Patent Information
- Application Number
- CN202210807709.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-11
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2042-07-11
Smart Images

Figure CN115309740B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of data processing, and in particular to a data archiving method, system, electronic device and storage medium. BACKGROUND
[0002] At present, the mysql database is generally used to store business data. With the expansion of the business scale, the mysql database stores massive data. For historical data, archiving operations are often required. However, due to the large amount of data, the database and table are generally divided for archiving.
[0003] In the prior art, data archiving is generally performed by backing up historical data to a sql file through mysqldump, backing up historical data through other backup tools, extracting historical data to a data warehouse, or extracting historical data to other databases (such as Tidb). However, after the above-mentioned methods are used to archive data, the data cannot be viewed and quickly queried at any time, and the maintenance cost is high.
[0004] At present, there is no effective solution to the problems of the archived data being unable to be quickly queried and the high maintenance cost in the related art. SUMMARY
[0005] The present application provides a data archiving method, system, electronic device and storage medium to solve the problems of the archived data being unable to be quickly queried and the high maintenance cost in the related art.
[0006] In a first aspect, the present application provides a data archiving method, comprising:
[0007] Obtaining data to be archived after database and table division, and shard meta information of the data to be archived;
[0008] Inserting a streaming service into the data to be archived, and performing streaming processing on the data to be archived according to at least the shard meta information and unique identifier information corresponding to the streaming service, to obtain first synchronization data;
[0009] Creating a distributed database according to the shard meta information, and synchronizing the first synchronization data to the distributed database for distributed storage, to obtain second synchronization data;
[0010] Taking the second synchronization data as a data archiving result.
[0011] In some embodiments, the streaming processing on the data to be archived according to at least the shard meta information and unique identifier information corresponding to the streaming service, to obtain first synchronization data, comprises:
[0012] acquire a time field of the data to be archived according to the shard meta information;
[0013] acquire a configuration parameter of the streaming service, and acquire the unique identifier information according to the configuration parameter;
[0014] obtain the first synchronization data according to at least the shard meta information, the time field of the data to be archived, and the unique identifier information.
[0015] In some embodiments, the synchronizing the first synchronization data into the distributed database for distributed storage to obtain second synchronization data comprises:
[0016] synchronizing the first synchronization data into the distributed database to obtain second synchronization data according to at least the shard meta information and field information of the distributed database.
[0017] In some embodiments, after the second synchronization data is obtained, before the second synchronization data is taken as the data archiving result, the method further comprises:
[0018] comparing the data to be archived and the second synchronization data to obtain a comparison result; in a case where the comparison result indicates that the data to be archived and the second synchronization data are different, acquiring and executing a transmission state statement of the distributed database according to at least the field information of the distributed database to obtain a transmission state result; in a case where the transmission state result indicates that the distributed database is in a synchronization state, acquiring a preset waiting time length, and after waiting for the preset waiting time length, obtaining a re-comparison result;
[0019] or, deleting the synchronized second synchronization data, performing streaming processing on the data to be archived to obtain third synchronization data, synchronizing the third synchronization data into the distributed database to obtain fourth synchronization data, and comparing the data to be archived and the fourth synchronization data to obtain a re-comparison result;
[0020] in a case where the comparison result or the re-comparison result indicates that the data to be archived and the second synchronization data are the same, deleting the data to be archived according to at least the shard meta information.
[0021] In some embodiments, after the second synchronization data is taken as the data archiving result, the method further comprises:
[0022] acquiring a recycle shard statement for the data archiving result according to the shard meta information, and executing the recycle shard statement to recycle disk fragments of the data to be archived.
[0023] In some embodiments, after the second synchronization data is taken as the data archiving result, the method further comprises:
[0024] Distributed meta information of the distributed database is acquired, and the data archiving result corresponding to the data to be archived is queried according to at least the distributed meta information.
[0025] In some embodiments, the distributed database is a StarRocks database, and / or the streaming service is a maxwell service.
[0026] In a second aspect, a data archiving system is provided in the present embodiment, comprising a terminal device, a transmission device and a server device; wherein the terminal device is connected to the server device through the transmission device;
[0027] The server device is configured to execute the data archiving method of the first aspect described above;
[0028] The transmission device is configured to transmit the data archiving result.
[0029] The terminal device is configured to display the data archiving result.
[0030] In a third aspect, an electronic device is provided in the present embodiment, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor executes the computer program to implement the data archiving method of the first aspect described above.
[0031] In a fourth aspect, a storage medium is provided in the present embodiment, and the storage medium stores a computer program executable by a processor to implement the data archiving method of the first aspect described above.
[0032] Compared with the related art, the data archiving method, system, electronic device and storage medium provided in the present embodiment solve the problem of high maintenance cost and slow query of archived data by acquiring the data to be archived after database and table splitting, and shard meta information of the data to be archived; inserting a streaming service into the data to be archived, and performing streaming processing on the data to be archived according to at least the shard meta information and unique identifier information corresponding to the streaming service, to obtain first synchronization data; creating a distributed database according to the shard meta information, and synchronizing the first synchronization data to the distributed database for distributed storage, to obtain second synchronization data; and taking the second synchronization data as a data archiving result, so as to realize fast storage and query of the data archiving result.
[0033] The details of one or more embodiments of the present application are set forth in the following drawings and description, so that other features, objects and advantages of the present application are more apparent. Attached Figure Description
[0034] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0035] Figure 1 This is an application scenario diagram of a data archiving method in one embodiment;
[0036] Figure 2 This is a flowchart illustrating a data archiving method in one embodiment;
[0037] Figure 3 This is a flowchart illustrating a data archiving method in another embodiment;
[0038] Figure 4 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0039] To better understand the purpose, technical solution, and advantages of this application, the application is described and illustrated below in conjunction with the accompanying drawings and embodiments.
[0040] Unless otherwise defined, the technical or scientific terms used in this application shall have the general meaning as understood by one of ordinary skill in the art to which this application pertains. Words such as “a,” “an,” “an,” “the,” “the,” and “these,” used in this application, do not indicate quantitative limitation and may be singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps or modules (units) is not limited to the listed steps or modules (units) but may include steps or modules (units) not listed, or may include other steps or modules (units) inherent to such processes, methods, products, or devices. The terms “connected,” “linked,” and “coupled,” used in this application, are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. The term “multiple” used in this application refers to two or more. The "and / or" operator describes the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: A alone, A and B simultaneously, and B alone. Typically, the character " / " indicates that the objects before and after it are in an "or" relationship. The terms "first," "second," and "third," etc., used in this application are merely for distinguishing similar objects and do not represent a specific ordering of the objects.
[0041] The data archiving method provided in the application can be applied to an application environment as shown in Figure 1 The terminal device 102 communicates with the server device 104 through a network. The server device 104 obtains the to-be-archived data after database and table splitting and the shard metadata of the to-be-archived data. The server device 104 inserts a streaming service into the to-be-archived data, at least performs streaming processing on the to-be-archived data according to the shard metadata and unique identifier information corresponding to the streaming service, to obtain first synchronization data. The server device 104 creates a distributed database according to the shard metadata and synchronizes the first synchronization data to the distributed database for distributed storage, to obtain second synchronization data. The server device 104 takes the second synchronization data as a data archiving result. The terminal device 102 is configured to display the data archiving result. The terminal device 102 can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices. The server device 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0042] In the embodiment, a data archiving method is provided, Figure 2 The flowchart of the data archiving method of the embodiment is shown in Figure 2 The flowchart includes the following steps:
[0043] In step S202, the to-be-archived data after database and table splitting and the shard metadata of the to-be-archived data are obtained. The to-be-archived data is massive data, and the to-be-archived data is at least one piece after database and table splitting. The to-be-archived data is stored in a relational database such as MySQL and NoSQL, and preferably stored in MySQL. The shard metadata of the to-be-archived data includes the IP address, port, database name, table name, time field, user name, user password and user identifier of the to-be-archived data.
[0044] Step S204, inserting a streaming service into the data to be archived, at least according to the shard element information, and the unique identifier information corresponding to the streaming service to stream process the data to be archived to obtain first synchronization data. The streaming service refers to a service of converting the data to be archived into streaming data, and has the characteristics of periodicity and real-time, and can periodically or in real time stream process at least one piece of the data to be archived. The streaming service can be Canal, maxwell or mysql_streamer, etc. There is at least one unique identifier information, and the unique identifier information corresponds to the data to be archived one by one. The first synchronization data is stored in a topic of a streaming database, and the streaming database can be kafka, Redis or Kinesis, etc. Preferably, the first synchronization data is stored in kafka. In particular, in the embodiment, the insertion of the streaming service into the data to be archived is based on the relational database shard table shard modulo implementation, and the data to be archived after one database, table and / or shard modulo corresponds to one streaming service. The insertion operation of the streaming service can be realized by a sql statement, for example, a bootstrap statement is used in the maxwell service.
[0045] Step S206, creating a distributed database according to the shard element information, and synchronizing the first synchronization data to the distributed database for distributed storage to obtain second synchronization data. The distributed database can be TiDB, Spanner or StarRocks, etc.
[0046] Step S208, taking the second synchronization data as a data archiving result. The data archiving method in the embodiment can be automatically run through a shell script.
[0047] Through the above steps, the data to be archived is first periodically or in real time synchronized into the first synchronization data in the streaming database, and then the first synchronization data is synchronized into the second synchronization data in the distributed database, the data to be archived is data archived, and the data archiving result is obtained. The data archiving result adopts distributed columnar storage, can easily support massive data storage, has high access performance, saves a large amount of disk space, thereby realizing fast storage and query of the data archiving result, and solving the problems of slow query and high maintenance cost of the archived data.
[0048] In some embodiments, the at least according to the shard element information, and the unique identifier information corresponding to the streaming service to stream process the data to be archived to obtain first synchronization data comprises:
[0049] According to the shard element information, obtaining the time field of the data to be archived;
[0050] Obtain the configuration parameters of the streaming service, and then obtain the unique identifier information based on the configuration parameters;
[0051] The first synchronized data can be obtained based at least on the fragment metadata, the time field of the data to be archived, and the unique identifier information.
[0052] The configuration parameters of the streaming service also include filter parameters. These filter parameters perform basic filtering on the data to be archived during the operation of the streaming service. The filter parameters can be userid>0, where userid refers to the user identifier that generated the data to be archived.
[0053] Specifically, in this embodiment, the start and end times are obtained, which are time conditions for filtering the time field of the data to be archived; the time field of the data to be archived is obtained based on the shard metadata; the configuration parameters of the streaming service are obtained, and the unique identifier information and filter parameters are obtained based on the configuration parameters; based on the start and end times, the time field of the data to be archived, the unique identifier information, and the shard name, table name, and user identifier of the shard metadata, the data to be archived is extracted and stored in the streaming database to obtain the first synchronized data. Taking Maxwell as the streaming service and Kafka as the streaming database as an example, firstly, the Kafka database is started and a topic is created to store the first synchronized data; secondly, for each shard of the data to be archived, a corresponding Maxwell service is synchronously created and integrated into the Maxwell management database, where all Maxwell services are managed; for each shard of the data to be archived, the following SQL is executed:
[0054] insert into bootstrap(database_name,table_name,where_clause,client_id)
[0055] values('database name', 'table name', "where condition", 'maxwell's unique identifier information');
[0056] Among them, the database name and table name refer to the database name and table name in the sharding metadata of the data to be archived; the WHERE condition refers to the condition for filtering the time field based on the start time and the end time; the Maxwell unique identifier information refers to the unique identifier information obtained from the configuration parameters of the Maxwell service.
[0057] After starting the maxwell service by executing the above sql statement, check the topic message of kafka, which is empty, then run normally, maxwell automatically extracts historical data to the corresponding topic of kafka, and generates the first synchronization data.
[0058] Through the above steps, according to the obtained start time and end time, the historical data in the specified time range can be periodically or real-time extracted to the stream database, so as to easily realize the extraction function of massive data to be archived, and support the fast extraction of hundreds of tables of data to be archived; By setting the filtering parameters, invalid data is filtered out, so as to improve the data extraction efficiency and quality, and solve the problems of slow query and high maintenance cost of archived data.
[0059] In some embodiments, the first synchronization data is synchronized to the distributed database for distributed storage to obtain second synchronization data, including:
[0060] At least according to the shard meta information and the field information of the distributed database, the first synchronization data is synchronized to the distributed database to obtain second synchronization data.
[0061] The field information of the distributed database includes the library name, table name and field name of the distributed database, and the field information of the distributed database corresponds to the shard meta information of the data to be archived.
[0062] Specifically, the first synchronization data to be archived is obtained from the stream database according to the shard meta information; the IP, port, username and user password of the distributed database are obtained, and the first synchronization data is synchronized to the distributed database according to the IP, port, username and user password of the distributed database and the field information of the distributed database, to obtain second synchronization data. Taking the stream database as kafka and the distributed database as StarRocks library as an example, the data archiving method in the embodiment first obtains the IP, port, username and user password of the deployed StarRocks, logs in the StarRocks service and creates a StarRocks library according to the shard meta information and the syntax requirements of StarRocks, and the StarRocks library is partitioned according to the time field; according to the topic format of the stream database kafka and the StarRocks field information, the following sql is executed to create the corresponding routine load task:
[0063]
[0064] In this embodiment, sr_db and sr_tab refer to the library name and table name of StarRocks; rl_name refers to the name of a specific routine load task; the where condition refers to the corresponding filter condition of c1, c2, c3, c4, c5, etc., for example, c1 is a partition field, and the where condition can be: c1>0; the kafka cluster link address and the topic name of kafka refer to the specific kafka address and topic name respectively, and if the topic is set with N partitions, the topic partition number is replaced with: 0, 1, 2,..., N-1.
[0065] Through the above steps, the first synchronization data is synchronized to the second synchronization data through distributed storage, the valid messages can be accurately distinguished according to the field information of the distributed database, massive data storage can be easily supported, and a large amount of disk space can be saved, so that the fast storage and query of the data archiving result are realized, and the problems of slow query and high maintenance cost of the archived data are solved.
[0066] In some embodiments, after obtaining the second synchronization data, before taking the second synchronization data as the data archiving result, the method further includes:
[0067] comparing the to-be-archived data and the second synchronization data to obtain a comparison result; in a case where the comparison result indicates that the to-be-archived data and the second synchronization data are different, obtaining and executing a transmission state statement of the distributed database according to at least the field information of the distributed database to obtain a transmission state result; in a case where the transmission state result indicates that the distributed database is in a synchronization state, obtaining a preset waiting time length, and after waiting for the preset waiting time length, obtaining a re-comparison result;
[0068] Alternatively, deleting the synchronized second synchronization data, performing stream processing on the to-be-archived data to obtain third synchronization data, synchronizing the third synchronization data to the distributed database to obtain fourth synchronization data, and comparing the to-be-archived data and the fourth synchronization data to obtain a re-comparison result.
[0069] in a case where the comparison result or the re-comparison result indicates that the to-be-archived data and the second synchronization data are the same, deleting the to-be-archived data according to at least the shard meta information.
[0070] The comparison result can be a result of comparing data volume or data similarity of the to-be-archived data and the second synchronization data in the same time range. Taking data volume comparison as an example, the comparison result includes two results of data volume difference and data volume same. In the case of data volume difference, the process of synchronizing the second synchronization data to the distributed database is ended, or the above-mentioned process of stream processing and distributed synchronization is re-executed to obtain third synchronization data and fourth synchronization data, so as to obtain again a comparison result of data volume. In the case of data volume same, the to-be-archived data is deleted to release the disk space.
[0071] Specifically, taking the stream service as maxwell, the stream database as kafka, the distributed database as StarRocks database, and the comparison result as data volume comparison result as an example, in the embodiment, the data volume of the to-be-archived data and the second synchronization data is compared to obtain a comparison result of data volume. In the case that the comparison result of data volume indicates that the data volume of the to-be-archived data and the second synchronization data is different, a sql statement is executed according to the field information of the distributed database StarRocks database to obtain the synchronization state of StarRocks, for example: show routine load for rl_name\G. In the case that the running result of the sql statement indicates that the transmission state result State of StarRocks is the synchronization state RUNNING and the value of the ReasonOfStateChanged field is empty, it indicates that the StarRocks is normally synchronizing the kafka message. A preset waiting time is obtained, and after waiting for the preset waiting time, a comparison result is obtained again.
[0072] Or, the second synchronization data that has been synchronized is deleted, the to-be-archived data is re-executed to perform maxwell stream processing transmission to kafka to obtain third synchronization data, the third synchronization data is synchronized to the distributed database StarRocks database to obtain fourth synchronization data, and the data volume of the to-be-archived data and the fourth synchronization data is compared to obtain a comparison result again.
[0073] In the case that the comparison result or the comparison result again indicates that the to-be-archived data and the second synchronization data are the same, the to-be-archived data is deleted according to the shard meta information, the start time and the end time, and the following sql statement is executed:
[0074] pt-archiver --source h=shard IP, P=shard port, D=shard name, t=shard name, u=mysql username, p=mysql password --where "where condition" --purge --limit=1000 --no-check-charset --txn-size=1000 --progress=1000 --max-lag=3600 --why-quit
[0075] The shard IP, the shard port, the shard name, the shard name, the mysql username, and the mysql password respectively refer to the IP address, the port, the shard name, the shard name, the username, and the user password in the shard information. The where condition refers to a condition for filtering a time field according to a start time and an end time.
[0076] Through the above steps, the comparison result is obtained by comparing the second synchronization data and the to-be-archived data, large-scale table splitting and StarRocks table data volume comparison can be realized, the data archiving accuracy is ensured, the to-be-archived data in a certain time range of the second synchronization data is deleted, the database service is not affected, a large amount of disk space can be released, and the data storage cost is saved, thereby improving the data archiving accuracy and efficiency, and solving the problems of slow query and high maintenance cost of the archived data.
[0077] In some embodiments, after the second synchronization data is taken as the data archiving result, the method further includes:
[0078] According to the shard information, a recycle shard statement for the data archiving result is obtained, and the recycle shard statement is executed to recycle the disk fragments of the to-be-archived data.
[0079] Specifically, the sql of the recycle shard statement is as follows:
[0080] perl pt-online-schema-change h=shard IP, P=shard port, D=shard name, t=shard name, u=mysql username, p=mysql password --alter "recycle shard sql" --execute --recursion-method="hosts" --charset=utf8mb4 --critical-load Threads_running=500 --no-check-alter --no-version-check
[0081] The "recycle fragment sql" in the above recycle fragment statement is in the form of: alter table shard db.shard name engine = innodb.
[0082] By the above steps, the disk fragments of the archived data are recycled, a large amount of disk space is released, the data storage cost is saved, and the problems of slow query and high maintenance cost of the archived data are solved.
[0083] In some embodiments, after recycling the disk fragments of the data to be archived, the method further comprises:
[0084] Obtaining distributed meta information of the distributed database, and querying the data archiving result corresponding to the data to be archived according to at least the distributed meta information.
[0085] The distributed meta information refers to the IP, port, username, user password, database name and table name of the distributed database.
[0086] By the above steps, the data archiving result after archiving is queried, the fast storage and query of the data archiving result are realized, and the problems of slow query and high maintenance cost of the archived data are solved.
[0087] In some embodiments, the distributed database is a StarRocks database, and / or the streaming service is a maxwell service.
[0088] The present embodiment also provides a data archiving method. Figure 3 The flowchart of another data archiving method of the present embodiment is shown in Figure 3 The flowchart of another data archiving method of the present embodiment is shown in
[0089] Step S302, obtaining data to be archived. The data to be archived after database and table splitting is obtained, and the shard meta information of the data to be archived is obtained.
[0090] Step S304, the streaming service synchronizes the data to be archived to obtain first synchronization data. The streaming service is inserted into the data to be archived, the time field of the data to be archived is obtained according to the shard meta information; the configuration parameters of the streaming service are obtained, and the unique identifier information is obtained according to the configuration parameters; the data to be archived is extracted and stored into the streaming database according to the start time, the end time, the time field of the data to be archived, the unique identifier information, and the database name, table name and user identifier of the shard meta information, to obtain the first synchronization data.
[0091] Step S306, the first synchronization data is distributed processed to obtain second synchronization data. A distributed database is created according to the sharding element information, the first synchronization data is synchronized into the distributed database according to the IP, port, user name and user password of the distributed database and the field information of the distributed database to obtain the second synchronization data.
[0092] Step S308, the second synchronization data is verified. It is judged whether the second synchronization data is same as the to-be-archived data. If yes, step S310 is executed; if no, step S304 is returned to be executed.
[0093] Step S310, the to-be-archived data of which the archiving is completed is deleted. The to-be-archived data of which the archiving is completed is deleted according to the sharding element information, the start time and the end time, and the second synchronization data is taken as the data archiving result.
[0094] Step S312, the sql disk fragments are recycled. The recycling fragment statement for the data archiving result is obtained according to the sharding element information, and the recycling fragment statement is executed to recycle the disk fragments of the to-be-archived data.
[0095] Through the above steps, the to-be-archived data is first periodically or real-timely synchronized as the first synchronization data in the stream database, and then the first synchronization data is synchronized as the second synchronization data in the distributed database, the to-be-archived data is archived to obtain the data archiving result, the data archiving result adopts the distributed column storage, can easily support mass data storage, has high access performance, and saves a large amount of disk space, so that the fast storage and query of the data archiving result are realized, and the problem that the archived data cannot be quickly queried and the maintenance cost is high is solved.
[0096] It should be understood that, although Figures 2-3 The steps in the flowchart are shown in sequence according to the arrows, but these steps are not necessarily executed in sequence according to the arrows. Unless otherwise specified in this article, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, Figures 2-3 At least part of the steps in the flowchart can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily sequential, but can be alternately executed with other steps or at least part of the sub-steps or stages of other steps.
[0097] In this embodiment, a data archiving system is also provided, characterized in that it comprises a terminal device 102, a transmission device and a server device 104; wherein the terminal device 102 is connected to the server device 104 through the transmission device;
[0098] The server device 104 is configured to perform the steps in any of the above method embodiments.
[0099] The transmission device is configured to transmit the data archiving result.
[0100] The terminal device 102 is configured to display the data archiving result.
[0101] In this embodiment, an electronic device is also provided, which includes a memory and a processor, the memory stores a computer program, and the processor is configured to execute the computer program to perform the steps in any of the above method embodiments.
[0102] Optionally, the electronic device described above can further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0103] Optionally, in this embodiment, the processor can be configured to perform the following steps through the computer program:
[0104] S1, obtaining the to-be-archived data after the database and table splitting, and shard metadata information of the to-be-archived data.
[0105] S2, inserting a streaming service into the to-be-archived data, and performing streaming processing on the to-be-archived data according to at least the shard metadata information and unique identifier information corresponding to the streaming service, to obtain first synchronization data.
[0106] S3, creating a distributed database according to the shard metadata information, and synchronizing the first synchronization data to the distributed database for distributed storage, to obtain second synchronization data.
[0107] S4, taking the second synchronization data as a data archiving result.
[0108] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, which will not be described herein again.
[0109] In addition, in combination with the data archiving method provided in the above embodiments, a storage medium can also be provided to implement the method in this embodiment. The storage medium stores a computer program; when the computer program is executed by a processor, any of the above data archiving methods is implemented.
[0110] In one embodiment, a computer device is provided, which can be a server, and its internal structure diagram can be as shown in Figure 4As shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium, an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The database of the computer device is used to store data archiving result data. The network interface of the computer device is used to communicate with the external terminal through the network connection. The computer program is executed by the processor to implement a data archiving method.
[0111] Those skilled in the art can understand that, Figure 4 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0112] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM) and the like.
[0113] It should be understood that the specific embodiments described herein are used to explain this application, but not to limit it. According to the embodiments provided by the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0114] It is apparent that the drawings depicted are only a few examples of embodiments of this application and that a person of ordinary skill in the art would be able to adapt these drawings to a variety of other situations without paying creative thought. It is also understood that, although the work done in developing this application can have been lengthy and expensive, certain modifications, such as design, manufacture, or production changes, made in light of the teachings of this application to an ordinarily skilled person in the art are deemed to be within the scope of the disclosure.
[0115] The word "example" is used herein to mean incorporating by reference the specific feature, structure, or characteristic being described with an embodiment. The appearance of the word "example" in various places in the specification is not necessarily intended to further enhance its significance as regards specific embodiments of the application. It should be understood that an embodiment described herein as including an example encompasses both alternative and / or "what if" embodiments that might be explicit as certifications in this Overview.
[0116] The above-described embodiments are merely some implementations of the present application, and the description is specific and detailed, but it should not be understood as a limitation on the scope of patent protection. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of the present application. Therefore, the scope of protection of the present application should be subject to the appended claims.
Claims
1. A data archiving method, characterized in that, include: Obtain the data to be archived after database sharding and table partitioning, as well as the fragment metadata of the data to be archived; The streaming service is inserted into the data to be archived, and the data to be archived is streamed based at least on the fragment metadata and the unique identifier information corresponding to the streaming service to obtain the first synchronized data. A distributed database is created based on the fragmentation metadata, and the first synchronized data is synchronized to the distributed database for distributed columnar storage to obtain the second synchronized data. The second synchronized data is used as the data archiving result; The fragmented metadata of the data to be archived includes the database name, table name, and user identifier of the data to be archived; the unique identifier information corresponds one-to-one with the data to be archived. The step of performing streaming processing on the data to be archived based at least on the fragment metadata and the unique identifier information corresponding to the streaming service to obtain the first synchronized data includes: Obtain the start time and end time, and obtain the time field of the data to be archived based on the shard metadata; obtain the configuration parameters of the streaming service, and obtain the unique identifier information based on the configuration parameters; extract and store the data to be archived into the streaming database based on the start time, end time, the time field of the data to be archived, the unique identifier information, and the database name, table name and user identifier of the shard metadata to obtain the first synchronization data; The streaming service is Maxwell, and the streaming database is Kafka.
2. The data archiving method according to claim 1, characterized in that, The step of synchronizing the first synchronized data to the distributed database for distributed storage to obtain the second synchronized data includes: Based at least on the fragmentation metadata and the field information of the distributed database, the first synchronization data is synchronized to the distributed database to obtain the second synchronization data.
3. The data archiving method according to claim 1, characterized in that, After obtaining the second synchronized data, and before using the second synchronized data as the data archiving result, the method further includes: The data to be archived and the second synchronized data are compared to obtain a comparison result. If the comparison result indicates that the data to be archived and the second synchronized data are different, the transmission status statement of the distributed database is obtained and executed based on the field information of the distributed database to obtain a transmission status result. If the transmission status result indicates that the distributed database is in a synchronized state, a preset waiting time is obtained, and after waiting for the preset waiting time, a comparison result is obtained again. Alternatively, delete the already synchronized second synchronized data, perform streaming processing on the data to be archived to obtain third synchronized data, synchronize the third synchronized data to the distributed database to obtain fourth synchronized data; compare the data to be archived and the fourth synchronized data to obtain a comparison result; If the comparison result or the second comparison result indicates that the data to be archived and the second synchronized data are the same, the data to be archived shall be deleted at least according to the fragment metadata.
4. The data archiving method according to claim 1, characterized in that, After using the second synchronized data as the data archiving result, the method further includes: Based on the fragment metadata, obtain the fragment recovery statement for the data archiving result, and execute the fragment recovery statement to recover the disk fragments of the data to be archived.
5. The data archiving method according to any one of claims 1 to 4, characterized in that, After using the second synchronized data as the data archiving result, the method further includes: Obtain the distributed metadata of the distributed database, and at least query the data archiving result corresponding to the data to be archived based on the distributed metadata.
6. The data archiving method according to claim 1, characterized in that, The distributed database is the StarRocks database.
7. A data archiving system, characterized in that, include: Terminal equipment, transmission equipment, and server equipment; wherein the terminal equipment is connected to the server equipment through the transmission equipment; The server device is used to execute the data archiving method according to any one of claims 1 to 6; The transmission device is used to transmit data archiving results; The terminal device is used to display the data archiving results.
8. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the data archiving method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the data archiving method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-source heterogeneous incremental data synchronization method and system
CN111723160A