Artificial Intelligence-Based Data Archiving Method, Device, Equipment, and Storage Medium

By initializing the data archiving system, loading archive information and determining the list of automatic archiving rules, the problem of inability to archive existing data in the prior art is solved, automatic archiving and performance monitoring of the entire library is realized, and the efficiency and flexibility of data archiving are improved.

CN113946543BActive Publication Date: 2025-07-22PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111273636.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-29
Publication Date
2025-07-22
Estimated Expiration
2041-10-29

AI Technical Summary

Technical Problem

Existing data archiving methods cannot archive existing data, and automatic archiving of the entire library cannot be achieved when the archived table type does not support the archive. This lacks performance monitoring, which affects the flexibility of the database and the efficiency of data archiving.

Method used

By initializing the data archiving system, loading archive information, determining the automatic archive rule list, and receiving archive requests based on the idle status information of the data archiving system, traversing the automatic archive rule list for data archiving processing, supporting the archiving and concurrent number control, recording the archive operation details and statistical data.

Benefits of technology

It realizes the automation of full-data archiving, improves the efficiency and flexibility of data archiving, supports secondary archiving, adapts to the progress of the source library, and provides global self-monitoring to improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113946543B_ABST
    Figure CN113946543B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and discloses a data archiving method based on artificial intelligence, including: initializing a data archiving system, and loading archiving information from a preset configuration source into the data archiving system; preprocessing the archiving information to determine an automatic archiving rule list; receiving an archiving request for archiving the data in the automatic archiving rule list based on the idle state information of the data archiving system; traversing the automatic archiving rule list based on the archiving request, and performing archiving processing on the data in the automatic archiving rule list. The present invention can improve the efficiency of data archiving based on artificial intelligence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and in particular, to a method, apparatus, electronic device, and computer-readable storage medium for data archiving based on artificial intelligence. Background Art

[0002] With the rapid development of the network, the data stored in its backend database and the occupied disk space are also increasing. To ensure the normal operation of online services, it is necessary to clean up historical data in a timely manner to free up the storage space of the database.

[0003] Currently, open-source distributed time-series databases are increasingly being applied to major network enterprises due to their high performance, automatic cleaning of expired data, support for various aggregation functions, etc. In the case of massive data updates, the original data is usually saved for a relatively short period of time. To reduce the storage pressure of the database and improve the query speed of data, etc., it is usually necessary to archive the original data.

[0004] The existing data archiving methods mainly have the following disadvantages: they cannot archive the stock data, and when there is a table type that does not support archiving, they cannot achieve full-library automatic archiving; in addition, there is a lack of performance monitoring, and automatic table addition or deletion operations cannot be realized, which affects the flexibility of the database, the efficiency of data archiving, and the user experience. Summary of the Invention

[0005] The present invention provides a method, apparatus, electronic device, and computer-readable storage medium for data archiving based on artificial intelligence, and its main purpose is to improve the efficiency and flexibility of data archiving based on artificial intelligence.

[0006] To achieve the above object, a method for data archiving based on artificial intelligence provided by the present invention includes: initializing a data archiving system, and loading archiving information from a preset configuration source into the data archiving system;

[0007] Preprocessing the archiving information to determine a list of automatic archiving rules for the archiving information;

[0008] Receiving an archiving request for archiving the data in the list of automatic archiving rules based on the idle state information of the data archiving system; wherein, the state information is acquired by collecting based on a preset frequency;

[0009] Traversing the list of automatic archiving rules based on the archiving request, and performing archiving processing on the data in the list of automatic archiving rules.

[0010] In addition, an optional technical solution is that the archiving information includes the concurrency number, original library information, and a list of archiving strategies; wherein, the archiving strategies in the list of archiving strategies include:

[0011] Source library, used to represent the data source for archiving;

[0012] Target library, used to represent the location where the archived data is stored;

[0013] Archiving period round, used to represent the time period for data aggregation;

[0014] List of aggregation functions, including average value, maximum value, P95;

[0015] Delay time delay, used to represent the delay time of the data set;

[0016] Initial archiving time, and the number of failure retries.

[0017] In addition, an optional technical solution is that preprocessing the archiving information in the data archiving system to determine the list of automatic archiving rules for the archiving information includes:

[0018] Summarize the archiving policies in the archiving information in the data archiving system to form an initial list;

[0019] Sort the archiving policies in the initial list according to preset rules to determine the list of automatic archiving rules for the archiving information; where,

[0020] The preset rules include: arranging the archiving policies with the source library being the original library in the front, and the remaining archiving policies satisfying the criterion that the target library of the previous archiving policy of the current archiving policy is the source library of the next archiving policy of the current archiving policy.

[0021] In addition, an optional technical solution is that after the archiving request is sent and before traversing the list of automatic archiving rules, it also includes a process of screening the data tables in the original library, and the screening process includes:

[0022] Obtain all the data tables in the original library at the current time T and form a table set MS;

[0023] Traverse the table set MS and obtain the latest value corresponding to each data table based on the last function;

[0024] Judge the type of each latest value and filter the data tables that do not meet the archiving conditions based on the type;

[0025] Among them, if the type of the latest value of the data table is not a numeric type, it indicates that the data table corresponding to the latest value does not meet the archiving conditions.

[0026] In addition, an optional technical solution is that the process of archiving the data in the automatic archiving rule list includes:

[0027] Obtain the archiving progress of the current archiving policy. If the archiving progress is empty, set the archiving progress to the initial archiving time of the current archiving policy;

[0028] Determine whether the source library of the current archiving policy is the original library. If not, continue to search for the second archiving policy and set the target library of the second archiving policy to the source library of the current archiving policy;

[0029] Obtain the archiving progress of the second archiving policy and determine whether the archiving progress of the second archiving policy is less than the current round of archiving time. If the archiving progress of the second archiving policy is less than the current round of archiving time, end the current process; otherwise, execute the next step;

[0030] Determine the relationship between the current round of archiving time and the current time T. If the current time T < the current round of archiving time + delay time, end the current process; otherwise, perform one round of archiving operation on the data corresponding to the current archiving policy.

[0031] In addition, an optional technical solution is that the process of performing one round of archiving operation on the data corresponding to the current archiving policy includes:

[0032] Create a task queue based on the control module, generate multiple tasks according to the table set MS, and put the multiple tasks into the task queue; wherein, the content of the task includes the data table, the archiving policy, the last archiving time, and the current round of archiving time;

[0033] Create a statistical queue based on the control module, and create and start multiple working threads according to the concurrency number in the archiving information;

[0034] The working threads obtain the corresponding task information from the task queue, and perform data archiving in the order of the archiving policy based on the task information and the automatic archiving rule list.

[0035] In addition, an optional technical solution is that the process of the working threads obtaining the corresponding task information from the task queue and performing data archiving in the order of the archiving policy based on the task information and the automatic archiving rule list includes:

[0036] Based on the task information, preprocess the data table in the source library corresponding to the current archiving policy to obtain the corresponding data points;

[0037] Save the data points to the target library corresponding to the source library to complete the archiving operation of the data table in the current source library;

[0038] Repeat the above steps in the order of the archiving policy until all data tables are archived.

[0039] To solve the above problems, the present invention also provides an artificial intelligence-based data archiving device, which includes: an initialization unit for initializing a data archiving system and loading archiving information from a preset configuration source into the data archiving system;

[0040] An automatic archiving rule list determination unit for preprocessing the archiving information to determine an automatic archiving rule list of the archiving information;

[0041] An archiving request sending unit for receiving an archiving request for archiving data in the automatic archiving rule list based on the idle state information of the data archiving system; wherein, the state information is collected and obtained based on a preset frequency;

[0042] A data archiving unit for traversing the automatic archiving rule list based on the archiving request and performing archiving processing on the data in the automatic archiving rule list.

[0043] To solve the above problems, the present invention also provides an electronic device, which includes:

[0044] A memory storing at least one instruction; and

[0045] A processor for executing the instructions stored in the memory to implement the above-mentioned artificial intelligence-based data archiving method.

[0046] To solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is executed by a processor in an electronic device to implement the above-mentioned artificial intelligence-based data archiving method.

[0047] In the embodiment of the present invention, by initializing a data archiving system, loading archiving information from a preset configuration source into the data archiving system, then preprocessing the archiving information to determine an automatic archiving rule list; and receiving an archiving request for archiving data in the automatic archiving rule list based on the idle state information of the initialized data archiving system collected at a preset frequency, and finally traversing the automatic archiving rule list based on the archiving request and performing archiving processing on the data in the automatic archiving rule list, it supports secondary archiving based on archived data, can adapt to the progress of the source library, and supports archiving of stock data and concurrency control, and records the operation details and statistical data of each round of archiving, providing a relatively complete global and table-level self-monitoring, improving the efficiency of data archiving and the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 Schematic flowchart of a data archiving method based on artificial intelligence according to an embodiment of the present invention;

[0049] Figure 2 Principle block diagram of a data archiving system according to an embodiment of the present invention;

[0050] Figure 3 Schematic structural diagram of an automatic archiving policy list according to an embodiment of the present invention;

[0051] Figure 4 Module schematic diagram of a data archiving device based on artificial intelligence according to an embodiment of the present invention;

[0052] Figure 5 Internal structural schematic diagram of an electronic device for implementing a data archiving method based on artificial intelligence provided by an embodiment of the present invention;

[0053] The realization, functional features and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the accompanying drawings. Detailed implementation manners

[0054] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0055] To solve the problems existing in the existing data archiving, such as the inability to archive the stock data, and when there is a table type that does not support archiving, the full-library automatic archiving cannot be realized; in addition, there is a lack of performance monitoring, and the automatic table addition or deletion operation cannot be realized, affecting the archiving efficiency, etc., the present invention provides a data archiving method based on artificial intelligence, which can perform data archiving on the full library based on an automatic archiving rule list and ignore the discarded data, with flexible operation and high efficiency.

[0056] The embodiments of the present invention can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, sense the environment, acquire knowledge and use the knowledge to obtain the best results of theory, method, technology and application system.

[0057] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0058] The present invention provides a data archiving method based on artificial intelligence. Refer to Figure 1 As shown in

[0059] In this embodiment, the data archiving method based on artificial intelligence includes:

[0060] S100: Initialize the data archiving system, and load archiving information from a preset configuration source into the data archiving system.

[0061] Among them, the data archiving system further includes an initialization module, a timing module, a control module, a working thread, and a statistics module; when initializing the data archiving system, archiving information such as the concurrency number, raw library information, and archiving policy list can be loaded from a preset configuration file or other preset configuration sources for subsequent automatic data archiving processing.

[0062] As a specific example, Figure 2 shows a schematic principle block diagram of a data archiving system according to an embodiment of the present invention.

[0063] As Figure 2 shown, the influxdb system includes various types of databases, that is, databases such as the source library and the target library are all in the influxdb system. The raw library mainly refers to the most primitive database, and the data in it has not been archived. However, the data in other databases has been archived based on the raw library or other libraries. Among them, the initialization module is used to initialize the data archiving system; the control module is used to create a working thread according to the archiving request and perform data archiving processing; the timing module is used to periodically judge the state of the control module so as to send an archiving request to it when the control module is in an idle state; the statistics module is used to obtain the statistical data of archiving and calculate corresponding data such as the total elapsed time, total write volume, minimum and maximum elapsed time, etc.

[0064] Specifically, the data structure of each archiving policy can include the following parameters:

[0065] 1. Source library, which is used to represent the data source for archiving;

[0066] 2. Target library, which is used to represent the location where the archived data is stored;

[0067] 3. Archiving period round, which is used for the time period of data aggregation;

[0068] 4. Aggregation function list, including average value, maximum value, P95, etc.;

[0069] 5. Delay time, which is used to represent the delay time of the data set;

[0070] 6. Initial archiving time;

[0071] 7. Number of failure retries.

[0072] In the above archiving strategy, the initial archiving time is mainly used to calculate the progress of data archiving, and the calculation of the progress of data archiving may include: first query the archiving progress of this archiving strategy. If the progress is empty, it means that this archiving strategy has not been run, and then further judge the initial archiving time; if the initial archiving time is also empty, it means that the stored data will not be archived, and then set the archiving progress to the current time, otherwise set the archiving progress to the initial archiving time.

[0073] S200: Preprocess the archiving information to determine the automatic archiving rule list of the archiving information.

[0074] Among them, the preprocessing of the archiving information to determine the automatic archiving rule list includes:

[0075] S210: Summarize the archiving strategies in the archiving information to form an initial list;

[0076] S220: Sort the archiving strategies in the initial list according to preset rules to determine the automatic archiving rule list; among them,

[0077] The preset rules include: arranging the archiving strategies whose source library is the original library in the front, and the remaining archiving strategies satisfy the criterion that the target library of the previous archiving strategy of the current archiving strategy is the source library of the next archiving strategy of the current archiving strategy. Sort all the archiving strategies based on the above preset rules.

[0078] In addition, if there is an archiving strategy whose source library does not belong to the original library and is not the target library of the other archiving strategies, then mark this archiving strategy as an invalid strategy. The invalid archiving strategy will be ignored in the subsequent archiving process. Furthermore, based on the sorted archiving strategies and the marked invalid strategies, determine the automatic archiving rule list.

[0079] As a specific example, Figure 3 shows the schematic structure of the automatic archiving strategy list according to an embodiment of the present invention.

[0080] As Figure 3As shown, DB1 is the original database. The archiving policy list contains 4 archiving policies. The meaning of archiving policy P1 is that for the data in the source database DB1, the average value and maximum value are calculated for every 5-minute time range, aggregated into one point, and saved to the target database DB2. The delay time refers to that the current time should be greater than the rightmost point of the time range for evaluation + the delay time. For example, the delay time in P1 is 10 minutes. For the time range [2020-06-01 12:00:00, 2020-06-01 12:05:00), the current time should be greater than 2020-06-01 12:15:00 to perform data aggregation for this time range. Regarding the sorting of the archiving policy list, since the original database must have data, the policies with the original database as the source are ranked first and archived preferentially. In the above example, policies P1 and P2 depend on DB1, P3 depends on DB2, and P4 depends on DB4, so P1 or P2 is archived preferentially.

[0081] S300: Receive an archiving request to archive the data in the automatic archiving rule list based on the idle state information of the data archiving system; wherein, the state information is collected and obtained based on a preset frequency.

[0082] Among them, the state of the control module of the data archiving system is judged based on a preset frequency, and when the state of the control module meets the archiving requirements, an archiving request is sent to the control module. The preset frequency can be flexibly set according to the business scenario or requirements. For example, the state of the control module is judged every minute regularly. If the control module is in an idle state, an archiving request can be further sent to the control module; otherwise, no archiving request is sent.

[0083] S400: Traverse the automatic archiving rule list based on the archiving request and perform archiving processing on the data in the automatic archiving rule list.

[0084] Specifically, the database to be processed can be uniformly set in influxDB, that is, this influxDB contains the original database and multiple other databases. The data in the original database is the most original data, such as the monitored video data obtained from the server, which can also be understood as the data that has not been archived. And other databases can be the databases after archiving the data based on the original database, or the databases after archiving other databases.

[0085] Since the data in the original database does not necessarily meet the archiving requirements, after the control module receives the archiving request and before traversing the automatic archiving rule list, it also includes a process of screening the tables in the original database. This process can further include:

[0086] S410: When obtaining all data tables in the original library at the current time T, form a table set MS;

[0087] S420: Traverse the table set MS, and obtain the latest value corresponding to each data table based on the last function;

[0088] S430: Judge the type of each latest value, and filter out the data tables that do not meet the archiving conditions based on the type. Among them, if the type of the latest value of the data table is not a numeric type, it indicates that the data table corresponding to this latest value does not meet the archiving conditions.

[0089] Specifically, since among all the databases in influxDB, only the data in the original library has not been archived, therefore, before archiving, it is necessary to screen the data in the original library, while other databases are all based on the data archived from the original library or other databases. Therefore, data screening operations can be not performed on such databases, and the data in them can be defaulted to meet the archiving requirements.

[0090] Furthermore, by using the last function to obtain the latest value of each table in the original library, if the type of this latest value is not a numeric type, it indicates that the corresponding table does not meet the archiving requirements, and such tables can be directly filtered to prevent them from affecting subsequent archiving operations.

[0091] Finally, during the process of archiving the data to be processed, the control module sequentially traverses the automatic archiving rule list in a serial manner to archive a single archiving policy.

[0092] In a specific embodiment of the present invention, the process of archiving the data in the automatic archiving rule list includes:

[0093] S440: Obtain the archiving progress of the current archiving policy. If the archiving progress is empty, set the archiving progress to the initial archiving time of the current archiving policy.

[0094] S450: Judge whether the source library of the current archiving policy is the original library. If not, continue to find the second archiving policy, and set the target library of the second archiving policy to the source library of the current archiving policy.

[0095] Among them, if the source library of the current archiving policy is not the original library, it indicates that the data in the source library is formed by other archiving policies. Therefore, it is necessary to find the corresponding second archiving policy and judge the progress of the second archiving policy to judge whether the source library data of the second archiving policy can perform the archiving operation of the current archiving policy.

[0096] In addition, if the source library of the current archiving policy is the original library, then perform the following step S460; if the second archiving policy does not exist, then end the current process; to ensure that the source library of the current archiving policy is archived first, so as to ensure the smooth progress of this archiving.

[0097] S460: Obtain the archiving progress of the second archiving policy, and determine whether the archiving progress of the second archiving policy is less than the current round of archiving time. If the archiving progress of the second archiving policy is less than the current round of archiving time, then end the current process; otherwise, perform the next step.

[0098] Among them, the current round of archiving time can be determined based on the last archiving time and the archiving period, and the current round of archiving time = the last archiving time + the archiving period.

[0099] S470: Judge the relationship between the current round of archiving time and the current time T. If the current time T < the current round of archiving time + the delay time, then end the current process; otherwise, perform an archiving operation on the data corresponding to the current archiving policy.

[0100] Among them, if the operation of the current archiving policy fails, then end this process; otherwise, record that the archiving progress of the current archiving policy = the current round of archiving time, and then repeat the above steps until all archiving policies are executed.

[0101] As an example, each round represents an archiving period, the initial archiving time is the start time of the first archiving period, and each archiving operation archives a period of time. If the current time > the current round of archiving time + the delay time, then archive the data within the time period [the last archiving time, the current round of archiving time).

[0102] Among them, the timing module and archiving progress in the data archiving system can ensure the timely archiving and abnormal recovery of the latest data. Under normal circumstances, for the newly configured automatic archiving, since there is no archiving progress, it will start from the initial archiving time and progress in sections of time to archive the data until the current time; furthermore, when the current time is greater than the archiving progress + the delay time again during the subsequent timed trigger archiving, archive the data within a new period of time; and the abnormal recovery mainly refers to performing abnormal recovery and continuing to process in the case of process termination or program exception.

[0103] Furthermore, the process of performing an archiving operation on the data corresponding to the current archiving policy includes:

[0104] S471: Create a task queue based on the control module, generate multiple tasks according to the table set MS, and place the multiple tasks into the task queue; wherein, the content of the task includes a data table, an archiving policy, the last archiving time, and the current round of archiving time.

[0105] S472: Create a statistics queue based on the control module, and create and start multiple worker threads according to the concurrency number in the archiving information.

[0106] S473: The worker threads obtain corresponding task information from the task queue, and perform data archiving in the order of the archiving policy based on the task information and the automatic archiving rule list.

[0107] In the above step S473, the process of the worker threads obtaining corresponding task information from the task queue and performing data archiving in the order of the archiving policy based on the task information and the automatic archiving rule list may include:

[0108] S4731: Based on the task information, preprocess the data table in the source library corresponding to the current archiving policy to obtain corresponding data points.

[0109] Among them, the process of obtaining data points may include: aggregating the monitoring items in the source library within the time period [last archiving time, current round of archiving time) into one point through an aggregation function, that is, forming corresponding data points.

[0110] S4732: Save the data points to the target library corresponding to the source library to complete the archiving operation of the data table in the current source library.

[0111] S4733: Repeat the above steps in the order of the archiving policy until all data tables are archived.

[0112] Specifically, if the archiving is successful, record the table name, time consumption, and write time amount of the current archiving, and place them into the statistics queue, then re-obtain tasks and execute; until the task queue is empty, send a completion signal to the control module and end the thread. Otherwise, if the archiving fails, judge the error type of the failure: a. If the error type is a data type error, ignore the current failure and re-obtain a new task; wherein, a data type error means that the data type reported by the user belongs to a non-numeric type and conflicts with the original data type, which will cause statistical anomalies; b. If it is a timeout error, execute the aggregation function separately. If it times out again, execute the monitoring items separately; if the execution is successful, continue with the archiving, otherwise send an exception message to the control module and end the thread. C. For other errors, send an exception message to the control module and end the thread.

[0113] Specifically, after a worker thread obtains a task, assuming the task is to calculate the average, maximum, minimum, and P95 for the disk write speed table in database DB1 within the time range [2020-06-01 12:00:00, 2020-06-01 12:05:00), the monitoring items here may be the sda disk of physical machine A, the sdb disk of physical machine A, the sdd disk of physical machine C, etc. Assuming the source library has one data point per minute, within this time range, each monitoring item has 5 data points, and aggregating to calculate 1 data point means calculating the average, maximum, minimum, and P95 within these 5 minutes. More specifically, by constructing a statement and using the source library, table, aggregation function, etc. as parameters, it is sent to influxdb for execution.

[0114] When a timeout error occurs, the aggregation functions can be executed separately, which means calculating the average, maximum, minimum, and P95 separately. For example, in the previous example where we need to calculate the average, maximum, minimum, and P95, the execution statement needs to be divided into 4 sentences for separate calculations; correspondingly, executing the monitoring items separately means generating and executing statements for each monitoring item in this table.

[0115] Furthermore, when the control module receives an exception message, all threads will end the archiving process and record the failure information simultaneously; if the control module receives the completion signals from all threads, it indicates that the current round of archiving is successful, updates the archiving progress of the current archiving policy, and notifies the statistics module to perform statistical summarization. The statistics module can pull statistical data from the statistical queue and calculate data such as the total elapsed time, total write volume, minimum elapsed time, maximum elapsed time, etc.

[0116] The data automatic archiving method for artificial intelligence provided by the present invention is based on influxdb and archives the entire database in the form of a policy; it can include the new monitoring metrics and monitoring items reported by users into the archiving scope, and ignore the obsolete monitoring metrics and monitoring items without having to reset the rules again; it supports secondary archiving based on the archived data and can adapt to the progress of the source library; it supports archiving of existing data, control of the number of concurrent operations, and records the operation details and statistical data of each round of archiving, providing a relatively complete global and table-level self-monitoring.

[0117] As Figure 4 shown, it is the functional module diagram of the data archiving device for artificial intelligence based on the present invention.

[0118] The AI-based data archiving device 100 according to the present invention can be installed in an electronic device. According to the functions achieved, the AI-based data archiving device may include an initialization unit 101, an automatic archiving rule list determination unit 102, an archiving request sending unit 103, and a data archiving unit 104. The units described in the present invention may also be referred to as modules, which mainly refer to a series of computer program segments that can be executed by an electronic device processor and can complete fixed functions, and are stored in the memory of the electronic device.

[0119] In this embodiment, the functions of each module / unit are as follows:

[0120] The initialization unit 101 is used to initialize the data archiving system and load archiving information from a preset configuration source into the data archiving system.

[0121] Among them, the data archiving system further includes an initialization module, a timing module, a control module, a working thread, and a statistics module; when initializing the data archiving system, archiving information such as the concurrency number, raw library information, and archiving policy list can be loaded from a preset configuration file or other preset configuration sources for subsequent automatic data archiving processing.

[0122] As a specific example, as Figure 2 shown, there are various types of databases in the influxdb system, that is, the source library, the target library, etc. are all in the influxdb system. The raw library mainly refers to the most primitive database, and the data in it has not been archived. However, the data in other databases has been archived based on the raw library or other libraries. Among them, the initialization module is used to initialize the data archiving system; the control module is used to create a working thread according to the archiving request and perform data archiving processing; the timing module is used to periodically judge the state of the control module so as to send an archiving request to it when the control module is in an idle state; the statistics module is used to obtain the statistical data of archiving and calculate corresponding data such as the total elapsed time, total write volume, minimum and maximum elapsed time, etc.

[0123] Specifically, the data structure of each archiving policy may include the following parameters:

[0124] 1. Source library, which is used to represent the data source for archiving;

[0125] 2. Target library, which is used to represent the location where the archived data is stored;

[0126] 3. Archiving period round, which is used for the time period of data aggregation;

[0127] 4. Aggregation function list, including average value, maximum value, P95, etc.;

[0128] 5. Delay time, which is used to represent the delay time of the data set;

[0129] 6. Initial archiving time;

[0130] 7. Number of failure retries.

[0131] In the above archiving strategy, the initial archiving time is mainly used to calculate the progress of data archiving, and the calculation of the progress of data archiving may include: first query the archiving progress of this archiving strategy. If the progress is empty, it means that this archiving strategy has not been run, and then further judge the initial archiving time; if the initial archiving time is also empty, it means that the stored data is not archived, and then set the archiving progress to the current time, otherwise set the archiving progress to the initial archiving time.

[0132] The automatic archiving rule list determination unit 102 is used to preprocess the archiving information to determine the automatic archiving rule list of the archiving information.

[0133] Among them, this unit may further include the following modules:

[0134] The initial list formation module is used to summarize the archiving strategies in the archiving information to form an initial list;

[0135] The automatic archiving rule list determination module is used to sort the archiving strategies in the initial list according to preset rules to determine the automatic archiving rule list; among them,

[0136] The preset rules include: arranging the archiving strategies whose source library is the original library in the front, and the remaining archiving strategies satisfy the criterion that the target library of the previous archiving strategy of the current archiving strategy is the source library of the next archiving strategy of the current archiving strategy. Sort all the archiving strategies based on the above preset rules.

[0137] In addition, if the source library of an archiving strategy does not belong to the original library and is not the target library of the remaining archiving strategies, then mark this archiving strategy as an invalid strategy. The invalid archiving strategy will be ignored in the subsequent archiving process. Furthermore, based on the sorted archiving strategies and the marked invalid strategies, determine the automatic archiving rule list.

[0138] As a specific example, such as Figure 3As shown in the figure, DB1 is the original database. The archive policy list contains 4 archive policies. The meaning of archive policy P1 is that for the data in the source database DB1, the average value and maximum value are calculated for every 5-minute time range, and one point is aggregated and saved to the target database DB2. The delay time refers to that the current time should be greater than the rightmost point of the time range for value calculation + the delay time. For example, the delay time in P1 is 10 minutes. For the time range [2020-06-01 12:00:00, 2020-06-01 12:05:00), the current time should be greater than 2020-06-01 12:15:00 to perform data aggregation for this time range. Regarding the sorting of the archive policy list, since the original database must have data, the policies with the original database as the source are arranged in the front and archived preferentially. In the above example, policies P1 and P2 depend on DB1, P3 depends on DB2, and P4 depends on DB4, so P1 or P2 is archived preferentially.

[0139] The archive request sending unit 103 is configured to receive an archive request for archiving the data in the automatic archive rule list based on the idle state information of the data archiving system; wherein, the state information is acquired by collecting based on a preset frequency.

[0140] Among them, the state of the control module of the data archiving system is judged based on a preset frequency, and when the state of the control module meets the archive requirements, an archive request is sent to the control module. The preset frequency can be flexibly set according to the business scenario or requirements. For example, the state of the control module is judged every minute at a fixed time. If the control module is in an idle state, an archive request can be further sent to the control module; otherwise, no archive request is sent.

[0141] The data archiving unit 104 is configured to traverse the automatic archive rule list based on the archive request and perform archive processing on the data in the automatic archive rule list.

[0142] Specifically, the database to be processed can be uniformly set in influxDB, that is, the influxDB contains the original database and multiple other databases. The data in the original database is the most original data, such as the monitored video data obtained from the server, which can also be understood as the data that has not been archived. And other databases can be the databases after archiving the data based on the original database, or the databases after archiving other databases.

[0143] Since not all the data in the original database necessarily meets the archive requirements, after the control module receives the archive request and before traversing the automatic archive rule list, it further includes a process of screening the tables in the original database. This process can further include:

[0144] A table set formation module, configured to obtain all data tables in the original library at the current time T and form a table set MS;

[0145] A latest value acquisition module, configured to traverse the table set MS and obtain the latest value corresponding to each data table based on the last function;

[0146] A filtering module, configured to determine the type of each latest value and filter the data tables that do not meet the archiving conditions based on the type. Among them, if the type of the latest value of the data table is not a numeric type, it indicates that the data table corresponding to the latest value does not meet the archiving conditions.

[0147] Specifically, since in each database of influxDB, only the data in the original library has not been archived, therefore, before archiving, the data in the original library needs to be screened. And other databases are all based on the data archived from the original library or other databases. Therefore, data screening operations can be not performed on such databases, and the data in them can be defaulted to meet the archiving requirements.

[0148] Furthermore, the latest value of each table in the original library is obtained through the last function. If the type of the latest value is not a numeric type, it indicates that the corresponding table does not meet the archiving requirements, and such tables can be directly filtered to prevent them from affecting subsequent archiving operations.

[0149] Finally, during the process of archiving the data to be processed, the control module sequentially traverses the automatic archiving rule list in a serial manner and archives a single archiving policy.

[0150] In a specific embodiment of the present invention, the process of archiving the data in the automatic archiving rule list includes:

[0151] An archiving progress acquisition module, configured to obtain the archiving progress of the current archiving policy. If the archiving progress is empty, set the archiving progress to the initial archiving time of the current archiving policy.

[0152] A judgment module, configured to judge whether the source library of the current archiving policy is the original library. If not, continue to find the second archiving policy and set the target library of the second archiving policy to the source library of the current archiving policy.

[0153] Among them, if the source library of the current archiving policy is not the original library, it indicates that the data in the source library is formed by other archiving policies. Therefore, it is necessary to find the corresponding second archiving policy and judge the progress of the second archiving policy to determine whether the source library data of the second archiving policy can perform the archiving operation of the current archiving policy.

[0154] In addition, if the source library of the current archiving policy is the original library, then perform the following step S460; if the second archiving policy does not exist, then end the current process; to ensure that the source library of the current archiving policy is archived first, so as to ensure the smooth progress of this archiving.

[0155] An archiving progress acquisition module, configured to acquire the archiving progress of the second archiving policy, and determine whether the archiving progress of the second archiving policy is less than the current round of archiving time. If the archiving progress of the second archiving policy is less than the current round of archiving time, then end the current process; otherwise, perform the next step.

[0156] Wherein, the current round of archiving time can be determined based on the last archiving time and the archiving period, and the current round of archiving time = the last archiving time + the archiving period.

[0157] An archiving module, configured to judge the relationship between the current round of archiving time and the current time T. If the current time T < the current round of archiving time + the delay time, then end the current process; otherwise, perform one round of archiving operation on the data corresponding to the current archiving policy.

[0158] Wherein, if the operation of the current archiving policy fails, then end this process; otherwise, record that the archiving progress of the current archiving policy = the current round of archiving time, and then repeat the above steps until all archiving policies are executed.

[0159] As an example, each round represents an archiving period, the initial archiving time is the start time of the first archiving period, and each archiving operation archives a period of time. If the current time > the current round of archiving time + the delay time, then archive the data within the time period [the last archiving time, the current round of archiving time).

[0160] Wherein, the latest data can be archived in a timely manner and abnormal recovery can be ensured through the timing module and archiving progress in the data archiving system. Under normal circumstances, for the newly configured automatic archiving, since there is no archiving progress, it will start from the initial archiving time and progress in segments of time to archive the data until the current time; furthermore, when the current time is greater than the archiving progress + the delay time again during subsequent timed trigger archiving, archive the data within a new period of time; and the abnormal recovery mainly refers to performing abnormal recovery and continuing to process in the case of process termination or program exception.

[0161] Furthermore, the above-mentioned archiving module further includes:

[0162] A task queue creation module, configured to create a task queue based on the control module, generate multiple tasks according to the table set MS, and put the multiple tasks into the task queue; wherein, the content of the task includes a data table, an archiving policy, the last archiving time, and the current round of archiving time.

[0163] A working thread creation module, configured to create a statistical queue based on the control module, and create and start a plurality of working threads according to the number of concurrency in the archiving information;

[0164] A data archiving module, configured to obtain corresponding task information from the task queue through the working thread, and perform data archiving in the order of the archiving strategy based on the task information and the automatic archiving rule list.

[0165] In the above data archiving module, the process of obtaining corresponding task information from the task queue through the working thread, and performing data archiving in the order of the archiving strategy based on the task information and the automatic archiving rule list may include:

[0166] First, based on the task information, preprocess the data table in the source library corresponding to the current archiving strategy to obtain corresponding data points;

[0167] Among them, the process of obtaining data points may include: aggregating the monitoring items in the source library within the time period [last archiving time, current round of archiving time) into one point through an aggregation function, that is, forming corresponding data points.

[0168] Secondly, save the data points to the target library corresponding to the source library to complete the archiving operation of the data table in the current source library;

[0169] Finally, in the order of the archiving strategy, repeat the above steps until all data tables are archived.

[0170] Specifically, if the archiving is successful, record the table name, time consumption, and write time amount of the current archiving, and put them into the statistical queue, then re-obtain tasks and execute; until the task queue is empty, send a completion signal to the control module and end the thread. Otherwise, if the archiving fails, judge the error type of the failure: a. If the error type is a data type error, ignore the current failure and re-obtain a new task; among them, a data type error means that the data type reported by the user belongs to a non-numeric type and conflicts with the original data type, which will cause statistical anomalies; b. If it is a timeout error, execute the aggregation function separately. If it times out again, execute the monitoring items separately; if the execution is successful, continue with the archiving, otherwise send an exception message to the control module and end the thread. C. For other errors, send an exception message to the control module and end the thread.

[0171] Specifically, after the worker thread obtains a task, assuming the task is to calculate the average, maximum, minimum, and P95 for the disk write speed table in database DB1 within the time range of [2020-06-01 12:00:00, 2020-06-01 12:05:00), the monitoring items here may be the sda disk of physical machine A, the sdb disk of physical machine A, the sdd disk of physical machine C, etc. Assuming the source database has one data point per minute, within this time range, each monitoring item has 5 data points, and aggregating to calculate 1 data point means calculating the average, maximum, minimum, and P95 within these 5 minutes. More specifically, by constructing a statement and using the source database, table, aggregation function, etc. as parameters, it is sent to influxdb for execution.

[0172] When a timeout error occurs, the aggregation functions can be executed separately, which means calculating the average, maximum, minimum, and P95 separately. For example, in the previous example where we need to calculate the average, maximum, minimum, and P95, the execution statement needs to be divided into 4 sentences for separate calculations; correspondingly, executing the monitoring items separately means generating and executing statements for each monitoring item in this table.

[0173] Furthermore, when the control module receives exception information, all threads end the archiving process and record the failure information simultaneously; if the control module receives the completion signals from all threads, it indicates that the current round of archiving is successful, updates the archiving progress of the current archiving policy, and notifies the statistics module to perform statistical summarization. The statistics module can pull statistical data from the statistical queue and calculate data such as total elapsed time, total write volume, minimum elapsed time, maximum elapsed time, etc.

[0174] As Figure 5 shown, it is a schematic structural diagram of an electronic device for implementing the data archiving method based on artificial intelligence according to the present invention.

[0175] The electronic device 1 may include a processor 10, a memory 11, and a bus, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as the data archiving program 12 based on artificial intelligence.

[0176] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device 1, such as the mobile hard disk of the electronic device 1. In some other embodiments, the memory 11 may also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 1. Further, the memory 11 may also include both an internal storage unit and an external storage device of the electronic device 1. The memory 11 can be used not only to store application software installed in the electronic device 1 and various types of data, such as the code of an artificial intelligence-based data archiving program, etc., but also to temporarily store data that has been output or will be output.

[0177] In some embodiments, the processor 10 may be composed of integrated circuits. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including a combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting all components of the entire electronic device through various interfaces and lines, and by running or executing programs or modules stored in the memory 11 (such as an artificial intelligence-based data archiving program, etc.), and calling data stored in the memory 11, to perform various functions of the electronic device 1 and process data.

[0178] The bus may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is set to enable connection and communication between the memory 11 and at least one processor 10, etc.

[0179] Figure 5 Only an electronic device with components is shown. Those skilled in the art can understand that Figure 5The structures shown do not constitute a limitation on the electronic device 1, and may include fewer or more components than those shown, or combine certain components, or have different component arrangements.

[0180] For example, although not shown, the electronic device 1 may further include a power source (such as a battery) for supplying power to each component. Preferably, the power source may be logically connected to the at least one processor 10 through a power management device, so as to implement functions such as charging management, discharging management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, a power status indicator, etc. The electronic device 1 may also include a variety of sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0181] Furthermore, the electronic device 1 may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device 1 and other electronic devices.

[0182] Optionally, the electronic device 1 may further include a user interface. The user interface may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device 1 and to display a visual user interface.

[0183] It should be understood that the embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.

[0184] The artificial intelligence-based data archiving program 12 stored in the memory 11 of the electronic device 1 is a combination of multiple instructions, and when running in the processor 10, can achieve:

[0185] Initialize the data archiving system and load archiving information from a preset configuration source into the data archiving system;

[0186] Preprocess the archiving information to determine a list of automatic archiving rules for the archiving information;

[0187] Receive an archiving request to archive the data in the automatic archiving rule list based on the idle state information of the data archiving system; wherein, the state information is collected and obtained based on a preset frequency.

[0188] Traverse the automatic archiving rule list based on the archiving request, and perform archiving processing on the data in the automatic archiving rule list.

[0189] Furthermore, if the modules / units integrated in the electronic device 1 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory).

[0190] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0191] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0192] In addition, in each embodiment of the present invention, the various functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware, or in the form of hardware plus software functional modules.

[0193] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-mentioned exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.

[0194] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to cover all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention. Any associated drawing marks in the claims should not be regarded as limiting the claimed rights.

[0195] In addition, it is obvious that the term "comprising" does not exclude other units or steps, and the singular does not exclude the plural. A plurality of units or devices stated in the system claims can also be implemented by one unit or device through software or hardware. The terms such as second are used to denote names and do not denote any particular order.

[0196] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A data archiving method based on artificial intelligence, characterized in that, The method includes: Initializing a data archiving system and loading archiving information from a preset configuration source into the data archiving system; the archiving information includes the concurrency number, original library information, and an archiving policy list, and the archiving policy list includes archiving policies; Preprocessing the archiving information to determine an automatic archiving rule list of the archiving information; Receiving an archiving request for archiving data in the automatic archiving rule list based on the idle state information of the data archiving system; wherein, the state information is collected and obtained based on a preset frequency; Traversing the automatic archiving rule list based on the archiving request and performing archiving processing on the data in the automatic archiving rule list; wherein, Obtaining the archiving progress of the current archiving policy, and if the archiving progress is empty, setting the archiving progress to the initial archiving time of the current archiving policy; Judging whether the source library of the current archiving policy is the original library, if not, continuing to find a second archiving policy and making the target library of the second archiving policy be the source library of the current archiving policy; Obtaining the archiving progress of the second archiving policy and judging whether the archiving progress of the second archiving policy is less than the current round of archiving time. If the archiving progress of the second archiving policy is less than the current round of archiving time, ending the current process, otherwise, performing the next step; Judging the relationship between the current round of archiving time and the current time T. If the current time T < the current round of archiving time + delay time, ending the current process; otherwise, performing a round of archiving operation on the data corresponding to the current archiving policy.

2. The data archiving method based on artificial intelligence according to claim 1, wherein The archiving policies in the archiving policy list include: A source library, which is used to represent the data source for archiving; A target library, which is used to represent the location where the archived data is stored; An archiving period round, which is used to represent the time period for data aggregation; A list of aggregation functions, including average value, maximum value, and P95; A delay time delay, which is used to represent the delay time of the data set; The initial archiving time, and the number of failure retries.

3. The data archiving method based on artificial intelligence according to claim 2, wherein The preprocessing of the archiving information to determine the automatic archiving rule list of the archiving information includes: Summarizing the archiving policies in the archiving information in the data archiving system to form an initial list; Sorting the archiving policies in the initial list according to a preset rule to determine the automatic archiving rule list of the archiving information; wherein, The preset rule includes: arranging the archiving policies with the source library being the original library in the front, and the remaining archiving policies satisfying the criterion that the target library of the previous archiving policy of the current archiving policy is the source library of the next archiving policy of the current archiving policy.

4. The data archiving method based on artificial intelligence according to claim 3, wherein, After the archiving request is sent and before traversing the automatic archiving rule list, it further includes a process of screening the data tables in the original library, and the screening process includes: Obtaining all the data tables in the original library at the current time T and forming a table set MS; Traversing the table set MS and obtaining the latest value corresponding to each data table based on the last function; Determine the types of each latest value, and filter the data tables that do not meet the archiving conditions based on the types; Among them, if the type of the latest value of the data table is not a numeric type, it indicates that the data table corresponding to the latest value does not meet the archiving conditions.

5. The data archiving method based on artificial intelligence according to claim 4, wherein, The process of performing a round of archiving operations on the data corresponding to the current archiving policy includes: Create a task queue based on the control module, generate multiple tasks according to the table set MS, and put the multiple tasks into the task queue; wherein, the content of the task includes the data table, the archiving policy, the last archiving time, and the current round of archiving time; Create a statistical queue based on the control module, and create and start multiple worker threads according to the concurrency number in the archiving information; The worker threads obtain the corresponding task information from the task queue, and perform data archiving in the order of the archiving policy based on the task information and the automatic archiving rule list.

6. The data archiving method based on artificial intelligence according to claim 5, wherein The process of the worker threads obtaining the corresponding task information from the task queue and performing data archiving in the order of the archiving policy based on the task information and the automatic archiving rule list includes: Based on the task information, preprocess the data table in the source library corresponding to the current archiving policy to obtain the corresponding data points; Save the data points to the target library corresponding to the source library to complete the archiving operation of the data table in the current source library; Repeat the above steps in the order of the archiving policy until all data tables are archived.

7. An artificial intelligence-based data archiving device, characterized in that, The device includes: An initialization unit for initializing the data archiving system and loading the archiving information from a preset configuration source into the data archiving system; the archiving information includes the concurrency number, the original library information, and the archiving policy list, and the archiving policy list includes archiving policies; An automatic archiving rule list determination unit for preprocessing the archiving information to determine the automatic archiving rule list of the archiving information; An archiving request sending unit for receiving an archiving request to archive the data in the automatic archiving rule list based on the idle state information of the data archiving system; wherein, the state information is collected and obtained based on a preset frequency; A data archiving unit for traversing the automatic archiving rule list based on the archiving request and performing archiving processing on the data in the automatic archiving rule list; wherein, Obtain the archiving progress of the current archiving policy, and if the archiving progress is empty, set the archiving progress to the initial archiving time of the current archiving policy; Judge whether the source library of the current archiving policy is the original library, if not, continue to find the second archiving policy, and make the target library of the second archiving policy the source library of the current archiving policy; Obtain the archiving progress of the second archiving policy, and judge whether the archiving progress of the second archiving policy is less than the current round of archiving time. If the archiving progress of the second archiving policy is less than the current round of archiving time, end the current process, otherwise, execute the next step; Determine the relationship between the current archiving time and the current time T. If the current time T < the current archiving time + the delay time, end the current process; otherwise, perform one round of archiving operation on the data corresponding to the current archiving policy.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the steps in the artificial intelligence-based data archiving method according to any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps in the artificial intelligence-based data archiving method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Data archiving method, server and computer readable storage medium

    CN110457255A

  • Data archiving processing method and device, computer equipment and storage medium

    CN112181945A