Data table cleaning method, system and equipment based on deep Q network and medium

Through the deep Q network-based method, the agent is constructed and DQN algorithm training is carried out, and the problem of lack of planning and judgment of data table cleaning is solved, efficient and accurate data table cleaning is achieved, and database performance and data availability are ensured.

CN120011355APending Publication Date: 2025-05-16CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510185782.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, data table cleaning lacks clear planning or judgment, resulting in inaccurate manual cleaning and error deletion, affecting database performance and data availability.

Method used

Using a deep Q network method, by collecting feature information of the data table, building an agent and training DQN algorithm, we learn how to select optimized cleaning actions based on the state of the database table, thereby maximizing cumulative rewards.

Benefits of technology

It realizes the importance of accurately predicting data tables during the data table cleaning process, ensures that key data is retained and redundant information is effectively removed, improves cleaning efficiency and accuracy, and ensures the accuracy and efficiency of database cleaning work.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011355A_ABST
    Figure CN120011355A_ABST
Patent Text Reader

Abstract

The invention provides a data table cleaning method and system based on a deep Q network, electronic equipment and a storage medium, and aims to solve the problems that table cleaning lacks planning or judgment, is inaccurate and is mistakenly deleted. The method comprises the following steps: collecting stock data feature information of related data tables in a database; intelligent agent parameters are determined according to the collected stock data feature information, an intelligent agent is newly built, and the intelligent agent parameters comprise a state space describing the state of the data table, an action space of all possible actions capable of being executed by the intelligent agent and a reward function for obtaining a reward after the intelligent agent executes a certain action; dQN algorithm training is carried out on the newly-built agent to obtain a trained agent, so that the agent learns how to select an optimized cleaning action according to a database table state based on agent parameters, and thus accumulated rewards are maximized; and cleaning the data table through the trained intelligent agent. The efficiency and accuracy in the data table cleaning process can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a data table cleaning method based on a deep Q network, a data table cleaning system based on a deep Q network, an electronic device, and a computer-readable storage medium. Background Art

[0002] With the rapid development of information technology, the popularization of the Internet and the Internet of Things, data-driven business models are developing rapidly, and the speed of data generation is also increasing at an unprecedented rate. By analyzing and mining these massive amounts of data, enterprises can make more accurate and effective business decisions in various fields. However, as the amount of data surges, the data will also be mixed with a large amount of repeated, redundant and useless junk data. When the junk data accumulates to a certain extent, it will affect the overall performance of the database. Then, cleaning up specific data tables becomes a major challenge for data maintainers. In order to meet this challenge, many companies currently set early warnings based on the utilization rate of data storage space, set regular data cleaning plans, database load analysis and data life cycle management, etc., to clean up database tables and reduce database load.

[0003] Most of the time, the existing methods for cleaning data tables rely on manual cleaning of data, and lack clear planning or judgment for table cleaning. Such manual cleaning often triggers a series of chain reactions, posing a potential threat to the operational efficiency and data security of enterprises. Without a clear table cleaning plan, storage space may be continuously occupied by a large amount of redundant and invalid data, which not only increases unnecessary hardware equipment costs, but also may affect the overall performance of the database due to tight storage space, such as slow query speed tables, ETL (Extract-Transform-Load) processing errors, etc., which interfere with the company's accurate decision-making process. More seriously, if the early warning process fails, long-term failure to clean the database may lead to increased data inconsistency, affecting the normal operation of the business and data availability. In addition, enterprises manually clear data table space based on monitoring database indicators. Due to the large number of database indicators that need to be monitored and the difficulty in grasping the importance of indicators, there may be problems such as inaccurate table cleaning and accidental deletion. Summary of the invention

[0004] In order to at least solve the problems of manual data table cleaning in the prior art, such as lack of clear planning or judgment for table cleaning, as well as inaccurate table cleaning and mistaken deletion, the present invention provides a data table cleaning method based on a deep Q network, a data table cleaning system based on a deep Q network, an electronic device and a computer-readable storage medium, which are continuously trained through reinforcement learning, and ultimately accurately predict the importance of each data table, thereby ensuring that key data can be retained during the cleaning process, while effectively removing redundant or no longer needed information, and improving the efficiency and accuracy of the data table cleaning process.

[0005] In a first aspect, the present disclosure provides a data table cleaning method based on a deep Q network, the method comprising:

[0006] Collect the characteristic information of stock data from relevant data tables in the database;

[0007] Determine agent parameters based on the collected stock data feature information and create a new agent. The agent parameters include a state space describing the state of the data table, an action space of all possible actions that the agent can perform, and a reward function for the agent to obtain a reward after performing a certain action.

[0008] The newly created agent is trained with the DQN (Deep Q-Network) algorithm to obtain a trained agent, so that the agent can learn how to select optimized cleaning actions according to the database table status based on the agent parameters, thereby maximizing the cumulative reward;

[0009] The data table is automatically cleaned through the trained intelligent agent.

[0010] Furthermore, the stock data characteristic information includes:

[0011] Structural information, including column name, data type, and whether there is an index;

[0012] Attribute information, including the size of the table and its relationships with other data tables.

[0013] Furthermore,

[0014] The state space includes the number of records in the table, creation time, access frequency, data table size, whether the belonging storage is invalid, daily growth, monthly growth, month-on-month growth, update time, and creator;

[0015] The action space includes a clearing table, a retaining table, and a migration table;

[0016] The reward function is designed based on the table size after the table cleanup action and factors affecting database performance improvement.

[0017] Furthermore, the reward function is calculated by the following formula:

[0018] R(s,a,s ‘ )=-m*Δdata integrity+o*I / O speed improvement+p*Δstorage space

[0019] Where R(s,a,s ‘ ) is the reward function value, Δdata integrity represents the change in the constraints on database integrity after table clearing, I / O (Input / Output) speed improvement represents the positive impact on database read and write performance after action a is executed, Δstorage space represents the change in storage space after data clearing, and m, o, and p are weight coefficients.

[0020] Furthermore, the DQN algorithm training of the newly created agent includes:

[0021] Initialization, initialize the following variables:

[0022] Q-Network: Initialize a deep neural network as Q-Network(s,a;θ), where θ is the parameter of the neural network,

[0023] Target network: Create a copy of the Q network as the target network, whose initial parameters θ' are consistent with those of the Q network;

[0024] Experience pool: Initialize an empty experience pool to store the experience data generated by the interaction between the agent and the environment;

[0025] The agent chooses actions, including:

[0026] In the early stages of training, the agent uses an ε-greedy strategy to explore the environment:

[0027] a = random action (with probability ε)

[0028] Where ε is the exploration rate;

[0029] After training to a certain stage, the agent chooses actions by relying on the Q network;

[0030] Perform actions, observe results, and store experiences, including:

[0031] The agent selects an action a in the database environment and observes the result, which contains the next state s of the database. ‘ , determine whether the reward r reaches the termination condition; the result in the current state (state s, action a, reward r, next state s ‘ ) is stored in the experience pool D;

[0032] Update the target network, including:

[0033] Extract a batch of experience data (s i ,a i ,r i ,s ‘ i ) as training samples, and use the target network Q for each training sample ′ (s ‘ i ,a ‘ θ ′ ) Calculate the target Q value y i :

[0034] y i =r i +γ*maxQ ′ (s ‘ i ,a ‘ θ ′ )

[0035] Where γ represents the discount factor, which is used to control the importance of future rewards, and then the mean square error function value L is used to measure the difference between the predicted Q value and the target Q value:

[0036]

[0037] Update the parameters in the Q network according to the back-propagation algorithm and gradient descent method to minimize the loss function;

[0038] Repeat training to get the best, including:

[0039] Repeat the training steps from agent selection to updating the target network until the preset number of training rounds is reached, or the performance of the agent and the loss function converge to a stable state, then the Q network can be used to select the optimal use action of the table.

[0040] Furthermore, the method further comprises:

[0041] Embed an alarm mechanism for the database. Before the agent cleans up the data table, it automatically sends the corresponding data table information to the corresponding auditor, so that the auditor can confirm the cleanup based on the data table information and return the confirmation information.

[0042] If the confirmation message is to confirm the cleanup, or if the confirmation message is not received within the preset time, the data table will be cleaned up;

[0043] If the confirmation information indicates that the data table is not to be cleaned, the cleanup field in the result table is changed to not perform any operation on the data table. The result table stores information of all data tables to be cleaned.

[0044] In a second aspect, the present disclosure further provides a data table cleaning system based on a deep Q network, the system comprising:

[0045] A collection module, which is configured to collect feature information of stock data from relevant data tables in a database;

[0046] A creation module is configured to determine agent parameters based on the collected stock data feature information and create a new agent, wherein the agent parameters include a state space describing the state of the data table, an action space of all possible actions that the agent can perform, and a reward function for the agent to obtain a reward after performing a certain action;

[0047] The training module is configured to perform DQN algorithm training on the newly created agent to obtain a trained agent, so that the agent can learn how to select optimized cleaning actions according to the database table status based on the agent parameters, thereby maximizing the cumulative reward;

[0048] A cleaning module is configured to automatically clean the data table through a trained agent.

[0049] Furthermore,

[0050] The state space includes the number of records in the table, creation time, access frequency, data table size, whether the belonging storage is invalid, daily growth, monthly growth, month-on-month growth, update time, and creator;

[0051] The action space includes a clearing table, a retaining table, and a migration table;

[0052] The reward function is designed based on the table size after the table cleanup action and the factors affecting database performance improvement, and is calculated using the following formula:

[0053] R(s,a,s ‘ )=-m*Δdata integrity+o*I / O speed improvement+p*Δstorage space

[0054] Where R(s,a,s ‘ ) is the reward function value, Δdata integrity represents the change in the constraints on database integrity after table clearing, I / O speed improvement represents the positive impact on database read and write performance after action a is executed, Δstorage space represents the change in storage space after data clearing, and m, o, and p are weight coefficients.

[0055] In a third aspect, the present disclosure provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor runs the computer program stored in the memory, the processor executes a data table cleaning method based on a deep Q network as described in any one of the first aspects.

[0056] In a fourth aspect, the present disclosure provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the data table cleaning method based on the deep Q network described in any one of the first aspects above is implemented.

[0057] Beneficial effects:

[0058] The data table cleaning method based on deep Q network, data table cleaning system based on deep Q network, electronic device and storage medium provided by the present disclosure; by training and analyzing the characteristic states of the data table in the database, the cleaning strategy of the data table in the database is optimized to the maximum extent, and the importance of each data table is accurately predicted through continuous training through reinforcement learning, so as to ensure that key data can be retained during the cleaning process, and at the same time, redundant or no longer needed information is effectively removed. The efficiency and accuracy of data engineers in the process of cleaning data tables are improved, and the accuracy and efficiency of database cleaning work are ensured. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 A flowchart of a data table cleaning method based on a deep Q network provided in the first embodiment of the present disclosure;

[0060] Figure 2 A schematic diagram of an overall process of data table cleaning provided by an embodiment of the present disclosure;

[0061] Figure 3 A schematic diagram of a process of training an agent using a DQN algorithm provided in an embodiment of the present disclosure;

[0062] Figure 4 An architecture diagram of a data table cleaning system based on a deep Q network provided in Embodiment 2 of the present disclosure;

[0063] Figure 5 This is an architecture diagram of an electronic device provided in Embodiment 3 of the present disclosure. DETAILED DESCRIPTION

[0064] In order to enable those skilled in the art to better understand the technical solution of the present disclosure, the present disclosure is further described in detail below in conjunction with the drawings and embodiments. It should be understood that the specific embodiments and drawings described herein are only used to explain the present invention, rather than to limit the present invention.

[0065] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence; and, in the absence of conflict, the embodiments in the present disclosure and the features in the embodiments can be arbitrarily combined with each other.

[0066] The terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure. The singular forms of "a", "said" and "the" used in the embodiments of the present disclosure and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings.

[0067] In the subsequent description, the suffixes such as "module", "component" or "unit" used to represent elements are only used to facilitate the description of the present disclosure, and have no specific meanings. Therefore, "module", "component" or "unit" can be used in a mixed manner.

[0068] The technical solution of the present invention and how the technical solution of the present invention solves the technical problems existing in the prior art are described in detail below with specific embodiments. It is understandable that in the embodiments of the present application, the execution subject can perform some or all of the steps in the embodiments of the present application, and these steps or operations are only examples. The embodiments of the present application can also perform other operations or deformations of various operations. In addition, each step can be performed in a different order presented in the embodiments of the present application, and it is possible not to perform all the operations in the embodiments of the present application. And, the following several specific embodiments can be combined with each other, and may not be repeated in some embodiments for the same or similar concepts or processes.

[0069] Figure 1 A flowchart of a data table cleaning method based on a deep Q network provided in the first embodiment of the present disclosure is shown in FIG. Figure 1 As shown, the method includes:

[0070] Step S101: Collecting feature information of stock data from relevant data tables in the database;

[0071] Step S102: Determine agent parameters based on the collected stock data feature information and create a new agent. The agent parameters include a state space describing the state of the data table, an action space of all possible actions that the agent can perform, and a reward function for the agent to obtain a reward after performing a certain action.

[0072] Step S103: Perform DQN algorithm training on the newly created agent to obtain a trained agent, so that the agent learns how to select optimized cleaning actions according to the database table status based on the agent parameters, thereby maximizing the accumulated reward;

[0073] Step S104: Automatically clean the data table through the trained intelligent agent.

[0074] The purpose of the disclosed embodiment is to propose a data table cleaning method, which uses reinforcement learning design and integrates user-defined reward functions to learn and optimize the exact indicators that the user wants, locate the data tables that should be cleaned up in the database, and design triggers to allow autonomous cleaning of data tables, so as to facilitate enterprises to manage complex and changeable databases.

[0075] To achieve this goal, Figure 2 As shown, the embodiment of the present disclosure specifically includes the following steps:

[0076] Collection of stock data characteristics;

[0077] Determine the agent parameters and create a new agent;

[0078] Deep Q network model training;

[0079] Automatically clear the table.

[0080] Specifically, when collecting features of existing data, in order to build and train the subsequent intelligent agent, it is first necessary to comprehensively and systematically collect feature information of all relevant data tables in the database, such as basic metadata, column-level information, keys and constraints, index information, storage and performance, etc. By carefully collecting and organizing this information, we can better understand the distribution status of the data tables, provide strong support for the subsequent intelligent agent design, and ensure the best performance in practical applications.

[0081] In order to achieve automatic and efficient cleaning of tables in the database, after obtaining the characteristic information of the data table, the characteristic information can be used to understand the state of the data table, all possible actions that can be performed in the data table, and the impact of the actions on the data table, so as to determine the parameters of the intelligent agent. The intelligent agent mainly includes three key characteristic parameters: state space, action space, and reward function. An intelligent agent system is built, and then the intelligent agent is trained using the DQN algorithm.

[0082] The DQN algorithm is an algorithm that combines deep learning and reinforcement learning to solve decision-making problems in high-dimensional state spaces. It is an extension of the Q-learning algorithm. It introduces a deep neural network to approximate the Q-value function, so that it can handle complex inputs. The DQN algorithm is based on Q-learning and aims to learn the optimal strategy through the interaction between the agent and the environment. After training the agent, the agent will learn how to select the best cleanup action based on the database table status based on three types of parameters to maximize the cumulative reward. The trained agent can use the Q network to select the optimal usage action for the data table.

[0083] The trained intelligent agent automatically executes the table clearing trigger to automatically clean up the data table.

[0084] The disclosed embodiment obtains data table features and constructs an intelligent agent, which is continuously trained through deep Q network reinforcement learning, and automatically selects the optimal cleaning action (such as filling, deleting, and correcting) according to data features (such as missing values, outliers, and duplicate data); the intelligent agent can process high-dimensional data (such as multi-column, multi-type data), extract complex features through deep neural networks, and adapt to structured data tables and unstructured data; and can dynamically adjust the cleaning strategy according to changes in data distribution. The cleaning effect is optimized through a reward mechanism, and the dependence on manual rules is reduced through autonomous learning, reducing labor costs; and it can handle complex and ambiguous cleaning tasks; DQN efficiently utilizes historical cleaning experience through an experience replay mechanism to accelerate the learning process. It can learn effective cleaning strategies from a small amount of labeled data. Finally, the importance of each data table is accurately predicted, thereby ensuring that key data can be retained during the cleaning process, while effectively removing redundant or no longer needed information, and improving the efficiency and accuracy of the data table cleaning process.

[0085] Furthermore, the stock data characteristic information includes:

[0086] Structural information, including column name, data type, and whether there is an index;

[0087] Attribute information, including the size of the table and its relationships with other data tables.

[0088] By querying the system table through SQL (Structured Query Language) or using database management tools, it is possible to collect the feature information of the stock data of each data table, including but not limited to the following information: structural information, such as: column name, data type, whether null values ​​are allowed, default values, comments, etc., whether there is an index, and the existing index information: index name, index type, columns included in the index, etc.; attribute information, such as table size (number of rows in the table, storage size), inter-table associations (primary key, foreign key, and associations with other data tables (such as whether there are foreign key constraints, etc.)), etc.; of course, it can also include other metadata, such as: table creation time, modification time, engine type, etc. Feature information is collected according to the actual situation of the data table to provide strong support for intelligent agent design.

[0089] Furthermore,

[0090] The state space includes the number of records in the table, creation time, access frequency, data table size, whether the belonging storage is invalid, daily growth, monthly growth, month-on-month growth, update time, and creator;

[0091] The action space includes a clearing table, a retaining table, and a migration table;

[0092] The reward function is designed based on the table size after the table cleanup action and factors affecting database performance improvement.

[0093] The key characteristic parameters of the agent include:

[0094] 1. State Space

[0095] The state space S mainly includes various characteristics of the data table, which is used to describe the state of the table. For example, the state space includes the number of records in the table, creation time, access frequency, data table size, whether the belonging storage is invalid, daily growth, monthly growth, month-on-month growth, update time, creator and other characteristic information; it can be expressed as:

[0096] S={s|s=[number of records, creation time...]}.

[0097] 2. Action Space

[0098] The action space mainly includes all possible actions that the agent can perform. In the process of data table cleaning, these action spaces include clearing tables, retention tables, migration tables (commonly used but with low update frequency), etc.; they can be expressed as:

[0099] A = {a|a = [clear table, keep table...]}.

[0100] 3. Reward Function

[0101] The reward function R defines the reward that the agent receives after performing an action. During the data table cleanup process, the reward function can be designed based on influencing factors such as the table size after cleanup and database performance improvement.

[0102] The reward function can be designed differently according to actual needs to dynamically adjust the cleaning strategy according to changes in data distribution. The cleaning effect can be optimized through reward function design (such as maximizing data quality, minimizing information loss, improving database performance, etc.).

[0103] Furthermore, the reward function is calculated by the following formula:

[0104] R(s,a,s ‘ )=-m*Δdata integrity+o*I / O speed improvement+p*Δstorage space

[0105] Where R(s,a,s ‘ ) is the reward function value, Δdata integrity represents the change in the constraints on database integrity after table clearing, I / O speed improvement represents the positive impact on database read and write performance after action a is executed, Δstorage space represents the change in storage space after data clearing, and m, o, and p are weight coefficients.

[0106] ΔData integrity can be measured by reducing the foreign key constraints of the table; I / O speed improvement can be reflected by query speed, and the importance of different factors can be balanced by setting corresponding weight coefficients m, o, p for each factor.

[0107] Furthermore, the DQN algorithm training of the newly created agent includes:

[0108] Initialization, initialize the following variables:

[0109] Q-Network: Initialize a deep neural network as Q-Network(s,a;θ), where θ is the parameter of the neural network,

[0110] Target network: Create a copy of the Q network as the target network, whose initial parameters θ' are consistent with those of the Q network;

[0111] Experience pool: Initialize an empty experience pool to store the experience data generated by the interaction between the agent and the environment;

[0112] The agent chooses actions, including:

[0113] In the early stages of training, the agent uses an ε-greedy strategy to explore the environment:

[0114] a = random action (with probability ε)

[0115] Where ε is the exploration rate;

[0116] After training to a certain stage, the agent chooses actions by relying on the Q network;

[0117] Perform actions, observe results, and store experiences, including:

[0118] The agent selects an action a in the database environment and observes the result, which contains the next state s of the database. ‘ , determine whether the reward r reaches the termination condition; the result in the current state (state s, action a, reward r, next state s ‘ ) is stored in the experience pool D;

[0119] Update the target network, including:

[0120] Extract a batch of experience data (s i ,a i ,r i ,s ‘ i ) as training samples, and use the target network Q for each training sample ′ (s ‘ i ,a ‘ θ′ ) Calculate the target Q value y i :

[0121] y i =r i +γ*maxQ ′ (s ‘ i ,a ‘ θ ′ )

[0122] Where γ represents the discount factor, which is used to control the importance of future rewards, i represents the i-th data, and then the mean square error function value L is used to measure the difference between the predicted Q value and the target Q value:

[0123]

[0124] Update the parameters in the Q network according to the back-propagation algorithm and gradient descent method to minimize the loss function;

[0125] Repeat training to get the best, including:

[0126] Repeat the training steps from agent selection to updating the target network until the preset number of training rounds is reached, or the performance of the agent and the loss function converge to a stable state, then the Q network can be used to select the optimal use action of the table.

[0127] The newly created agent is trained through the DQN algorithm, including environment initialization, hyperparameter setting, interaction and experience collection, sampling and training, target network update and exploration rate decay. Through repeated iterations, the agent gradually learns the optimal strategy and can eventually make efficient decisions in complex environments.

[0128] The training process, such as Figure 3 As shown, by determining the state space, action space and reward function, constructing, and training the agent with the DQN algorithm, by initializing the DQN algorithm, initializing the variables Q network, target network and experience pool, the agent selects actions, executes actions, observes results, stores experience, updates the target network, measures the difference between the predicted Q value and the target Q value, updates the parameters in the Q network, and selects actions again to minimize the loss function, thereby obtaining the optimal agent action for the data table.

[0129] During the initialization process, MLP (Multilayer Perceptron) can be selected as the initial deep neural network; after the target network is created, the parameters of the target network will not change for a period of time, which is used to stabilize the training process. When the agent selects an action, the exploration rate ε controls the balance between the agent's exploration and utilization (the initial value is high and gradually decays). As the training progresses, when it reaches a certain stage, if the ε value drops to a lower level (such as 0.01~0.1), the agent selects actions by relying on the Q network. The agent performs action a, and the environment returns the reward r and the next state s ‘ , and store the experience in the replay buffer D. In updating the target network, the discount factor γ weighs the importance of current rewards and future rewards (usually 0.9 to 0.99); after predicting the difference between the Q value and the target Q value, the gradient of the loss function to the network parameter θ is calculated through the back propagation algorithm to minimize the loss function.

[0130] Repeat the training steps until the preset number of training rounds is reached, or the performance of the agent and the loss function converge to a stable state, then use the Q network to select the optimal action for the table. The agent will input the state s of the current data table into the Q network, obtain the Q value after each action, and then select the action a with the optimal Q value. ‘ As the best action for the data table in its current state.

[0131] During the training process, by paying attention to the following parameters and steps:

[0132] Target Q value: calculated by the target network to ensure training stability.

[0133] Current Q value: calculated by the main network and used to compare with the target Q value.

[0134] Loss function: Use mean-square error (MSE) to measure the prediction error.

[0135] Backpropagation: Calculate the gradient of the loss function with respect to the network parameters.

[0136] Gradient descent: Use gradients to update network parameters and minimize the loss function.

[0137] DQN can gradually optimize the Q network parameters, allowing the agent to learn to make the best decision in a complex environment. It enables the agent to significantly improve the efficiency and effectiveness of data cleaning, and better complete the automated cleaning tasks of large-scale and complex data tables.

[0138] Furthermore, the method further comprises:

[0139] Embed an alarm mechanism for the database. Before the agent cleans up the data table, it automatically sends the corresponding data table information to the corresponding auditor, so that the auditor can confirm the cleanup based on the data table information and return the confirmation information.

[0140] If the confirmation message is to confirm the cleanup, or if the confirmation message is not received within the preset time, the data table will be cleaned up;

[0141] If the confirmation information indicates that the table data is not to be cleaned, the cleanup field in the result table is changed to not perform any operation on the table data. The result table stores information on all data tables to be cleaned.

[0142] When the intelligent agent is continuously optimized and iterated, and the performance tends to be stable, the database can automatically clean up invalid tables. However, the importance of some tables may not be observed from the table name, frequency of use, etc. This may mislead the intelligent agent and cause the table to be deleted by mistake. In order to improve this mechanism, an alarm mechanism is embedded in the database before the table is deleted, and the result table of the data table to be cleaned is obtained (storing all the information such as the name of the data table that needs to be cleaned), and the corresponding table is sent to the corresponding reviewer (person in charge) through the system's automatic sending of information (SMS, email or other information) function.

[0143] Automatic deletion

[0144] After the information alarm, the corresponding person in charge confirms the table name. If the corresponding person in charge finds the table useful, he can change the cleanup field in the result table to not operate the table. However, if the data table maintainer still does not operate the table after the corresponding time period, which is generally set to 3-4 days, the table clearing trigger will be automatically executed to clean up the table.

[0145] The disclosed embodiment performs training and analysis on the characteristic states of the data tables in the database, maximizes and optimizes the cleanup strategy for the data tables in the database, and continuously trains through reinforcement learning to accurately predict the importance of each data table, thereby ensuring that key data can be retained during the cleanup process, while effectively removing redundant or no longer needed information. This improves the efficiency and accuracy of data engineers in the process of cleaning data tables, and ensures the accuracy and efficiency of database cleanup work.

[0146] The second embodiment of the present disclosure also provides a data table cleaning system based on a deep Q network, such as Figure 4 As shown, the system comprises:

[0147] A collection module 11, which is configured to collect feature information of stock data from relevant data tables in a database;

[0148] A creation module 12 is configured to determine agent parameters based on the collected stock data feature information and create a new agent, wherein the agent parameters include a state space describing the state of the data table, an action space of all possible actions that the agent can perform, and a reward function for the agent to obtain a reward after performing a certain action;

[0149] The training module 13 is configured to perform DQN algorithm training on the newly created agent to obtain a trained agent, so that the agent learns how to select an optimized cleaning action according to the database table state based on the agent parameters, thereby maximizing the accumulated reward;

[0150] The cleaning module 14 is configured to automatically clean the data table through the trained intelligent agent.

[0151] Furthermore, the stock data characteristic information includes:

[0152] Structural information, including column name, data type, and whether there is an index;

[0153] Attribute information, including the size of the table and its relationships with other data tables.

[0154] Furthermore,

[0155] The state space includes the number of records in the table, creation time, access frequency, data table size, whether the belonging storage is invalid, daily growth, monthly growth, month-on-month growth, update time, and creator;

[0156] The action space includes a clearing table, a retaining table, and a migration table;

[0157] The reward function is designed based on the table size after the table cleanup action and the factors affecting database performance improvement, and is calculated using the following formula:

[0158] R(s,a,s ‘ )=-m*Δdata integrity+o*I / O speed improvement+p*Δstorage space

[0159] Where R(s,a,s ‘ ) is the reward function value, Δdata integrity represents the change in the constraints on database integrity after table clearing, I / O speed improvement represents the positive impact on database read and write performance after action a is executed, Δstorage space represents the change in storage space after data clearing, and m, o, and p are weight coefficients.

[0160] Furthermore, the training module 13 is specifically configured as follows:

[0161] Initialization, initialize the following variables:

[0162] Q-Network: Initialize a deep neural network as Q-Network(s,a;θ), where θ is the parameter of the neural network,

[0163] Target network: Create a copy of the Q network as the target network, whose initial parameters θ' are consistent with those of the Q network;

[0164] Experience pool: Initialize an empty experience pool to store the experience data generated by the interaction between the agent and the environment;

[0165] The agent chooses actions, including:

[0166] In the early stages of training, the agent uses an ε-greedy strategy to explore the environment:

[0167] a = random action (with probability ε)

[0168] Where ε is the exploration rate;

[0169] After training to a certain stage, the agent chooses actions by relying on the Q network;

[0170] Perform actions, observe results, and store experiences, including:

[0171] The agent selects an action a in the database environment and observes the result, which contains the next state s of the database. ‘ , determine whether the reward r reaches the termination condition; the result in the current state (state s, action a, reward r, next state s ‘ ) is stored in the experience pool D;

[0172] Update the target network, including:

[0173] Extract a batch of experience data (s i ,a i ,r i ,s ‘ i ) as training samples, and use the target network Q for each training sample ′ (s ‘ i ,a ‘ θ ′ ) Calculate the target Q value y i :

[0174] y i =r i +γ*maxQ ′ (s ‘ i ,a ‘ θ ′ )

[0175] Where γ represents the discount factor, which is used to control the importance of future rewards, and then the mean square error function value L is used to measure the difference between the predicted Q value and the target Q value:

[0176]

[0177] Update the parameters in the Q network according to the back-propagation algorithm and gradient descent method to minimize the loss function;

[0178] Repeat training to get the best, including:

[0179] Repeat the training steps from agent selection to updating the target network until the preset number of training rounds is reached, or the performance of the agent and the loss function converge to a stable state, then the Q network can be used to select the optimal use action of the table.

[0180] Furthermore, the system also includes an alarm module 15;

[0181] The alarm module 15 is configured to embed an alarm mechanism into the database, and automatically sends the information of the corresponding data table to the corresponding auditor before the agent cleans up the data table, so that the auditor can confirm the cleanup according to the information of the data table and return the confirmation information;

[0182] If the confirmation message is to confirm the cleanup, or if the confirmation message is not received within the preset time, the cleanup module 14 cleans up the data table;

[0183] If the confirmation information indicates that the data table is not to be cleaned, the cleaning module 14 is enabled to not operate the data table by changing the field "Whether to clean" in the result table, wherein the result table stores information of all data tables to be cleaned.

[0184] The data table cleaning system based on deep Q network in the embodiment of the present disclosure is used to implement the data table cleaning method based on deep Q network in method embodiment 1, so the description is relatively simple. For details, please refer to the relevant description in the previous method embodiment, which will not be repeated here.

[0185] In addition, if Figure 5 As shown, the third embodiment of the present disclosure further provides an electronic device, including a memory 100 and a processor 200, wherein the memory 100 stores a computer program, and when the processor 200 runs the computer program stored in the memory 100, the processor 200 executes the above-mentioned various possible methods.

[0186] The memory 100 is connected to the processor 200. The memory 100 may be a flash memory, a read-only memory or other memory, and the processor 200 may be a central processing unit or a single-chip microcomputer.

[0187] In addition, an embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to perform the above-mentioned various possible methods.

[0188] The computer-readable storage medium includes volatile or non-volatile, removable or non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, computer program modules or other data). Computer-readable storage media include, but are not limited to, RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable read only memory), flash memory or other memory technology, CD-ROM (Compact Disc Read-Only Memory), Digital Video Disc (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer.

[0189] It is to be understood that the above embodiments are merely exemplary embodiments used to illustrate the principles of the present disclosure, but the present disclosure is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and substance of the present disclosure, and these modifications and improvements are also considered to be within the scope of protection of the present disclosure.

Claims

1. A data table cleaning method based on deep Q network, characterized in that: The method comprises: Collect the characteristic information of stock data from relevant data tables in the database; Determine agent parameters based on the collected stock data feature information and create a new agent. The agent parameters include a state space describing the state of the data table, an action space of all possible actions that the agent can perform, and a reward function for the agent to obtain a reward after performing a certain action. The newly created agent is trained with the deep Q network DQN algorithm to obtain a trained agent, so that the agent can learn how to select optimized cleaning actions according to the database table status based on the agent parameters, thereby maximizing the cumulative reward; The data table is automatically cleaned through the trained intelligent agent.

2. The method according to claim 1, characterized in that: The stock data characteristic information includes: Structural information, including column name, data type, and whether there is an index; Attribute information, including the size of the table and its relationships with other data tables.

3. The method according to claim 1, characterized in that: The state space includes the number of records in the table, creation time, access frequency, data table size, whether the belonging storage is invalid, daily growth, monthly growth, month-on-month growth, update time, and creator; The action space includes a clearing table, a retaining table, and a migration table; The reward function is designed based on the table size after the table cleanup action and factors affecting database performance improvement.

4. The method according to claim 3, characterized in that The reward function is calculated by the following formula: R(s,a,s ‘ )=-m*Δdata integrity+o*I / O speed improvement+p*Δstorage space Where R(s,a,s ‘ ) is the reward function value, Δdata integrity represents the change in the constraints on database integrity after table clearing, input / output I / O speed improvement represents the positive impact on database read and write performance after action a is executed, Δstorage space represents the change in storage space after data clearing, and m, o, p are weight coefficients.

5. The method according to claim 4, characterized in that The DQN algorithm training of the newly created agent includes: Initialization, initialize the following variables: Q-Network: Initialize a deep neural network as Q-Network(s,a;θ), where θ is the parameter of the neural network, Target network: Create a copy of the Q network as the target network, whose initial parameters θ' are consistent with those of the Q network; Experience pool: Initialize an empty experience pool to store the experience data generated by the interaction between the agent and the environment; The agent selects actions, including: In the early stages of training, the agent uses an ε-greedy strategy to explore the environment: a = random action (with probability ε) Where ε is the exploration rate; After training to a certain stage, the agent chooses actions by relying on the Q network; Perform actions, observe results, and store experiences, including: The agent selects an action a in the database environment and observes the result, which contains the next state s of the database. ‘ , determine whether the reward r reaches the termination condition; the result in the current state (state s, action a, reward r, next state s ‘ ) is stored in the experience pool D; Update the target network, including: Extract a batch of experience data (s i ,a i ,r i ,s ‘ i ) as training samples, and use the target network Q for each training sample ′ (s ‘ i ,a ‘ θ ′ ) Calculate the target Q value y i : and i =r i +γ*maxQ ′ (s ‘ i ,to ‘ ;θ ′ ) Where γ represents the discount factor, which is used to control the importance of future rewards, and then the mean square error function value L is used to measure the difference between the predicted Q value and the target Q value: Update the parameters in the Q network according to the back-propagation algorithm and gradient descent method to minimize the loss function; Repeat training to get the best, including: Repeat the training steps from agent selection to updating the target network until the preset number of training rounds is reached, or the performance of the agent and the loss function converge to a stable state, then the Q network can be used to select the optimal use action of the table.

6. The method according to claim 1, characterized in that The method further comprises: Embed an alarm mechanism for the database. Before the agent cleans up the data table, it automatically sends the corresponding data table information to the corresponding auditor, so that the auditor can confirm the cleanup based on the data table information and return the confirmation information. If the confirmation message is to confirm the cleanup, or if the confirmation message is not received within the preset time, the data table will be cleaned up; If the confirmation information indicates that the data table is not to be cleaned, the cleanup field in the result table is changed to not perform any operation on the data table. The result table stores information of all data tables to be cleaned.

7. A data table cleaning system based on deep Q network, characterized in that: The system comprises: A collection module, which is configured to collect feature information of stock data from relevant data tables in a database; A creation module is configured to determine agent parameters based on the collected stock data feature information and create a new agent, wherein the agent parameters include a state space describing the state of the data table, an action space of all possible actions that the agent can perform, and a reward function for the agent to obtain a reward after performing a certain action; The training module is configured to perform DQN algorithm training on the newly created agent to obtain a trained agent, so that the agent can learn how to select optimized cleaning actions according to the database table status based on the agent parameters, thereby maximizing the cumulative reward; A cleaning module is configured to automatically clean the data table through a trained agent.

8. The system according to claim 7, characterized in that The state space includes the number of records in the table, creation time, access frequency, data table size, whether the belonging storage is invalid, daily growth, monthly growth, month-on-month growth, update time, and creator; The action space includes a clearing table, a retaining table, and a migration table; The reward function is designed based on the table size after the table cleanup action and the factors affecting database performance improvement, and is calculated using the following formula: R(s,a,s ‘ )=-m*Δdata integrity+o*I / O speed improvement+p*Δstorage space Where R(s,a,s ‘ ) is the reward function value, Δdata integrity represents the change in the constraints on database integrity after table clearing, I / O speed improvement represents the positive impact on database read and write performance after action a is executed, Δstorage space represents the change in storage space after data clearing, and m, o, and p are weight coefficients.

9. An electronic device, characterized in that: It includes a memory and a processor, wherein a computer program is stored in the memory, and when the processor runs the computer program stored in the memory, the processor executes the data table cleaning method based on the deep Q network as described in any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the data table cleaning method based on the deep Q network according to any one of claims 1 to 6 is implemented.