A data processing method, apparatus, electronic device, and storage medium
By de-private and clustering database query statements, generate statement clusters and output processing tasks, the slow query problem in the database is solved, and efficient processing and user information protection is achieved.
Patent Information
- Application Number
- CN202210412172.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-19
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-04-19
AI Technical Summary
There are slow query problems in database queries, resulting in degradation in database performance and impact on user experience. It is difficult for the existing technology to effectively handle slow query online, and may lead to user information leakage.
By determining the target query statement with the execution time exceeding the predetermined time threshold from the query statement, de-privacy processing is performed, statement clusters are clustered, and processing tasks are output for each statement cluster, so as to achieve efficient processing of slow queries.
While taking into account the user information is not leaked, it efficiently processes slow query statements, reduces the number of processing tasks, and improves database performance and user experience.
Smart Images

Figure CN114911817B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing, in particular to the field of databases, and specifically relates to a data processing method, apparatus, electronic device, and storage medium. Background Art
[0002] A business system can generate a query statement for a database, such as an SQL (Structured Query Language) statement, based on query information given by a user, and send the generated query statement to the database system so that the database system performs a data query.
[0003] However, for database queries, there is a problem of slow queries, such as slow SQL (Structured Query Language) problems. Since the execution time of slow queries is too long, it has always been an important factor affecting the performance of the database, and greatly affects the user experience. Summary of the Invention
[0004] The present disclosure provides a data processing method, apparatus, electronic device, and storage medium.
[0005] According to one aspect of the present disclosure, a data processing method is provided, including:
[0006] Determining at least one target query statement that meets a specified condition from at least one query statement; wherein the specified condition includes that the execution duration exceeds a predetermined duration threshold;
[0007] Performing de-privatization processing on the at least one target query statement respectively to obtain at least one statement to be utilized;
[0008] Clustering the at least one statement to be utilized according to the text content of the at least one statement to be utilized to obtain at least one statement cluster;
[0009] Outputting a corresponding processing task for each statement cluster in the at least one statement cluster.
[0010] According to a second aspect of the present disclosure, a data processing apparatus is provided, including:
[0011] A first determination module, configured to determine at least one target query statement that meets a specified condition from at least one query statement; wherein the specified condition includes that the execution duration exceeds a predetermined duration threshold;
[0012] A first processing module, configured to perform de-privatization processing on the at least one target query statement respectively to obtain at least one statement to be utilized;
[0013] A clustering module, configured to cluster the at least one statement to be utilized according to the text content of the at least one statement to be utilized, so as to obtain at least one statement cluster;
[0014] An output module, configured to output a corresponding processing task for each statement cluster in the at least one statement cluster.
[0015] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0016] At least one processor; and
[0017] A memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute any one of the data processing methods.
[0019] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any one of the data processing methods.
[0020] According to a fifth aspect of the present disclosure, there is further provided a computer program product, including a computer program, where the computer program, when executed by a processor, implements any one of the data processing methods.
[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0022] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0023] Figure 1 is a flowchart of the data processing method provided by the embodiment of the present disclosure;
[0024] Figure 2 is another flowchart of the data processing method provided by the embodiment of the present disclosure;
[0025] Figure 3 is a schematic diagram of a processing interface for a processing task provided by the embodiment of the present disclosure;
[0026] Figure 4 is a schematic diagram of a system architecture applicable to the data processing method provided by the embodiment of the present disclosure;
[0027] Figure 5It is a schematic structural diagram of the data processing device provided by an embodiment of the present disclosure;
[0028] Figure 6 It is another schematic structural diagram of the data processing device provided by an embodiment of the present disclosure;
[0029] Figure 7 It is a block diagram of an electronic device for implementing the data processing method according to an embodiment of the present disclosure. Detailed implementation manners
[0030] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0031] As a user-oriented system, the business system can generate a query statement for the database based on the query information given by the user, such as an SQL (Structured Query Language) statement, and send the generated query statement to the database system so that the database system can perform data query. Exemplarily, the business system can be a library platform, a cloud storage platform, various information processing systems, and so on.
[0032] However, for database queries, due to factors such as insufficient optimization of the code content of the business system, problems such as slow queries may occur, such as slow SQL (Structured Query Language) problems. Since the execution time of slow queries is too long, it has always been an important factor affecting database performance, and it greatly affects the user experience. Moreover, with the complexity of the database architecture and the increase in the amount of data, the problems brought by slow queries have become increasingly prominent: resulting in synchronization delays in the entire database cluster and master-slave delays in the business; large consumption of host and database resources, affecting the performance of other queries, and a decrease in the database throughput capacity; query requests piling up, ultimately resulting in slow response or even complete unresponsiveness of the database business, and so on.
[0033] Currently, there are simple processing methods for slow SQL, but they only identify slow SQL in the development environment and then notify the corresponding handler for corresponding processing, without dealing with the slow SQL problems that already exist in the online business system. Moreover, since the identified slow SQL may contain user information, it may lead to the leakage of user information.
[0034] Based on the above problems, embodiments of the present disclosure provide a data processing method, apparatus, electronic device, and storage medium to achieve the purpose of efficiently processing slow query statements while taking into account the non-disclosure of user information.
[0035] First, a data processing method provided by embodiments of the present disclosure will be introduced below.
[0036] Among them, a data processing method provided by embodiments of the present disclosure can be applied to an electronic device, which can be a terminal device or a server, and the present disclosure does not limit the specific form of the electronic device. In addition, a data processing method provided by embodiments of the present disclosure can be applied to a cluster, distributed, or any other application scenario with a slow query statement processing requirement, and embodiments of the present disclosure do not limit the specific scenario.
[0037] In addition, the execution subject of a data processing method provided by embodiments of the present disclosure can be a data processing apparatus. Exemplarily, when the data processing method is applied to a terminal device, the data processing apparatus can be a functional software running on the terminal device, such as: a tool software for processing slow query statements; of course, the data processing apparatus can also be a plug-in in an existing client, such as: a plug-in in a client for software processing. Exemplarily, when the data processing method is applied to a server, the data processing apparatus can be a computer program running on the server, and the computer program can be used to initiate slow query statement processing.
[0038] Moreover, the slow query involved in the present disclosure refers to a query duration exceeding a predetermined time threshold. Correspondingly, a slow query statement is a query statement whose query duration exceeds the predetermined duration threshold. Among them, the predetermined duration threshold can be set according to the actual situation. For example: for a query scenario that requires immediate response, the predetermined duration threshold can be, for example, 0.5 seconds, 1 second, 2 seconds, etc.; while for other scenarios with lower requirements for response speed, the predetermined duration threshold can be, for example, 0.5 minutes, 1 minute, etc.
[0039] In addition, exemplarily, for a database that uses SQL statements for queries, the so-called slow query problem can be referred to as the slow SQL problem. The present disclosure does not specifically limit the type of query statement used for slow queries.
[0040] A data processing method provided by embodiments of the present disclosure may include the following steps:
[0041] Determine at least one target query statement that meets the specified conditions from at least one query statement; wherein, the specified conditions include an execution duration exceeding the predetermined duration threshold;
[0042] Perform de - anonymization processing on the at least one target query statement respectively to obtain at least one statement to be utilized;
[0043] Cluster the at least one statement to be utilized according to the text content of the at least one statement to be utilized to obtain at least one statement cluster;
[0044] For each statement cluster in the at least one statement cluster, output a corresponding processing task.
[0045] In this solution, after determining at least one target query statement that meets the specified conditions in at least one query statement, by performing de - anonymization processing on the at least one target query statement, information is prevented from being leaked; and, considering that the at least one statement to be utilized after de - anonymization processing contains statements with common content, therefore, cluster the at least one statement to be utilized, and output processing tasks respectively according to each statement cluster obtained by clustering. In this way, compared with setting processing tasks for each target query statement, the number of processing tasks is greatly reduced. It can be seen that through this solution, while taking into account the non - leakage of user information, slow query statements can be processed efficiently.
[0046] Next, in combination with the accompanying drawings, an exemplary introduction to a data processing method provided by the present disclosure will be given.
[0047] As Figure 1 shown, a data processing method provided by the present disclosure may include the following steps:
[0048] S101: Determine at least one target query statement that meets the specified conditions from at least one query statement;
[0049] Wherein, the specified condition includes that the execution duration exceeds a predetermined duration threshold;
[0050] In the embodiments of the present disclosure, when the target business system runs, log data will be generated, and the log data includes at least one query statement; then, at least one target query statement that meets the specified conditions can be determined from at least one query statement, so as to perform subsequent processing on at least one target query statement that meets the specified conditions. Wherein, the specified condition includes that the execution duration exceeds a predetermined duration threshold. Correspondingly, the target query statement can be a slow query statement. In addition, in practical applications, a slow query statement can also be referred to as a statement with a slow query status or a statement with a slow query problem. Exemplarily, the business system can be a library platform, a cloud storage platform, various information processing systems, etc.
[0051] Among them, the specific implementation manner of determining at least one target query statement that meets the specified conditions from at least one query statement is not limited in the embodiments of the present disclosure. Any manner capable of identifying whether a query statement meets the specified conditions can be applied to the embodiments of the present disclosure. Exemplarily, the manner of determining at least one target query statement that meets the specified conditions from at least one query statement may include: obtaining at least one query statement; for each query statement, if it is identified that the query statement meets the specified conditions based on the execution duration and / or EXPLAIN result of the query statement, determining the query statement as a target query statement. Among them, if it is identified whether the query statement meets the specified conditions based on the execution duration of the query statement, a predetermined duration threshold may be set in advance. When identifying the query statement, the execution duration of the query statement is compared with the predetermined duration threshold. If it exceeds the predetermined duration threshold, it is determined that the query statement meets the specified conditions; if it is identified whether the query statement meets the specified conditions based on the EXPLAIN result of the query statement, the EXPLAIN command (used to detect the query execution plan of the query statement) may be executed on the query statement to obtain the syntax analysis result, that is, the EXPLAIN result, and the EXPLAIN result is scored, so as to determine whether the query statement meets the specified conditions based on whether the scoring score meets the score conditions set for slow queries, that is, based on whether the scoring score meets the score conditions set for the specified conditions; and if it is identified whether the query statement meets the specified conditions based on the execution duration and EXPLAIN result of the query statement, it may be determined that the query statement meets the specified conditions when the execution duration exceeds the predetermined duration threshold and the evaluation score meets the score conditions set for slow queries. It can be understood that the EXPLAIN result may include multiple columns of content: query type (select_type), scanning method (type), number of scanned rows (rows), actually used index (key), etc. Different evaluation scores may be set for different values of the columns affecting query efficiency, so as to identify whether it meets the specified conditions based on the total evaluation score.
[0052] In addition, in actual applications, considering that the target business system can call the database system in various operating environments to generate log data, and the generated log data includes at least one query statement. Therefore, at least one query statement of the required operating environment can be flexibly selected to process slow queries in the specified operating environment. Then, the method for determining at least one target query statement that meets the specified conditions from at least one query statement may include: determining at least one target query statement that meets the specified conditions from at least one query statement generated in the specified operating environment; wherein, the specified operating environment includes at least one of a research and development environment, a test environment, a heterogeneous environment, and a production environment. It should be noted that in the embodiments of the present disclosure, the log data generated in the specified operating environment can be collected, which can be real-time collection or scheduled collection, etc., all of which are reasonable; the so-called research and development environment, that is, the development environment, at this time is for the development or programming at the code level; the so-called test environment is the environment for testing the code content generated in the research and development environment; the so-called heterogeneous environment, also known as the access environment, is for further testing of the test environment. At this time, it completely simulates the production environment and further tests the code content generated in the research and development environment; the production environment, at this time the developed code content has been put on the market for users to use.
[0053] Based on the above implementation method, the target query statements that meet the specified conditions in any specified operating environment can be flexibly processed, and finally the business code can be flexibly optimized. Moreover, multiple operating environments can be selected to intercept and optimize slow queries layer by layer to achieve the effect of comprehensive processing.
[0054] It should be noted that the above description of the specified operating environment and the method for determining each target query statement that meets the specified conditions is only an example and should not constitute a limitation to the present disclosure.
[0055] S102: Perform de-privatization processing on at least one target query statement to obtain at least one statement to be utilized;
[0056] Among them, each target query statement may contain content related to user information. At this time, in order to ensure information security and protect personal privacy, the target query statement can be de-privatized, that is, the user information in the target query statement is de-privatized. It can be understood that by de-privatizing the user information in the target query statement, data desensitization of the target query statement can be achieved.
[0057] Exemplarily, the method for de - anonymizing the target query statement can be: key information hiding or key information replacement, where the key information here is the user information that needs to be de - anonymized. For example, if the target query statement contains information such as personal name and address, these information can be hidden, that is, this part of the content can be not displayed, or this part of the content can be replaced with fixed characters, such as "?" or "*".
[0058] In addition, the target fields related to user information can be determined in advance, and when de - anonymizing, the field content corresponding to the target fields is de - anonymized. For example, a slow query statement is SELECT*FROM Persons WHEREcity=‘Beijing’; the target fields are preset as: FROM, WHERE, then after de - anonymization, it becomes SELECT*FROM?WHERE%?=‘?’.
[0059] It should be emphasized that the above - mentioned specific implementation methods for de - anonymizing the target query statement are only examples and should not constitute a limitation to the embodiments of the present disclosure. In addition, in the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0060] S103: Cluster at least one statement to be utilized according to the text content of the at least one statement to be utilized, to obtain at least one statement cluster;
[0061] Since a query statement is a statement formed after assigning specific query information to a certain type of query statement, and the query information usually involves user information, then, in order to prevent user information leakage, and considering that at least one target query statement after de - anonymization contains statements with common content, at least one statement to be utilized can be clustered to obtain at least one statement cluster, so as to perform subsequent data processing according to the clustering result, and finally achieve efficient processing of slow query statements.
[0062] It can be understood that if all of the at least one statement to be utilized are the same, one statement cluster can be obtained; if there are identical statements among the at least one statement to be utilized, the number of statement clusters obtained is less than the number of the at least one statement to be utilized.
[0063] It should be noted that there can be various ways to cluster at least one statement to be utilized according to the text content of the at least one statement to be utilized.
[0064] Exemplarily, in one implementation, the strings of at least one statement to be utilized may be compared, and the target query statements with the same strings may be grouped into one cluster, so as to obtain at least one statement cluster.
[0065] Exemplarily, in another implementation, clustering the at least one statement to be utilized according to the text content of the at least one statement to be utilized to obtain at least one statement cluster may include: performing a hash operation on each statement to be utilized in the at least one statement to be utilized to obtain a signature corresponding to the statement to be utilized; and obtaining at least one statement cluster according to the signature corresponding to the statement to be utilized.
[0066] The signature obtained by performing a hash operation on the statement to be utilized may be an 8-bit, 16-bit, or 32-bit string, and this signature may be used as an identifier of the target query statement, so as to retrieve and analyze the statement to be utilized. Through the hash operation, a string of a fixed size may be obtained, so that according to the results of each hash operation, the statements to be utilized can be clustered more efficiently. Among them, the condition for obtaining at least one statement cluster according to the signature corresponding to the statement to be utilized may be: the signatures are the same, or the signature similarity is greater than a predetermined threshold, etc.
[0067] S104: Output a corresponding processing task for each statement cluster in the at least one statement cluster;
[0068] Wherein, the processing task is a task used to indicate a specified processing for this statement cluster.
[0069] After clustering the de-privatized target query statements, a processing task may be output for each statement cluster obtained by clustering. Compared with the prior art, the present solution does not output tasks for each slow query statement, which avoids a large amount of manpower for processing each slow query statement, thereby realizing efficient processing of slow query statements. It should be noted that in practical applications, the number of statement clusters obtained after clustering is much less than the number of statements before clustering.
[0070] Wherein, the specified processing may be a processing for optimizing the business code content corresponding to the statements in this statement cluster. Correspondingly, the processing task may be a task used to indicate a processor to optimize the business code content corresponding to the statements in this statement cluster. In this way, by outputting the processing task for the statement cluster, the processor can locate the code content corresponding to the statements in this statement cluster, so as to complete the optimization of the code content. Of course, the specified processing may also be other types of processing, for example: a processing for merely locating the business code content corresponding to the statements in this statement cluster, etc. The embodiments of the present disclosure do not make any limitations thereto.
[0071] Exemplarily, in one implementation, for each statement cluster among at least one statement cluster, outputting a corresponding processing task includes: for each statement cluster, determining the statement identifier of the statement cluster as the task content of the statement cluster. For example, using the signature of the statements in the statement cluster as the task content; for each statement cluster, outputting a processing task including the statement identifier of the statement cluster. Among them, the statement identifier of each statement cluster is associated with a statement to be utilized in the statement cluster. In this way, the processing personnel can, through the statement identifier of each statement cluster, locate a statement to be utilized in the statement cluster from a data table containing the mapping relationship between the statement identifiers and each statement cluster, and then use the located statement to perform subsequent code content location and optimization.
[0072] Exemplarily, in one implementation, for each statement cluster among at least one statement cluster, outputting a corresponding processing task includes: for each statement cluster, determining the corresponding processing object of the statement cluster and outputting a corresponding processing task; where the processing object is a statement to be utilized in the statement cluster. It should be noted that, for the purpose of achieving efficient and intuitive processing of slow query statements, when outputting the processing task for each statement cluster, the statement content involved in each statement cluster can be output. And since the statements included in the same statement cluster are the same, therefore, for each statement cluster, any statement to be utilized in the statement cluster, that is, any query statement after de-identification, can be used to form the task content corresponding to the statement cluster.
[0073] In addition, for the case of determining at least one target query statement that meets the specified conditions for at least one query statement generated under a specified operating environment, for each obtained statement cluster, outputting a corresponding processing task includes:
[0074] For each statement cluster among the at least one statement cluster, outputting the processing task corresponding to the statement cluster to the processing end corresponding to the specified operating environment. The processing end corresponding to the specified operating environment is the processing end used by the processing personnel of the specified operating environment. By outputting the processing task to the processing end corresponding to the specified operating environment, the processing personnel of the specified operating environment can be informed of the processing task, so as to perform corresponding processing on the processing task.
[0075] In this solution, after determining at least one target query statement that meets the specified conditions among at least one query statement, the at least one target query statement is de-privatized to prevent information leakage. Moreover, considering that the at least one statement to be utilized after de-privatization contains statements with common content, the at least one statement to be utilized is clustered, and processing tasks are output separately for each of the at least one statement clusters obtained by clustering. In this way, compared with setting processing tasks for each target query statement, the number of processing tasks is greatly reduced. It can be seen that through this solution, slow query statements can be efficiently processed while ensuring that user information is not leaked.
[0076] Based on the above data processing method, in another embodiment of the present disclosure, another data processing method is further provided. As Figure 2 shown, it may include the following steps:
[0077] S201: Determine at least one target query statement that meets the specified conditions from at least one query statement;
[0078] wherein, the specified conditions include that the execution duration exceeds a predetermined duration threshold;
[0079] S202: De-privatize each of the at least one target query statement to obtain at least one statement to be utilized;
[0080] S203: Cluster the at least one statement to be utilized according to the text content of the at least one statement to be utilized to obtain at least one statement cluster;
[0081] S204: Output a corresponding processing task for each of the at least one statement clusters;
[0082] wherein, the processing task is a task for indicating to perform a specified processing on this statement cluster;
[0083] The content of steps S201 - S204 is similar to that of S101 - S104 above and will not be elaborated here.
[0084] S205: Determine the query statement to be executed;
[0085] Since during the operation of the system, there are not only executed query statements but also new, unexecuted query statements generated; therefore, when any unexecuted query statement is received, this unexecuted query statement can be used as the query statement to be executed, and subsequent data processing steps can be performed for this query statement to be executed.
[0086] It should be noted that the query statement to be executed can be a query statement that has never been executed, or a query statement improved based on a previously executed query statement, which is not limited here.
[0087] S206: Perform de-privacy processing on the query statement to be executed to obtain the statement to be analyzed;
[0088] When the target business system generates a business request carrying a query statement, that is, a query statement to be executed is generated. Before the query statement to be executed is sent to the database system, the user information in the query statement to be executed can be de-privatized to obtain the statement to be analyzed, and the subsequent processing method steps are executed on the statement to be analyzed, so as to achieve real-time control of slow query statements.
[0089] The method of de-privatizing the user information in the query statement to be executed is the same as the method of de-privatizing the user information in the target query statement described above. Exemplarily, in one implementation, the method of de-privatizing the query statement to be executed is: hiding or information replacement of the part of the query statement to be executed that contains user information; for example: replacing the part containing information such as personal name and address with "?" or "*", etc. Of course, it is not limited to this.
[0090] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of user information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0091] S207: If it is detected that the statement to be analyzed belongs to any one of at least one statement cluster, intercept and / or alarm the query statement to be executed;
[0092] It should be noted that at least one statement cluster is obtained by clustering the above target query statement. The target query statement is a slow query statement, that is, the target query statement is a statement that meets the specified conditions. Therefore, if the statement to be analyzed belongs to any one of the at least one statement cluster, the query statement to be executed corresponding to the statement to be analyzed is a slow query statement, that is, the query statement to be executed is a statement that meets the specified conditions, and the query statement to be executed can be directly intercepted and / or alarmed. It can be understood that the so-called interception means intercepting the query statement to be executed outside the database system so that the database system cannot receive the query statement to be executed; and, in practical applications, the alarm methods can include: sending an alarm notification, an alarm signal, emitting an alarm sound, or any alarm notification that can remind the corresponding processing personnel. Exemplarily, a method for detecting whether the statement to be analyzed belongs to any one of the at least one statement cluster can be: detecting whether the string of the statement to be analyzed is the same as the string of any one of the at least one statement cluster, or detecting whether the result obtained after hashing the statement to be analyzed is the same as the hashing result of any one of the at least one statement cluster. If they are the same, it is detected that the statement to be analyzed belongs to any one of the at least one statement cluster.
[0093] In addition, in order to avoid the problem of possible losses to the target business system caused by intercepting and / or alarming any slow query statement, therefore, the statement to be analyzed can be re-identified and then decide whether to intercept and / or alarm, so as to achieve flexible filtering of slow query statements and prevent misjudgment of the production environment.
[0094] Based on this processing idea, exemplarily, in one implementation, if it is detected that the statement to be analyzed belongs to any one of the at least one statement cluster, intercepting and / or alarming the query statement to be executed may include:
[0095] If it is detected that the statement to be analyzed belongs to any one of the at least one statement cluster, and the specified evaluation index corresponding to any one of the statement clusters meets the predetermined threshold condition, then intercept and / or alarm the query statement to be executed.
[0096] Exemplarily, the specified evaluation metrics can be execution duration, lock duration (the time when a slow query statement is locked), EXPLAIN result, etc. These results can be pre-recorded, or statistically calculated, which are all reasonable. For example, for a slow query, the specified evaluation metric is the execution duration. At this time, if the execution duration corresponding to any one statement cluster exceeds a preset duration threshold, the query statement to be executed can be intercepted and / or an alarm can be issued. It can be understood that the execution duration corresponding to any one statement cluster can be obtained by averaging the execution durations of the target query statements included in any one statement cluster. Additionally, it can be understood that if the statement to be analyzed does not belong to any one of the at least one statement cluster, or the statement to be analyzed belongs to any one of the at least one statement cluster but the specified evaluation metric for the slow query corresponding to any one statement cluster does not meet the predetermined threshold condition, the query statement to be executed corresponding to the statement to be analyzed can be issued, that is, the statement to be analyzed is issued to the database system so that the database system can respond to the statement to be analyzed.
[0097] The solution provided in this embodiment can not only efficiently process slow query statements while taking into account the protection of user information from leakage, but also filter newly generated slow query statements in the business system, achieving the purpose of real-time control of slow query statements. Moreover, by setting predetermined threshold conditions, slow query statements can be reasonably filtered to prevent damage to the production environment.
[0098] Optionally, based on the above embodiment, in the data processing method provided in another embodiment of the present disclosure, after outputting corresponding processing tasks for each of the at least one statement clusters, it may further include:
[0099] When receiving a specified control instruction for any one processing task, the processing task responds to the specified control instruction.
[0100] Among them, the specified control instruction for any one processing task is an instruction for controlling the processing personnel, status information, etc. of the processing task. Exemplarily, the specified control instruction can be: an instruction for indicating transferring the task to others, an instruction for marking the task as an exempt task, an instruction for setting the status of the task to the resolved state, an instruction for setting note information for the task, etc. Among them, marked as exempt: after marking as exempt, such signatures will no longer be alarmed and notified, and need to be reviewed by the processing personnel.
[0101] It can be understood that on the output interface of the processing task, a trigger button corresponding to each specified control instruction can be set. In this way, the processing personnel can issue the corresponding specified control instruction by clicking the trigger button, so that the data processing device can respond to the specified control instruction after receiving the specified control instruction.
[0102] For the convenience of understanding the solution, Figure 3 a schematic diagram of the processing interface for a processing task is given. As Figure 3 shown, for a processing task, the task content of the processing task and other information can be displayed in the output interface. The other information includes the current assignee, status, type, and trigger buttons corresponding to each specified control instruction. Among them, if the processing task is a task determined based on the log data generated in the production environment, the type is the online type; if the processing task is a task determined based on the log data generated in other operating environments, the type is the offline type.
[0103] It can be seen that by responding to the specified control instruction for any processing task, the processing task can be flexibly processed by the processing personnel.
[0104] Next, in combination with another specific embodiment, the principle content of a slow query statement processing method provided by the present disclosure will be introduced in detail.
[0105] As Figure 4 shown in the system architecture diagram, Figure 4 the following five parts are provided:
[0106] Simple database architecture 410: The database system includes a master node and slave nodes (master node and slave nodes), a primary and standby node, and a proxy instance; among them, the proxy instance transfers the external write traffic to the master node to enable the master node to complete the data writing process; transfers the external read traffic to each slave node slave-01 to slave-n to enable the slave node to complete the data reading process. Among them, during the process of the business system accessing the database, a log file can be generated (which contains at least one of the above query statements).
[0107] Collection and processing of slow query statements 420: Real-time log collection can be performed from the log file generated in the simple database architecture, and then slow query statement recognition can be performed on the collected logs (corresponding to determining at least one target query statement that meets the specified conditions from at least one of the above query statements), and then signature information generation (corresponding to de-privatizing the target query statement and then performing a hash operation to obtain the signature corresponding to the target query statement), and finally relevant information is stored. This part can store information such as the execution duration of the query statement, the number of parsed lines: the number of code lines for parsing the query statement, the number of returned lines: the number of lines of data returned for the query statement, the source IP: the IP of the source of the query statement, and the request time: the request time of the query statement.
[0108] Slow query statement governance 430: Receive the data pushed by the slow query statement collection and processing part, aggregate similar slow query statements (corresponding to clustering at least one to-be-utilized statement above to obtain at least one statement cluster), and store the aggregated information: that is, store each statement cluster obtained after aggregation, and then process each statement cluster according to the aggregated result (corresponding to outputting the corresponding processing task for each statement cluster in the at least one statement cluster above). The so-called processing may include: notification and alarm: alarm for the corresponding slow query statement and notify the corresponding handler; claim function: the handler claims the processing task belonging to himself; handover function: hand over the processing task that does not belong to himself to the corresponding handler; allocation function: allocate the corresponding processing task to the corresponding handler; set schedule: set the deadline of the processing task; set remarks: set the remarks of the processing task, for example: special requirements for a certain processing task: it must be processed on the same day, etc.
[0109] Slow query statement identification and interception 440: When any business request arrives (corresponding to determining the to-be-executed query statement above), perform de-privacy processing on the user information in the query statement of the business request to obtain the to-be-analyzed statement (corresponding to performing de-privacy processing on the to-be-executed query statement to obtain the to-be-analyzed statement above); request the data of the slow query statement governance part, and then identify whether the to-be-analyzed statement is a slow query statement according to the data of the slow query statement governance part (corresponding to if it is detected that the to-be-analyzed statement belongs to any one of the at least one statement cluster, then the to-be-executed query statement corresponding to the to-be-analyzed statement is a slow query statement). If it is not a slow query statement, it is normally sent to the database system; if it is a slow query statement, judge again whether the threshold is reached (corresponding to if it is detected that the to-be-analyzed statement belongs to any one of the at least one statement cluster, and the specified evaluation index regarding slow query corresponding to this any one statement cluster meets the predetermined threshold condition). If the threshold is reached, intercept and alarm the slow query statement; if not, normally send the slow query statement to the database.
[0110] Slow query statement interception 450: The first interception in the development environment; the second interception in the test environment; the third interception in the heterogeneous environment; the fourth interception in the production environment; this part expands multiple interception environments, and can intercept slow query statements as much as possible before the production environment and optimize the business code in a timely manner. The multi-environment interception can be reasonably allocated according to business requirements. For example: it can be applied only to the development environment and the test environment; or it can be applied to the development environment, the test environment, the heterogeneous environment, and the production environment, but the proportion of the development environment and the test environment is relatively large.
[0111] In this embodiment, it is possible to dynamically monitor the log data generated by the service, with stronger real-time performance. Then, perform de-privacy processing on the slow query statements. It can not only cluster slow query statements belonging to the same category, but also ensure the confidentiality of user data, avoiding potential security risks caused by user data leakage. It can efficiently process each category of slow query statements in any environment, avoiding the situation where no one cares about them online. Moreover, by setting thresholds, it is more flexible to reasonably filter slow query statements to prevent accidental damage to the production environment. The multi-channel interception strategies in multiple environments can intercept slow query statements before the production environment and optimize the business code in a timely manner to avoid affecting the business. And, for complex master-slave architectures or distributed architectures, the improvement in the efficiency of slow query processing by this solution is particularly obvious.
[0112] According to an embodiment of the present disclosure, the present disclosure also provides a data processing device, as Figure 5 shown. The device includes:
[0113] A first determination module 510, configured to determine at least one target query statement that meets specified conditions from at least one query statement; wherein, the specified conditions include that the execution duration exceeds a predetermined duration threshold;
[0114] A first processing module 520, configured to perform de-privacy processing on the at least one target query statement respectively to obtain at least one statement to be utilized;
[0115] A clustering module 530, configured to cluster the at least one statement to be utilized according to the text content of the at least one statement to be utilized to obtain at least one statement cluster;
[0116] An output module 540, configured to output a corresponding processing task for each statement cluster in the at least one statement cluster.
[0117] In this solution, after determining at least one target query statement that meets specified conditions in at least one query statement, by performing de-privacy processing on the at least one target query statement, information is prevented from being leaked. And, considering that the at least one statement to be utilized after de-privacy processing contains statements with common content, therefore, the at least one statement to be utilized is clustered, and a processing task is output respectively according to each statement cluster obtained by clustering. In this way, compared with setting processing tasks for each target query statement, the number of processing tasks is greatly reduced. It can be seen that through this solution, slow query statements can be efficiently processed while taking into account that user information is not leaked.
[0118] Optionally, the output module is specifically configured to:
[0119] For each statement cluster, determine the processing object corresponding to the statement cluster and output the corresponding processing task; wherein, the processing object is a to-be-exploited statement in the statement cluster.
[0120] Optionally, the determining module is specifically configured to:
[0121] Determine at least one target query statement that meets the specified conditions from at least one query statement generated in the specified operating environment;
[0122] The output module is specifically configured to:
[0123] For each statement cluster in the at least one statement cluster, output the processing task corresponding to the statement cluster to the processing end corresponding to the specified operating environment.
[0124] Optionally, the clustering module is specifically configured to:
[0125] Perform a hash operation on each to-be-exploited statement in the at least one to-be-exploited statement to obtain a signature corresponding to the to-be-exploited statement;
[0126] Obtain at least one statement cluster according to the signature corresponding to the to-be-exploited statement.
[0127] Optionally, as Figure 6 shown, the apparatus further includes:
[0128] A second determining module 650, configured to determine a query statement to be executed;
[0129] A second processing module 660, configured to perform a de-privatization process on the query statement to be executed to obtain a statement to be analyzed;
[0130] An interception module 670, configured to intercept and / or alarm the query statement to be executed if it is detected that the statement to be analyzed belongs to any one of the at least one statement cluster.
[0131] Optionally, the interception module is specifically configured to:
[0132] If it is detected that the statement to be analyzed belongs to any one of the at least one statement cluster and the specified evaluation index corresponding to the any one statement cluster meets the predetermined threshold condition, intercept and / or alarm the query statement to be executed.
[0133] Optionally, the apparatus further includes:
[0134] A response module, configured to respond to the specified control instruction for the processing task when receiving the specified control instruction for any processing task.
[0135] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0136] An embodiment of the present disclosure provides an electronic device, including:
[0137] at least one processor; and
[0138] a memory communicatively connected to the at least one processor; wherein,
[0139] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any data processing method.
[0140] An embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any data processing method.
[0141] The present disclosure provides a computer program product, including a computer program, where the computer program, when executed by a processor, implements any data processing method.
[0142] Figure 7 FIG. shows a schematic block diagram of an exemplary electronic device 700 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0143] As Figure 7 shown, the device 700 includes a computing unit 701, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0144] Multiple components in device 700 are connected to I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, an optical disc, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0145] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the data processing method. For example, in some embodiments, the data processing method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the data processing method described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute the data processing method in any other suitable manner (e.g., by means of firmware).
[0146] The various embodiments of the systems and technologies described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0147] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0148] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0149] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0150] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0151] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain.
[0152] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0153] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A data processing method, comprising: determining at least one target query statement that meets specified conditions from at least one query statement in log data; wherein, the specified conditions include that the execution duration exceeds a predetermined duration threshold; any one of the at least one query statement in the log data is an executed query statement; performing de - anonymization processing on the at least one target query statement respectively to obtain at least one statement to be utilized; clustering the at least one statement to be utilized according to the text content of the at least one statement to be utilized to obtain at least one statement cluster; for each statement cluster in the at least one statement cluster, outputting a corresponding processing task; wherein, the processing task is a task used to indicate specified processing for this statement cluster, and the specified processing is processing for optimizing and / or locating the business code content corresponding to the statements in this statement cluster; determining a query statement to be executed; performing de - anonymization processing on the query statement to be executed to obtain a statement to be analyzed; if it is detected that the statement to be analyzed belongs to any one of the at least one statement cluster, and the specified evaluation index corresponding to this any one statement cluster meets a predetermined threshold condition, then intercepting and / or alarming the query statement to be executed.
2. The method according to claim 1, wherein, the outputting a corresponding processing task for each statement cluster in the at least one statement cluster includes: for each statement cluster, determining a processing object corresponding to this statement cluster and outputting a corresponding processing task; wherein, the processing object is a statement to be utilized in this statement cluster.
3. The method according to claim 1, wherein, the determining at least one target query statement that meets specified conditions from at least one query statement in log data includes: determining at least one target query statement that meets specified conditions from at least one query statement in log data generated under a specified operating environment; the outputting a corresponding processing task for each statement cluster in the at least one statement cluster includes: for each statement cluster in the at least one statement cluster, outputting the processing task corresponding to this statement cluster to the processing end corresponding to the specified operating environment.
4. The method according to claim 1, wherein, the clustering the at least one statement to be utilized according to the text content of the at least one statement to be utilized to obtain at least one statement cluster includes: performing a hash operation on each statement to be utilized in the at least one statement to be utilized to obtain a signature corresponding to the statement to be utilized; obtaining at least one statement cluster according to the signature corresponding to the statement to be utilized.
5. The method according to claim 1, wherein, after the outputting a corresponding processing task for each statement cluster in the at least one statement cluster, further comprising: when receiving a specified control instruction for any processing task, responding to the specified control instruction for this processing task.
6. A data processing device, comprising: A first determination module, configured to determine at least one target query statement that meets specified conditions from at least one query statement of log data; wherein, the specified conditions include that the execution duration exceeds a predetermined duration threshold; any one of the at least one query statement of the log data is an executed query statement; A first processing module, configured to perform de-privatization processing on the at least one target query statement respectively to obtain at least one statement to be utilized; A clustering module, configured to cluster the at least one statement to be utilized according to the text content of the at least one statement to be utilized to obtain at least one statement cluster; An output module, configured to output a corresponding processing task for each statement cluster in the at least one statement cluster; wherein, the processing task is a task for indicating to perform specified processing on this statement cluster, and the specified processing is processing for optimizing and / or locating the business code content corresponding to the statements in this statement cluster; An interception module, configured to determine a query statement to be executed; perform de-privatization processing on the query statement to be executed to obtain a statement to be analyzed; if it is detected that the statement to be analyzed belongs to any one of the at least one statement cluster, and the specified evaluation index corresponding to this any one statement cluster meets a predetermined threshold condition, then intercept and / or alarm the query statement to be executed.
7. The apparatus according to claim 6, wherein, the output module is specifically configured to: For each statement cluster, determine a processing object corresponding to this statement cluster and output a corresponding processing task; wherein, the processing object is a statement to be utilized in this statement cluster.
8. The apparatus according to claim 6, wherein, the first determination module is specifically configured to: Determine at least one target query statement that meets specified conditions from at least one query statement of log data generated under a specified operating environment; the output module is specifically configured to: For each statement cluster in the at least one statement cluster, output the processing task corresponding to this statement cluster to the processing end corresponding to the specified operating environment.
9. The apparatus according to claim 6, wherein, the clustering module is specifically configured to: Perform a hashing operation on each statement to be utilized in the at least one statement to be utilized to obtain a signature corresponding to the statement to be utilized; Obtain at least one statement cluster according to the signature corresponding to the statement to be utilized.
10. The apparatus according to claim 6, wherein, the apparatus further includes: A response module, configured to, when receiving a specified regulation instruction for any processing task, respond to this processing task with the specified regulation instruction.
11. An electronic device, including: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-5.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-5.
13. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-5.
Citation Information
Patent Citations
Statement early warning method and device, equipment and computer readable storage medium
CN110019349A
Database slow query log processing method, server, computing device and system
CN112506951A