A Method for Calculating the Sensitivity of SQL Statements in a Multi-Source Heterogeneous Database

Through the multi-source heterogeneous database SQL statement sensitivity calculation method, the privacy protection problem of complex SQL statements in the multi-source heterogeneous database environment is solved, and the balance of data analysis value is achieved while ensuring data privacy, which is suitable for databases in multiple SQL dialects.

CN119166671BActive Publication Date: 2025-07-29INSPUR SOFTWARE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411687174.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-25
Publication Date
2025-07-29
Estimated Expiration
2044-11-25

AI Technical Summary

Technical Problem

In a multi-source heterogeneous database environment, how to effectively protect data privacy when executing complex SQL statements, while maintaining the practicality of data and the accuracy of analysis results, especially how to accurately calculate and add appropriate amounts of noise to achieve a balance between data utility and privacy protection.

Method used

The sensitivity calculation method of SQL statements of multi-source heterogeneous database is adopted. By receiving and analyzing SQL query requests, an abstract syntax tree is generated, key syntax features are identified and modified, and the sensitivity calculation is performed based on the metadata information of the target database, guiding the addition of differential privacy noise.

Benefits of technology

It realizes that in a multi-source heterogeneous database environment, compatible with different types of databases, accurately calculates the sensitivity of complex SQL query statements, reduces usage costs, guarantees user query experience, and ensures privacy protection effects while retaining data analysis value to the greatest extent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119166671B_ABST
    Figure CN119166671B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for calculating the sensitivity of SQL statements in a multi-source heterogeneous database, which relates to the technical field of SQL queries. In order to accurately calculate and add appropriate levels of noise to complex SQL statements, the following solution is adopted: receiving and preprocessing an SQL query request, and disassembling it to generate an abstract syntax tree; traversing the abstract syntax tree, identifying the syntax features to be processed, defining rewrite rules, and modifying and optimizing the corresponding nodes in the abstract syntax tree; selecting a serialization format and converting the optimized abstract syntax tree into corresponding serialized data; performing semantic analysis and logical analysis on the serialized data, binding the semantic information contained in the request to the metadata information of the target database, and calling a sensitivity evaluation algorithm to calculate the sensitivity of the request. The sensitivity calculated by the present invention is used to guide the addition of data noise in the SQL query results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of SQL queries, and specifically to a method for calculating the sensitivity of SQL statements in a multi-source heterogeneous database. Background Art

[0002] In the current era of the prevalence of big data and cloud computing, the multi-source heterogeneous database environment has become the norm for data collection, storage, and analysis, which contains data sets from different sources, different formats, and structures, such as relational databases, NoSQL databases, file systems, etc. Data queries in such an environment, especially the execution of complex SQL statements, pose a severe challenge to data privacy protection. Complex SQL statements, usually involving multi-table joins, subqueries, aggregate functions, etc., not only increase the semantic complexity of the query but also make it extremely difficult to accurately evaluate the impact of the query on sensitive data, amplifying the potential risk of privacy leakage. For example, a query containing multiple JOIN and GROUP BY operations may inadvertently disclose a large amount of private information. How to effectively protect data privacy while maintaining the usability of the data and the accuracy of the analysis results when executing such statements has become a key problem that urgently needs to be solved.

[0003] Differential privacy, as an effective privacy protection technology, ensures that the query results do not change significantly whether or not an individual data is included by adding an appropriate amount of random noise to the query results, thereby protecting individual privacy. However, how to accurately calculate and add an appropriate amount of noise in complex SQL statements to achieve the optimal privacy protection effect while minimizing the impact on data utility is a technical difficulty. Especially in a multi-source heterogeneous database environment, because it is necessary to consider the distribution characteristics of the data, the complexity of the query, and the heterogeneity between data sources, measuring the sensitivity of the query results to changes in individual data records, that is, sensitivity calculation, becomes more complex. Therefore, developing a calculation method that can accurately evaluate the sensitivity of complex SQL statements and guide the addition of noise accordingly is crucial for achieving the balance between data utility and privacy protection. This requires the technical solution to not only deeply understand the structure and semantics of SQL statements but also be able to effectively integrate data fusion, query optimization, and privacy protection technologies to meet the data processing requirements in a multi-source heterogeneous environment. Summary of the Invention

[0004] In view of the requirements and deficiencies of the current technological development, the present invention provides a method for calculating the sensitivity of SQL statements in a multi-source heterogeneous database. This method can be compatible with the query dialects of different types of databases, accurately calculate the sensitivity of complex SQL query statements, reduce the usage cost, ensure the user query experience, and guide the reasonable addition of differential privacy noise, while ensuring the privacy protection effect and maximizing the retention of the analysis value of the data.

[0005] The present invention provides a method for calculating the sensitivity of SQL statements in a multi-source heterogeneous database. The technical solution adopted to solve the above technical problems is as follows:

[0006] A method for calculating the sensitivity of SQL statements in a multi-source heterogeneous database, the implementation of which includes the following four stages:

[0007] S1. SQL Receiving and Parsing Stage: Receive the SQL query request from the user, check and preprocess the SQL query request, then disassemble it into a set of lexical units, and then combine them to generate an abstract syntax tree, and verify the abstract syntax tree;

[0008] S2. Abstract Syntax Tree Rewriting Stage: Traverse the abstract syntax tree, identify the syntax features that need to be specified for processing; define a set of rewriting rules for the syntax features that need to be specified for processing; according to the defined rewriting rules, modify and optimize the corresponding nodes in the abstract syntax tree;

[0009] S3. Serialization Stage: Select the serialization format of the abstract syntax tree, and convert the optimized abstract syntax tree into the corresponding serialized data;

[0010] S4. Sensitivity Calculation Stage: Perform semantic analysis and logical analysis on the serialized data, bind the semantic information contained in the SQL query request to the metadata information of the target database, call the preset sensitivity evaluation algorithm, calculate the sensitivity of the SQL query request, and the calculated sensitivity is used to guide the addition of data noise in the SQL query result.

[0011] Optionally, the specific steps of step S1 include:

[0012] S1.1. Receive the SQL query request from the user, and the SQL query request supports multiple SQL dialects, including standard SQL and the dialects of databases MySQL, Oracle, SQL Server, and PostgreSQL;

[0013] S1.2. Check the SQL query request. The checking process includes verifying whether the SQL query request is a valid SQL statement, whether it contains potential malicious code or SQL injection attempts through regular expression matching, blacklist / whitelist filtering, or a preset SQL injection prevention mechanism;

[0014] S1.3. After the check passes, perform preprocessing operations on the SQL query request, including removing comments, normalizing whitespace, and unifying case;

[0015] S1.4. Disassemble the preprocessed SQL query request into a set of lexical units, and combine this set of lexical units into an abstract syntax tree according to the grammar rules corresponding to the SQL query request;

[0016] S1.5. Verify the abstract syntax tree, including: verifying whether the table names and column names involved in the abstract syntax tree nodes exist in the target database; checking whether the user has the permissions required to execute the query.

[0017] Further optionally, the involved step S1.3 performs preprocessing operations of removing comments, normalizing whitespace, and unifying case on the SQL query request, where:

[0018] The removing of comments includes: traversing the SQL query string and removing single-line comments and multi-line comments;

[0019] The whitespace normalization includes: standardizing the whitespace characters in the SQL query request to ensure consistency and accuracy during parsing;

[0020] The case unification includes: performing case unification on the keywords and identifiers in the SQL query request according to the type of the SQL query request.

[0021] Further optionally, the specific operations for performing step S1.4 to obtain the abstract syntax tree are as follows:

[0022] S1.4.1. Construct a lexical analyzer that splits the preprocessed SQL query string into a set of lexical units according to five categories: keywords, identifiers, operators, numbers, and string constants; during the splitting process, when the lexical analyzer discovers unrecognized characters, it generates an error message and indicates the error location;

[0023] S1.4.2. According to the grammar rules corresponding to the SQL query request, construct a syntax analyzer that recognizes and parses subqueries, JOIN operations, and GROUP BY statements, thereby identifying the relationships between lexical units, generating corresponding abstract syntax tree nodes, and generating meta-information including the query type, involved table names, and column names;

[0024] S1.4.3. During the process of generating the abstract syntax tree, perform syntax checking, and when a syntax error is found, generate a detailed error message and indicate the specific location and reason of the error.

[0025] Further optionally, the involved step S2 specifically includes:

[0026] S2.1. Traverse the abstract syntax tree to identify the syntactic features that need to be specified for processing, and the syntactic features that need to be specified for processing include aggregate functions, multi-layer subqueries, complex JOIN operations, GROUP BY, and ORDER BY clauses;

[0027] S2.2. Define a set of rewrite rules for the syntactic features that need to be specified for processing;

[0028] S2.3. Modify the corresponding nodes in the abstract syntax tree according to the defined rewriting rules. The specific modification operations include: ① Replace the nodes, replacing the old function nodes with new function expressions; ② Add nodes to the nodes, adding additional nodes required for the noise generation logic; ③ Delete nodes from the nodes, deleting redundant nodes that do not affect the query results but may disclose privacy.

[0029] S2.4. After modifying the abstract syntax tree by applying all necessary rewriting rules, further optimize the abstract syntax tree to improve the execution efficiency of the query. Among them, the optimization measures include: identifying and removing duplicate or unnecessary calculation logics, selecting relatively efficient indexes, adjusting the order of join operations, and simplifying expressions.

[0030] Further optionally, the specific steps involved in step S3 include:

[0031] S3.1. Select the serialization format of the abstract syntax tree according to the specific requirements and performance considerations of the system. The serialization formats of the abstract syntax tree include JSON, XML, Protocol Buffers, YAML, Avro, MessagePack, and CBOR.

[0032] S3.2. Use the selected serialization format to convert the optimized abstract syntax tree into corresponding serialized data to ensure the integrity and consistency of the query request during the transmission across systems or components.

[0033] Further optionally, when performing step S3.2, it is necessary to ensure that all information in the abstract syntax tree can be accurately encoded into the serialized data, including node types, node attributes, and child node relationships.

[0034] Further optionally, after performing step S3 and converting the optimized abstract syntax tree into corresponding serialized data:

[0035] Add necessary meta-information to the serialized data to assist subsequent parsing and processing. The necessary meta-information includes SQL dialect information and the list of rules applied during the rewriting process.

[0036] Verify the generated serialized data to ensure that the serialized data completely contains all information of the abstract syntax tree and does not introduce any unexpected errors or inconsistencies.

[0037] Further optionally, the specific steps involved in step S4 include:

[0038] S4.1. Perform semantic analysis and logical analysis on the serialized data to obtain the semantic information and logical structure contained in the SQL query request. Among them, performing semantic analysis on the serialized data includes extracting identifiers, operators, functions, as well as table names and field names involved in the SQL query request; performing logical analysis on the serialized data includes extracting node types, node attributes, and child node relationships contained in the SQL query request;

[0039] S4.2. Obtain the metadata information of the target database, and bind the semantic information contained in the SQL query request to the metadata of the target database to ensure that the objects involved in the query exist in the target database and meet the expectations;

[0040] S4.3. Invoke a preset sensitivity evaluation algorithm, calculate the sensitivity based on the logical structure of the SQL query request and the bound metadata information, and obtain the sensitivity of the SQL query request. The obtained sensitivity is used to add data noise to the SQL query result.

[0041] Preferably, the preset sensitivity evaluation algorithm is at least one of rule-based evaluation, statistics-based evaluation, machine learning-based evaluation, graph structure-based evaluation, risk matrix-based evaluation, and context-based evaluation.

[0042] A method for calculating the sensitivity of SQL statements in a multi-source heterogeneous database according to the present invention has the following beneficial effects compared with the prior art:

[0043] 1. The present invention can be compatible with the query dialects of different types of databases, accurately calculate the sensitivity of complex SQL query statements, reduce the usage cost, ensure the user query experience, and thus facilitate the subsequent guidance on the reasonable addition of differential privacy noise, while ensuring the privacy protection effect, maximizing the analysis value of the data;

[0044] 2. By calculating the sensitivity of complex SQL query statements, the present invention facilitates the subsequent reasonable addition of differential privacy noise, and can solve the problem of privacy leakage in query statistics and user data collection scenarios. On the basis of protecting user privacy during query and collection, the high availability goal of the data is achieved;

[0045] 3. By introducing the abstract syntax tree rewriting and serialization mechanism, the present invention can rewrite the complex feature structure into a form convenient for calculating sensitivity, and ensure the integrity and consistency of the query request during cross-system or component transmission, improving the sensitivity calculation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Attached Figure 1 is a flowchart of the method in Embodiment 1 of the present invention;

[0047] Attached Figure 2It is the flowchart of step S1 in Embodiment 1 of the present invention;

[0048] Appendix Figure 3 It is the flowchart of step S2 in Embodiment 1 of the present invention;

[0049] Appendix Figure 4 It is the flowchart of step S3 in Embodiment 1 of the present invention;

[0050] Appendix Figure 5 It is the flowchart of step S4 in Embodiment 1 of the present invention. Detailed implementation manner

[0051] To make the technical solutions, problems to be solved and technical effects of the present invention clearer and more understandable, the following describes the technical solutions of the present invention clearly and completely in combination with specific embodiments.

[0052] Embodiment 1: In combination with Appendix Figures 1-5 , this embodiment proposes a method for calculating the sensitivity of SQL statements in a multi-source heterogeneous database, and its implementation includes the following four stages:

[0053] S1. SQL receiving and parsing stage:

[0054] S1.1. Receive the SQL query request from the user. The SQL query request supports multiple SQL dialects, including standard SQL (such as ANSI SQL) and the dialects of databases MySQL, Oracle, SQL Server, and PostgreSQL;

[0055] S1.2. Check the SQL query request. The checking process includes verifying whether the SQL query request is a valid SQL statement, whether it contains potential malicious code or SQL injection attempts through regular expression matching, blacklist / whitelist filtering, or a preset SQL injection prevention mechanism;

[0056] After passing the check, perform preprocessing operations on the SQL query request, including removing comments, normalizing whitespace, and unifying case. Specifically: Removing comments includes traversing the SQL query string and removing single-line comments (such as lines starting with -- or #, depending on the SQL dialect) and multi-line comments (such as / *... * / ); Whitespace normalization includes standardizing the whitespace characters in the SQL query request to ensure consistency and accuracy during parsing; Case unification includes unifying the case of keywords and identifiers in the SQL query request according to the type of the SQL query request;

[0057] S1.4. Decompose the preprocessed SQL query request into a set of lexical units, and combine these lexical units into an abstract syntax tree according to the corresponding grammar rules of the SQL query request. The specific operations include:

[0058] S1.4.1. Construct a lexical analyzer which splits the preprocessed SQL query string into a set of lexical units according to five categories: keywords (such as SELECT, FROM, WHERE), identifiers (such as table names, field names), operators (such as =, <>, AND, OR), numbers, and string constants. During the splitting process, when the lexical analyzer encounters unrecognized characters, it generates an error message and indicates the error location.

[0059] S1.4.2. According to the grammar rules corresponding to the SQL query request, construct a syntax analyzer which identifies and parses subqueries, JOIN operations, and GROUP BY statements, thereby identifying the relationships between lexical units, generating corresponding abstract syntax tree nodes, and generating meta-information including the query type, involved table names, and column names.

[0060] S1.4.3. During the process of generating the abstract syntax tree, perform syntax checking, and when a syntax error is found, generate a detailed error message and indicate the specific location and reason of the error.

[0061] S1.5. Verify the abstract syntax tree, including: verifying whether the table names and column names involved in the abstract syntax tree nodes exist in the target database; checking whether the user has the permissions required to execute the query.

[0062] S2. Abstract Syntax Tree Rewriting Phase:

[0063] S2.1. Traverse the abstract syntax tree to identify the syntax features that need to be processed specifically. The syntax features that need to be processed specifically include aggregate functions (such as SUM, AVG, COUNT, etc.), multi-layer subqueries, complex JOIN operations, GROUP BY, and ORDER BY clauses.

[0064] S2.2. Define a set of rewriting rules for the syntax features that need to be processed specifically. For example, for the AVG function, define a rule to rewrite it as a division expression of the SUM function and the COUNT function, so as to utilize the metadata information and sensitivity algorithm involved in the SUM function and the COUNT function to calculate the overall sensitivity in subsequent processing.

[0065] S2.3. According to the defined rewriting rules, modify the corresponding nodes in the abstract syntax tree. The specific modification operations include: ① replacing the node, replacing the old function node with a new function expression; ② adding nodes to the node, adding additional nodes required for the noise generation logic; ③ deleting nodes from the node, deleting redundant nodes that do not affect the query result but may disclose privacy.

[0066] S2.4. After modifying the abstract syntax tree by applying all necessary rewrite rules, further optimize the abstract syntax tree to improve the execution efficiency of the query. The optimization measures include: identifying and removing duplicate or unnecessary calculation logic, selecting relatively efficient indexes, adjusting the order of join operations, and simplifying expressions.

[0067] S3. Serialization stage:

[0068] S3.1. According to the specific requirements and performance considerations of the system, select the serialization format of the abstract syntax tree. The serialization formats of the abstract syntax tree include JSON, XML, Protocol Buffers, YAML, Avro, MessagePack, and CBOR.

[0069] S3.2. Use the selected serialization format to convert the optimized abstract syntax tree into corresponding serialized data, ensuring the integrity and consistency of the query request during cross-system or component transmission. At the same time, ensure that all information in the abstract syntax tree can be accurately encoded into the serialized data, including node types, node attributes, and child node relationships.

[0070] S3.3. Add necessary meta-information to the serialized data to assist subsequent parsing and processing. The necessary meta-information includes SQL dialect information and the list of rules applied during the rewrite process.

[0071] S3.4. Verify the generated serialized data to ensure that the serialized data completely contains all information of the abstract syntax tree and does not introduce any unexpected errors or inconsistencies.

[0072] S4. Sensitivity calculation stage:

[0073] S4.1. Perform semantic analysis and logical analysis on the serialized data to obtain the semantic information and logical structure contained in the SQL query request. Among them, performing semantic analysis on the serialized data includes extracting identifiers, operators, functions, and table names and field names involved in the SQL query request; performing logical analysis on the serialized data includes extracting node types, node attributes, and child node relationships contained in the SQL query request.

[0074] S4.2. Obtain the metadata information of the target database, bind the semantic information contained in the SQL query request to the metadata of the target database, and ensure that the objects involved in the query exist in the target database and meet the expectations.

[0075] S4.3. Invoke the preset sensitivity evaluation algorithm, calculate the sensitivity based on the logical structure of the SQL query request and the bound metadata information, obtain the sensitivity of the SQL query request, and use the obtained sensitivity to add data noise to the SQL query result.

[0076] The preset sensitivity evaluation algorithm is at least one of the rule-based evaluation, weight-based evaluation, historical data-based evaluation, and graph structure-based evaluation. Among them: ① The rule-based evaluation algorithm defines a set of rules and performs matching and scoring based on features such as the operation type, table name, and field name included in the SQL query request. For example, some operations (such as DELETE, UPDATE) may be more sensitive than SELECT; ② The statistics-based evaluation algorithm uses historical data or metadata information to perform statistical analysis on the SQL query request. For example, it statistically analyzes the access frequency and update frequency of a certain table to evaluate its sensitivity; ③ The machine learning-based evaluation algorithm uses a machine learning model to classify and score the SQL query request. Through the training dataset, the machine learning model can learn which feature combinations correspond to high-sensitivity or low-sensitivity SQL queries; ④ The graph structure-based evaluation algorithm represents the SQL query request as a graph structure, where nodes represent tables and fields, and edges represent relationships. By analyzing the structural features of the graph (such as path length, connection complexity, etc.), its sensitivity is evaluated; ⑤ The risk matrix-based evaluation algorithm defines a risk matrix, looks up the corresponding risk values according to the tables, fields, and their operation types involved in the SQL query request, and calculates the total risk score; ⑥ The context-based evaluation algorithm considers the context information of the SQL query request, such as user role, current time, business logic, etc., and dynamically adjusts the sensitivity evaluation result.

[0077] In summary, by using the method for calculating the sensitivity of SQL statements in a multi-source heterogeneous database of the present invention, it is possible to efficiently be compatible with the query dialects of different types of databases, accurately calculate the sensitivity of complex SQL query statements, solve the problem of privacy leakage in query statistics and user data collection scenarios, and achieve the goal of high availability of data on the basis of protecting user privacy during the query and collection processes.

[0078] The above application of specific examples has elaborated in detail the principle and implementation manner of the present invention. These embodiments are only used to help understand the core technical content of the present invention. Based on the above specific embodiments of the present invention, those skilled in the art of this technology, without departing from the principle of the present invention, any improvements and modifications made to the present invention shall fall within the scope of patent protection of the present invention.

Claims

1. A method for calculating the sensitivity of SQL statements in a multi-source heterogeneous database, characterized in that, Its implementation includes the following four stages: S1. SQL Receiving and Parsing Stage: Receive the user's SQL query request, check and preprocess the SQL query request, then disassemble it into a set of lexical units, and then combine them to generate an abstract syntax tree, and verify the abstract syntax tree; This process specifically includes: S1.

1. Receive the user's SQL query request, and the SQL query request supports multiple SQL dialects, including standard SQL and the dialects of databases MySQL, Oracle, SQL Server, and PostgreSQL; S1.

2. Check the SQL query request. The checking process includes verifying whether the SQL query request is a valid SQL statement, whether it contains potential malicious code or SQL injection attempts through regular expression matching, blacklist / whitelist filtering, or a preset SQL injection prevention mechanism; S1.

3. After the check passes, perform preprocessing operations on the SQL query request, such as removing comments, normalizing whitespace, and unifying case; S1.

4. Disassemble the preprocessed SQL query request into a set of lexical units, and combine this set of lexical units into an abstract syntax tree according to the corresponding grammar rules of the SQL query request; S1.

5. Verify the abstract syntax tree, including: verifying whether the table names and column names involved in the abstract syntax tree nodes exist in the target database; checking whether the user has the permissions required to execute the query; S2. Abstract Syntax Tree Rewriting Stage: Traverse the abstract syntax tree, identify the syntax features that need to be specified for processing; define a set of rewriting rules for the syntax features that need to be specified for processing; modify and optimize the corresponding nodes in the abstract syntax tree according to the defined rewriting rules; S3. Serialization Stage: Select the serialization format of the abstract syntax tree, and convert the optimized abstract syntax tree into the corresponding serialized data; S4. Sensitivity Calculation Stage: Perform semantic analysis and logical analysis on the serialized data, bind the semantic information contained in the SQL query request to the metadata information of the target database, call a preset sensitivity evaluation algorithm, calculate the sensitivity of the SQL query request, and the calculated sensitivity is used to guide the addition of data noise in the SQL query result; This process specifically includes: S4.

1. Perform semantic analysis and logical analysis on the serialized data to obtain the semantic information and logical structure contained in the SQL query request. Among them, performing semantic analysis on the serialized data includes extracting the identifiers, operators, functions, and the involved table names and field names contained in the SQL query request; performing logical analysis on the serialized data includes extracting the node types, node attributes, and child node relationships contained in the SQL query request; S4.

2. Obtain the metadata information of the target database, and bind the semantic information contained in the SQL query request to the metadata of the target database to ensure that the objects involved in the query exist in the target database and meet the expectations; S4.

3. Call the preset sensitivity evaluation algorithm to calculate the sensitivity of the SQL query request based on the logical structure of the SQL query request and the bound metadata information, and use the obtained sensitivity to add data noise to the SQL query result.

2. A method for calculating the sensitivity of SQL statements in a multi-source heterogeneous database according to claim 1, characterized in that, In step S1.3, preprocessing operations of removing comments, normalizing spaces, and unifying case for the SQL query request are performed, where: The comment removal includes: traversing the SQL query string and removing single-line comments and multi-line comments; The space normalization includes: standardizing the whitespace characters in the SQL query request to ensure consistency and accuracy during parsing; The case unification includes: unifying the case of keywords and identifiers in the SQL query request according to the type of the SQL query request.

3. A method for calculating the sensitivity of SQL statements in a multi-source heterogeneous database according to claim 1, characterized in that, The specific operation of executing step S1.4 to obtain the abstract syntax tree is as follows: S1.4.

1. Construct a lexical analyzer that splits the preprocessed SQL query string into a set of lexical units according to five categories: keywords, identifiers, operators, numbers, and string constants; during the splitting process, when the lexical analyzer finds an unrecognized character, it generates an error message and indicates the error location. S1.4.

2. According to the grammar rules corresponding to the SQL query request, construct a syntax analyzer that recognizes and parses subqueries, JOIN operations, and GROUP BY statements, thereby identifying the relationships between lexical units, generating corresponding abstract syntax tree nodes, and generating meta-information including the query type, table names involved, and column names. S1.4.

3. During the process of generating the abstract syntax tree, perform syntax checking, and when a syntax error is found, generate a detailed error message and indicate the specific location and reason of the error.

4. A method for calculating the sensitivity of SQL statements in a multi-source heterogeneous database according to claim 1, characterized in that Step S2 specifically includes: S2.

1. Traverse the abstract syntax tree to identify the syntax features that need to be specified for processing, and the syntax features that need to be specified for processing include aggregate functions, multi-layer subqueries, complex JOIN operations, GROUP BY, and ORDER BY clauses; S2.

2. Define a set of rewrite rules for the syntax features that need to be specified for processing; S2.

3. According to the defined rewrite rules, modify the corresponding nodes in the abstract syntax tree. The specific modification operations include: ① replacing the node, replacing the old function node with a new function expression; ② adding nodes to the node, adding additional nodes required for the noise generation logic; ③ deleting nodes from the node, deleting redundant nodes that do not affect the query result but may disclose privacy. S2.

4. After applying all necessary rewrite rules to modify the abstract syntax tree, further optimize the abstract syntax tree to improve the execution efficiency of the query. Among them, the optimization measures include: identifying and removing duplicate or unnecessary calculation logic, selecting relatively efficient indexes, adjusting the order of join operations, and simplifying expressions.

5. A method for calculating the sensitivity of SQL statements in a multi-source heterogeneous database according to claim 1, characterized in that, Step S3 specifically includes: S3.

1. Select the serialization format of the abstract syntax tree according to the specific requirements and performance considerations of the system. The serialization formats of the abstract syntax tree include JSON, XML, Protocol Buffers, YAML, Avro, MessagePack, and CBOR; S3.

2. Use the selected serialization format to convert the optimized abstract syntax tree into corresponding serialized data to ensure the integrity and consistency of the query request during cross-system or component transmission.

6. A method for calculating the sensitivity of SQL statements in a multi-source heterogeneous database according to claim 5, characterized in that, When performing step S3.2, it is necessary to ensure that all information in the abstract syntax tree can be accurately encoded into the serialized data, including node types, node attributes, and child node relationships.

7. A method for calculating the sensitivity of SQL statements in a multi-source heterogeneous database according to claim 5, characterized in that After performing step S3 and converting the optimized abstract syntax tree into corresponding serialized data: Add necessary meta-information to the serialized data to assist subsequent parsing and processing. The necessary meta-information includes SQL dialect information and a list of rules applied during the rewriting process; Verify the generated serialized data to ensure that the serialized data completely contains all information of the abstract syntax tree and does not introduce any unexpected errors or inconsistencies.

8. A method for calculating the sensitivity of SQL statements in a multi-source heterogeneous database according to claim 1, characterized in that, The preset sensitivity assessment algorithm is at least one of rule-based assessment, statistics-based assessment, machine learning-based assessment, graph structure-based assessment, risk matrix-based assessment, and context-based assessment.

Citation Information

Patent Citations

  • Data desensitization method and device, computer readable storage medium and electronic equipment

    CN112560100A

  • Differential privacy noise adding method and device, medium and electronic equipment

    CN118246055A

  • SQL statement generation method and device, computer equipment and storage medium

    CN118568123A