Artificial intelligence-based database optimization method

By combining query tree construction, semi-connection traversal, aggregation query and noise addition technologies, a privacy protection query optimization solution is designed, which solves the problems of low efficiency and insufficient privacy protection in traditional database query methods, and realizes efficient and secure database optimization, which is especially suitable for large-scale and high-complex employment planning data queries.

CN120336368APending Publication Date: 2025-07-18菏泽职业学院 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510450582.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Traditional database query methods are inefficient when facing complex queries and large-scale data, consume high resources, and insufficient privacy protection, making it difficult to balance the security of privacy and query efficiency.

Method used

Using a database optimization method based on artificial intelligence, combining query tree construction, semi-connection traversal, aggregation query and noise addition technologies, a privacy protection query optimization solution is designed, and through precise feature extraction and hierarchical decision tree optimization, query methods are intelligently selected.

Benefits of technology

It significantly improves query efficiency, ensures data privacy, improves the accuracy and response speed of query results. It is suitable for large-scale and high-complex employment planning data query, providing efficient and secure technical support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336368A_ABST
    Figure CN120336368A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of database optimization, provides a database optimization method based on artificial intelligence, and aims to solve the problems of low efficiency, high resource consumption and insufficient privacy protection in a traditional query method. According to the method, query tree construction, semi-connection traversal, aggregation query and noise addition technologies are combined, the query efficiency is remarkably improved, the Laplace noise and confusion circuit privacy protection technologies are utilized, the security of sensitive information is protected, and the effectiveness and reliability of a query result are ensured; through accurate feature extraction and intelligent decision tree optimization, a query method is automatically selected and optimized, and high precision and low resource consumption are realized in a complex query scene with a large data volume; the method provides efficient and safe technical support for intelligent analysis and decision-making of college employment planning data, and has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of database optimization, and particularly to a database optimization method based on artificial intelligence. Background Art

[0002] With the advent of the information age, the collection and analysis of college career planning data have become particularly important. However, with the continuous expansion of the data scale, how to efficiently and securely query these large datasets through a database has become an urgent problem to be solved. The traditional database query methods have the following problems: First, the traditional methods usually rely on techniques such as sequential scanning and index query. When facing complex queries and large-scale data, the query efficiency drops significantly, affecting the real-time performance and overall efficiency of the query. Second, the traditional methods often need to perform multiple intermediate calculations, generating a large number of irrelevant intermediate results, which not only fails to effectively shorten the query time but also increases the burden on the system. Finally, when processing career planning data, a large amount of personal privacy information is involved. The traditional methods often ignore privacy protection and it is difficult to strike a balance between ensuring privacy security and query efficiency, thus increasing the risk of data leakage. Summary of the Invention

[0003] The present invention provides a database optimization method based on artificial intelligence, aiming to solve the problems of low database query efficiency, high resource consumption, and insufficient privacy protection in the prior art. Specifically, the present invention first designs a privacy protection query optimization scheme by combining query tree construction, semi-join traversal, aggregation query, and noise addition techniques, thereby improving the query efficiency while ensuring the privacy of data. Secondly, with the help of accurate feature extraction and hierarchical decision tree judgment, the present invention can intelligently select the most suitable query method to improve the accuracy, response speed, and efficiency of query results. In summary, the present invention provides an efficient, secure, and database optimization method, which is particularly suitable for querying large-scale and high-complexity career planning data, providing solid technical support for college student career guidance.

[0004] The present invention provides a database optimization method based on artificial intelligence, which specifically includes the following steps:

[0005] Step S1: Dataset preparation and enhancement: Collect the career planning query dataset, perform data enhancement on the career planning query dataset to generate an enhanced query dataset;

[0006] Step S2: Feature extraction: Extract the query structure features, query statistical features, and query execution plan features of the enhanced query dataset, and integrate them to obtain a comprehensive feature set;

[0007] Step S3: Design query method: Construct a privacy-preserving query method by combining query tree construction technology, semi-join traversal, aggregation query, and noise addition technology, and perform database queries through the privacy-preserving query method;

[0008] Step S4: Query judgment: Construct a hierarchical decision tree by combining decision tree depth control, Lagrange multiplier method optimization, and greedy splitting method; Analyze the comprehensive feature set through the hierarchical decision tree to generate a trained hierarchical decision tree for selecting query methods and optimizing database query performance. The query methods include traditional query methods and privacy-preserving query methods;

[0009] Optimizing the database query performance through the hierarchical decision tree includes: query path optimization, high-concurrency query optimization, complex query optimization, privacy-preserving query optimization, and database resource scheduling optimization.

[0010] Further, step S2 specifically includes the following steps:

[0011] Step S21: Extract the query structure features of the enhanced query dataset through the Earley parsing algorithm, KMP algorithm, and maximum common subtree algorithm;

[0012] Step S22: Extract the query statistical features of the enhanced query dataset through the hash algorithm and K-means clustering algorithm;

[0013] Step S23: Extract the query execution plan features of the enhanced query dataset through depth-first search, breadth-first search, and simulated annealing algorithm.

[0014] Further, in the process of constructing the privacy-preserving query optimization method in step S3, it specifically includes the following steps:

[0015] Step S31: Receive the query statement and query target, judge the query statement, and construct a query connection tree. The query connection tree includes leaf nodes and root nodes;

[0016] Step S32: Start from the leaf nodes of the query connection tree and traverse the query connection tree in a bottom-up order; Perform semi-joins step by step from the leaf nodes to the root nodes; Optimize data transmission and retain the data related to the query target;

[0017] Step S33: Start from the root node of the query connection tree and traverse the query connection tree in a top-down order; Perform semi-joins step by step from the root node to the leaf nodes;

[0018] Step S34: Integrate all nodes through connection operations, and then generate an aggregated query result through aggregation operations; Protect the sensitive information in the aggregated query result through Laplace noise and garbled circuits to generate a privacy query result.

[0019] Further, the process of constructing the query connection tree in step S31 specifically includes the following steps:

[0020] Step S311: Receive the query statement, analyze the structure of the query statement, and determine whether the query statement is an acyclic join query;

[0021] Step S312: For the query statement that meets the acyclic join query, construct the query connection tree; for the query statement that does not meet the acyclic join query, convert it into an acyclic join query through the loop detection and decomposition strategy, and construct the query connection tree;

[0022] Step S313: During the process of constructing the query connection tree, introduce selective join and data partitioning to optimize the structure of the query connection tree; Selective join improves the execution efficiency of the query by optimizing the join order and conditions. Especially in complex queries, by reducing the number of joins and optimizing the join order, unnecessary calculations are avoided, and the execution time of the query is reduced; Data partitioning improves the execution efficiency of the query by splitting large datasets into small pieces for parallel processing; effectively reduces the burden on a single computing unit, reduces the scan of the entire dataset, and shortens the query execution time.

[0023] Further, the process of constructing the hierarchical decision tree in step S4 specifically includes the following steps:

[0024] Step S41: Construct a decision tree, define the current depth of the decision tree, and the decision tree includes a root node and leaf nodes;

[0025] Step S42: Introduce the foresight depth. When the foresight depth ≥ the current depth, in the part of the decision tree close to the root node, simultaneously control the complexity and accuracy of the decision tree through the Lagrange multiplier method, thereby splitting the hierarchical decision tree;

[0026] Step S43: When the foresight depth < the current depth, perform splitting in the part of the decision tree close to the leaf nodes through the greedy method;

[0027] Step S44: After the splitting in step S42 and step S43 is completed, generate the hierarchical decision tree.

[0028] Adopting the above solution, the beneficial effects obtained by the present invention are as follows:

[0029] The present invention provides a database optimization method based on artificial intelligence, which combines artificial intelligence and privacy protection technology to achieve efficient and secure optimization of database queries. First, by introducing query tree construction, semi-join traversal, aggregation query, and noise addition technology, the database query process is optimized, significantly improving the efficiency of data query. Based on the traditional database query method, the privacy protection query optimization scheme of the present invention can avoid the efficiency loss existing in the traditional privacy protection technology while ensuring data privacy. By using Laplace noise and garbled circuit privacy protection technology, it further ensures that sensitive information in the query process is not leaked, while maintaining the validity and reliability of the query results. The technology that combines privacy protection and query optimization not only enhances the security of data query, but also provides a solid guarantee for the long-term use of university employment planning data, laying a foundation for intelligent analysis and decision support of data.

[0030] Secondly, the present invention realizes the automatic selection and optimization of query methods through precise feature extraction and intelligent decision tree optimization. By deeply analyzing the query structure, statistical features, and execution plan features of the enhanced query dataset, the query strategy is optimized to ensure high query accuracy and low resource consumption in complex and large-scale data query scenarios. In terms of intelligent decision tree optimization, the present invention adopts depth control, Lagrange multiplier method, and greedy splitting method to control the depth and splitting method of the decision tree, enabling the decision tree to be flexibly adjusted in different query scenarios, thereby improving the adaptability and overall performance of the database query method.

[0031] In summary, the present invention provides a database optimization method integrating artificial intelligence and privacy protection technology, which significantly improves the efficiency and security of data query, and has broad application prospects and important practical value. Brief Description of the Drawings

[0032] Figure 1 It is a schematic flow chart of a database optimization method based on artificial intelligence provided by the present invention;

[0033] Figure 2 It is a schematic structural diagram of a hierarchical decision tree in Embodiment 6. Detailed Embodiments

[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0035] Embodiment 1, according to Figure 1, the present invention provides an artificial intelligence-based database optimization method, which specifically includes the following steps:

[0036] Step S1: Dataset preparation and enhancement: Collect the employment planning query dataset, perform data enhancement on the employment planning query dataset, and generate an enhanced query dataset;

[0037] Step S2: Feature extraction: Extract the query structure features, query statistical features, and query execution plan features of the enhanced query dataset, and integrate them to obtain a comprehensive feature set;

[0038] Step S3: Design query method: Combine query tree construction technology, semi-join traversal, aggregation query, and noise addition technology to construct a privacy-preserving query method, and perform database query through the privacy-preserving query method;

[0039] Step S4: Query judgment: Combine decision tree depth control, Lagrange multiplier method optimization, and greedy splitting method to construct a hierarchical decision tree; Analyze the comprehensive feature set through the hierarchical decision tree to generate a trained hierarchical decision tree for selecting query methods and optimizing database query performance. The query methods include traditional query methods and privacy-preserving query methods;

[0040] Train the hierarchical decision tree through the comprehensive feature set to obtain a trained hierarchical decision tree;

[0041] Receive a query statement and a query target, and judge whether to use a traditional query method or a privacy-preserving query method through the trained hierarchical decision tree;

[0042] If it is a traditional query method, directly use it for query to generate a query result;

[0043] If it is a privacy-preserving query method, judge and perform privacy protection on the query statement to generate a privacy query result;

[0044] Optimizing the database query performance through the hierarchical decision tree includes: query path optimization, high-concurrency query optimization, complex query optimization, privacy-preserving query optimization, and database resource scheduling optimization;

[0045] Query path optimization (select index scan and full table scan to reduce query overhead);

[0046] High-concurrency query optimization (divide query tasks to improve concurrent processing ability);

[0047] Complex query optimization (optimize JOIN association query and aggregation query to improve execution efficiency);

[0048] Privacy-preserving query optimization (switch between traditional query and privacy-preserving query to ensure data security);

[0049] Database resource scheduling optimization (optimize CPU and IO resource allocation to improve query throughput).

[0050] Embodiment 2: This embodiment is based on Embodiment 1. In this embodiment, step S2 specifically includes the following steps:

[0051] Step S21: Extract the query structure features of the enhanced query dataset through the Earley parsing algorithm, KMP algorithm, and maximum common subtree algorithm;

[0052] Step S22: Extract the query statistical features of the enhanced query dataset through the hash algorithm and K-means clustering algorithm;

[0053] Step S23: Extract the query execution plan features of the enhanced query dataset through depth-first search, breadth-first search, and simulated annealing algorithm;

[0054] The query structure features include table join type, join tree depth, query field type, aggregation operation type, and inter-table redundancy;

[0055] The query statistical features include cache hit rate, data filtering rate, structural complexity, and data mobility;

[0056] The query execution plan features include scan type, Join method, and parallel execution feature.

[0057] Embodiment 3: This embodiment is based on Embodiment 2. In this embodiment, the process of performing database queries through the privacy-preserving query method in step S3 specifically includes the following steps:

[0058] Step S31: Receive the query statement and query target, and judge the query statement to construct a query connection tree, where the query connection tree includes leaf nodes and root nodes;

[0059] Step S32: Bottom-up semi-join traversal: Start from the leaf nodes of the query connection tree and traverse the query connection tree in a bottom-up order; perform semi-joins step by step from the leaf nodes to the root node; optimize data transmission and retain the data related to the query target;

[0060] Step S33: Top-down semi-join traversal: Start from the root node of the query connection tree and traverse the query connection tree in a top-down order; perform semi-joins step by step from the root node to the leaf nodes, eliminate irrelevant data, reduce computational overhead, and improve query efficiency;

[0061] Step S34: Connection traversal: Integrate all nodes through connection operations to generate a preliminary query result; perform an aggregation operation on the preliminary query result to generate an aggregated query result; protect the sensitive information in the aggregated query result through Laplace noise and garbled circuits to generate a privacy query result. The formula used is as follows:

[0062] ;

[0063] Among them, represents the query result after adding Laplace noise, represents the aggregated query result, represents the scale parameter for adding Laplace noise, represents the differential privacy parameter;

[0064] ;

[0065] Among them, represents the privacy query result, represents the operations of encryption and obfuscation;

[0066] Sensitive information includes health information, financial information, personal behavior patterns, social identity data, and social media data.

[0067] Example 4. This example is based on Example 3. In this example, step S31 specifically includes the following steps:

[0068] Step S311: Receive a query statement, analyze the structure of the query statement, and determine whether the query statement is an acyclic join query;

[0069] Step S312: For a query statement that conforms to an acyclic join query, construct a query join tree; for a query statement that does not conform to an acyclic join query, convert it into an acyclic join query through a loop detection and decomposition strategy, and construct a query join tree;

[0070] Step S313: During the process of constructing the query join tree, introduce selective join and data partitioning to optimize the structure of the query join tree; Selective join improves the execution efficiency of the query by optimizing the join order and conditions. Especially in complex queries, by reducing the number of joins and optimizing the join order, unnecessary calculations are avoided, and the execution time of the query is reduced; Data partitioning improves the execution efficiency of the query by splitting a large dataset into small chunks for parallel processing; effectively reduces the burden on a single computing unit, reduces the scan of the entire dataset, and shortens the query execution time.

[0071] Example 5. This example is based on Example 3. In this example, step S31 specifically includes the following steps:

[0072] Step Q1: Receive a query statement, analyze the structure of the query statement, and determine whether the query statement is an acyclic join query;

[0073] Step Q2: For a query statement that conforms to an acyclic join query, construct a query join tree; for a query statement that does not conform to an acyclic join query, convert it to an acyclic join query through a loop detection and decomposition strategy, and construct a query join tree.

[0074] Example 6. According to Figure 2 , this example is based on Example 4. In this example, in step S4, the process of constructing a hierarchical decision tree specifically includes the following steps:

[0075] Step S41: Construct a decision tree, define the current depth of the decision tree, and the decision tree includes a root node and leaf nodes;

[0076] Step S42: Introduce a foresight depth. When the foresight depth ≥ the current depth, in the part of the decision tree close to the root node, use the Lagrange multiplier method to simultaneously control the complexity and accuracy of the decision tree, thereby splitting the hierarchical decision tree. The formula used is as follows:

[0077] ;

[0078] Among them, represents the data set, that is, the enhanced query data set, represents the current depth of the decision tree, that is, the number of layers of the tree, represents the Lagrange multiplier, which is used to constrain the depth and computational cost of the decision tree, represents the optimization objective function, which controls the splitting process of the hierarchical decision tree; represents the sparsity penalty term, represents the layer index of the decision tree, represents the negative samples of the data set, represents the positive samples of the data set; represents the depth of the decision tree at layer , represents the data balance term, which avoids overfitting of the decision tree; represents the comprehensive feature set, represents the features of the comprehensive feature set, including query structure features, query statistical features, and query execution plan features, represents the weight of the feature cost, represents the feature cost term; represents the Lagrange multiplier, which adjusts the penalty intensity of the depth constraint, represents the Lagrange multiplier, which adjusts the penalty intensity of the computational cost, represents the constraint for controlling the depth of the decision tree, represents the constraint for controlling the computational cost;

[0079] Step S43: When the foresight depth < the current depth, perform splitting in the part of the decision tree close to the leaf nodes through a greedy method;

[0080] Step S44: After the splitting in Steps S42 and S43 is completed, generate a hierarchical decision tree.

[0081] Example 7: This example is based on Example 6. In this example, in Step S4, the process of training a hierarchical decision tree specifically includes the following steps:

[0082] Step W1: Divide the comprehensive feature set into a training set and a test set;

[0083] Among them, 80% is the training set and 20% is the test set;

[0084] Step W2: Use the training set to train the hierarchical decision tree and use the test set to evaluate the performance of the decision tree.

[0085] Example 8: This example is based on Example 7. In this example, receive a query statement and a query target, and use the trained hierarchical decision tree to determine whether to use a traditional query method or a privacy-preserving query method;

[0086] Example 1:

[0087] SELECT major, AVG(salary)

[0088] FROM employment data

[0089] WHERE graduation year = 2023

[0090] GROUP BY major;

[0091] Analysis results of the hierarchical decision tree:

[0092] Query structure: Standard GROUP BY statistical query;

[0093] Data sensitivity: No direct personal privacy data (average salary statistics, not involving individuals);

[0094] Query complexity: Moderate computational complexity, suitable for traditional SQL processing;

[0095] Query method selection: Traditional query method;

[0096] Example 2:

[0097] SELECT student ID, name, employment unit

[0098] FROM employment data

[0099] WHERE major = 'Computer Science' AND salary > 30000;

[0100] Analysis results of the hierarchical decision tree:

[0101] Query structure: SELECT to directly query sensitive information (student ID, name, employer);

[0102] Data sensitivity: Involves personal information and privacy (name, student ID, salary);

[0103] Query complexity: Medium complexity, but involves privacy protection;

[0104] Query method selection: Privacy-preserving query method.

[0105] The above describes the present invention and its implementation manners. Such description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and design, without creative efforts, a structural manner and an embodiment similar to the technical solution without departing from the gist of the present invention, they shall fall within the protection scope of the present invention.

Claims

1. An artificial intelligence-based database optimization method, characterized in that: The method specifically includes the following steps: Step S1: Dataset preparation and enhancement: Collect an employment planning query dataset, perform data enhancement on the employment planning query dataset to generate an enhanced query dataset; Step S2: Feature extraction: Extract the query structure features, query statistical features, and query execution plan features of the enhanced query dataset, and integrate them to obtain a comprehensive feature set; Step S3: Design a query method: Combine query tree construction techniques, semi-join traversal, aggregation queries, and noise addition techniques to construct a privacy-preserving query method, and perform database queries through the privacy-preserving query method; Step S4: Query judgment: Combine decision tree depth control, Lagrange multiplier method optimization, and greedy splitting methods to construct a hierarchical decision tree; Analyze the comprehensive feature set through the hierarchical decision tree to generate a trained hierarchical decision tree for selecting query methods and optimizing database query performance. The query methods include traditional query methods and privacy-preserving query methods.

2. The database optimization method based on artificial intelligence according to claim 1, characterized in that: In step S3, the process of performing database queries through the privacy-preserving query method specifically includes the following steps: Step S31: Receive a query statement and a query target, judge the query statement, and construct a query connection tree. The query connection tree includes leaf nodes and a root node; Step S32: Start from the leaf nodes of the query connection tree and traverse the query connection tree in a bottom-up order; Perform semi-joins step by step from the leaf nodes to the root node; Optimize data transmission and retain data related to the query target; Step S33: Start from the root node of the query connection tree and traverse the query connection tree in a top-down order; Perform semi-joins step by step from the root node to the leaf nodes; Step S34: Integrate all nodes through a connection operation, and then generate an aggregated query result through an aggregation operation; Protect sensitive information in the aggregated query result through Laplace noise and garbled circuits to generate a privacy query result.

3. The database optimization method based on artificial intelligence according to claim 2, characterized in that: In the process of constructing the query connection tree in step S31, it specifically includes the following steps: Step S311: Receive a query statement, analyze the structure of the query statement, and judge whether the query statement is an acyclic join query; Step S312: For query statements that meet the acyclic join query, construct a query connection tree; For those that do not meet the acyclic join query, convert them into acyclic join queries through loop detection and decomposition strategies and construct query connection trees; Step S313: During the process of constructing the query connection tree, introduce selective joins and data partitioning to optimize the query connection tree structure.

4. The database optimization method based on artificial intelligence according to claim 1, wherein: In step S4, the process of constructing the hierarchical decision tree specifically includes the following steps: Step S41: Construct a decision tree, define the current depth of the decision tree, and the decision tree includes a root node and leaf nodes; Step S42: Introduce a foresight depth. When the foresight depth ≥ the current depth, in the part of the decision tree close to the root node, simultaneously control the complexity and accuracy of the decision tree through the Lagrange multiplier method to split the hierarchical decision tree; Step S43: When the foresight depth < the current depth, perform splitting in the part of the decision tree close to the leaf nodes through a greedy method; Step S44: After the splitting in steps S42 and S43 is completed, generate a hierarchical decision tree.