Intelligent database interaction method and intelligent chain database system
By integrating open source large models and databases, the database query capabilities and data understanding are enhanced, and the existing database is difficult to deal with natural language query in large-scale data and complex query scenarios, and an intelligent and humanized interactive interface is realized, simplifying the data access process and lowering the threshold for use.
Patent Information
- Application Number
- CN202510209666.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-30
AI Technical Summary
In large-scale data and complex query scenarios, existing databases are difficult to efficiently handle users' natural language query needs, and the threshold for use is high and cannot meet the growing user needs.
By integrating open source models and databases, we can enhance the query capabilities and data understanding of the database, and provide an intelligent and humanized interactive interface. Specific steps include selecting and deploying large models, database connection and data preparation, model and database interaction, security control and performance optimization, user interface and feedback loops, deployment and maintenance, and continuous learning and updates.
It realizes the understanding of users' natural language instructions and automatically converts them into accurate SQL query statements, simplifies the data access process, lowers the threshold for database use, and provides a more intelligent and efficient data query and analysis experience.
Smart Images

Figure CN120067136A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and more specifically, to an intelligent database interaction method and an intelligent chain database system. Background Art
[0002] With the wide application of databases, for data query, especially in the scenarios of large-scale data and complex queries, the data access process is complex. The database has high requirements for the standardization and accuracy of query statements, has poor ability to understand natural language, cannot efficiently process users' query requirements, and has a high threshold for database use, which cannot meet the growing user needs. Summary of the Invention
[0003] The technical task of the present invention is to address the above deficiencies and provide an intelligent database interaction method and an intelligent chain database system, which can significantly improve the ability to quickly locate and solve problems, provide a more intelligent and user-friendly interaction interface, and provide users with a more intelligent and efficient data query and analysis experience.
[0004] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0005] An intelligent database interaction method, by integrating an open-source large model and a database, enhances the query ability and data understanding of the database, and provides an intelligent and user-friendly interaction interface; the implementation of this method includes the following steps:
[0006] 1) Preparation of the development environment;
[0007] 2) Selection and deployment of the large model: Based on the open-source large model for model deployment, including model selection, model deployment, training and fine-tuning, rule definition, sequence-to-sequence conversion, and reinforcement learning;
[0008] 3) Database connection and data preparation, including database initialization and data import;
[0009] 4) Interaction between the model and the database;
[0010] 5) Security control and performance optimization;
[0011] 6) User interface and feedback loop, create an intuitive user interface, allowing users to input queries in natural language; establish a user feedback loop to collect users' evaluations of query results for continuous improvement of model performance;
[0012] 7) Deployment and maintenance;
[0013] 8) Continuous learning and update.
[0014] This method provides users with a more intelligent and efficient data query and analysis experience by integrating open-source large models and databases. Through deep learning and natural language processing technologies, it can understand users' natural language instructions, simplify the data access process, lower the threshold of database use, and provide continuous monitoring and updates to keep the functions highly available and intelligent.
[0015] Furthermore, the development environment for implementing this method needs to meet the following conditions:
[0016] Python and pip are the latest versions;
[0017] Docker is installed for containerized deployment;
[0018] MongoDB or other compatible database systems are ready.
[0019] Furthermore, the open-source large models include Bert, RoBERTa, or GPT series with the Transformers architecture, and these models perform excellently in NLP tasks;
[0020] The deployment model: Deploy the model using Docker Compose; refer to the docker-compose.yml file to start the model service and adjust it according to the actual production environment selected by the user;
[0021] The training and fine-tuning: Use the labeled dataset to pre-train and fine-tune the model to recognize and understand the natural language query intent input by the user;
[0022] The rule definition: Define the mapping rules from natural language to SQL, including keyword matching, semantic parsing, and context understanding;
[0023] The sequence-to-sequence (Seq2Seq) conversion: Use the sequence-to-sequence conversion Seq2Seq model to convert the semantic representation output by the NLU module into an SQL statement, which usually involves an encoder-decoder architecture;
[0024] The reinforcement learning (RL): Optimize the accuracy and efficiency of the generated SQL statement through reinforcement learning RL to ensure consistency with the user's intention Figure 1 .
[0025] Furthermore, for step 3),
[0026] The database initialization: Initialize the database using postgresql or other database services as needed;
[0027] The data import: Import the user's production data into the database to ensure good data quality and clear structure.
[0028] Furthermore, the interaction between the model and the database is specifically implemented as follows:
[0029] Build an intermediate layer: Create an intermediate layer or interface to enable the large model to read from and write to the database; this may involve using database adapters of LangChain or similar libraries;
[0030] Obtain the database table structure: Use a parser developed based on the get_table_schema() function or other methods to extract table structure information from the database, read the table structure and relationships of the database so that the model can understand the data layout;
[0031] Generate SQL queries: When receiving a natural language query, use the large model to parse the query intent and generate corresponding SQL statements;
[0032] Execute the query: Build a secure execution environment for executing the SQL statements generated by the NLU and SQL generation modules, while preventing SQL injection attacks; Send the generated SQL statements to the database for execution and capture the results;
[0033] Result parsing: Convert the query results back into a natural language form for easy understanding by users.
[0034] Furthermore, the security control: Implement strict permission control and SQL statement verification to avoid risks of malicious operations and data leakage; Ensure that all database operations follow appropriate security policies to avoid attacks such as SQL injection;
[0035] The performance optimization: Integrate query optimization algorithms, including cost baseline optimization, index usage analysis, etc., to improve query efficiency; Monitor query performance and adjust the configuration of the large model or database indexes when necessary to improve the response speed.
[0036] Furthermore, the deployment and maintenance include:
[0037] Cloud / local deployment options: Provide flexible deployment methods, which can be used as cloud services or run on local servers;
[0038] Continuous monitoring and update: Regularly check the system performance, collect runtime data for model iteration and software upgrade;
[0039] The continuous learning and update include:
[0040] Model update: Regularly update the large model to reflect the latest progress in language understanding and database technology;
[0041] Data update: Keep the database data up-to-date and accurate, and regularly perform data cleaning and maintenance.
[0042] The present invention also claims to protect an intelligent chain database system, which enhances the query ability and data understanding of the database through the above method and establishes an intelligent and user-friendly interaction interface; the system includes:
[0043] A large model deployment module for deploying a model based on an open-source large model, including model selection, model deployment, training and fine-tuning, rule definition, sequence-to-sequence conversion, and reinforcement learning;
[0044] A database connection and data preparation module, including database initialization and data import;
[0045] An interaction module between the model and the database;
[0046] A security control and performance optimization module;
[0047] A user interface and feedback loop module;
[0048] A deployment and maintenance module;
[0049] A continuous learning and updating module.
[0050] The present invention also claims to protect an intelligent database interaction device, which is characterized by including at least one memory and at least one processor;
[0051] The at least one memory is used to store machine-readable programs;
[0052] The at least one processor is used to call the machine-readable program to implement the above method.
[0053] The present invention also claims to protect a computer-readable medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the processor implements the above method.
[0054] Compared with the prior art, an intelligent database interaction method and an intelligent chain database system of the present invention have the following beneficial effects:
[0055] The present invention ingeniously integrates an advanced open-source large model with a traditional database system, providing users with a more intelligent and efficient data query and analysis experience. Through deep learning and natural language processing technologies, the intelligent chain assistant can understand users' natural language instructions and automatically convert them into accurate SQL query statements, greatly simplifying the data access process and lowering the threshold for database use.
[0056] Combining a large model (such as GPT series models or other advanced language models) with a database can significantly improve the ability to quickly locate and solve problems, especially in scenarios of large-scale data and complex queries. At the same time, a more intelligent and user-friendly interaction interface can be provided. Description of the Drawings
[0057] Figure 1 It is a flowchart of the intelligent database interaction method provided by the embodiments of the present invention. Detailed Embodiments
[0058] The present invention will be further described below in conjunction with specific embodiments.
[0059] The embodiments of the present invention provide an intelligent database interaction method, which enhances the query ability and data understanding of the database by integrating open-source large models and databases, and provides an intelligent and user-friendly interaction interface; the implementation of this method includes the following steps:
[0060] 1. Preparation of the development environment:
[0061] First of all, the development environment needs to meet the following conditions:
[0062] The latest versions of Python and pip.
[0063] Docker is installed for containerized deployment.
[0064] MongoDB or other compatible database systems are ready.
[0065] 2. Selection and deployment of the large model:
[0066] Based on open-source large models for model deployment, including model selection, model deployment, training and fine-tuning, rule definition, sequence-to-sequence conversion, and reinforcement learning.
[0067] The model selection: Based on open-source large models, such as Bert, RoBERTa, or GPT series of the Transformers architecture, which perform excellently in NLP tasks.
[0068] The model deployment: Use Docker Compose to deploy the model; refer to the docker-compose.yml file to start the model service, which needs to be adjusted according to the actual production environment selected by the user;
[0069] The training and fine-tuning: Use the labeled dataset to pre-train and fine-tune the model to identify and understand the natural language query intent of the user input.
[0070] The rule definition: Define the mapping rules from natural language to SQL, including keyword matching, semantic parsing, and context understanding.
[0071] The sequence-to-sequence (Seq2Seq) conversion: Use the Seq2Seq model to convert the semantic representation output by the NLU module into an SQL statement, which usually involves an encoder-decoder architecture.
[0072] The reinforcement learning (RL): Optimize the accuracy and efficiency of the generated SQL statement through RL to ensure consistency with the user's intention. Figure 1 Consistency.
[0073] 3. Database connection and data preparation, including database initialization and data import:
[0074] The database initialization: Initialize the database using postgresql or other database services as needed.
[0075] The data import: Import the user-generated data into the database to ensure good data quality and clear structure.
[0076] 4. Interaction between the model and the database:
[0077] Build an intermediate layer: Create an intermediate layer or interface to enable the large model to read and write to the database; this may involve using a database adapter of LangChain or a similar library.
[0078] Obtain the database table structure: Use a parser developed based on the get_table_schema() function or other methods to extract the table structure information from the database, read the table structure and relationships of the database so that the model can understand the data layout.
[0079] Generate an SQL query: When receiving a natural language query, use the large model to parse the query intention and generate the corresponding SQL statement.
[0080] Execute the query: Build a secure execution environment to execute the SQL statements generated by the NLU and SQL generation modules, while preventing SQL injection attacks. Send the generated SQL statement to the database for execution and capture the results;
[0081] Result parsing: Convert the query result back to a natural language form for easy user understanding.
[0082] 5. Security control and performance optimization:
[0083] Security control: Implement strict permission control and SQL statement verification to avoid risks of malicious operations and data leakage; ensure that all database operations follow appropriate security policies to avoid attacks such as SQL injection.
[0084] Performance Optimization: Integrate query optimization algorithms, including cost baseline optimization, index usage analysis, etc., to improve query efficiency; monitor query performance and adjust the configuration of the large model or database index when necessary to improve the response speed.
[0085] 6. User Interface and Feedback Loop:
[0086] UI Design: Create an intuitive user interface that allows users to enter queries in natural language.
[0087] Feedback Mechanism: Establish a user feedback loop to collect user evaluations of query results for continuous improvement of model performance.
[0088] 7. Deployment and Maintenance:
[0089] Cloud / Local Deployment Options: Provide flexible deployment methods, which can be used as cloud services or run on local servers.
[0090] Continuous Monitoring and Update: Regularly check system performance and collect runtime data for model iteration and software upgrade.
[0091] 8. Continuous Learning and Update:
[0092] Model Update: Regularly update the large model to reflect the latest progress in language understanding and database technology;
[0093] Data Update: Keep the database data up-to-date and accurate, and regularly perform data cleaning and maintenance.
[0094] This method provides a more intelligent and efficient data query and analysis experience for users by integrating open-source large models and databases. Through deep learning and natural language processing technologies, it can understand users' natural language instructions, simplify the data access process, lower the threshold of database use, and provide continuous monitoring and update to keep the functions highly available and intelligent.
[0095] The embodiment of the present invention also provides an intelligent chain database system, which enhances the query ability and data understanding of the database through the intelligent database interaction method described in the above embodiment and establishes an intelligent and user-friendly interaction interface. The system includes:
[0096] 1. Large Model Deployment Module, which deploys models based on open-source large models, including model selection, model deployment, training and fine-tuning, rule definition, sequence-to-sequence conversion, and reinforcement learning.
[0097] The model selection: Based on open-source large models, such as Bert, RoBERTa, or GPT series of the Transformers architecture, which perform excellently in NLP tasks.
[0098] The deployment model: Use the Docker Compose deployment model; refer to the docker-compose.yml file to start the model service, which needs to be adjusted according to the actual production environment selected by the user;
[0099] The training and fine-tuning: Use the labeled dataset to pre-train and fine-tune the model to recognize and understand the natural language query intention input by the user.
[0100] The rule definition: Define the mapping rules from natural language to SQL, including keyword matching, semantic parsing, and context understanding.
[0101] The sequence-to-sequence (Seq2Seq) conversion: Use the Seq2Seq model of sequence-to-sequence conversion to convert the semantic representation output by the NLU module into an SQL statement, which usually involves an encoder-decoder architecture.
[0102] The reinforcement learning (RL): Optimize the accuracy and efficiency of the generated SQL statement through RL to ensure consistency with the user's intention. Figure 1 consistent.
[0103] 2. Database connection and data preparation module, including database initialization and data import;
[0104] The database initialization: Initialize the database using postgresql or other database services as needed.
[0105] The data import: Import the user's production data into the database to ensure good data quality and clear structure.
[0106] 3. Model and database interaction module, including:
[0107] Build the middle layer: Create a middle layer or interface to enable the large model to read and write to the database; this may involve using a database adapter of LangChain or a similar library.
[0108] Obtain the database table structure: Use a parser developed based on the get_table_schema() function or other methods to extract the table structure information from the database, read the table structure and relationships of the database so that the model can understand the data layout.
[0109] Generate SQL queries: When receiving a natural language query, use the large model to parse the query intention and generate the corresponding SQL statement.
[0110] Execute the query: Build a secure execution environment to execute the SQL statements generated by the NLU and SQL generation modules, while preventing SQL injection attacks. Send the generated SQL statement to the database for execution and capture the results;
[0111] Result analysis: Convert the query result back to a natural language form for easy user understanding.
[0112] 4. Security control and performance optimization module, including:
[0113] Security control: Implement strict permission control and SQL statement verification to avoid malicious operations and data leakage risks; Ensure that all database operations follow appropriate security policies to avoid attacks such as SQL injection.
[0114] Performance optimization: Integrate query optimization algorithms, including cost baseline optimization, index usage analysis, etc., to improve query efficiency; Monitor query performance and adjust the configuration of the large model or database index when necessary to improve the response speed.
[0115] 5. User interface and feedback loop module, including:
[0116] UI design: Create an intuitive user interface that allows users to enter queries in natural language.
[0117] Feedback mechanism: Establish a user feedback loop to collect user evaluations of query results for continuous improvement of model performance.
[0118] 6. Deployment and maintenance module, including:
[0119] Cloud / local deployment options: Provide flexible deployment methods, which can be used as cloud services or run on local servers.
[0120] Continuous monitoring and update: Regularly check system performance and collect runtime data for model iteration and software upgrade.
[0121] 7. Continuous learning and update module, including:
[0122] Model update: Regularly update the large model to reflect the latest progress in language understanding and database technology;
[0123] Data update: Keep the database data up-to-date and accurate, and regularly perform data cleaning and maintenance.
[0124] An embodiment of the present invention also provides an intelligent database interaction device, characterized by including at least one memory and at least one processor;
[0125] The at least one memory is used to store machine-readable programs;
[0126] The at least one processor is used to call the machine-readable program to implement the intelligent database interaction method described in the above embodiment.
[0127] An embodiment of the present invention further provides a computer-readable medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the processor is caused to execute the intelligent database interaction method described in the above embodiment. Specifically, a system or device equipped with a storage medium can be provided, on which software program codes for implementing the functions of any one of the above embodiments are stored, and the computer (or CPU or MPU) of the system or device is caused to read and execute the program codes stored in the storage medium.
[0128] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments. Therefore, the program code and the storage medium storing the program code constitute a part of the present invention.
[0129] Embodiments of the storage medium for providing program codes include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer via a communication network.
[0130] In addition, it should be clear that not only can the functions of any one of the above embodiments be implemented by executing the program code read by the computer, but also by causing an operating system or the like operating on the computer based on the instructions of the program code to complete part or all of the actual operations.
[0131] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU or the like installed on the expansion board or the expansion unit is caused to execute part and all of the actual operations, thereby implementing the functions of any one of the above embodiments.
[0132] The present invention has been described in detail above with reference to the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above-mentioned multiple embodiments, those skilled in the art can know that more embodiments of the present invention can be obtained by combining the code review means in the above different embodiments, and these embodiments are also within the protection scope of the present invention.
Claims
1. An intelligent database interaction method, characterized in that: By integrating the open source big model with the database, the query capability and data understanding of the database are enhanced, and an intelligent and humanized interactive interface is provided; the implementation of this method includes the following steps: 1) Development environment preparation; 2) Select and deploy large models: Model deployment based on open source large models, including model selection, model deployment, training and fine-tuning, rule definition, sequence-to-sequence conversion, and reinforcement learning; 3) Database connection and data preparation, including database initialization and data import; 4) Interaction between model and database; 5) Safety control and performance optimization; 6) User interface and feedback loop: Create an intuitive user interface that allows users to enter queries in natural language; establish a user feedback loop to collect user evaluation of query results for continuous improvement of model performance; 7) Deployment and maintenance; 8) Continuous learning and updating.
2. The intelligent database interaction method according to claim 1, characterized in that: The development environment to implement this method needs to meet the following conditions: Python and pip are the latest versions; Docker is installed for containerized deployment; MongoDB or other compatible database systems are ready.
3. The intelligent database interaction method according to claim 1, characterized in that: The open source large models include Bert, RoBERTa or GPT series of Transformers architecture; The deployment model described: Use Docker Compose to deploy the model; refer to the docker-compose.yml file to start the model service and adjust it according to the production environment actually selected by the user; The training and fine-tuning: using a labeled dataset to pre-train and fine-tune the model to recognize and understand the natural language query intent input by the user; The rule definition: defines the mapping rules from natural language to SQL, including keyword matching, semantic parsing and context understanding; The sequence-to-sequence conversion: using the sequence-to-sequence conversion model to convert the semantic representation output by the NLU module into SQL statements; The reinforcement learning: optimizes the accuracy and efficiency of the generated SQL statements through reinforcement learning to ensure consistency with user intentions.
4. The intelligent database interaction method according to claim 1, characterized in that: For step 3), Initializing the database: using postgresql or other database services to initialize the database as needed; The data import is to import the user's production data into the database to ensure that the data is of good quality and has a clear structure.
5. The intelligent database interaction method according to claim 1, characterized in that: The interaction between the model and the database is specifically implemented as follows: Build a middle layer: Create a middle layer or interface that allows the large model to read and write to the database; Get the database table structure: Use a parser or other method developed based on the get_table_schema() function to extract table structure information from the database, read the database table structure and relationships so that the model can understand the data layout; Generate SQL queries: When receiving natural language queries, use the big model to parse the query intent and generate the corresponding SQL statements; Execute queries: Build a secure execution environment to execute SQL statements generated by the NLU and SQL generation modules while preventing SQL injection attacks; send the generated SQL statements to the database for execution and capture the results; Result parsing: Convert query results back into natural language form for easy understanding by users.
6. The intelligent database interaction method according to claim 1, characterized in that: The security controls described above: implement strict permission control and SQL statement verification to avoid malicious operations and data leakage; ensure that all database operations follow appropriate security policies to avoid SQL injection attacks; The performance optimization: integrates query optimization algorithms, including cost baseline optimization, index usage analysis, monitors query performance, and adjusts large model configurations or database indexes when necessary to improve response speed.
7. The intelligent database interaction method according to claim 1, characterized in that: The deployment and maintenance include: Cloud / local deployment options: Provides flexible deployment options, can be used as a cloud service or run on a local server; Continuous monitoring and updating: Regularly check system performance and collect runtime data for model iteration and software upgrades; The continuous learning and updating include: Model updates: Regularly update the big model to reflect the latest advances in language understanding and database technology; Data update: Keep database data up-to-date and accurate, and perform data cleanup and maintenance regularly.
8. A Zhilian database system, characterized in that: The system enhances the query capability and data understanding of the database through the method described in any one of claims 1 to 7, and establishes an intelligent and humanized interactive interface; the system includes: Large model deployment module, which performs model deployment based on open source large models, including model selection, model deployment, training and fine-tuning, rule definition, sequence-to-sequence conversion, and reinforcement learning; Database connection and data preparation module, including database initialization and data import; Model and database interaction module; Security control and performance optimization modules; User interface and feedback loop module; Deployment and maintenance modules; Continuous learning and updating modules.
9. An intelligent database interaction device, characterized in that: comprising at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is used to call the machine-readable program to implement the method described in any one of claims 1 to 7.
10. A computer-readable medium, characterized in that The computer readable medium stores computer instructions, which, when executed by a processor, enable the processor to implement the method according to any one of claims 1 to 7.
Citation Information
Cited By
Large model driven efficient NL2SQL conversion system and method based on big data training
CN120872999A