Database query statement generation method, system and equipment and storage medium

By acquiring the user's natural language and using a large model to generate database query statements, the problems of database query complexity and low accuracy in the existing technology are solved, and efficient and accurate database query statement generation is achieved.

CN120670447APending Publication Date: 2025-09-19NORTH CHINA DIGITAL HEALTH TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510548877.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In the existing technology, database query statement generation relies on users to manually write SQL statements, which is highly professional and error-prone. The accuracy of the generative large model is not high in different database scenarios, and the existing improvement methods increase computing resources and time costs.

Method used

By obtaining user natural language, extracting keywords and obtaining database location information with natural language annotations, using large models to generate query statements, and improving accuracy through review and optimization processes, including database information acquisition, natural language annotation, query statement generation and error statement processing.

Benefits of technology

It improves the efficiency of large models in understanding database information, reduces data processing volume, reduces irrelevant interference, improves the accuracy of query statements, has a simple structure and has broad application prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670447A_ABST
    Figure CN120670447A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and particularly provides a database query statement generation method, system and device and a storage medium, and the method comprises the steps: obtaining a user natural language, and extracting keywords from the user natural language; data position information is obtained according to the keyword, the data position information comprises database position description information with a natural language annotation, and the natural language annotation and the keyword have an association relationship; and writing the data position information into a prompt, and generating a query statement based on the prompt by using a large model. By reducing the data processing amount of the large model, the interference of irrelevant data is reduced, and the accuracy of generating the query statement by the large model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and specifically relates to a database query statement generation method, system, device and storage medium. Background Art

[0002] In traditional data interaction models, database queries often rely on users manually writing SQL statements. This approach has significant limitations. It not only requires high professional skills but is also prone to errors during actual operation, making database queries a significant obstacle for non-professional users.

[0003] In recent years, the rapid development of natural language processing (NLP) technology has brought new hope to database querying, making natural language querying a highly sought-after solution. However, current mainstream NLP methods mostly rely on predefined grammatical rules and templates. This rigid model is inadequate for complex query scenarios and fails to meet diverse practical needs.

[0004] The emergence of large generative models, such as large language models (LLMs), has provided a new approach for automatically generating queries. In theory, users can describe their requirements in natural language, and the LLM will automatically convert them into corresponding queries. However, in reality, queries generated by different databases in different application scenarios vary significantly. This results in unsatisfactory accuracy of LLM-generated queries, significantly reducing their effectiveness in practical applications.

[0005] To address the shortcomings of LLM in query generation, related fields have explored several improvement techniques. For example, deploying knowledge graphs or combining BERT with LLM to generate more standardized prompts aims to improve its performance. However, these approaches also introduce new challenges, significantly increasing the amount of model training and resource overhead. While improving query accuracy, they also place higher demands on computing resources and time, limiting their widespread application and promotion. Summary of the Invention

[0006] In view of the above-mentioned deficiencies in the prior art, the present invention provides a method, system, device and storage medium for generating a database query statement to solve the above-mentioned technical problems.

[0007] In a first aspect, the present invention provides a method for generating a database query statement, comprising: Acquire user natural language and extract keywords from the user natural language; Acquiring data location information according to the keyword, the data location information including database location description information with natural language annotations, wherein the natural language annotations are associated with the keyword; The data location information is written into a prompt, and a query statement is generated based on the prompt using a large model.

[0008] In an optional embodiment, the method further comprises: Confirm that there is an update in the database and obtain database information through database connection parameters. The database information includes database type, database name, table name, table description, field name, field description, field type, and some non-empty sample data for each field; Using the large model to gradually generate natural language annotations for the database information; The natural language annotations are reviewed, and the reviewed annotations and the database information to which the annotations belong are saved as database location description information in JSON format.

[0009] In an optional embodiment, using a large model to gradually generate natural language annotations for the database information includes: Encapsulating the database information into a first prompt, and using a large model to generate a natural language annotation of the field name of each field based on the field name and sample data of the first prompt; Encapsulate the database information, the field name, and the natural language annotation of each field into a second prompt, and use the large model to generate a natural language annotation of the table name for each data table based on the table name, table description information, and the natural language annotation of the field name of each field in the table in the second prompt; The database information, the table name and natural language label of each table, and the field name and natural language label of each field are encapsulated into a third prompt. The large model is used to generate a natural language label for the database name for each database based on the database name, database description information, the natural language label of the table name of each table in the database, and the natural language label of the field name of each field in each table in the third prompt.

[0010] In an optional embodiment, the natural language annotations are reviewed, and the reviewed annotations and the database information to which the annotations belong are saved as database location description information in JSON format, including: Use the natural language processing library to check the sentence structure, part of speech, and grammatical rules of natural language annotations; Verify the logical consistency of the natural language annotation of the database name with the natural language annotation of the subordinate table name and the natural language annotation of the field name; Analyze the characteristics of some non-empty sample data for each field, compare the annotation content with the characteristics of the sample data, and verify whether the two are consistent; Verify whether the natural language annotations conform to the preset standard template; The reviewed annotations and the database information to which the annotations belong are saved as database location description information in JSON format, and the database location description information in JSON format is updated to the database architecture description file.

[0011] In an optional embodiment, the data location information is written into a prompt, and a query statement is generated based on the prompt using a large model, including: Retrieving code data associated with the data location information; The code data and the data location information are written into a prompt, and the prompt is input into the large model to obtain a query statement.

[0012] In an optional embodiment, the method further comprises: Collect the erroneous query statements output by the large model that do not meet the query expectations, and save the data location information corresponding to the erroneous query statements and the corresponding correct query statements to the code library.

[0013] In an optional embodiment, the data location information is written into a prompt, and a query statement is generated based on the prompt using a large model, including: retrieving code data associated with the data location information from the code repository; Writing the code data and the data location information into a prompt, and inputting the prompt into the large model to obtain multiple query statements and predicted probabilities output by the large model; Counting the number of occurrences of multiple query statements contained in the code data, and calculating the occurrence frequencies of the multiple query statements based on the number of occurrences of the multiple query statements, and using the occurrence frequencies as predicted probabilities of the query statements; For any possible query statement, calculate its comprehensive probability as:

[0014] Where k is the weight factor, is the probability output by the large model, is the probability obtained by statistics;

[0015] in, It is a hyperparameter used to control the influence of error rate on the overall probability. e represents the error rate of the large model. Determine the minimum value of the hyperparameter and maximum value , and the minimum error rate and maximum value ; The mapping formula between hyperparameters and error rate is:

[0016] Calculate hyperparameters based on the current error rate of the large model.

[0017] In a second aspect, the present invention provides a database query statement generation system, comprising: An acquisition module, configured to acquire user natural language and extract keywords from the user natural language; a retrieval module, configured to obtain data location information according to the keyword, wherein the data location information includes database location description information with natural language annotations, wherein the natural language annotations are associated with the keyword; A generation module is used to write the data location information into a prompt and generate a query statement based on the prompt using a large model.

[0018] According to a third aspect, a device is provided, comprising: A memory, used for storing a database query statement generating program; A processor is used to implement the steps of the database query statement generation method provided in the first aspect when executing the database query statement generation program.

[0019] In a fourth aspect, a computer-readable storage medium is provided, on which a database query statement generation program is stored. When the database query statement generation program is executed by a processor, the steps of the database query statement generation method provided in the first aspect are implemented.

[0020] The beneficial effects of the present invention lie in the fact that the database query statement generation method, system, device, and storage medium provided herein generate natural language annotations for database information using a large model. The natural language annotations improve the efficiency of the large model's understanding of database information and reduce the large model's data processing load. Furthermore, by reducing the large model's data processing load and reducing interference from irrelevant data, the accuracy of query statements generated by the large model is improved.

[0021] In addition, the present invention has a reliable design principle, a simple structure and a very broad application prospect. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0023] Figure 1 is a schematic flow chart of a method according to an embodiment of the present invention.

[0024] Figure 2 FIG. 4 is a schematic block diagram of a system according to an embodiment of the present invention.

[0025] Figure 3 A schematic structural diagram of a device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0026] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0028] The key terms appearing in the present invention are explained below.

[0029] Large Language Models (LLMs), or Large Language Models, are a technology that has made significant progress in the field of artificial intelligence in recent years. These AI models are trained on large amounts of text data. They aim to learn the patterns, structure, and semantics of language to generate natural, fluent, and logically coherent text responses. They have a large number of parameters and are capable of handling a wide range of natural language processing tasks.

[0030] In the context of artificial intelligence, especially large language models (LLMs), "prompt" is usually translated as "prompt word". It is the input text provided to the model when interacting with the model, used to guide the model to generate output that meets specific requirements.

[0031] The database query statement generation method provided by the embodiments of the present invention is executed by a computer device. Correspondingly, the database query statement generation system runs in the computer device.

[0032] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention. Among them, Figure 1 The execution subject can be a database query statement generation system. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted.

[0033] Such as Figure 1 shown, this method includes: S1. Obtain the user's natural language and extract keywords from the user's natural language.

[0034] The user's natural language can be a question entered by the user in the search interface. For example, on a data analysis platform, the user enters "Query the top three products with the highest sales volume in the past month"; it can also be text information obtained from channels such as documents and voice transcripts, such as the demand description in a market research report or the data analysis requirements in a meeting record.

[0035] In order to accurately extract keywords from these natural languages, various methods can be used. A common way is to use natural language processing (NLP) toolkits, such as NLTK (Natural Language Toolkit) and SpaCy in Python. First, preprocess the obtained natural language, including removing stop words (such as words with no actual semantics like "of", "is", "in", etc.), stemming (还原单词为其基本形式) and词性标注. For example, for the sentence "Query the top three products with the highest sales volume in the past month", after removing stop words, it becomes "Query past one month sales highest three products", and then combined with词性标注, identify keywords with key semantics such as "Query", "sales volume", "products", "one month", etc. In addition, the TF-IDF (Term Frequency - Inverse Document Frequency) algorithm can also be used to evaluate the importance of each word in the text and extract words with higher importance as keywords.

[0036] S2. Obtain data location information according to the keywords. The data location information includes database location description information with natural language annotations, and the natural language annotations are associated with the keywords.

[0037] After obtaining keywords, we need to use them to locate relevant data locations. This requires a pre-built database indexing system that contains natural language-annotated database location descriptions. This indexing system can associate keywords with database location information, such as the database server address, database name, table name, and related fields.

[0038] When the system receives keywords, it searches for matches within this indexing system. For example, if the keywords are "sales" and "product," the system searches for all natural language annotations related to these terms. This annotation might describe the server where the database storing product sales data is located, specifically which table within that database. Furthermore, this annotation is described in natural language, such as "In the sales database, the product sales table stores detailed sales information for each product," facilitating subsequent interaction with the large model.

[0039] During the search and matching process, semantic matching technology can be used to improve matching accuracy. This goes beyond simple string matching and also considers the semantic similarity of terms. For example, "sales volume" and "sales amount" may have similar semantics in certain contexts. The system will treat them as related keywords for matching, thereby obtaining more comprehensive data location information associated with the keywords.

[0040] S3. Write the data location information into a prompt, and use the large model to generate a query statement based on the prompt.

[0041] After obtaining the data location information, incorporate this information into the prompt. The design of the prompt is crucial; it needs to clearly convey the specific requirements of the larger model and the relevant data location information. For example, the prompt could be: "In the sales database located at [server address], the product sales table stores detailed sales information for each product. Based on this information, please generate a SQL query to find the three products with the highest sales in the past month." After accurately entering the data location information into the prompt word, you can input this prompt word into the big model. The big model will understand and analyze the prompt word based on its learned language patterns and knowledge. It will combine database structure information, query requirements, and SQL syntax rules to generate the corresponding query statement. During the generation process, the big model will continuously adjust and optimize the statement to ensure that the generated query statement accurately retrieves the required data from the specified database location.

[0042] In one embodiment of the present invention, a method for generating natural language annotations includes: (1) Confirm that there is an update in the database and obtain database information through database connection parameters. The database information includes database type, database name, table name, table description, field name, field description, field type, and some non-empty sample data for each field.

[0043] First, establish a connection to the target database using a database connection library (e.g., Python's psycopg2 for PostgreSQL, mysql.connector for MySQL, pyodbc for SQL Server, etc.). You must provide database connection parameters, including the host name, port number, username, password, and database name. Then, use SQL queries (e.g., SELECT * FROM information_schema.tables or SELECT * FROM information_schema.columns ) to retrieve database schema information, including the database name, table name, table description, field names, field descriptions, and field types. To obtain sample data, execute SQL queries such as SELECT * FROM table_name LIMIT 10 for each table to obtain a small amount of non-empty sample data for subsequent analysis and model training.

[0044] (2) Using the large model to gradually generate natural language annotations for the database information.

[0045] 1) Encapsulate the database information into a first prompt, and use a large model to generate a natural language annotation of the field name of each field based on the field name and sample data of the first prompt.

[0046] Implementation example: First, construct a prompt: "You are a database expert. Please infer the Chinese meaning of the field based on the following information. Output the JSON representation in the format [field name: Chinese meaning]: \n\s\sField name: user_id \n\s\sField type: INT \n\s\sSample data: [1, 2, 3, 4, 5]." Next, submit the prompt to LLM to infer the Chinese meaning of the field. LLM will return: {"user_id": "User ID"}. Finally, store "User ID" in the schema and associate it with the user_id field.

[0047] 2) Encapsulate the database information, the field name, and the natural language annotation of each field into a second prompt, and use the large model to generate a natural language annotation of the table name for each data table based on the table name, table description information, and the natural language annotation of the field name of each field in the table in the second prompt.

[0048] Specific implementation example: First, construct a prompt: "You are a database expert. Please infer the Chinese meaning of the data table based on the following information. The output format is [field name: Chinese meaning] JSON expression: \n\s\sTable name: users\n\s\sField information: \n\s\s - user_id (user ID, INT) \n\s\s - username (user name, VARCHAR(255)) \n\s\s - email (email address, VARCHAR(255)) \n\s\s - created_at (creation time, TIMESTAMP)". Second, submit the prompt to LLM and let it infer the Chinese meaning of the field. LLM will return: {"users": "User table"}. Finally, we store the "user table" in the schema and associate it with the users table.

[0049] 3) Encapsulate the database information, the table name and natural language label of each table, and the field name and natural language label of each field into a third prompt, and use the large model to generate a natural language label for the database name for each database based on the database name, database description information, the natural language label of the table name of each table in the database, and the natural language label of the field name of each field in each table in the third prompt.

[0050] Specific implementation example: First, construct a prompt: "You are a database expert. Please infer the Chinese meaning of the database based on the following information. The output format is a JSON expression of [database name: Chinese meaning]: \n\s\sDatabase name: his_info \n\s\sTable information: \n\s\s - users (user information table) \n\s\s - visit_infos (visit information table) \n\s\s - diagnose_records (diagnostic record table)". Secondly, submit the prompt to LLM and let it infer the Chinese meaning of the field. LLM will return: {"his_info": "Hospital Information Management Database"}. Finally, we store "Hospital Information Management Database" in the schema and associate it with the his_info table.

[0051] (3) Reviewing the natural language annotations, and saving the reviewed annotations and the database information to which the annotations belong as database location description information in JSON format.

[0052] Through steps (1) and (2), we can obtain information such as the database library name, the Chinese meaning of the library, the table name, the Chinese meaning of the table, the primary key within a single table, the internal and external keys within a single table, the index of a single table, the name of each field, the Chinese meaning of each field, the type of each field, and the sample data for each field. Then, we manually annotate the schema information and review the Chinese meaning of the library, table, and field inferred by the computer. This annotation information can help the model better understand the semantics of the database.

[0053] Specific marking review methods include: 1) Use the natural language processing library to check the sentence structure, part of speech, and grammatical rules of natural language annotations.

[0054] 2) Verify the logical consistency of the natural language annotation of the database name with the natural language annotation of the subordinate table name and the natural language annotation of the field name; Verify the logical consistency of the natural language annotation of the database name with the natural language annotation of the subordinate table name and the natural language annotation of the field name to ensure that the entire database architecture is semantically coherent.

[0055] 3) Analyze the characteristics of some non-null sample data for each field, compare the annotations with the characteristics of the sample data, and verify their consistency. Perform statistical analysis on the sample data for each field, such as data type, value range, and distribution patterns. Compare the field annotations with the characteristics of the sample data to check for consistency. Check that the annotations and sample data for each field match.

[0056] 4) Verify that natural language annotations conform to pre-set standard templates. For example, the database annotation format is specified as "[Specific business domain] Database," the table annotation format is specified as "[Business object] Table," and the field annotation format is specified as "[Business attribute] [Data type]." Check that the annotations conform to the pre-set format.

[0057] The reviewed annotations and the database information to which the annotations belong are saved as database location description information in JSON format, and the database location description information in JSON format is updated to the database architecture description file.

[0058] Convert the annotated Schema information into JSON format. Finally, load the pre-trained LLM and use the converted JSON data for fine-tuning. The following is an example of JSON data: {"database_name": "his_info", "database_chinese_meaning": "Hospital Information Management Database", "tables": [ {"table_name": "users", "table_chinese_meaning": "User Information Table", "columns": [ {"column_name": "user_id", "column_chinese_meaning": "User ID", "column_type": "INT"}, "table_name": "visit_info", "table_chinese_meaning": "Medical Visit Information Table", "columns": [ {"column_name": "user_id", "column_chinese_meaning": "User ID", "column_type": "INT"}, {"column_name": "visit_id", "column_chinese_meaning": "Visit ID", "column_type": "INT"}, {"column_name": "visit_time", "column_chinese_meaning": "Visit Time", "column_type": "datetime"}, "table_name": "diagnose_records", "table_chinese_meaning": "Diagnosis Record Table", "columns": [ { "column_name": "visit_id", "column_chinese_meaning": "Visit ID", "column_type": "INT" }, {"column_name": "diagnose_name", "column_chinese_meaning": "Diagnosis Name", "column_type": "varchar(64)"}, / / ... additional fields]}]}.

[0059] In an embodiment of the present invention, based on step S2, a possible embodiment will be given below to illustrate its specific implementation scheme in a non-limiting manner.

[0060] If data location information is directly written into a prompt and then fed into a large model, the large model, after extensive training on a public database, can generate corresponding SQL queries. However, for some specialized application scenarios, such as medical databases, the SQL queries directly generated by the large model may not be appropriate due to the database's unique characteristics.

[0061] Therefore, this application discloses the following query statement query method: S301. Retrieve code data related to the data location information.

[0062] Receive input data location information, which should include the target data's possible database name, table name, field name, and so on. Extract and initially verify this information to ensure its integrity and legitimacy. For example, check that the database name, table name, and field name comply with the corresponding database naming conventions.

[0063] Based on the data location or other contextual information, determine the target database type, such as MySQL, PostgreSQL, Oracle, etc. Different database types may have differences in SQL syntax and features, so accurately identifying the database type is crucial.

[0064] If the code repository is located locally, specify the root directory of the code repository. You can use the os module to traverse this directory and its subdirectories. If the code repository is stored in a remote repository (such as GitHub or GitLab), clone it using a version control system (such as Git).

[0065] Build search keywords based on database type, library name, table name, and field name. Search for code snippets containing these keywords in code files, prioritizing SQL generation code related to the database type.

[0066] For more complex matching requirements, use regular expressions to precisely match SQL generated code in a specific format, such as matching code for operations such as creating tables, inserting data, and querying data.

[0067] S302. Write the code data and the data location information into a prompt, and input the prompt into the large model to obtain a query statement.

[0068] Clearly state the task in the prompt: generate a query statement suitable for a specific database type based on the given data location information and related code data. Add the data location information and retrieved code data to the prompt line by line to ensure that the large model has a comprehensive understanding of the context.

[0069] Use the API provided by the big model to interact and send the constructed prompt as input to the big model.

[0070] Perform syntax checks on query statements generated by large models to ensure they comply with the target database's syntax rules. Simple verification can be performed using the database's command-line tools or client.

[0071] If a query statement performs poorly during execution, optimize it based on the recommendations provided by the database's performance analysis tool, such as adding indexes or adjusting the query order. These poorly performing queries are also marked as incorrect and the corrected query statements, along with the corresponding data location information and code data, are saved to the code repository. The error rate of the large model is also calculated.

[0072] To further improve the adaptability of SQL query statements generated by the large model and make up for the lack of understanding of some highly specialized databases by the large model, the following improvement methods are provided: (1) Retrieving code data associated with the data location information from the code library.

[0073] In a specific example, the input data location information is parsed to identify key information contained within, such as the library name, table name, and field name of the target data. This information can be stored in a dictionary or object for later use. The code repository's storage location is determined. This could be a directory in the local file system or a remote code repository (such as a Git repository). If it's a remote repository, it must be cloned first. Using the parsed data location information, search keywords are constructed and the code repository is searched for code files containing these keywords. This can be achieved using file traversal and text matching.

[0074] (2) Writing the code data and the data location information into a prompt, and inputting the prompt into the large model to obtain multiple query statements and prediction probabilities output by the large model.

[0075] (3) Counting the number of occurrences of the plurality of query statements contained in the code data, and calculating the frequency of occurrence of the plurality of query statements based on the number of occurrences of the plurality of query statements, and using the frequency of occurrence as the predicted probability of the query statement.

[0076] For a query statement that appears in the code data, its statistical probability is: P2=n / N Where n is the number of occurrences of the query statement, and N is the total number of query statements contained in the code data.

[0077] (4) For any possible query statement, calculate its comprehensive probability as:

[0078] Where k is the weight factor, is the probability output by the large model, is the probability obtained by statistics;

[0079] in, It is a hyperparameter used to control the influence of error rate on the overall probability. e represents the error rate of the large model. Determine the minimum value of the hyperparameter and maximum value , and the minimum error rate and maximum value ; The mapping formula between hyperparameters and error rate is:

[0080] Calculate hyperparameters based on the current error rate of the large model.

[0081] By combining the historical performance of the big model, when the big model predicts and generates query statements, the historical error predictions are used to correct the prediction results of the big model, thereby improving the performance of the big model when processing special database codes.

[0082] In some embodiments, the database query statement generation system may include multiple functional modules composed of computer program segments. The computer program of each program segment in the database query statement generation system may be stored in a memory of a computer device and executed by at least one processor to perform (see Figure 1 Description) Function to generate database query statements.

[0083] In this embodiment, the database query statement generation system can be divided into multiple functional modules according to the functions it performs, such as Figure 2 As shown. The functional modules of the system may include: an acquisition module, a retrieval module, and a generation module. A module as referred to in the present invention refers to a series of computer program segments that can be executed by at least one processor and can perform fixed functions, and is stored in a memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0084] An acquisition module, configured to acquire user natural language and extract keywords from the user natural language; a retrieval module, configured to obtain data location information according to the keyword, wherein the data location information includes database location description information with natural language annotations, wherein the natural language annotations are associated with the keyword; A generation module is used to write the data location information into a prompt and generate a query statement based on the prompt using a large model.

[0085] Figure 3 The database query statement generation method provided for the embodiment of the present application can be applied to a device. Those skilled in the art will understand that the device structure involved in the embodiment of the present invention does not constitute a limitation on the device, and the device may include more or fewer components than shown, or combine certain components, or arrange different components. In an embodiment of the present invention, the device includes but is not limited to a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described and / or required herein.

[0086] The device 300 may include a processor 310, a memory 320, and a communication unit 330. These components communicate via one or more buses. Those skilled in the art will appreciate that the server structure shown in the figure does not limit the present invention. The server structure may be a bus structure or a star structure, and may include more or fewer components than shown, or combine certain components, or arrange the components differently.

[0087] The memory 320 can be used to store execution instructions of the processor 310. The memory 320 can be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk. When the execution instructions in the memory 320 are executed by the processor 310, the device 300 can perform some or all of the steps in the above-described method embodiments.

[0088] The processor 310 is the control center of the storage device, which uses various interfaces and lines to connect various parts of the entire electronic device. It executes various functions of the electronic device and / or processes data by running or executing software programs and / or modules stored in the memory 320, and calling data stored in the memory. The processor can be composed of an integrated circuit (IC), for example, it can be composed of a single packaged IC, or it can be composed of multiple packaged ICs with the same or different functions. For example, the processor 310 can only include a central processing unit (CPU). In an embodiment of the present invention, the CPU can be a single computing core or multiple computing cores.

[0089] The communication unit 330 is configured to establish a communication channel so that the storage device can communicate with other devices, receive user data sent by other devices, or send user data to other devices.

[0090] The present invention also provides a computer storage medium, wherein the computer storage medium may store a program that, when executed, may include some or all of the steps of each embodiment provided by the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0091] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software and a necessary general-purpose hardware platform. Based on this understanding, the technical solutions in the embodiments of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, among other media capable of storing program code, and includes instructions for causing a computer device (which can be a personal computer, a server, or a second device, a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention.

[0092] In this specification, the same or similar parts between the various embodiments can be referred to each other. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.

[0093] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of systems or modules, and can be electrical, mechanical or other forms.

[0094] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected to achieve the purpose of the present embodiment according to actual needs.

[0095] In addition, each functional module in each embodiment of the present invention may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0096] Although the present invention has been described in detail with reference to the accompanying drawings and in conjunction with preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, persons of ordinary skill in the art may make various equivalent modifications or substitutions to the embodiments of the present invention, and such modifications or substitutions shall be within the scope of the present invention. Any changes or substitutions that can be easily conceived by persons skilled in the art within the technical scope disclosed in the present invention shall be within the scope of protection of the present invention.

Claims

1. A method for generating a database query statement, characterized in that: include: Acquire user natural language and extract keywords from the user natural language; Acquiring data location information according to the keyword, the data location information including database location description information with natural language annotations, wherein the natural language annotations are associated with the keyword; The data location information is written into a prompt, and a query statement is generated based on the prompt using a large model.

2. The method according to claim 1, characterized in that The method further comprises: Confirm that there is an update in the database and obtain database information through database connection parameters. The database information includes database type, database name, table name, table description, field name, field description, field type, and some non-empty sample data for each field; Using the large model to gradually generate natural language annotations for the database information; The natural language annotations are reviewed, and the reviewed annotations and the database information to which the annotations belong are saved as database location description information in JSON format.

3. The method according to claim 2, characterized in that Using the large model to gradually generate natural language annotations for the database information, including: Encapsulating the database information into a first prompt, and using a large model to generate a natural language annotation of the field name of each field based on the field name and sample data of the first prompt; Encapsulate the database information, the field name, and the natural language annotation of each field into a second prompt, and use the large model to generate a natural language annotation of the table name for each data table based on the table name, table description information, and the natural language annotation of the field name of each field in the table in the second prompt; The database information, the table name and natural language label of each table, and the field name and natural language label of each field are encapsulated into a third prompt. The large model is used to generate a natural language label for the database name for each database based on the database name, database description information, the natural language label of the table name of each table in the database, and the natural language label of the field name of each field in each table in the third prompt.

4. The method according to claim 3, characterized in that The natural language annotations are reviewed, and the reviewed annotations and the database information to which the annotations belong are saved as database location description information in JSON format, including: Use the natural language processing library to check the sentence structure, part of speech, and grammatical rules of natural language annotations; Verify the logical consistency of the natural language annotation of the database name with the natural language annotation of the subordinate table name and the natural language annotation of the field name; Analyze the characteristics of some non-empty sample data for each field, compare the annotation content with the characteristics of the sample data, and verify whether the two are consistent; Verify whether the natural language annotations conform to the preset standard template; The reviewed annotations and the database information to which the annotations belong are saved as database location description information in JSON format, and the database location description information in JSON format is updated to the database architecture description file.

5. The method according to claim 1, wherein The data location information is written into a prompt, and a query statement is generated based on the prompt using a large model, including: Retrieving code data associated with the data location information; The code data and the data location information are written into a prompt, and the prompt is input into the large model to obtain a query statement.

6. The method according to claim 5, characterized in that The method further comprises: Collect the erroneous query statements output by the large model that do not meet the query expectations, and save the data location information corresponding to the erroneous query statements and the corresponding correct query statements to the code library.

7. The method according to claim 6, characterized in that The data location information is written into a prompt, and a query statement is generated based on the prompt using a large model, including: retrieving code data associated with the data location information from the code repository; Writing the code data and the data location information into a prompt, and inputting the prompt into the large model to obtain multiple query statements and predicted probabilities output by the large model; Counting the number of occurrences of multiple query statements contained in the code data, and calculating the occurrence frequencies of the multiple query statements based on the number of occurrences of the multiple query statements, and using the occurrence frequencies as predicted probabilities of the query statements; For any possible query statement, calculate its comprehensive probability as: Where k is the weight factor, is the probability output by the large model, is the probability obtained by statistics; in, It is a hyperparameter used to control the influence of error rate on the overall probability. e represents the error rate of the large model. Determine the minimum value of the hyperparameter and maximum value , and the minimum error rate and maximum value ; The mapping formula between hyperparameters and error rates is: Calculate hyperparameters based on the current error rate of the large model.

8. A database query statement generation system, characterized in that: include: An acquisition module, configured to acquire user natural language and extract keywords from the user natural language; a retrieval module, configured to obtain data location information according to the keyword, wherein the data location information includes database location description information with natural language annotations, wherein the natural language annotations are associated with the keyword; A generation module is used to write the data location information into a prompt and generate a query statement based on the prompt using a large model.

9. A device, characterized in that include: A memory, used for storing a database query statement generating program; A processor, configured to implement the steps of the database query statement generation method according to any one of claims 1 to 7 when executing the database query statement generation program.

10. A computer-readable storage medium storing a computer program, characterized in that: The readable storage medium stores a database query statement generation program, and when the database query statement generation program is executed by a processor, the steps of the database query statement generation method according to any one of claims 1 to 7 are implemented.