Database statement generation method based on natural language processing and related equipment

By natural language processing, generating and extending statements on the SQL statements entered by users, the problem of traditional SQL editing tools lacking intelligent recommendations is solved, and SQL writing efficiency and accuracy are improved.

CN120067128APending Publication Date: 2025-05-30BEIJING CHINA POWER INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510108734.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional SQL editing tools lack intelligent recommendation completion functions, which leads to users who need to frequently review documents or perform manual input when writing SQL, which reduces work efficiency and is prone to syntax errors.

Method used

A database statement generation method based on natural language processing is adopted to analyze and standardize the user's input statements, generate preprocessing statements, and expand them to generate target statements, including the original input statement information.

Benefits of technology

Improves the efficiency of writing database statements, reduces manual input and syntax errors, and enhances the readability and maintenance of the code.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067128A_ABST
    Figure CN120067128A_ABST
Patent Text Reader

Abstract

The invention provides a database statement generation method based on natural language processing and related equipment. The method comprises the steps of obtaining an input statement of a user; analyzing and standardizing the input statement to obtain a preprocessed statement; and expanding the preprocessed statement to generate at least one target statement, wherein the target statement comprises the input statement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing, and in particular, to a method for generating database statements based on natural language processing and related devices. Background Art

[0002] With the continuous growth of the data volume, database queries have become increasingly complex. Writing efficient SQL query statements is a challenge for database administrators and developers. Traditional SQL editing tools lack the intelligent recommendation and completion function for large segments of code, resulting in users having to frequently consult documents or perform manual input when writing SQL, which not only reduces work efficiency but also easily leads to syntax errors. Summary of the Invention

[0003] The present disclosure provides a method for generating database statements based on natural language processing and related devices, so as to solve, to a certain extent, the technical problems such as low work efficiency and easy occurrence of syntax errors in the generation of database statements by manual input.

[0004] In the first aspect of the present disclosure, a method for generating database statements based on natural language processing is provided, including:

[0005] Obtaining an input statement of a user;

[0006] Parsing and standardizing the input statement to obtain a preprocessed statement;

[0007] Expanding the preprocessed statement to generate at least one target statement, where the target statement includes the input statement.

[0008] In the second aspect of the present disclosure, a device for generating database statements based on natural language processing is provided, including:

[0009] An obtaining module, configured to obtain an input statement of a user;

[0010] A preprocessing module, configured to parse and standardize the input statement to obtain a preprocessed statement;

[0011] An expanding module, configured to expand the preprocessed statement to generate at least one target statement, where the target statement includes the input statement.

[0012] In the third aspect of the present disclosure, an electronic device is provided, including one or more processors, a memory; and one or more programs, where the one or more programs are stored in the memory and are executed by the one or more processors, and the programs include instructions for executing the method according to the first aspect.

[0013] In a fourth aspect of the present disclosure, there is provided a non-volatile computer-readable storage medium including a computer program, which, when executed by one or more processors, causes the processors to execute the method described in the first aspect.

[0014] In a fifth aspect of the present disclosure, there is provided a computer program product including computer program instructions, which, when executed on a computer, cause the computer to execute the method described in the first aspect.

[0015] As can be seen from the above, a method for generating database statements based on natural language processing and related devices provided by the present disclosure parse and standardize the input statement of the user to obtain a preprocessed statement, and expand the preprocessed statement to generate at least one target statement containing the information of the original input statement. This not only improves the writing efficiency of database statements and reduces the manual input of developers through the automatic completion function, but also reduces problems such as syntax errors caused by humans. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following description are only embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 It is a schematic diagram of the architecture for generating database statements based on natural language processing according to an embodiment of the present disclosure.

[0018] Figure 2 It is a schematic diagram of the hardware structure of an exemplary electronic device according to an embodiment of the present disclosure.

[0019] Figure 3 It is a schematic diagram of the flowchart of the method for generating database statements based on natural language processing according to an embodiment of the present disclosure.

[0020] Figure 4 It is a schematic diagram of the method for generating database statements based on natural language processing according to an embodiment of the present disclosure.

[0021] Figure 5 It is a schematic diagram of the code structure according to an embodiment of the present disclosure.

[0022] Figure 6 It is a schematic diagram of the device for generating database statements based on natural language processing according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] To make the objectives, technical solutions, and advantages of the present disclosure clearer and more understandable, the present disclosure will be further described in detail below with reference to specific embodiments and the accompanying drawings.

[0024] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the ordinary meanings understood by those of ordinary skill in the field to which the present disclosure belongs. The "first", "second", and similar terms used in the embodiments of the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this term cover the elements or objects listed after this term and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0025] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0026] For example, when receiving a user's active request, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.

[0027] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0028] Figure 1 shows a schematic diagram of the database statement generation architecture based on natural language processing according to an embodiment of the present disclosure. Refer to Figure 1, the database statement generation architecture 100 based on natural language processing may include a server 110, a terminal 120, and a network 130 providing a communication link. The server 110 and the terminal 120 may be connected through the wired or wireless network 130. Among them, the server 110 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, security services, and CDN.

[0029] The terminal 120 may be implemented by hardware or software. For example, when the terminal 120 is implemented by hardware, it may be various electronic devices with a display screen and supporting page display, including but not limited to smartphones, tablets, e-book readers, laptop computers, and desktop computers, etc. When the terminal 120 device is implemented by software, it may be installed in the above-listed electronic devices; it may be implemented as multiple software or software modules (such as software or software modules for providing distributed services), or it may be implemented as a single software or software module, which is not specifically limited here.

[0030] It should be noted that the database statement generation method based on natural language processing provided by the embodiments of the present application may be executed by the terminal 120 or by the server 110. It should be understood that Figure 1 the numbers of the terminals, networks, and servers in are only for illustration and are not intended to limit them. According to the implementation needs, there may be any number of terminals, networks, and servers.

[0031] Figure 2 shows a schematic hardware structure diagram of an exemplary electronic device 200 provided by an embodiment of the present disclosure. As Figure 2 shown, the electronic device 200 may include: a processor 202, a memory 204, a network module 206, a peripheral interface 208, and a bus 210. Among them, the processor 202, the memory 204, the network module 206, and the peripheral interface 208 are communicatively connected to each other inside the electronic device 200 through the bus 210.

[0032] The processor 202 can be a Central Processing Unit (CPU), a Neural Network Processor (NPU), a Microcontroller Unit (MCU), a Programmable Logic Device, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits. The processor 202 can be used to execute functions related to the technologies described in this disclosure. In some embodiments, the processor 202 may also include multiple processors integrated as a single logic component. For example, as Figure 2 shown, the processor 202 can include multiple processors 202a, 202b, and 202c.

[0033] The memory 204 can be configured to store data (e.g., instructions, computer code, etc.). As Figure 2 shown, the data stored in the memory 204 can include program instructions (e.g., program instructions for implementing the natural language processing-based database statement generation method of the embodiments of this disclosure) and data to be processed (e.g., the memory can store configuration files of other modules, etc.). The processor 202 can also access the program instructions and data stored in the memory 204 and execute the program instructions to operate on the data to be processed. The memory 204 can include a volatile storage device or a non-volatile storage device. In some embodiments, the memory 204 can include a Random Access Memory (RAM), a Read Only Memory (ROM), an optical disc, a magnetic disk, a hard disk, a Solid State Drive (SSD), a flash memory, a memory stick, etc.

[0034] The network module 206 can be configured to provide communication with other external devices to the electronic device 200 via a network. The network can be any wired or wireless network capable of transmitting and receiving data. For example, the network can be a wired network, a local wireless network (e.g., Bluetooth, WiFi, Near Field Communication (NFC), etc.), a cellular network, the Internet, or a combination of the above. It can be understood that the type of the network is not limited to the above specific examples. In some embodiments, the network module 206 can include any combination of any number of Network Interface Controllers (NICs), radio frequency modules, transceivers, modems, routers, gateways, adapters, cellular network chips, etc.

[0035] The peripheral interface 208 can be configured to connect the electronic device 200 to one or more peripheral devices to achieve information input and output. For example, the peripheral devices can include input devices such as a keyboard, a mouse, a touchpad, a touch screen, a microphone, various sensors, etc. and output devices such as a display, a speaker, a vibrator, an indicator light, etc.

[0036] The bus 210 can be configured to transfer information between various components of the electronic device 200, such as the processor 202, the memory 204, the network module 206, and the peripheral interface 208, such as an internal bus (e.g., a processor-memory bus), an external bus (USB port, PCI-E bus), etc.

[0037] It should be noted that although the architecture of the above-mentioned electronic device 200 only shows the processor 202, the memory 204, the network module 206, the peripheral interface 208, and the bus 210, in the specific implementation process, the architecture of the electronic device 200 may further include other components necessary for normal execution. In addition, those skilled in the art can understand that the architecture of the above-mentioned electronic device 200 may also only include the components necessary to implement the solution of the embodiments of the present disclosure, and do not necessarily include all the components shown in the figure.

[0038] With the continuous growth of the data volume, database queries have become increasingly complex. Writing efficient SQL query statements is a challenge for database administrators and developers. Traditional SQL editing tools lack the intelligent recommendation and completion function for large segments of code, resulting in users having to frequently consult documents or perform manual input when writing SQL, which not only reduces work efficiency but also easily leads to syntax errors. In addition, in large projects, understanding and maintaining complex SQL scripts is also a major problem. Therefore, the function of being able to quickly and accurately extract the SQL outline is particularly important. Therefore, how to improve the writing efficiency of database statements, reduce errors, increase the readability of the code, and reduce the maintenance cost of the code has become a technical problem that needs to be solved urgently.

[0039] In view of this, the embodiments of the present disclosure provide a method for generating database statements based on natural language processing and related devices. By parsing and standardizing the input statement of the user, a preprocessing statement is obtained, and the preprocessing statement is extended to generate at least one target statement containing the information of the original input statement. It not only improves the writing efficiency of database statements, reduces the manual input of developers through the automatic completion function, but also reduces problems such as syntax errors caused by humans.

[0040] See Figure 3 , Figure 3 shows a schematic flowchart of a method for generating database statements based on natural language processing according to an embodiment of the present disclosure. The method for generating database statements based on natural language processing according to an embodiment of the present disclosure can be deployed on a terminal or a server side. Figure 3 In, the method 300 for generating database statements based on natural language processing may further include the following steps.

[0041] In step S310, obtain the input statement of the user.

[0042] Among them, an intuitive and fully functional user interface can be provided so that users can conveniently input statements and view the processing results. The statement can be a database statement, such as an SQL statement. Specifically, the user interface can include an input area, for example, providing a text box or a code editor as the input area for users to input statements. The input area can have basic text editing functions, such as copy, paste, undo, and redo. The user interface can also include a display area for displaying at least one of the statement input by the user, the preprocessed statement after processing, and the final generated target statement. The display area can also clearly present the differences and connections between different statements to help users understand the processing process. For error or warning messages, clear prompts can also be given in the display area. The user interface can also include operation buttons, such as adding a "submit" or "execute" button on the user interface. After the user clicks, the system will start processing the input SQL statement. Specifically, it can receive the SQL query statement input by the user through the input area and capture the input SQL statement when the user clicks the "submit" or "execute" button. Through the user interface, users can conveniently input and view the processing process and results of SQL query statements. This not only improves the readability and execution efficiency of SQL statements but also provides users with a more intuitive and rich query experience.

[0043] In step S320, the input statement is parsed and standardized to obtain a preprocessed statement.

[0044] Among them, parsing and standardizing the input statement makes the preprocessed statement clearer and more readable visually, which helps to better understand and analyze the SQL statement. The standardized preprocessed statement is easier to be parsed and executed by the database system, thus improving the query performance. The execution results of the standardized preprocessed statement are more consistent on different database systems, improving the cross-database compatibility of SQL statements.

[0045] In some embodiments, parsing and standardizing the input statement to obtain a preprocessed statement includes:

[0046] Performing data cleaning and data filtering on the input statement to obtain an intermediate statement;

[0047] Performing word segmentation and standardization on the intermediate statement to obtain the preprocessed statement.

[0048] Among them, through data cleaning and filtering, irrelevant characters and errors in the input statement are removed, improving the integrity and accuracy of the statement. Specifically, redundant spaces and line breaks in the SQL statement can be removed to make it more compact and readable. Keywords in the SQL statement (such as SELECT, FROM, WHERE, etc.) can be converted to uppercase to make it easier to distinguish keywords from table names and column names. Ensure that string constants use consistent quotation marks (usually single quotes). Comments in the SQL statement can be removed for further processing. Comments include single-line comments (--) and multi-line comments ( / *...* / ). It can be split into multiple simple statements or its structure can be reorganized to make it clearer. SQL parsing tools or database management tools can be used to check and correct syntax errors in the SQL statement. The SQL statement can be formatted into a standard format to improve readability. For example, use tools or manually add appropriate indentation and line breaks.

[0049] The intermediate statement is split into independent lexical units, and these units are standardized to ensure their consistency and readability. Specifically, a tokenization algorithm can be used to split the intermediate statement into a series of independent lexical units, which may be keywords, identifiers (such as table names and column names), operators, constant values, etc. Converting all SQL keywords to uppercase helps to identify them more easily in subsequent parsing and execution processes. According to the database configuration or user selection, identifiers (such as table names and column names) are converted to lowercase or uppercase to maintain consistency. Escape the string constant values to ensure they are correctly interpreted and executed. Ensure that all operators and delimiters conform to the SQL syntax specification and no non-standard alternative characters are used. After tokenization and standardization, a preprocessed statement that is more standardized and unified in syntax and format can be obtained, providing a good foundation for subsequent processing and execution.

[0050] In step S330, the preprocessed statement is expanded to generate at least one target statement, and the target statement includes the input statement.

[0051] Among them, after obtaining the preprocessed statement, it can be further expanded to generate one or more target statements. The target statement not only contains all the information of the original input statement, but may also contain additional information or optimized statements.

[0052] In some embodiments, expanding the preprocessed statement to generate at least one target statement includes:

[0053] Feature extraction is performed on the preprocessed statement to obtain syntactic features and context information;

[0054] Expand the preprocessed statement based on the syntax features and the context information to obtain at least one of the target statements.

[0055] Among them, performing syntax analysis on the preprocessed statement to extract its structural features and components, such as keywords, identifiers (table names, column names), operators, constant values, subqueries, etc., helps to understand the basic structure and logical structure of the statement. Analyzing the context environment where the preprocessed statement is located, including the database schema (such as table structure, field type, etc.), the user's historical query records, business logic, etc. Through the context information, the possible intentions and potential requirements of the preprocessed statement can be inferred, such as the data range that the user may want to query, the degree of aggregation, etc.

[0056] In some embodiments, expanding the preprocessed statement based on the syntax features and the context information to obtain at least one of the target statements includes:

[0057] Expanding the preprocessed statement based on the trained statement generation model according to the syntax features and the context information to obtain at least one of the target statements; wherein, the trained statement generation model is obtained by training an initial model based on training samples, and the training samples include training statements and extended statements including the training statements.

[0058] Among them, through the trained statement generation model, the preprocessed statement can be automatically expanded to generate target statements that meet the user's needs. The statement generation model can perform feature extraction and context analysis on the preprocessed statement, and using the trained statement generation model for expansion can generate target statements that meet the user's needs. This not only improves the efficiency and accuracy of query generation, but also provides greater flexibility and scalability.

[0059] In some embodiments, the trained statement generation model is obtained by training an initial model based on training samples, including:

[0060] Generating predicted statements based on the training statements;

[0061] Obtaining a loss function based on the cross-entropy function between the predicted statements and the extended statements;

[0062] Adjusting the model parameters of the initial model to minimize the loss function to obtain the trained statement generation model.

[0063] Among them, the training statements and their corresponding extended statements can be used as training samples. Using these training samples to train the initial model, a statement generation model that can generate target statements based on syntactic features and context information is obtained. During the training process, the model learns how to extract key information from the given preprocessed statements and generate reasonable extended statements according to this information. This reduces the workload of manually writing SQL queries and improves the efficiency of query generation. It makes the generated queries more in line with the user's expectations, improving the accuracy and practicality of the queries. Since the model is trained based on training samples, the generation strategy of the model can be changed by adjusting the training samples, which enables the model to adapt to different application scenarios and user requirements, providing greater flexibility. With the increase of training samples and the continuous optimization of the model, the performance of the statement generation model can be continuously improved, enabling the system to continuously learn and evolve to better meet the user's query needs.

[0064] Specifically, an initial model (which may be a deep learning model, such as a sequence-to-sequence (Seq2Seq) model, a Transformer model, etc.) can be used to process the training statements in the training samples. The model generates one or more predicted statements based on the syntactic features and context information of the training statements. For each training sample, the cross-entropy function between the predicted statement and the extended statement (i.e., the true target statement) is calculated to evaluate the matching degree between the predicted statement and the true target statement. The smaller the cross-entropy value, the closer the predicted statement is to the true target statement, and the better the performance of the model.

[0065] Based on the calculated loss function (i.e., the cross-entropy value), the model parameters of the initial model can be adjusted using the backpropagation algorithm. The backpropagation algorithm calculates the gradient of the loss function with respect to the model parameters and then updates the model parameters in the opposite direction of the gradient to minimize the loss function. This process is iterated multiple times until the loss function converges to a smaller value or reaches a preset number of iterations. When the loss function converges to a satisfactory level or reaches the preset number of iterations, the training process is stopped. The model obtained at this time is the trained statement generation model, which can generate accurate target statements based on syntactic features and context information.

[0066] In some embodiments, method 300 further includes:

[0067] Obtaining the confidence corresponding to the target statement based on the statement generation model;

[0068] Sorting the target statements in descending order based on the confidence to obtain a sorting result;

[0069] Displaying the top n target statements of the sorting result, where n is a positive integer.

[0070] Among them, while using the statement generation model to generate the target statement, the model will also calculate the corresponding confidence for each generated target statement. The confidence is a measure representing the accuracy of the generated statement by the model and can be a numerical value between 0 and 1. The closer the numerical value is to 1, the higher the confidence of the model in this statement. According to the calculated confidence, the generated target statements are sorted. The sorting rule can be to sort from the highest confidence to the lowest, that is, the target statement with the highest confidence is ranked first. From the sorted list of target statements, the first n (n is a positive integer) are selected as the final recommended results. The value of n can be adjusted according to user needs or system configuration to balance the diversity and accuracy of the results. The first n target statements of the sorting result can be displayed in a user-friendly manner. For example, the display method can include a list form, a tree structure, or any other form that helps users understand and select. By displaying the first n target statements with the highest confidence, users can find the query results that meet their needs faster. This reduces the time for users to browse and filter a large number of query results and improves the query efficiency. Different numbers of recommended results can be provided according to user needs, which not only ensures the accuracy of the results but also provides users with a diverse selection space. It can be seen that by calculating the confidence of the target statement and sorting and displaying according to the confidence, more accurate, efficient, and reliable query results can be provided for users. This not only improves the user experience but also enhances the reliability and resource utilization efficiency of the system.

[0071] In some embodiments, method 300 further includes:

[0072] Determine the statement to be processed for code structure analysis from at least one of the target statements;

[0073] Perform lexical analysis on the statement to be processed to obtain a syntax tree;

[0074] Based on preset clauses, determine the key nodes in the syntax tree;

[0075] Generate a code structure for the statement to be processed based on the key nodes and display the code structure in a preset format.

[0076] One or more target statements generated previously can be selected for code structure analysis, and these selected statements can be referred to as statements to be processed. Lexical analysis is performed on each statement to be processed, and the statement is decomposed into a series of lexical units (such as keywords, identifiers, operators, etc.). Based on these lexical units, a syntax tree of the statement is constructed to represent the syntax structure of the statement, where each node represents a lexical unit or a sub-structure. In the syntax tree, according to the preset clause patterns or rules, key nodes that have important impacts on the statement structure and semantics are identified. These key nodes may include selection lists, table names, conditional expressions, aggregate functions, query targets, data sources, filtering conditions, etc. Based on the identified key nodes, a code structure description of the statement to be processed is generated, which may include the hierarchical structure of the statement, the connection relationships of each part, the detailed information of the key nodes, etc. The generated code structure can be displayed in a preset format (such as a tree diagram, a text description, a table, etc.) to facilitate the user to understand and analyze the structure of the statement. Specifically, existing lexical analyzers (such as ANTLR, Flex / Bison, etc.) can be used to decompose the statement and generate the syntax tree. Text descriptions, tables, tree diagrams, or graphical interfaces can be used to display the hierarchical structure and connection relationships. By generating and displaying the code structure, the user can more intuitively understand the structure and semantics of the query statement, and thus more easily identify errors and optimize the query. Code structure analysis can reduce the debugging time and cost, and can be used as a teaching aid to help beginners better understand the SQL syntax rules and their application scenarios, and help beginners better understand the construction and syntax rules of query statements. It helps non-professionals quickly master the SQL syntax structure, reducing the learning curve. It also facilitates developers and maintainers to quickly browse the main logical framework of long SQL scripts, improving the code review efficiency.

[0077] See Figure 4 , Figure 4 shows a schematic diagram of a database statement generation method according to an embodiment of the present disclosure. Figure 4 In, user interface: provides an input and display area, receives the user's SQL input and displays completion suggestions and an outline. Receives the user's SQL input and passes the input to the preprocessing module. Preprocessing module: performs preliminary parsing and standardization on the SQL statement input by the user, and then passes the processed data to the deep learning model. Deep learning model: used to predict and generate a complete SQL statement, and passes the result to the post-processing module. Outline extraction module: extracts key information from the SQL statement, generates a code outline, and passes the outline to the post-processing module. Post-processing module: optimizes and formats the generated result, and finally displays the completion suggestions and the outline on the user interface deep learning model.

[0078] Among them, for model selection of the deep learning model, the Transformer model can be used, such as BERT or GPT-2. A large number of SQL statements are collected as training data, including common query, insert, update, and delete operations. Syntactic features and context information are extracted from the SQL statements for model training. The supervised learning method is used to train the model through a large amount of labeled data so that it can predict complete SQL statements. The partial SQL statements input by the user are converted into a format that the model can understand. The trained deep learning model is utilized to predict possible complete SQL statements. The results are sorted according to the predicted probability values, and the most likely completion suggestions are preferentially displayed. The sorted completion suggestions are presented to the user for selection. Specifically, the real SQL execution records from the production environment logs or other sources can be collected, and after cleaning to remove duplicates, a training set is formed. An encoder-decoder model adopting the Transformer architecture enables the model to learn to infer the unknown based on the known through sequence-to-sequence learning. When the user enters new characters in the editor, the background service is triggered to calculate the best match and instantaneously update the options in the interface prompt box. The top several candidate answers with the highest scores are displayed to the user in descending order for selection, and keyboard shortcuts are also supported for quick selection or skipping.

[0079] Perform syntactic analysis on the input SQL statements to construct a syntax tree. Extract key nodes from the syntax tree, such as clauses like SELECT, FROM, WHERE, etc. Generate an outline of the SQL code based on the key nodes and display it in a tree structure. Format the generated outline to make it easy to read and understand. Specifically, the full text of the SQL file to be processed can be read, and the text is initially segmented using regular expressions or other rule engines to distinguish basic units such as keywords, identifiers, and operators. An abstract syntax tree (AST) is generated based on the SQL syntax rules defined by the BNF paradigm. Recursively visit each node in the AST to extract important components such as query targets, data sources, filtering conditions, etc. The information extracted above is organized in a certain format and then returned to the front end for display or saved as an external file, as Figure 5 shown Figure 5 is a schematic diagram of the code structure of the embodiment of the present disclosure.

[0080] It can be seen that according to the method of the embodiment of the present disclosure, the SQL writing speed and accuracy are significantly improved, and the errors caused by typos are reduced. It helps non-professionals quickly master the SQL syntax structure and reduces the learning curve. It is convenient for developers and maintainers to quickly browse the main logical framework of long SQL scripts, improving the code review efficiency. It can be used as a teaching aid to help beginners better understand the SQL syntax rules and their application scenarios.

[0081] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In such a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.

[0082] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the above embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.

[0083] Based on the same technical concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a database statement generation device based on natural language processing. See Figure 6 , the database statement generation device based on natural language processing, the device includes:

[0084] An acquisition module, configured to acquire an input statement of a user;

[0085] A preprocessing module, configured to parse and standardize the input statement to obtain a preprocessed statement;

[0086] An extension module, configured to extend the preprocessed statement to generate at least one target statement, where the target statement includes the input statement.

[0087] For the convenience of description, when describing the above device, it is divided into various modules according to functions and described separately. Of course, when implementing the present disclosure, the functions of each module can be implemented in one or more software and / or hardware.

[0088] The device of the above embodiment is used to implement the corresponding database statement generation method based on natural language processing in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0089] Based on the same technical concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium, where the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the database statement generation method based on natural language processing as described in any of the foregoing embodiments.

[0090] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0091] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the database statement generation method based on natural language processing described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0092] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the concept of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present disclosure as described above. For the sake of brevity, they are not provided in detail.

[0093] In addition, for the sake of simplicity of description and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. In addition, the device may be shown in block diagram form to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be completely within the understanding of those skilled in the art). In the case where specific details (such as circuits) are set forth to describe the exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0094] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0095] Embodiments of the present disclosure are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Accordingly, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the embodiments of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A method for generating database statements based on natural language processing, comprising: Get the user's input statement; Parsing and standardizing the input sentence to obtain a preprocessed sentence; The preprocessing statement is expanded to generate at least one target statement, wherein the target statement includes the input statement.

2. The method according to claim 1, wherein: Expanding the prepared statement to generate at least one target statement includes: Performing feature extraction on the preprocessed sentence to obtain grammatical features and context information; The preprocessing sentence is expanded based on the grammatical features and the context information to obtain at least one target sentence.

3. The method according to claim 2, wherein: Expanding the preprocessing statement based on the grammatical features and the context information to obtain at least one target statement includes: Based on the trained sentence generation model, the preprocessed sentence is expanded according to the grammatical features and the context information to obtain at least one target sentence; wherein the trained sentence generation model is obtained by training an initial model based on training samples, and the training samples include training sentences and extended sentences including the training sentences.

4. The method according to claim 3, further comprising: Obtaining a confidence level corresponding to the target sentence based on the sentence generation model; Sort the target sentences from large to small based on the confidence level to obtain a sorting result; The first n target sentences of the sorting result are displayed, where n is a positive integer.

5. The method according to claim 3, wherein: The trained sentence generation model is obtained by training the initial model based on the training samples, including: Generate a prediction sentence based on the training sentence; Obtaining a loss function based on a cross entropy function between the predicted sentence and the extended sentence; The model parameters of the initial model are adjusted to minimize the loss function to obtain a trained sentence generation model.

6. The method according to claim 1, further comprising: Determining a statement to be processed for code structure analysis from at least one of the target statements; Performing lexical analysis on the sentence to be processed to obtain a syntax tree; Determine key nodes in the syntax tree based on preset clauses; A code structure about the statement to be processed is generated based on the key nodes and the code structure is displayed in a preset format.

7. The method according to claim 1, wherein: The input sentence is parsed and standardized to obtain a preprocessed sentence, including: Performing data cleaning and data filtering on the input sentence to obtain an intermediate sentence; The intermediate sentence is segmented and standardized to obtain the preprocessed sentence.

8. A database statement generation device based on natural language processing, comprising: The acquisition module is used to obtain the user's input statement; A preprocessing module, used for parsing and standardizing the input sentence to obtain a preprocessed sentence; An expansion module is used to expand the preprocessing statement to generate at least one target statement, wherein the target statement includes the input statement.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the program. 10 . A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method according to claim 1 .