Database Code Extraction Method and System Based on Rapid Screening of Large Language Models
By obtaining database type and version, word segmentation processing and building large language model prompt words, the problem of inaccuracy and efficiency of database code extraction in the prior art is solved, and efficient and accurate database code extraction is achieved.
Patent Information
- Application Number
- CN202411775270.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-12-05
AI Technical Summary
The existing technology has problems in database code extraction with low code recognition rate, lack of the latest database version corpus, and large files exceeding the token length limit of large language models, resulting in inefficiency.
By obtaining the database type and version in the user code base file, finding the corresponding database code keywords and performing word segmentation enhancement processing, filtering out files containing keywords, building basic prompt words for large language models and adding database query syntax, generating new large model prompt word information, and activate the large language model for database code extraction.
It realizes efficient and accurate extraction of database code in the software file code library, improves the efficiency of database code review, release and type conversion, and solves the problem of low accuracy and efficiency of large language models in database code extraction.
Smart Images

Figure CN119226380B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software development data management, and particularly to a method for extracting database code based on rapid screening by a large language model, a system for extracting database code based on rapid screening by a large language model, an electronic device, and a computer-readable storage medium. Background Art
[0002] During the software R & D process, code reviews are often required. There are relatively many review tools and methods for language codes such as Java and C++. However, database codes are usually written in SQL language, and the written SQL codes are not independent files. They are usually scattered in other language code files such as Java, C++, and Python. And database codes usually need to be submitted to database administrators for review and operation separately. Therefore, it is necessary to extract these SQL codes scattered in different code files. In the past, traditional database code extraction work was all manually operated by engineers one file at a time, which required a large amount of time, resulting in low efficiency of database code review and release.
[0003] In addition, although with the emergence of AI technology and the AI large language model, software code files can be sent to the large language model, and the large language model can automatically extract database codes, there are still several problems:
[0004] There are many types of databases. Without knowing the database type, the code recognition rate of the AI large language model is not high, and incorrect content may be extracted;
[0005] The AI large language model lacks the corpus of the latest database version, and the extraction success rate for new databases or new version codes is not high;
[0006] Database SQL codes are scattered in different program code files. A software may have hundreds or thousands of files. If all program code files are sent to the large language model, it will consume a large number of tokens of the large language model, and the generation efficiency is also low. In addition, large files may exceed the token length limit of the large language model and accurate results cannot be obtained. Summary of the Invention
[0007] In order to solve the technical problems existing in the prior art, the present invention provides the following technical solutions:
[0008] On the one hand, a method for extracting database code based on rapid screening by a large language model is provided. This method is implemented by an electronic device and includes:
[0009] S1. Obtain the database type and version in the user code library file;
[0010] S2. According to the database type and version, search for the corresponding database keywords from the preset database code keyword list and perform word segmentation enhancement processing;
[0011] S3. Based on the enhanced database code keywords, screen out the files containing the keywords from the user code library file and mark them as program files containing database code;
[0012] S4. Build a basic prompt word for the large language model based on the database code keywords, and add new database query syntax to the basic prompt word to generate new large model prompt word information and input it into the preset large language model;
[0013] S5. Activate the large language model, and the large language model extracts the corresponding database code from the program files containing database code based on the large model prompt word information and outputs it.
[0014] Further, S1. Obtain the database type and version in the user code library file, including:
[0015] Input the user code library file;
[0016] Parse the user code library file to obtain all program files;
[0017] Traverse each program file to check whether the program file contains a database configuration file:
[0018] If not, traverse to the next one until all program files are checked;
[0019] If so, read the database configuration file and parse the database configuration file to obtain the basic data corresponding to the database type and version, and obtain the database type and version.
[0020] Further, S2. According to the database type and version, search for the corresponding database keywords from the preset database code keyword list and perform word segmentation enhancement processing, including:
[0021] Pre - construct the database code keyword list composed of different versions of databases and corresponding keywords;
[0022] According to the database type and version, search for the corresponding database code keywords from the preset database code keyword list and input them into the preset Bloom filter to generate a Bloom filter data structure for data retrieval;
[0023] Use word segmentation processing technology to identify and extract the word segments in the code file content of the user code library file, and write the word segments into the Bloom filter data structure and input them into the Bloom filter.
[0024] Further, S3. Based on the enhanced database code keywords, screen out the files containing the keywords from the user code library files and mark them as program files containing database code, including:
[0025] Traverse all program files;
[0026] Perform word segmentation and parsing on each program file to obtain the word segmentation of the corresponding program file;
[0027] Use a Bloom filter to perform database keyword matching on each program file:
[0028] If it is found that the keywords are matched in the program file, mark the program file with the matched keywords as a program file containing database code;
[0029] If no keywords are found to be matched in the program file, proceed to the next step;
[0030] Determine whether all the program files have been traversed:
[0031] If so, all traversals are completed, and proceed to the next step;
[0032] If there are still other program files that have not been traversed, continue to traverse the program files;
[0033] Save the program files containing database code to the software file code library.
[0034] Further, S4. Build a basic prompt for the large language model based on the database code keywords, and add new database query syntax to the basic prompt to generate new large model prompt information and input it into the preset large language model, including:
[0035] Prepare the basic prompt of the large language model, where the basic prompt is constructed based on the database type and version in the database code keywords;
[0036] Based on the basic prompt, construct the corresponding database query syntax;
[0037] Organize the database query syntax, generate the large model prompt information and input it into the large language model;
[0038] When subsequently using the large language model to traverse program files, check in real time whether new database types and versions appear in the dynamically input user code library files:
[0039] If so, based on the new database type and version, construct new database query syntax and add it to the large language model;
[0040] If not, give up.
[0041] Further, before activating the large language model in S5 and extracting and outputting the corresponding database code from the program file containing the database code by the large language model based on the large model prompt word information, it further includes:
[0042] Traverse all program files marked as containing database code;
[0043] Calculate the file size of each program file containing database code: token length:
[0044] ,
[0045] B is the token length,
[0046] n is the number of text characters in the program file,
[0047] Lsi is the word length of each text;
[0048] Judge whether the file size: token length of the program file containing database code exceeds the token limit of the large language model:
[0049] If so, perform sharding processing on the program file containing database code to obtain several program shard files and input them into the large language model;
[0050] If not, directly input the program file containing database code into the large language model.
[0051] Further, in S5, activate the large language model, and extract and output the corresponding database code from the program file containing database code by the large language model based on the large model prompt word information, including:
[0052] Send an extraction instruction to the large language model to activate the large language model;
[0053] Extract and return the corresponding database code from the program file containing database code by the large language model based on the large model prompt word information;
[0054] Parse the database code and save it as the database code in the corresponding structured format.
[0055] On the other hand, a database code extraction system based on fast screening by a large language model is provided. The database code extraction system based on fast screening by a large language model is used to implement the above-mentioned database code extraction method based on fast screening by a large language model. It is characterized in that the system includes:
[0056] A file acquisition module, configured to acquire the database type and version in the user code library file;
[0057] A keyword search module, configured to search for corresponding database keywords from a preset database code keyword list according to the database type and version, and perform word segmentation enhancement processing;
[0058] A file screening module, configured to screen out files containing keywords from the user code library file based on the enhanced database code keywords, and mark them as program files containing database code;
[0059] An intelligent screening module, configured to build a basic prompt for the large language model based on the database code keywords, add new database query syntax to the basic prompt, generate new large model prompt information, and input it into a preset large language model;
[0060] A code extraction module, configured to activate the large language model, and the large language model extracts corresponding database code from the program files containing database code based on the large model prompt information and outputs it.
[0061] On the other hand, an electronic device is provided, and the electronic device includes: a processor; a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, any one of the methods in the above-mentioned database code extraction method based on rapid screening by a large language model is implemented.
[0062] On the other hand, a computer-readable storage medium is provided, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement any one of the methods in the above-mentioned database code extraction method based on rapid screening by a large language model.
[0063] The beneficial effects brought by the technical solution provided by the embodiments of the present invention at least include:
[0064] Based on the implementation of the present invention, the present invention proposes a database code extraction method based on an AI large language model, and users can automatically extract database code from a software file code library more accurately and efficiently, and apply it to database code review, release, and database type code conversion scenarios.
[0065] The present invention enables the large language model to better understand the database syntax supplement format, an automatic database type and version accurate identification scheme, provides a method for quickly screening and filtering database code files, and can effectively solve the problems of low accuracy and efficiency of database code extraction by the large language model. Description of the Drawings
[0066] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0067] Figure 1 is an operation flowchart of a method for extracting database code based on rapid screening by a large language model provided by an embodiment of the present invention;
[0068] Figure 2 is a schematic flowchart of a process for obtaining database configuration information from a user code library file provided by an embodiment of the present invention;
[0069] Figure 3 is a list of database code keywords provided by an embodiment of the present invention;
[0070] Figure 4 is a schematic flowchart of a process for screening out program files containing keywords provided by an embodiment of the present invention;
[0071] Figure 5 is a schematic flowchart of a process for extracting database code by a large language model provided by an embodiment of the present invention;
[0072] Figure 6 is a block diagram of a system for extracting database code based on rapid screening by a large language model provided by an embodiment of the present invention;
[0073] Figure 7 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0074] The following describes the technical solutions in the present invention with reference to the accompanying drawings.
[0075] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of the word "example" is intended to present concepts in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.
[0076] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, their intended meanings are the same. "of", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, their intended meanings are the same.
[0077] In the embodiments of the present invention, sometimes subscripts such as W 1 may be miswritten as non-subscript forms such as W1. When the difference is not emphasized, their intended meanings are the same.
[0078] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0079] The embodiments of the present invention provide a method for extracting database code based on rapid screening of large language models. This method can be implemented by an electronic device, which can be a terminal or a server.
[0080] The present invention is applicable to mainstream large language models in the market, such as large language models like ChatGPT, Llama, Hundred Rivers Intelligence, ERNIE Bot, Tongyi Qianwen, etc., or AI large language models. A large language model (LLM) is a deep learning model trained based on a vast amount of text data. It not only has the ability to generate natural language text but also can deeply understand the meaning of the text and handle various natural language tasks, such as text summarization, question answering, translation, etc.
[0081] Through deep learning technology, large language models use multi-layer neural networks to model the statistical laws and latent semantic information of language. During the training process, large language models will learn and abstract a large amount of text data, so as to be able to generate logical and coherent language outputs. At the same time, in order to ensure that the model has good generalization ability, it is also necessary to collect and organize large-scale data sets for training.
[0082] Large language models have demonstrated powerful capabilities and potential in the fields of natural language processing, machine translation, dialogue systems, text generation, etc. It can understand human natural language input and generate semantically relevant outputs according to the input content. By learning a large amount of text data, large language models can obtain an in-depth understanding of aspects such as language structure, grammar, and semantics.
[0083] In this embodiment, the AI large language model can be selected by the user, such as Chat GPT, Llama, ERNIE Bot, Tongyi Qianwen.
[0084] Regarding the semantic interaction of the AI large language model, it can be understood in combination with the logical architecture of the AI large language model, which will not be elaborated in this embodiment.
[0085] The present invention proposes a flowchart of a database code extraction method based on rapid screening of a large language model. The processing flow of this method can include the following steps:
[0086] S1. Obtain the database type and version in the user code library file;
[0087] S2. According to the database type and version, find the corresponding database keywords from the preset database code keyword list and perform word segmentation enhancement processing;
[0088] S3. Based on the enhanced database code keywords, screen out the files containing the keywords from the user code library file and mark them as program files containing database code;
[0089] S4. Build a basic prompt word for the large language model based on the database code keywords, and add new database query syntax to the basic prompt word to generate new large model prompt word information and input it into the preset large language model;
[0090] S5. Activate the large language model, and let the large language model extract the corresponding database code from the program files containing database code based on the large model prompt word information and output it.
[0091] As Figure 1 shown in the operation flowchart, the present invention mainly inputs the user code library file into the system, trains the large language model through the constructed large model prompt word information, and the large language model performs an automatic database query operation based on the prompt word, and the large model returns the database code (file) obtained from the database query and outputs it.
[0092] Specifically: The present invention obtains the user code library file, extracts the database configuration information, connects to the database to obtain the database type and version. Then, according to the database type and version, it extracts the database keywords and supplementary syntax definitions. It uses the large language model to traverse all program files, and quickly screens out the program files containing database code through the database keywords. It constructs the corresponding large model prompt word information through the keywords and inputs it into the large language model, allowing the large language model to traverse all files containing database code, sending the database type, version, and supplementary syntax to the large model, and extracting the database code content. Therefore, based on the database code extraction method of the AI large language model, users can accurately and efficiently automate the extraction of database code from the software file code library and apply it to database code review, release, and database type code conversion scenarios.
[0093] The following will specifically describe the implementation steps of the present invention.
[0094] Further, S1: Obtain the database type and version in the user code library file, including:
[0095] Input the user code library file;
[0096] Parse the user code library file to obtain all program files;
[0097] Traverse each program file to check whether the program file contains a database configuration file:
[0098] If not, traverse to the next one until all program files have been checked;
[0099] If so, read the database configuration file and parse the database configuration file to obtain the basic data corresponding to the database type and version, and obtain the database type and version.
[0100] As Figure 2 shown, extract all program code files according to the program code package uploaded by the user or the provided code library address. Parse the files to obtain the configuration files therein.
[0101] Traverse each code package to find the database configuration file (the information in the configuration file includes: type, connection protocol, connection address, login name, password), extract the database configuration file information, parse the configuration file format, extract the database type and connection information of the database link address, and then write code according to the information to connect to the database and obtain the database type and version used by the program code (that is, obtain the basic data corresponding to the database type and version, including the corresponding keywords and characteristic syntax).
[0102] The format of the user code library file is determined by the user. If it is a code library address, link based on the address to obtain the program code file.
[0103] Further, S2: Search for the corresponding database keywords from the preset database code keyword list according to the database type and version and perform word segmentation enhancement processing, including:
[0104] Pre-construct the database code keyword list composed of different versions of databases and the corresponding keywords;
[0105] According to the database type and version, search for the corresponding database code keywords from the preset database code keyword list and input them into the preset Bloom filter to generate a Bloom filter data structure for data retrieval;
[0106] Using the word segmentation processing technology, identify and extract the word segments in the code file content of the user code library file, and write the word segments into the Bloom filter data structure and input them into the Bloom filter.
[0107] Here, keywords need to be combined to retrieve the target file. For accurate and wider retrieval, the Bloom filter is used for retrieval, and files with database code are quickly identified through the Bloom filter method.
[0108] As Figure 3 shown, in order to quickly match keywords corresponding to different database types and versions, a database code keyword list (composed of database types, versions, and keywords, with several keywords set for different database types and versions, which can be user-defined) is created here, and the keywords corresponding to the database version can be quickly found. Therefore, according to the database type and version, the corresponding database code keywords can be found from the preset database code keyword list.
[0109] The word segmentation processing technology, such as the lex lexical analysis algorithm, has the following main steps for word segmentation extraction:
[0110] Lexical Analysis;
[0111] Input Buffer;
[0112] Token Initialization;
[0113] Character Reading;
[0114] Token Recognition;
[0115] Token Conversion;
[0116] Token Output;
[0117] Error Handling.
[0118] Therefore, here the lex lexical analysis algorithm can be used to identify and extract the word segments in the code file content of the user code library file, including:
[0119] Traverse the user code library file and scan each character;
[0120] Match the longest vocabulary according to the dictionary;
[0121] Add the matched vocabulary to the word segment list;
[0122] Remove the matched words and continue to tokenize the remaining text;
[0123] Repeat steps 3 - 5 until the entire text is tokenized.
[0124] The generation of the Bloom filter data structure can be automatically generated by the data format and template of the Bloom filter.
[0125] Further, S3. Based on the enhanced database code keywords, screen out the files containing the keywords from the user code library files and mark them as program files containing database code, including:
[0126] Traverse all program files;
[0127] Perform tokenization and parsing on each program file to obtain the tokens of the corresponding program file;
[0128] Use the Bloom filter to perform database keyword matching on each program file:
[0129] If a match is found, mark the matched file as a program file containing database code; determine whether all traversals are completed: if so, traverse all program files marked as containing database code and proceed to the next step; if not, continue traversing;
[0130] If no match is found, determine whether all traversals are completed: if so, traverse all program files marked as containing database code and proceed to the next step; if there are still other un-traversed program files, continue traversing;
[0131] Save the program files containing database code to the software file code library.
[0132] Specifically, when operating:
[0133] Find the pre-prepared database code keyword list by database type and version. For various database keyword lists, refer to Figure 3 Example; generate the Bloom filter data structure for all keywords corresponding to the database version.
[0134] Such as Figure 4 As shown, first start traversing all program files, perform tokenization and parsing on each traversed program file to obtain the tokens of the corresponding program file (perform tokenization on the code file content, and here you can use a similar lex lexical analysis algorithm or tool to obtain all tokens), and quickly match each token with the database code keywords through the Bloom filter; if any database keyword is found to be matched, mark this file as a file containing database code and proceed to the next step; if no match is found, continue traversing until all traversals are completed;
[0135] If there is no match, it is determined whether all traversals are completed: if so, all program files marked as containing database code are traversed, and the next step is entered; if not all traversals are completed, continue traversing;
[0136] Save the program file containing the database code to the software file code library.
[0137] Next, the construction of large model prompt words will be combined with keywords.
[0138] Further, S4. Based on the database code keywords, construct basic prompt words for the large language model, and add new database query syntax to the basic prompt words, generate new large model prompt word information and input it into the preset large language model, including:
[0139] Prepare the basic prompt words of the large language model, where the basic prompt words are constructed based on the database type and version in the database code keywords;
[0140] Based on the basic prompt words, construct the corresponding database query syntax;
[0141] Organize the database query syntax, generate the large model prompt word information and input it into the large language model.
[0142] This application needs to use a large language model (or large model) to traverse all files marked as containing database code, organize large model prompt word information and send it to the large language model to extract database code. Specifically:
[0143] Prepare the basic prompt words to be sent to the large language model. The prompt words contain database type and version information, which can make the large language model extract content more accurately (the specific information that can be included in the prompt words can refer to, for example, Figure 5 shown in "database type, version, new syntax, program file, extraction instruction, return format requirement", etc.).
[0144] For new database types or versions, append the additional database syntax to the prompt words to tell the large language model the new corpus information. The additional syntax for various databases is prepared in advance. See the following description of adding new database query syntax to the large language model for details.
[0145] The large language model can screen files that meet the characteristics of the prompt words from all program files based on the database query syntax constructed from keywords such as database type and version in the prompt words, and use them as target files (program files containing database code). Therefore, the model can be used to screen files instead of manual work, improving work efficiency.
[0146] During the software development process, the database type and version change relatively quickly. Therefore, the present invention proposes a supplementary format for database syntax that can be better understood by large language models, automatically updates the prompt words according to the database type and version, and accurately identifies them. Specifically, when subsequently traversing the program files using the large language model, it is checked in real time whether new database types and versions appear in the dynamically input user code library file:
[0147] If they exist, new database query syntax is constructed based on the new database type and version and added to the large language model;
[0148] If not, give up.
[0149] Example of the supplementary format for database syntax definition:
[0151] {
[0152] "COMMAND_KEY":"CREATE EVENT",
[0153] "COMMAND_DATABASE_TYPE":"MYSQL",
[0154] "COMMAND_DATABASE_VERSION":"5.1"
[0155] "COMMAND_GRAMMAR":"CREATE EVENT {event name} ON SCHEDULE
[0156] {schedule} Do {event body}"
[0157] },{
[0158] "COMMAND_KEY":"CREATE_PLUGGIN_DATABASE",
[0159] "COMMAND_DATABASE_TYPE":"ORACLE",
[0160] "COMMAND_DATABASE_VERSION":"12"
[0161] "COMMAND_GRAMMAR":"CREATE_PLUGGABLE_DATABASE{pdb_database_name}
[0162] FILES {file_ path1,file path2,...}"}
[0163] .
[0164] During the process of the large language model extracting database code, the system can, through step S1, monitor in real time and dynamically, automatically identify the database type and version, find the corresponding database keywords through the database version and type, efficiently identify the program files containing database code from all program files, and add database type version and additional database syntax information through the combined large language model prompt words to make the large language model more accurately extract the database code content.
[0165] The large language model has a token length limit. For extremely large program files, the present invention calculates the token length, performs lexical segmentation, and sends the file in slices to the large language model, so as to achieve the ability to completely extract extremely large program files.
[0166] Further, before activating the large language model in S5 and extracting the corresponding database code from the program files containing database code and outputting it by the large language model based on the large model prompt word information, it further includes:
[0167] Traverse all program files marked as containing database code;
[0168] Calculate the file size: token length of each program file containing database code:
[0169] ,
[0170] B is the token length,
[0171] n is the number of text characters in the program file,
[0172] Lsi is the word length of each text;
[0173] Judge whether the file size: token length of the program file containing database code exceeds the token limit of the large language model:
[0174] If so, perform slicing processing on the program file containing database code to obtain several program slice files and input them into the large language model;
[0175] If not, directly input the program file containing database code into the large language model.
[0176] Such as Figure 5As shown, when sending each program file to the large language model, it is necessary to perform smooth transmission of the program file according to the token limit of the large language model. If the token size of the program file exceeds the limit of the large language model, it is necessary to fragment the oversized program file. Several program fragment files of the program file are combined and input into the large language model (each program fragment file needs to be marked with the attributes, numbers, etc. of the original program file for subsequent collection based on the data of the large language model).
[0177] In the calculation regulations of the size of the program file: the token length in the present invention, due to different types of large language models, the algorithms will be different. Specifically, it is necessary to determine the definition in combination with the selected large language model. For example:
[0178] For English large models such as ChatGPT\Llama, the algorithm is as follows:
[0179] 1 English word = 1,
[0180] 1 Chinese character = 1.5,
[0181] 1 punctuation mark = 1,
[0182] 1 digit = 0.5;
[0183] For Chinese large models such as Wenxin Yiyan\Tongyi Qianwen, the algorithm is as follows:
[0184] 1 English word = 1.5,
[0185] 1 Chinese character = 1,
[0186] 1 punctuation mark = 1,
[0187] 1 digit = 0.5.
[0188] The steps for fragmenting and sending the file are as follows:
[0189] Reference for the file fragmentation algorithm steps:
[0190] 1. Read the token length limit for each interaction of the large language model. The limits of each large model are different, usually 4K~128K;
[0191] 2. Read the file content;
[0192] 3. Calculate the token length of each word in the file according to the previous token length calculation method, and then cut it into several fragments according to the token length limit for each interaction of the large model. The size of each fragment does not exceed the token length limit for each interaction of the large model;
[0193] 4. The large language model performs extraction operations on each shard in sequence until all shards are completed.
[0194] Through calculating the token length and performing lexical segmentation, the present invention sends the file shards to the large language model, thereby achieving the ability to completely extract an ultra-large program file.
[0195] Further, in S5, activate the large language model, and the large language model extracts the corresponding database code from the program file containing the database code based on the large model prompt word information and outputs it, including:
[0196] Send an extraction instruction to the large language model to activate the large language model;
[0197] Based on the large model prompt word information, extract and return the corresponding database code from the program file containing the database code through the large language model;
[0198] Parse the database code and save it as the database code in the corresponding structured format.
[0199] As Figure 5 shown, send an extraction instruction to the large language model, and require the large language model to return the database code content in the program file according to the large model prompt word information prepared above (the organized large model prompt words, such as "database type, version, new syntax, program file, extraction instruction, return format requirement"). After the large language model retrieves the corresponding extraction result, according to the requirements in the prompt words, parse the content returned by the large language model, generate content in a similar json structured format, with each database SQL code as a record, and save it to the result.sql file. This is the final file content of the extracted database code. It can be used for database code review, release, or code rewriting and replacement.
[0200] As shown in the following general Java language code example:
[0201] import java.sql.*;
[0202] public class demo{
[0203] public static void updateproduct(connection connection,int id,stringname,string type){String sql="update product p set name=?,type=?";
[0204] if(type!=null){
[0205] sql = sql + ", type = ?"
[0206] }
[0207] sql = sql + " where id = ?"
[0208] PrepareStatement statement = connection.prepareStatement(sql);
[0209] statement.setString(1, name);
[0210] if (type != null) {
[0211] statement.setString(2, type);
[0212] statement.setInt(3, id);
[0213] } else {
[0214] statement.setInt(2, id);
[0215] }
[0216] statement.execute(sql);
[0217] statement.close();
[0218] }
[0219] public static void deleteProduct(Connection connection, int id) {
[0220] String sql = "delete product p";
[0221] if (id != null) {
[0222] sql = sql + " where id =?"
[0223] } else {
[0224] throw new Exception("error, id is empty!")
[0225] } It should be noted that there are some potential issues in the original code, such as the incorrect spelling "nu1l" which should be "null", and the incorrect use of "statement.setInt(3,id);" and "statement.setInt(2,id);" in different logical branches which might lead to incorrect SQL execution. Also, the "connection" variable should be of the correct type (presumably a JDBC `Connection` object). And the "where id.?" in line 53 should be "where id =?". These are just for code review purposes and not directly related to the translation task.
[0226] Preparestatement statement = connection.preparestatement(sql);
[0227] statement.setInt(1,id);
[0228] statement.execute(sql);
[0229] statement.close();
[0230] }.
[0231] Example of database code results quickly screened by the large language model of the present invention:
[0233] {"sql":"update product p set name=?,type=? where id=?","line":5},
[0234] {"sql":"delete product p where id=?","line":22}
[0235] 。
[0236] Therefore, in contrast, the present invention can provide a database syntax supplement format that can be better understood by the large language model, an automated database type and version accurate identification scheme, and provides a method for quickly screening and filtering database code files, which can effectively solve the problems of low accuracy and efficiency in extracting database code by the large language model.
[0237] Figure 6 It is a block diagram of a database code extraction system for quick screening based on a large language model shown according to an exemplary embodiment. This system is used for the method of extracting database code for quick screening based on a large language model. Refer to Figure 6 , this system includes a file acquisition module 610, a keyword search module 620, a file screening module 630, an intelligent screening module 640, and a code extraction module 650. Among them:
[0238] The file acquisition module is used to acquire the database type and version in the user code library file;
[0239] The keyword search module is used to search for corresponding database keywords from the preset database code keyword list according to the database type and version and perform word segmentation enhancement processing;
[0240] A file screening module, configured to screen out files containing keywords from the user code library files based on the enhanced database code keywords and mark them as program files containing database code;
[0241] An intelligent screening module, configured to build basic prompt words for a large language model based on the database code keywords, add new database query syntax to the basic prompt words, generate new large model prompt word information and input it into a preset large language model;
[0242] A code extraction module, configured to activate the large language model, and the large language model extracts corresponding database code from the program files containing database code based on the large model prompt word information and outputs it.
[0243] For the functional services and interaction processes of the above-mentioned various modules, please understand them in combination with the corresponding steps and principles of the above method, and they will not be elaborated in this embodiment.
[0244] Figure 7 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 7 shown, the electronic device may include the above-mentioned Figure 6 database code extraction system for rapid screening based on a large language model shown. Optionally, the electronic device 410 may include a first processor 2001.
[0245] Optionally, the electronic device 410 may further include a memory 2002 and a transceiver 2003.
[0246] Among them, the first processor 2001 is connected to the memory 2002 and the transceiver 2003, such as through a communication bus.
[0247] Next, in combination with Figure 7 each component of the electronic device 410 will be specifically introduced:
[0248] Among them, the first processor 2001 is the control center of the electronic device 410, which may be a processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or may be a specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0249] Optionally, the first processor 2001 may execute various functions of the electronic device 410 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0250] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 7 CPU0 and CPU1 shown in
[0251] In a specific implementation, as an embodiment, the electronic device 410 may also include multiple processors, such as Figure 7 the first processor 2001 and the second processor 2004 shown in
[0252] Among them, the memory 2002 is used to store software programs for implementing the solution of the present invention and is controlled by the first processor 2001 for execution. The specific implementation manner may refer to the above method embodiments and will not be elaborated here.
[0253] Optionally, the memory 2002 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently and be coupled to the first processor 2001 through an interface circuit ( Figure 7 not shown in
[0254] The transceiver 2003 is used to communicate with network devices or with terminal devices.
[0255] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 7 not shown separately). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function.
[0256] Optionally, the transceiver 2003 may be integrated with the first processor 2001, or may exist independently, and is coupled to the first processor 2001 through an interface circuit ( Figure 7 not shown) of the electronic device 410. The embodiments of the present invention do not make specific limitations on this.
[0257] It should be noted that Figure 7 the structure of the electronic device 410 shown does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0258] In addition, the technical effects of the electronic device 410 may refer to the technical effects of the database code extraction method based on rapid screening by a large language model described in the above method embodiments, and will not be elaborated here.
[0259] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or this processor may also be any conventional processor, etc.
[0260] It should also be understood that the memory in the embodiments of the present invention can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0261] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable systems. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0262] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be understood specifically with reference to the context before and after.
[0263] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0264] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0265] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled artisans may use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0266] Those skilled in the art can clearly understand that for the sake of convenience and brevity of description, the specific working processes of the devices, systems, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0267] In several embodiments provided by the present invention, it should be understood that the disclosed devices, systems, and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of systems or units can be in electrical, mechanical, or other forms.
[0268] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0269] In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0270] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0271] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A database code extraction method based on rapid screening of a large language model, characterized in that: The method comprises: S1. Obtain the database type and version in the user code library file; S2. According to the database type and version, the corresponding database keywords are searched from the preset database code keyword list and word segmentation enhancement is performed; S3. Based on the enhanced database code keywords, files containing keywords are screened out from the user code library files and marked as program files containing database codes; S4, constructing basic prompt words for the large language model based on the database code keywords, wherein the basic prompt words include database type, version, new syntax, program file, extraction instruction, and return format requirements, and adding new database query syntax to the basic prompt words, generating new large model prompt word information and inputting it into the preset large language model; S5. activating the large language model, and having the large language model extract corresponding database codes from the program file containing the database codes based on the large model prompt word information, and outputting the codes.
2. The database code extraction method based on large language model rapid screening according to claim 1 is characterized in that: S1. Get the database type and version in the user code library file, including: Input the user code library file; Parse the user code library file to obtain all program files; Traverse each program file to see if it contains a database configuration file: If not, traverse the next one until all program files are checked; If yes, the database configuration file is read and parsed to obtain basic data corresponding to the database type and version, thereby obtaining the database type and version.
3. The database code extraction method based on large language model rapid screening according to claim 1 is characterized in that: S2. According to the database type and version, the corresponding database keywords are searched from the preset database code keyword list and word segmentation enhancement processing is performed, including: Pre-building the database code keyword list constructed by different versions of databases and corresponding keywords; According to the database type and version, the corresponding database code keyword is searched from the preset database code keyword list and input into the preset Bloom filter to generate a Bloom filter data structure for data retrieval; The word segmentation processing technology is used to identify and extract the word segments in the code file content of the user code library file, and the word segments are written into the Bloom filter data structure and input into the Bloom filter.
4. The database code extraction method based on large language model rapid screening according to claim 3 is characterized in that: S3. Based on the enhanced database code keywords, files containing keywords are screened out from the user code library files and marked as program files containing database codes, including: Traverse all program files; Perform word segmentation analysis on each program file to obtain the word segmentation of the corresponding program file; Use Bloom filters to match database keywords for each program file: If it is found that the program file matches the keyword, the program file matching the keyword is identified as a program file containing the database code; If no matching keywords are found in the program file, proceed to the next step; Determine whether all the program files have been traversed: If so, all traversals are completed and proceed to the next step; If there are other program files that have not been traversed, continue to traverse the program files; The program file containing the database code is saved in a software file code library.
5. The method for extracting database codes based on rapid screening of a large language model according to claim 4 is characterized in that: S4, constructing basic prompt words for the large language model based on the database code keywords, adding new database query syntax to the basic prompt words, generating new large model prompt word information and inputting it into the preset large language model, including: Preparing basic prompt words of a large language model, wherein the basic prompt words are constructed based on the database type and version in the database code keywords; Based on the basic prompt words, construct corresponding database query grammar; Organizing the database query grammar, generating the large model prompt word information and inputting it into the large language model; When the large language model is used to traverse the program files later, it is checked in real time whether a new database type and version appears in the dynamically input user code library file: If so, based on the new database type and version, a new database query grammar is constructed and added to the large language model; If it doesn't exist, give up.
6. The method for extracting database codes based on rapid screening of a large language model according to claim 4, characterized in that: Before S5, activating the large language model, and extracting and outputting the corresponding database code from the program file containing the database code based on the large model prompt word information by the large language model, the method further includes: Traverse all program files identified as containing database code; Calculate the file size of each program file containing the database code: token length: , B is the token length, n is the number of text characters in the program file, Lsi is the word length of each text; Determine whether the file size of the program file containing the database code: token length exceeds the token limit of the large language model: If yes, the program file containing the database code is fragmented to obtain a plurality of program fragment files and input them into the large language model; If not, the program file containing the database code is directly input into the large language model.
7. The method for extracting database codes based on rapid screening of a large language model according to claim 6, characterized in that: S5, activating the large language model, and having the large language model extract corresponding database codes from the program file containing the database codes based on the large model prompt word information and outputting the codes, including: issuing an extraction instruction to the large language model to activate the large language model; Extracting and returning the corresponding database code from the program file containing the database code based on the large model prompt word information through the large language model; The database code is parsed and saved as a database code in a corresponding structured format.
8. A database code extraction system based on rapid screening of a large language model, the database code extraction system based on rapid screening of a large language model is used to implement the database code extraction method based on rapid screening of a large language model as described in any one of claims 1 to 7, characterized in that: The system comprises: The file acquisition module is used to obtain the database type and version in the user code library file; The keyword search module is used to search for corresponding database keywords from a preset database code keyword list and perform word segmentation enhancement processing according to the database type and version; A file screening module, for screening out files containing keywords from the user code library files based on the enhanced database code keywords and marking them as program files containing database codes; An intelligent screening module, used to construct basic prompt words for a large language model based on the database code keywords, and to add new database query syntax to the basic prompt words, generate new large model prompt word information and input it into a preset large language model; The code extraction module is used to activate the large language model, and the large language model extracts and outputs the corresponding database code from the program file containing the database code based on the large model prompt word information.
9. An electronic device, characterized in that: The electronic device comprises: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Data batch generation method for limiting large language model based on token training
CN118331890A
Code retrieval method and device based on large language model
CN118643120A