Text generation method and device, electronic equipment and storage medium
By extracting database object identification, recalling and filtering metadata information, and building prompt text, the problem of insufficient accuracy of the general code completion model in database operation scenarios is solved, and a better understanding of vertical knowledge is achieved and the accuracy of database operation text completion is improved.
Patent Information
- Application Number
- CN202510240392.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
AI Technical Summary
The code text generated by the general code completion model is insufficient in database operation scenarios and cannot effectively utilize vertical knowledge.
By obtaining the currently entered database operation text and cursor position, extracting the database object identity, recalling candidate metadata information of multiple data tables, sorting and filtering to obtain target metadata information, building prompt text with context fragments, and inputting it into the text completion model for prediction to generate complete text.
It improves the understanding of vertical knowledge by the text completion model and improves the accuracy of generating database operation text completion text.
Smart Images

Figure CN120144017A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a text generation method, apparatus, electronic device, and storage medium. Background Art
[0002] With the development of artificial intelligence technology, developers have begun to use code assistants to assist in improving the efficiency of code writing. In related technologies, a general completion model is generally used for code completion. However, the information that can be obtained by the general completion model is limited, and in the scenario of database operations, the accuracy of the generated code text needs to be improved. Summary of the Invention
[0003] The following is an overview of the subject matter described in detail in the present disclosure. This overview is not intended to limit the scope of protection of the claims.
[0004] Embodiments of the present disclosure provide a text generation method, apparatus, electronic device, and storage medium, which can improve the accuracy when generating a completion text.
[0005] On the one hand, embodiments of the present disclosure provide a text generation method, including:
[0006] Obtain the currently input database operation text and the current cursor position, and extract the database object identifier in the database operation text based on the cursor position;
[0007] Recall the candidate metadata information of multiple data tables based on the database object identifier, sort the multiple candidate metadata information and then perform screening to obtain the target metadata information;
[0008] Obtain the context fragment adjacent to the cursor position, and construct a prompt text based on the target metadata information and the context fragment, where the prompt text is used to prompt a text completion model to complete the database operation text;
[0009] Input the prompt text into the text completion model for text prediction to generate a completion text of the database operation text.
[0010] On the other hand, embodiments of the present disclosure further provide a text generation apparatus, including:
[0011] An extraction module, configured to obtain the currently input database operation text and the current cursor position, and extract the database object identifier in the database operation text based on the cursor position;
[0012] A recall module, configured to recall the candidate metadata information of multiple data tables based on the database object identifier, sort the multiple candidate metadata information and then perform screening to obtain the target metadata information;
[0013] A building block for obtaining a context segment adjacent to the cursor position and constructing a prompt text based on the target metadata information and the context segment, where the prompt text is used to prompt a text completion model to complete the database operation text;
[0014] An inference module for inputting the prompt text into the text completion model for text prediction to generate a completion text of the database operation text.
[0015] Furthermore, the extraction module is further configured to:
[0016] Extract the table name text after performing a normalization transformation on the database operation text;
[0017] Divide the database operation text into multiple statement blocks, and respectively extract the field name text from each of the statement blocks based on a preset parsing library;
[0018] Merge the table name text and the field name text belonging to the same statement block to obtain the database object identifiers in each of the statement blocks of the database operation text.
[0019] Furthermore, the extraction module is further configured to:
[0020] Perform field name extraction on each of the statement blocks based on a preset parsing library. When the extraction based on the preset parsing library fails, traverse each line of the statement block;
[0021] If the database syntax keyword is missing at the beginning of the currently traversed line, add the database syntax keyword at the beginning of the currently traversed line and then extract the field name text.
[0022] Furthermore, the extraction module is further configured to:
[0023] Perform a prefix judgment on the table name text according to the positional relationship between the cursor position and the table name text to obtain a first judgment result, and perform a type division on the table name text according to the first judgment result, where the first judgment result is used to indicate that the table name text is a table name prefix or a complete table name;
[0024] Perform a prefix judgment on the field name text according to the positional relationship between the cursor position and the field name to obtain a second judgment result, and perform a type division on the field name text according to the second judgment result, where the second judgment result is used to indicate that the field name text is a field name prefix or a complete field name;
[0025] Merge the table name text and the field name text that belong to the same statement block and have been type-divided to obtain the database object identifiers in each statement block of the database operation text.
[0026] Further, the recall module is also used for:
[0027] Determine the business space associated with the target account for which the database operation text is input, and recall the metadata information of multiple recall paths based on the business space, the cursor position, and the database object identifier;
[0028] Merge the recall results of each recall path to obtain the candidate metadata information of multiple data tables.
[0029] Further, the recall module is also used for:
[0030] For the recall path corresponding to the business space, sort the data tables in the business space by popularity, and determine the first target data table from the data tables in the business space according to the sorting result, and recall the metadata information of the first target data table;
[0031] For the recall path corresponding to the cursor position and the database object identifier, determine the adjacent identifier of the cursor position in the database object identifier, and determine the data table containing the adjacent identifier as the second target data table, and recall the metadata information of the second target data table.
[0032] Further, the recall module is also used for:
[0033] Determine the second target data table as the data table that simultaneously contains the adjacent complete field name, the adjacent field name prefix, and the adjacent table name prefix.
[0034] Further, the recall module is also used for:
[0035] Obtain the path weights of each recall path, determine the sorting weights of the candidate metadata information according to the path weights, and perform an initial sorting on the multiple candidate metadata information based on the sorting weights to obtain an initial sorting result;
[0036] Based on at least one of the confidence level of the candidate metadata information, the access permission of the target account to the candidate metadata information, or the integrity verification result of the candidate metadata information, re-sort the initial sorting result to obtain a re-sorting result;
[0037] Screen and obtain the target metadata information from the re-sorting result.
[0038] Further, the recall module is also used for:
[0039] For any one of the candidate metadata information, splice the database operation text and the candidate metadata information and input them into a weight prediction model for prediction to obtain the feature weight of the candidate metadata information;
[0040] Perform a weighted sum of the feature weight and the path weight to obtain the sorting weight of the candidate metadata information.
[0041] Furthermore, the construction module is further configured to:
[0042] Determine the dialect information of the context segment, where the dialect information is used to indicate the dialect used in the context segment;
[0043] Obtain the current time when the database operation text is input, and construct a prompt text based on the dialect information, the context segment, the target metadata information, and the current time.
[0044] Furthermore, the construction module is further configured to:
[0045] Construct a dialect information text based on the dialect information and a first identification text for identifying the dialect information;
[0046] Construct a table creation text in the form of a data definition language based on the target metadata information, and construct a metadata information text based on the table creation text and a second identification text for identifying the table creation text;
[0047] Construct a context information text based on the context segment and a third identification text for identifying the context segment;
[0048] Construct a prompt text based on the dialect information text, the metadata information text, the current time, the context information text, and a fourth identification text for identifying the completion text.
[0049] Furthermore, the construction module is further configured to:
[0050] Retrieve a historical reference text similar to the database operation text;
[0051] Construct a metadata information text based on the table creation text after adding the enumeration value and a second identification text for identifying the table creation text.
[0052] Furthermore, the construction module is further configured to:
[0053] Retrieve a historical reference text similar to the database operation text;
[0054] Construct a prompt text based on the dialect information text, the metadata information text, the current time, the historical reference text, the context information text, and a fourth identification text for identifying the completed text.
[0055] Further, the inference module is further configured to:
[0056] Input the prompt text into the text completion model for streaming text prediction, and cache the predicted words output by the text completion model in each round;
[0057] When the text completion model ends, obtain a prediction result based on the predicted words cached in multiple rounds, perform an inhibition process on the prediction result, and generate a completed text of the database operation text, where the inhibition process is used to delete abnormal content in the prediction result.
[0058] On the other hand, an embodiment of the present disclosure further provides an electronic device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the above-mentioned text generation method is implemented.
[0059] On the other hand, an embodiment of the present disclosure further provides a computer-readable storage medium, where the storage medium stores a computer program, and the computer program is executed by a processor to implement the above-mentioned text generation method.
[0060] On the other hand, an embodiment of the present disclosure further provides a computer program product, which includes a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes to implement the above-mentioned text generation method.
[0061] The embodiments of the present disclosure at least include the following beneficial effects: By obtaining the currently input database operation text and the current cursor position, extracting the database object identifier in the database operation text based on the cursor position, it is possible to recall candidate metadata information of multiple data tables based on the database object identifier, thereby mining more vertical knowledge, sorting and screening multiple candidate metadata information, and then obtaining target metadata information with higher relevance. Subsequently, obtain the context fragment adjacent to the cursor position, construct a prompt text based on the target metadata information and the context fragment, and input the prompt text into the text completion model for text prediction, which can achieve metadata enhancement during the inference of the text completion model, improve the text completion model's understanding ability of vertical knowledge, and effectively improve the accuracy when the text completion model generates the completed text of the database operation text.
[0062] Other features and advantages of the present disclosure will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present disclosure. Description of the Drawings
[0063] The drawings are used to provide a further understanding of the technical solutions of the present disclosure, and constitute a part of the description. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure, and do not constitute a limitation to the technical solutions of the present disclosure.
[0064] Figure 1 Schematic diagram of an optional implementation environment provided for an embodiment of the present disclosure;
[0065] Figure 2 Optional flowchart of the text generation method provided for an embodiment of the present disclosure;
[0066] Figure 3 Optional schematic diagram of dividing a statement block provided for an embodiment of the present disclosure;
[0067] Figure 4 Optional flowchart of the preprocessing of database operation text provided for an embodiment of the present disclosure;
[0068] Figure 5 Optional flowchart of the recall of metadata information provided for an embodiment of the present disclosure;
[0069] Figure 6 Optional flowchart of the sorting of metadata information provided for an embodiment of the present disclosure;
[0070] Figure 7 Optional schematic diagram of the prompt text provided for an embodiment of the present disclosure;
[0071] Figure 8 Optional flowchart of constructing the prompt text provided for an embodiment of the present disclosure;
[0072] Figure 9 Optional schematic diagram of the post-processing of the prediction result of the text completion model provided for an embodiment of the present disclosure;
[0073] Figure 10 Optional flowchart of the inference process of the text completion model provided for an embodiment of the present disclosure;
[0074] Figure 11 Optional application overall framework of the text generation method provided for an embodiment of the present disclosure;
[0075] Figure 12 Provided for an embodiment of the present disclosure Figure 11 Optional overall schematic framework of the algorithm side in
[0076] Figure 13 Schematic structural diagram of the text generation device provided by an embodiment of the present disclosure;
[0077] Figure 14 Partial structural block diagram of the terminal provided by an embodiment of the present disclosure;
[0078] Figure 15 Partial structural block diagram of the server provided by an embodiment of the present disclosure. Detailed implementation manners
[0079] In order to make the objectives, technical solutions and advantages of the present disclosure more clear and understandable, the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not used to limit the present disclosure.
[0080] It should be noted that in each specific implementation manner of the present disclosure, when it comes to performing relevant processing based on data related to the characteristics of the target object, such as target object attribute information or a set of attribute information, the permission or consent of the target object will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. Among them, the target object may be a user. In addition, when an embodiment of the present disclosure needs to obtain target object attribute information, it will obtain the separate permission or separate consent of the target object through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the separate permission or separate consent of the target object, the necessary data related to the target object for enabling the normal operation of the embodiment of the present disclosure will be obtained.
[0081] In an embodiment of the present disclosure, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0082] To facilitate the understanding of the technical solutions provided by the embodiments of the present disclosure, some key terms used in the embodiments of the present disclosure are explained here first:
[0083] Metadata: Refers to data used to describe data, mainly used to record relevant information of the data, including the definition, source, format, update time, etc. of the data, to help understand, manage, and use the data.
[0084] Vertical knowledge: Refers to a professional knowledge system that has been systematically organized and summarized in a specific field, including the basic principles, knowledge structure, development context, etc. of that field.
[0085] With the development of artificial intelligence technology, developers have begun to use code-based assistants to assist in improving the efficiency of code writing. In related technologies, a general completion model is generally used for code completion. However, the information that can be obtained by the general completion model is limited, and in the scenario of database operations, the accuracy of the generated code text needs to be improved.
[0086] With the development of artificial intelligence technology, developers have begun to use code-based assistants to assist in improving the efficiency of code writing. Code completion processing, as an important application of code-based assistants, plays a crucial role in code editing. In related technologies, enhancing the code completion function usually involves training and optimizing the general code completion model to obtain rich context information, and based on information such as the editing environment of the current text editor, running files, and code reference relationships, prompting the general code completion model to perform code completion. However, the content provided by this method that relies on information such as the editing environment of the text editor is limited, and it is difficult to effectively supplement relevant vertical knowledge, resulting in the general code completion model being unable to adapt to the completion scenario of query languages, leading to situations where the general code completion model may generate fields and tables that do not exist in the database, unable to associate context semantic information, generate mismatched enumeration values, etc., affecting the accuracy of the generated code text.
[0087] Based on this, the embodiments of the present disclosure provide a text generation method, device, electronic device, and storage medium, which can improve the understanding ability of the text completion model for vertical knowledge, so that the text completion model can effectively improve the accuracy when generating the completion text of database operation texts.
[0088] Refer to Figure 1 , Figure 1 FIG. is a schematic diagram of an optional implementation environment provided by the embodiments of the present disclosure. The implementation environment includes a terminal 101 and a server 102, where the terminal 101 and the server 102 are connected through a communication network.
[0089] Exemplarily, in the terminal 101, the currently entered database operation text and the current cursor position are obtained, and the database operation text and the current cursor position are sent to the server 102. The server 102 extracts the database object identifier in the database operation text based on the cursor position, recalls the candidate metadata information of multiple data tables based on the database object identifier, sorts and filters the multiple candidate metadata information to obtain the target metadata information, obtains the context fragment adjacent to the cursor position in the database operation text, constructs a prompt text based on the target metadata information and the context fragment, and the prompt text is used to prompt the text completion model to complete the database operation text. Finally, the prompt text is input into the text completion model for text prediction to generate the completion text of the database operation text, and the completion text is sent to the terminal 101 to automatically complete the database operation text.
[0090] It can be understood that the text generation method provided in the embodiments of the present disclosure can also be executed alone in the terminal 101.
[0091] The server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In addition, the server 102 can also be a node server in a blockchain network.
[0092] The terminal 101 can be a mobile phone, a computer, a smart voice interaction device, a smart wearable device, a smart home appliance, a vehicle-mounted terminal, etc., but is not limited thereto. The terminal 101 and the server 102 can be directly or indirectly connected through wired or wireless communication methods, and the embodiments of the present disclosure do not limit this.
[0093] Refer to Figure 2 , Figure 2 FIG. is an optional flowchart of the text generation method provided in the embodiments of the present disclosure. The text generation method can be executed by the terminal, or by the server, or by the cooperation of the terminal and the server. The text generation method includes but is not limited to the following steps S201 to step S204.
[0094] Step S201: Obtain the currently entered database operation text and the current cursor position, and extract the database object identifier in the database operation text based on the cursor position.
[0095] Among them, the database operation text and the current cursor position are both obtained from a text editor, which is used to edit the database operation text. The database operation text is a multi-line code statement or instruction for interacting with the database or performing database operations, and is used to manage and use the data in the database. For example, when the database is MySQL, its database operation text can be SQL text, and based on the SQL text, the database can be instructed to perform operations such as creating, reading, updating, and deleting data. The current cursor position is located after the currently entered database operation text and is used to indicate the position where the next database operation text will be entered. For example, in the text editor, if the currently entered database operation text is "langu|", then the current cursor position is the position after "u", that is, the position where the vertical bar "|" is located. When the letter "a" is entered next, the database operation text and the cursor position displayed in the text editor can be "langua|". The database object identifier is used to identify a database object, which is the basic structure and entity for storing and managing data, such as a table or a field. Therefore, the database object identifier can be at least one of a table name text and a field name text.
[0096] In a possible implementation, during the process of extracting the database object identifier from the database operation text based on the cursor position, specifically, the database operation text can be normalized and transformed to extract the table name text, the database operation text can be divided into multiple statement blocks, and the field name text can be extracted from each statement block based on a preset parsing library. The table name text and the field name text belonging to the same statement block are merged to obtain the database object identifiers in each statement block of the database operation text. Among them, the table name text is the name of the data table called in the database operation text, which can be composed of one or more of letters, numbers, and symbols. The data table is a table in the database for storing data, with the field name as the table header and consisting of the field name and the data corresponding to the field name. The field name can be obtained by summarizing based on the data meaning or data attributes; the division operation can be based on a text delimiter, which can be a specific syntax keyword, such as "select"; a statement block usually starts with a syntax keyword and includes at least one operation statement, which is used to perform a database operation and can include a syntax keyword and a table name text, or can also include a syntax keyword and a field name text; the preset parsing library is used to extract the field name in the database operation text and can be SQL_META; the field name text is at least one field name selected from the data table called in the database operation text. For example, if the database operation text is "select a,b from table1", then "table1" is the table name text, and "a" and "b" are the field name texts.
[0097] Specifically, preprocess the database operation text. Using the database operation text and the current cursor position as input, perform a normalization transformation on the database operation text to unify the text format. For example, the text case can be unified, converting the identifiers in the database operation text to uppercase and other text content to lowercase; or the text format can be unified, renaming the table names according to the preset naming convention so that all table names follow a unified naming style, or clarifying the rules for using quotation marks, uniformly using single quotes to represent strings and double quotes to represent identifiers. Then, construct a first regular expression based on the editing rules of the table names, extract the table name text from the normalized database operation text based on the first regular expression, and divide the normalized database operation text into multiple statement blocks according to the text delimiter. Refer to Figure 3 , Figure 3 which is an optional schematic diagram for dividing statement blocks provided by an embodiment of the present disclosure. Figure 3 The database operation text in includes multiple operation statements. For the syntax keywords in the multiple operation statements, divide the database operation text using the syntax keyword B as the text delimiter to obtain statement block A, statement block B, and statement block C. Statement block A contains two operation statements, namely "syntax keyword B + field name text A" and "syntax keyword C + table name text A". Statement block B contains "syntax keyword B + field name text B" and "syntax keyword D + table name text B". Statement block C contains "syntax keyword B + field name text C" and "syntax keyword C + table name text C", while the operation statement "syntax keyword A + table name text A" does not belong to any statement block.
[0098] Next, preset a truncation parameter. The truncation parameter is used to limit the number of statement lines in each statement block. Truncate each statement block based on the truncation parameter, that is, retain the first several lines of statements in each statement block. Based on the preset parsing library, extract the field name text from each truncated statement block respectively, and associate and merge the table name text and the field name text belonging to the same statement block to obtain the database object identifiers in each statement block of the database operation text. By preprocessing the database operation text, the database operation text can be made consistent in text representation, avoiding misjudging the same operation text as different content due to format differences, and reducing the occurrence of errors and abnormal situations. In addition, by dividing the database operation text into blocks, parallel processing of multiple statement blocks can be achieved, improving the data preprocessing efficiency. On this basis, extracting the field name text from the truncated statement blocks reduces the extraction range, making the extracted field name text more accurate and also reducing the time wasted in processing overly long statement blocks, so that more accurate database object identifiers can be obtained quickly.
[0099] It should also be noted that when dividing the database operation text, if there are nested statements, the nested statements are regarded as the same statement block. For example, if the database operation text is "select a,b from table1 select c(*)from(select d from table2)", "select a,b from table1" is a statement block, and the structure of "select c(*)from(select d from table2)" is a nested query and is regarded as another statement block.
[0100] In a possible implementation manner, during the process of extracting the field name text from each statement block based on the preset parsing library, specifically, the field name extraction can be performed on each statement block based on the preset parsing library. When the extraction fails based on the preset parsing library, each line of the statement block is traversed. If the database syntax keyword is missing at the beginning of the currently traversed line, the database syntax keyword is added at the beginning of the currently traversed line and then the field name text is extracted.
[0101] Specifically, a second regular expression is constructed based on the editing rule of field name selection, and the preset parsing library is called to extract the field name text from each statement block based on the second regular expression. When the extraction fails with the preset parsing library, the statement block is traversed. If the database syntax keyword is missing at the beginning of the currently traversed line, after adding the database syntax keyword at the beginning of the currently traversed line, each line is parsed to extract the field name text. By executing the line-by-line extraction strategy for supplementary extraction of the field name text, the parsing rate of the statement block can be increased. Especially when the structure of the statement block is incomplete, using this strategy can make the syntax tree of the incomplete statement block more perfect, which is beneficial to the parsing of the incomplete statement block and further improves the parsing rate of the preset parsing library and the accuracy of extracting the field name text.
[0102] In a possible implementation, during the process of merging the table name text and the field name text belonging to the same statement block to obtain the database object identifiers in each statement block of the database operation text, specifically, a prefix judgment can be made on the table name text according to the positional relationship between the cursor position and the table name text to obtain a first judgment result, and the table name text can be classified according to the first judgment result. A prefix judgment is made on the field name text according to the positional relationship between the cursor position and the field name to obtain a second judgment result, and the field name text is classified according to the second judgment result. The table name text and the field name text that belong to the same statement block and have been classified are merged to obtain the database object identifiers in each statement block of the database operation text. Among them, the first judgment result is used to indicate that the table name text is a table name prefix or a complete table name, and the second judgment result is used to indicate that the field name text is a field name prefix or a complete field name; the classification of the table name text can be understood as dividing the table name prefix and the complete table name into two different categories, and the extracted table name text can be distinguished as a table name prefix or a complete table name by adding a new field. For example, if a new field "Table Name Text Type" is added, "T" can be used to represent the table name prefix, and "A" can be used to represent the complete table name. The classification of the field name text is the same.
[0103] Specifically, when extracting the table name text, record the start and end positions of the table name text in the database operation text, that is, record the first start position where the first character of the table name text appears in the database operation text and the first end position where the last character appears in the database operation text. Then, obtain the cursor position, and make a prefix judgment on the table name text according to the positional relationship between the cursor position and the table name text. If the cursor position is adjacent to the last character of the table name text, the type of the table name text is a table name prefix. If the cursor position is not within the table name text and is at least one character away from the last character of the table name text, the type of the table name text is a complete table name. Similarly, when extracting the field name text, record the second start position where the first character of the field name text appears in the database operation text and the second end position where the last character appears in the database operation text. If the cursor position is adjacent to the last character of the field name text, the type of the field name text is a field name prefix. If the cursor position is not within the field name text and is at least one character away from the last character of the field name text, the type of the field name text is a complete field name.
[0104] Next, obtain the third starting position of the first character of the statement block in the database operation text and the third ending position of the last character in the database operation text. When the first starting position and the second starting position are after the third starting position and the first starting position is after the second starting position, the first ending position and the second ending position are before the third ending position and the first ending position is after the second ending position, and the first ending position is before the second starting position, it indicates that the field name text is derived from the table name text, and the field name text and the table name text belong to the statement block. Combine the field name text and the table name text to obtain the database object identifier in the statement block. Similarly, the database object identifiers in all statement blocks in the database operation text can be obtained. By classifying the table name text and the field name text, it is possible to quickly determine whether the table name text and the field name text in the database operation text are complete, providing a basis for subsequent text completion. At the same time, combining the table name text and the field name text of the same statement block to obtain the database object identifier tightens the connection between the data table and the fields in the data table, thereby enhancing the consistency and reliability of the database object identifier. On this basis, the text completion model can quickly locate the required metadata information based on the database object identifier, effectively improving the efficiency of the text completion model in generating the completed text.
[0105] In addition, in order to quickly and accurately combine the table name text and the field name text of the same statement block, a table can be created and saved based on the extracted table name text, field name text, and their respective types. Referring to Table 1, Table 1 is an optional schematic table of the database operation text extraction table provided by the present disclosure embodiment. Combine each row of the database operation text extraction table in Table 1 to obtain a data entry, and use this data entry as the database object identifier of the statement block in the database operation text. The structure of the database object identifier can be table name text - table name text type - field name text - field name text type.
[0106] Table 1 Database Operation Text Extraction Table
[0107] Table Name Text Table Name Text Type Field Name Text Field Name Text Type user.table A name A user.table A age A music.table A ti(tle) T time.ta(ble) T time A
[0108] In Table 1, when the table name text type is A, it indicates that the table name text is a complete table name; when it is T, it indicates that the table name text is a table name prefix. The bracketed parts in the table name text time.ta(ble) and the field name text ti(tle) are the uncompleted parts. In a statement block, the table name text is user.table and the field name text is name. At this time, the database object identifier can be expressed as user.table-A-name-A. In another statement block, the table name text is music.table and the field name text is ti(tle). At this time, the database object identifier can be expressed as music.table-A-ti-T.
[0109] According to the description in step S201, the preprocessing process of the database operation text can refer to Figure 4 , Figure 4 which is an optional process schematic diagram for preprocessing the database operation text provided by the embodiments of the present disclosure. Using the database operation text and the current cursor position as inputs, the table name text is extracted and the type of the table name text is classified. The database operation text is segmented (statement block 1, statement block 2, statement block 3) according to the text delimiter, and each statement block is truncated. The field name text is extracted from each truncated statement block, the type of the field name text is classified, and the table name text and the field name text that belong to the same statement block and have been classified by type are associated and merged to obtain the database object identifiers in each statement block of the database operation text.
[0110] Step S202: Recall the candidate metadata information of multiple data tables based on the database object identifier, sort the multiple candidate metadata information and then perform screening to obtain the target metadata information.
[0111] Among them, the candidate metadata information is the metadata information obtained by recall. The metadata information is data about tabular data, providing descriptions of the table structure, content, and related attributes. The recall operation can be regarded as an operation of screening and sorting the obtained metadata information, which can be recalled from the data tables in the database or from the business space of the associated application. The target metadata information is the metadata information after sorting and screening, and is used to construct the prompt text in the subsequent steps.
[0112] In a possible implementation, during the process of recalling candidate metadata information of multiple data tables based on the database object identifier, specifically, it may be to determine the business space associated with the target account that inputs the database operation text, and recall the metadata information of multiple recall paths based on the business space, cursor position, and database object identifier, and merge the recall results of each recall path to obtain the candidate metadata information of multiple data tables. Among them, the target account is the management account of the text editor, and the database operation text can be edited in the text editor; the business space is the storage space of a specific business area served by the database operation text edited in the text editor, including business processes, business rules, business data, etc. The business space is associated with the target account, and its services include music applications, video applications, game applications, etc.; the recall path is the way and method to obtain metadata information, and can be recalled based on aspects such as popularity, permissions, fields, tables, etc.
[0113] Specifically, obtain the target account that currently inputs the database operation text in the text editor, determine the business space associated with the target account, and obtain the metadata information from the business space. At the same time, obtain the metadata information from the data table marked by the database object identifier based on the cursor position and the database object identifier. After preprocessing the obtained metadata information, recall the metadata information through multiple recall paths to obtain the recall results of each recall path. Check the results of each recall result. If there are situations such as partial metadata information missing or metadata information being damaged in the recall result, call the recall path corresponding to the recall result to recall the corresponding data table, and query in the data table to obtain the recall result. Merge the recall results of each recall path to obtain the candidate metadata information of multiple data tables. Refer to Figure 5 , Figure 5 FIG. is an optional flowchart of metadata information recall provided by an embodiment of the present disclosure. Taking the business space, database operation text, cursor position, and database object identifier as inputs to obtain pre-metadata information, and after processing the pre-metadata information, perform multi-path recall. The recall paths include global popularity recall, personal popularity recall, permission recall, exact field recall, exact table recall, prefix table recall, behavior recall, clipboard recall, and collaborative recall. Merge the recall results of each recall path to obtain the recall result of the metadata information, that is, the candidate metadata information. It should also be noted that for the database operation text, similar text fragments can be recalled from the database operation text as supplementary metadata information, and the supplementary metadata information is finally merged into the candidate metadata information.
[0114] By performing multi-path recall, the candidate metadata information can cover different data regions such as the business space and database object identifiers. This coverage greatly expands the information scope of recall, enabling the maximization of the recall of metadata information highly relevant to the requirements, thereby mining more vertical knowledge. Additionally, multi-path recall can be regarded as obtaining metadata information from multiple data sources. When a problem occurs in one recall path, other paths can still continuously and stably provide effective recall results, ensuring the continuity and stability of the overall recall. Meanwhile, the multi-path recall strategy can flexibly adjust the number and weight of recall paths according to specific requirements, making the obtained candidate metadata information more targeted and further enhancing the reliability of the candidate metadata information.
[0115] In a possible implementation, during the process of recalling metadata information for multiple recall paths based on the business space, cursor position, and database object identifier, specifically, for the recall path corresponding to the business space, the data tables in the business space are sorted by popularity, and the first target data table is determined from the data tables in the business space according to the sorting result, and the metadata information of the first target data table is recalled. For the recall paths corresponding to the cursor position and database object identifier, the adjacent identifier of the cursor position is determined in the database object identifier, and the data table containing the adjacent identifier is determined as the second target data table, and the metadata information of the second target data table is recalled. Among them, the first target data table is determined based on the metadata information in the business space and is used to store the popularity sorting result; the second target data table is determined based on the data tables in the database object identifier and is used to store the data tables containing the adjacent identifier, which can be regarded as a data table set; the adjacent identifier of the cursor position is the database object identifier obtained by associating and merging the table name text and field name text near the current cursor position, and the adjacent identifier includes at least one of the table name text and field name text, and the adjacent identifier can be determined based on the distance between the current cursor position and the table name text and field name text. For example, the adjacent identifier can be set to have a distance less than or equal to a preset distance threshold from the cursor position.
[0116] Specifically, for the recall path corresponding to the business space, there are multiple data tables in the business space. For all the data tables in the business space, obtain the table metadata of all the data tables from the business space, sort the table metadata in descending order of popularity, build a table based on the sorting result to obtain the first target data table, and perform global popularity recall based on the first target data table to obtain the metadata information of the recalled first target data table. For any data table managed by a target account in the business space, obtain the table metadata corresponding to the data table managed by the target account from the business space, sort the table metadata in descending order of popularity, build a table based on the sorting result to obtain the first target data table, and perform personal popularity recall based on the first target data table to obtain the metadata information of the recalled first target data table. Additionally, it is also possible to obtain the table metadata corresponding to the data tables for which the target account has operation permissions from the business space, sort the table metadata to obtain the first target data table, and perform permission recall based on the first target data table to obtain the metadata information of the recalled first target data table. It should be noted that the data tables managed by the target account are the data tables that the target account currently needs to process; the operation permissions are used to indicate the types of operations that the target account can perform on the data tables, including viewing, editing, deleting, etc. The data tables for which the target account has operation permissions are the data tables on which the target account can perform operations such as viewing, editing, deleting, etc., including the data tables managed by the target account and the data tables that the target account does not need to manage but has operation permissions for.
[0117] Next, for the recall path corresponding to the cursor position and the database object identifier, a character interval threshold can be set, and the field name text of adjacent identifiers can be determined based on the character interval threshold. Determine the field name text in the database object identifier whose distance from the current cursor position is less than the character interval threshold and whose second end position is not adjacent to the cursor position. At this time, the field name text is a complete field name, and the data table containing this field name text is determined as the second target data table. Perform accurate field recall based on the second target data table to obtain the metadata information of the recalled second target data table. For example, when the current input is "select a,bfrom|", the cursor position "|" is after the syntax keyword "from", indicating that the field name has been entered completely, and the data table containing "a" and "b" can be recalled based on the complete field name. Or, determine the field name text in the database object identifier whose distance from the current cursor position is less than the character interval threshold and whose second end position is adjacent to the cursor position. At this time, the field name text is a field name prefix, and the data table containing this field name prefix is determined as the second target data table. Perform prefix field recall based on the second target data table to obtain the metadata information of the recalled second target data table. For example, when the current input is "select ab|from", the cursor position "|" is adjacent to the field name text being entered, and the current field name text "ab" is regarded as a field name prefix. Recall the data table containing this field name prefix based on the field name prefix "ab", and the field names in the recalled data table can be "abs", "abstract", etc.
[0118] Next, determine the table name text of adjacent identifiers based on the character interval threshold. Determine the table name text in the database object identifier whose distance from the current cursor position is less than the character interval threshold and the first end position is not adjacent to the cursor position. At this time, the table name text is the complete table name. Determine the data table corresponding to the table name text as the second target data table, and perform accurate table recall based on the second target data table to obtain the metadata information of the recalled second target data table. For example, when the current input is "select a from music.talbe where|", the cursor position "|" is after the syntax keyword "from", indicating that the table name has been entered completely. Recall the corresponding data table directly based on this table name. Or, determine the table name text in the database object identifier whose distance from the current cursor position is less than the character interval threshold and the first end position is adjacent to the cursor position. At this time, the table name text is the table name prefix. Determine the data table containing this table name prefix as the second target data table, and perform prefix table recall based on the second target data table to obtain the metadata information of the recalled second target data table. For example, when the current input is "select a from music.ta|", the cursor position "|" is adjacent to the table name text being entered. Consider the current table name text "music.ta" as the table name prefix, and recall the data tables containing this table name prefix based on the table name prefix "music.ta". The recalled data tables can be "music.table", "music.task", etc.
[0119] By setting the recall paths for two data sources, namely the data tables in the business space or the database (obtained from the database object identifier), the multi-party metadata information can be integrated, thereby improving the comprehensiveness and richness of the recalled metadata. In addition, recall based on popularity can ensure that the recalled metadata information has strong relevance and timeliness. Moreover, the metadata information with high popularity usually undergoes a large number of interactions and retrievals, and has high accuracy and integrity. Recall based on adjacent identifiers can also ensure that the recalled metadata information has strong relevance to the currently entered database operation text, guaranteeing the high relevance of the recalled metadata information.
[0120] In a possible implementation, the adjacent identifier can also be determined according to the statement block where the current cursor position is located. Obtain the current cursor position, locate the statement block where the current cursor position is located according to the chunking strategy of the database operation text, determine the field name text and table name text of the adjacent identifier in the current statement block, use the data table recalled based on the field name text and containing this field name text as the second target data table, and obtain the table name text in the previous statement block and the next statement block based on this current statement block. Use the corresponding three data tables recalled based on the table name text in the previous statement block, the current statement block, and the next statement block as the second target data tables.
[0121] In a possible implementation manner, during the process of recalling metadata information for multiple recall paths, behavior recall, clipboard recall, and collaborative recall can also be performed. Behavior recall can be performed based on the operation behavior of the target account, clipboard recall can be performed based on the copy operation of the target account, and collaborative recall can be performed based on the data tables frequently used by other target accounts. Most of the metadata information recalled by these three recall paths is metadata information with high usage frequency and recently used, and can be cross-recalled with one or more of the above recall methods based on popularity, permissions, field name text, and table name text. The recall methods based on popularity, permissions, field name text, and table name text can be regarded as precise recalls because they have clear recall conditions; while the metadata information based on behavior, clipboard, and collaborative recall has a certain degree of randomness, but since the metadata information recalled by these three methods is usually highly correlated with the operation behavior of the target account, the recalled metadata information also has a certain reference value. For example, a table is created based on the metadata information sorted by popularity in the business space, and at the same time, the data tables frequently used by other target accounts are obtained through collaborative recall, and the data tables containing popularity sorting information and the data tables frequently used by other target accounts are determined as the first target data tables.
[0122] In a possible implementation manner, the proximity identifier includes a proximity complete field name, a proximity field name prefix, and a proximity table name prefix. In the process of determining the data table containing the proximity identifier as the second target data table, specifically, it can be to determine the data table that simultaneously contains the proximity complete field name, the proximity field name prefix, and the proximity table name prefix as the second target data table.
[0123] Specifically, when the proximity identifier includes a proximity complete field name, a proximity field name prefix, and a proximity table name prefix, cross-recall can be performed. First, the initial data table containing the proximity table name prefix is obtained through the proximity table name prefix, the candidate data tables containing the proximity complete field name are searched in the initial data table, and then the data tables containing the proximity field name prefix are searched in the candidate data tables, and the candidate data tables containing the proximity field name prefix are determined as the second target data tables. By combining multiple field name or table name conditions for cross-recall, the situation where a single recall condition may recall irrelevant data can be avoided, so as to more accurately locate the target data table, reduce the interference of irrelevant information, and improve the accuracy and efficiency of recall. In addition, the cross-recall of multiple conditions makes good use of the internal association between the field name and the table name, making the obtained second target data more logical and reliable.
[0124] In a possible implementation, during the process of sorting multiple candidate metadata information and then filtering to obtain the target metadata information, specifically, the path weights of each recall path can be obtained, the sorting weights of the candidate metadata information can be determined according to the path weights, the multiple candidate metadata information can be initially sorted based on the sorting weights to obtain an initial sorting result, and the initial sorting result can be re-sorted based on at least one of the confidence of the candidate metadata information, the access right of the target account to the candidate metadata information, or the integrity verification result of the candidate metadata information to obtain a re-sorting result, and the target metadata information can be filtered from the re-sorting result. Among them, the target metadata information is the information with relatively high relevance filtered from the candidate metadata information and is used for subsequent template construction; the path weight is used to indicate the relative importance of the recall path, and the higher the path weight, the higher the importance of the corresponding recall path; the sorting weight is used to initially sort the candidate metadata and can directly be the path weight. For example, if the path weight of the recall path is w, then the sorting weight of the metadata information recalled through the recall path is also w, or it can be obtained by weighted summation based on the path weight.
[0125] Specifically, obtain the historical database object identifier, historical business space, and historical recall metadata information. Based on the historical database object identifier, historical business space, and historical recall metadata information, obtain the path weights corresponding to each recall path. Determine the sorting weights of the candidate metadata information according to the path weights. Based on the sorting weights, initially sort multiple metadata information in a certain order to obtain an initial sorting result. Then, further adjust the initial sorting result according to the permission configuration policy. Determine the confidence level of the candidate metadata information based on the recall characteristics of each recall path. The higher the confidence level, the higher the weight of the candidate metadata information. Obtain the access permission of the target account for the candidate metadata information. The weight can be set according to the access scope or permission priority of the access permission. The larger the access scope or the higher the permission priority, the higher the weight of the candidate metadata information. Verify the integrity of the candidate metadata information or verify the field name text corresponding to the candidate metadata information to obtain a verification result. When the verification result indicates that the candidate metadata is incomplete or the field name text verification fails, reduce the weight of the candidate metadata information. Rearrange the initial sorting result based on at least one of the confidence level, access permission, or verification result to obtain a rearranged result. Filter out the candidate metadata information with low weights (i.e., the metadata information with weak relevance) in the rearranged result, and screen out the candidate metadata information with high weights (i.e., the metadata information with strong relevance) as the target metadata information. By performing multiple sorts on the candidate metadata information and considering the path weights, confidence levels, access permissions, and integrity of the candidate metadata information during the sorting process, it is possible to quickly filter out the metadata information with low relevance, reduce the interference of unreliable or incorrect information, thereby improving the reliability and efficiency of the sorting result, making the obtained target metadata information have high accuracy and strong relevance, and further improving the accuracy of the text completion model when generating the completion text of the database operation text.
[0126] Refer to Figure 6 , Figure 6 FIG. is an optional flowchart of metadata information sorting provided by an embodiment of the present disclosure. Taking the business space, candidate metadata information, and database object identifier as inputs, initially sort the candidate metadata according to the information priority (the feature weight of the metadata information, see below) and the information weight (the path weight of the recall path) to obtain an initial sorting result. Then, perform operations such as filtering, increasing the weight, or decreasing the weight on the initial sorting result according to information such as the confidence level and access permission to obtain the target metadata information.
[0127] In a possible implementation manner, in the process of determining the sorting weight of candidate metadata information according to the path weight, specifically, for any candidate metadata information, the database operation text and the candidate metadata information are concatenated and then input into a weight prediction model for prediction to obtain the feature weight of the candidate metadata information. The feature weight and the path weight are weighted and summed to obtain the sorting weight of the candidate metadata information. Among them, the weight prediction model is used to obtain the feature weight of the candidate metadata information, which can be a convolutional neural network, a recurrent neural network, or a Transformer network. The present application does not make specific limitations.
[0128] Specifically, for any candidate metadata information, the database operation text and the candidate metadata information are concatenated and then input into a weight prediction model for prediction. The scores of the corresponding data table and field are determined according to the types of the table name text and the field name text in the database operation text. The feature weight of the candidate metadata information is obtained according to the scores. For example, when the table name text music.ta is the table name prefix, the data tables music.table, music.task, and music.tal containing the table name prefix are recalled. The higher the proportion of the prefix characters in the recalled table names, the higher the score. At this time, 8 out of 11 characters in the table name of the data table music.table are prefix characters, and its score can be 0.7. 8 out of 10 characters in the table name of the data table music.task are prefix characters, and its score can be 0.8. 8 out of 9 characters in the table name of the data table music.tal are prefix characters, and its score can be 0.9. Based on the scores, the feature weight of the candidate metadata recalled from the data table music.table is 0.7, the feature weight of the candidate metadata recalled from the data table music.task is 0.8, and the feature weight of the candidate metadata recalled from the data table music.tal is 0.9. Then, the feature weight and the path weight are weighted and summed to obtain the sorting weight of the candidate metadata information. Continuing with the above example, the field name texts corresponding to the candidate metadata information A and the candidate metadata information B are complete field names. The path weight of the recall path A is w 1 , and the path weight of the recall path B is w 2 . The candidate metadata information A is obtained from the data table music.task through the recall path A, and the candidate metadata information B is obtained from the data table music.table through the recall path B. Then, the sorting weight of the candidate metadata information A is 0.8w 1 , and the sorting weight of the candidate metadata information B is 0.7w 2。The recall path indicates the source reliability of the candidate metadata information, and the characteristics of the candidate metadata information reflect the importance of the metadata itself. By performing a weighted sum of the path weight and the characteristic weight of the candidate metadata information, the advantages of the two dimensions of the recall path and the candidate metadata information can be balanced, avoiding the one-sidedness of sorting in a single dimension, thereby improving the accuracy of sorting.
[0129] In addition, the path weights corresponding to different recall paths can be obtained and input into the weight prediction model for prediction to obtain the characteristic weights of the candidate metadata information.
[0130] Step S203: Obtain the context fragments adjacent to the cursor position, and construct a prompt text based on the target metadata information and the context fragments.
[0131] Among them, the context fragments are the operation statements adjacent to the current cursor position before and after in the database operation text. The prompt text is used to prompt the text completion model to complete the database operation text, and the text completion model is used to complete the table name prefix and field name prefix in the database operation text.
[0132] In a possible implementation manner, in the process of constructing the prompt text based on the target metadata information and the context fragments, specifically, the dialect information of the context fragments can be determined, the current time when the database operation text is input is obtained, and the prompt text is constructed based on the dialect information, the context fragments, the target metadata information, and the current time. Among them, the dialect information is used to indicate the dialect used by the context fragments. The dialect can be understood as a kind of code-like language with unique grammar and rules developed based on a certain code language. For example, databases or applications such as MySQL, Hive, Presto, ClickHouse, and Impala support standard SQL and its extensions, so the code languages used by these databases or applications can be called dialects.
[0133] Specifically, when inputting the database operation text in the text editor, the current cursor position is obtained, the context fragments near the cursor position and the dialect information of the context fragments are determined, the current time when the database operation text is being input is obtained, and the prompt text is constructed based on the dialect information, the context fragments, the target metadata information, and the current time. By constructing the prompt text, the text completion model can accurately grasp the generation direction, enabling the text completion model to generate the completion text more pertinently. At the same time, the prompts of the dialect information and the current time also standardize the format standard and timeliness of the generated completion text, effectively improving the accuracy and professionalism of the text completion model in generating the completion text.
[0134] In a possible implementation manner, in the process of constructing the prompt text based on the dialect information, context fragment, target metadata information, and current time, specifically, the dialect information text may be constructed based on the dialect information and the first identification text for identifying the dialect information, the table creation text may be constructed in the form of a data definition language based on the target metadata information, the metadata information text may be constructed based on the table creation text and the second identification text for identifying the table creation text, the context information text may be constructed based on the context fragment and the third identification text for identifying the context fragment, and the prompt text may be constructed based on the dialect information text, metadata information text, current time, context information text, and the fourth identification text for identifying the completion text. Among them, the first identification text is used to identify the dialect information in the prompt text, the second identification text is used to represent the metadata information text in the prompt text, the third identification text is used to represent the extracted context fragment in the prompt text, and the fourth identification text is used to identify the completion text in the prompt text. The above identification texts can all be composed of one or more of letters, numbers, and symbols; the dialect information text is the part in the prompt text that represents the dialect information, the table creation text can be regarded as a table creation statement, the metadata information text is the part in the prompt text that represents the metadata information, and the context information text is the part of the extracted context fragment in the prompt text.
[0135] Specifically, the first identification text is added according to the dialect information, and the dialect information text is constructed based on the dialect information and the first identification text. The table creation text is constructed in the form of a data definition language based on the target metadata information. The table creation text includes information such as the table creation table name, table annotation, field name text, field type, field annotation, and partitioning. In order to avoid data redundancy caused by the excessive length of the target metadata information, the fields can be filtered according to the popularity information of the fields, and the fields with high popularity are retained. Then, the second identification text is added according to the table creation text, and the metadata information text is constructed based on the table creation text and the second identification text. The third identification text is added according to the context fragment, and the context information text is constructed based on the context fragment and the third identification text to perceive the input state of the current database operation text. The completion text is obtained, and the completion text contains the text content of the part to be completed. The fourth identification text is added according to the completion text, and the prompt text is constructed based on the dialect information text, metadata information text, current time, context information text, and the fourth identification text. By constructing the table creation text in the form of a data definition language, the text completion model can more easily understand the table creation text, thereby improving the effect of metadata enhancement.
[0136] According to the above description, referring to Figure 7 , Figure 7 is an optional schematic diagram of the prompt text provided by the embodiments of the present disclosure, including a dialect information text, a metadata information text, a related information text, and an operation statement text. The operation statement text includes a context information text. Assume that the first identification text is<language>, if the dialect information is PRESTO, then the dialect information text can be <language>PRESTO. Assume that the second identification text is <extra_info>, and the table creation text is CREATE TABLE db.table (col1 string COMMENT 'aaa', col2 bigint COMMENT 'bbb',...). This table creation text means to create a data table named table in the database db. In the data table table, the data type of the first column col1 is string. COMMENT can be regarded as a placeholder where the metadata information conforming to the first column can be written. '' is the comment for the first column. The data type of the second column col2 is big integer bigint. Similarly, the metadata information of the second column can be written in the position where COMMENT is located, and 'bbb' is the comment for the second column. According to the above description, the metadata information text can be <extra_info>CREATE TABLE db.table (col1 string COMMENT 'aaa', col2 bigint COMMENT 'bbb',...). The relevant information text includes the current time and the query statement. The current time can be expressed as current_date:xxxxxxxx, and the query statement can be select * from db.table where col2 = c or col3 = d. This query statement means to traverse the second and third columns in the data table table of the database db and find all rows where the value in the second column is equal to c or all rows where the value in the third column is equal to d. The final operation statement is used to find more similar code snippets. When the third identification prefix is <fim_prefix>, the third identification suffix is <fim_suffix>, the statement keyword 1 is select, the statement keyword 2 is from, and the fourth identification text is <fim_suffix>, the context information text can be <fim_prefix>select col1,col2<fim_suffix>, where <fim_prefix> represents the above text, that is, the part of the database operation text before the current cursor position, and <fim_suffix> represents the following text, that is, the part of the database operation text after the current cursor position. Further, in order to conform to the generation process of the text completion model, the operation statement can be constructed in the order of the third identification prefix, the third identification suffix, the fourth identification text or the third identification suffix, the third identification prefix, the fourth identification text to ensure that the text to be completed is at the end, which helps the generation of the completed text. Therefore, the operation statement can be <fim_prefix>select col1,col2<fim_suffix> from db.table <fim_suffix>, where <fim_suffix> represents the text to be completed.
[0137] Further, in order to clarify the role of the operation statement text in the prompt text, a fifth identification text is newly added based on this. The fifth identification text is used to indicate that the operation statement is used to extract similar code segments from the database. The fifth identification text can be <code_content>. Based on this, the operation statement text can also be <code_content><fim_prefix>select col1, col2<fim_suffix>from db.table<fim_suffix>. Since the dialect information provides the code language habit, the metadata information provides the metadata features, the context fragment provides the coherent semantic information, the current time provides the timeliness of the prompt text, and the similar code segment provides an effective reference paradigm, by constructing the prompt text from aspects such as dialect information, metadata information, context fragment, current time, and similar code segments, a prompt text with clear requirements and clear directions can be obtained, enabling the text completion model to generate accurate completion text based on the prompt text, effectively avoiding the problem of deviation of the completion text caused by vague requirements, thereby improving the accuracy and efficiency of the text completion model in generating the completion text.
[0138] In a possible implementation manner, the target metadata information includes the field name text. In the process of constructing the metadata information text based on the table creation text and the second identification text for identifying the table creation text, specifically, the corresponding enumeration value can be obtained according to the field name text, and the enumeration value is added to the table creation text in a manner parallel to the field name text, and the metadata information text is constructed based on the table creation text after adding the enumeration value and the second identification text for identifying the table creation text.
[0139] Specifically, an open data source is accessed through an external interface, the corresponding enumeration value is obtained according to the field name text in the open data source, the enumeration value is added to the table creation text in a manner parallel to the field name text, and the metadata information text is constructed based on the table creation text after adding the enumeration value and the second identification text for identifying the table creation text. For example, the field name text in the table creation text is "age", and the metadata information corresponding to the field name text "age" is {15, 4, 5, 12, 75}. At this time, the enumeration values can be age values other than the metadata information, such as 10 and 20. At this time, the metadata information added to the prompt text is {15, 4, 5, 12, 75, 10, 20}. By obtaining the corresponding enumeration value according to the field name text, the metadata information becomes more diverse, thereby improving the possibility and comprehensiveness of obtaining metadata information.
[0140] In addition, in the process of constructing the metadata information text based on the table creation text and the second identification text for identifying the table creation text, specifically, it may also be to retrieve the enumeration text similar to the field name text, add the enumeration text to the table creation text in a way parallel to the field name text, and construct the metadata information text based on the table creation text after adding the enumeration text and the second identification text for identifying the table creation text. Among them, the enumeration text is the text with semantic features similar to the field name text. For example, if the field name text is "age", the enumeration text may be "year of age", "number of years", etc.
[0141] Specifically, access the open data source through an external interface, retrieve the enumeration text similar to the field name text in the open data source, add the enumeration text to the table creation text in a way parallel to the field name text, and construct the metadata information text based on the table creation text after adding the enumeration text and the second identification text for identifying the table creation text. For example, the field name text in the table creation text is "age", and the metadata information corresponding to the field name text "age" is {15, 4, 5, 12, 75}. At this time, if the input fields are "year of age" or "number of years", the metadata information {15, 4, 5, 12, 75} cannot be obtained according to the table creation text. When the two enumeration texts "year of age" and "number of years" are added to the table creation text in parallel with the field name text "age", any one of the fields "age", "year of age", and "number of years" can be used to obtain the metadata information {15, 4, 5, 12, 75}. By adding enumeration texts with similar semantic features to the field name text, it is possible to cover diverse expressions of the same semantic feature, avoid omissions or inability to obtain relevant metadata information that may occur due to the use of a single field, lower the threshold for obtaining metadata information, and thus improve the possibility and comprehensiveness of obtaining metadata information.
[0142] In a possible implementation manner, in the process of constructing the prompt text based on the dialect information text, the metadata information text, the current time, the context information text, and the fourth identification text for identifying the completion text, specifically, it may be to retrieve the historical reference text similar to the database operation text, and construct the prompt text based on the dialect information text, the metadata information text, the current time, the historical reference text, the context information text, and the fourth identification text for identifying the completion text. Among them, the historical reference text is the code text in other data sources that is similar to the database operation text in terms of content and structure, and the other data sources are other databases, web resources, etc. except for the database currently in use.
[0143] Specifically, an open data source is accessed through an external interface, and historical reference texts that are similar in structure and content to the database operation text are retrieved from the open data source. A sixth identification text is added based on the historical reference text. The sixth identification text is used to identify the historical reference text in the prompt text, and the sixth identification text belongs to the historical reference text. For example, if the obtained historical reference text is START xxxxxx, when the sixth identification text is <history_code>, the final historical reference text can be <history_code>START xxxxxx. Finally, a prompt text is constructed based on the dialect information text, metadata information text, current time, historical reference text, context information text, and fourth identification text for identifying the completed text. By obtaining historical reference texts to provide similar code snippets, more effective reference paradigms can be further provided for the text completion model, enabling the text completion model to generate more standardized completed texts, effectively improving the accuracy and standardization of the completed texts generated by the text completion model.
[0144] According to the description in step S203, the process of constructing the prompt text can refer to Figure 8 , Figure 8 which is an optional flowchart for constructing the prompt text provided by an embodiment of the present disclosure. Using the database operation text, the current cursor position, and the target metadata information as inputs, the dialect information of the context fragment near the cursor position is obtained. A table creation text is constructed for the target metadata information in the form of a data definition language, and the fields in the table creation text are filtered according to the field heat information. Then, the current time of the input data operation text is obtained, similar enumeration texts are obtained based on the field name texts in the table creation text, historical reference texts are obtained based on the database operation text, and retrieval enhancement is achieved according to the enumeration texts and historical reference texts. Finally, the context fragments before and after the current cursor position in the database operation text are obtained, the context fragments are truncated, and a prompt text is constructed based on the dialect information, metadata information, current time, enumeration texts, historical reference texts, and context fragments.
[0145] Step S204: Input the prompt text into the text completion model for text prediction to generate a completed text for the database operation text.
[0146] Among them, the completed text is the part to be completed in the database operation text. For example, when "select a from" is input in the database operation text, the corresponding completed text can be "db.table where n>0", and the complete database operation text after completion is "select a from db.table where n>0".
[0147] In a possible implementation, during the process of inputting the prompt text into the text completion model for text prediction to generate the completion text of the database operation text, specifically, the prompt text can be input into the text completion model for streaming text prediction, and the predicted words output by the text completion model in each round are cached. When the text completion model ends, the prediction result is obtained based on the predicted words cached in multiple rounds, and the prediction result is subjected to suppression processing to generate the completion text of the database operation text. Among them, the suppression processing is used to delete abnormal content in the prediction result. The suppression processing includes stop word truncation, deletion of extreme repetitions, suppression of abnormal output by the text completion model, deletion of non-existent table name text or field name text, line number truncation, block truncation, etc.; the predicted word is the prediction result input by the text completion model, and the text completion model outputs one predicted word in each round, and the prediction result is composed of multiple predicted words.
[0148] Specifically, the prompt text and the database operation text are input into the text completion model for streaming text prediction, and the predicted words output by the text completion model in each round are sent to the buffer for caching. When the text completion model finishes prediction, the predicted words cached in multiple rounds form the prediction result in the buffer, and the prediction result is subjected to suppression processing line by line in the buffer. One processing result is output for each processed line until all lines of the prediction result are processed, and the final completion text of the database operation text is output. When it is not necessary to output the processing result of each line, the prompt text and the database operation text can also be input into the text completion model for non-streaming text prediction. After the text completion model finishes prediction and obtains the prediction result, suppression processing is performed based on the overall prediction result, and the completion text of the database operation text is directly output. The processes of the above two text prediction methods can refer to Figure 9 , Figure 9 which is an optional schematic diagram of the post-processing of the prediction result of the text completion model provided by this embodiment of the present disclosure. By inputting the prompt text into the text completion model for text prediction, metadata enhancement can be achieved during the inference of the text completion model, improving the text completion model's understanding ability of vertical knowledge. Further, by using the streaming text prediction method, the prediction result is cached in the buffer and subjected to suppression processing line by line, enabling the model to output the completion text line by line, which can reduce the occupation of computing resources. In addition, when an error or interruption occurs during the generation process, the already generated completion text can still be used, improving the fault tolerance of the text completion model.
[0149] Before post-processing the processing result, the inference process of the text completion model can refer to Figure 10 , Figure 10 which is an optional flowchart of the inference process of the text completion model provided by this embodiment of the present disclosure. Figure 10 The provided inference process involves the interaction process of the server, taking the prompt text, request information, and model configuration as inputs. Among them, the request information is to request the server to perform full-text generation, and the model configuration is the parameter configuration of the text completion model. Specifically, the prompt text, request information, and model configuration are sent to the server. After the server responds to the request information, it calls the text completion model based on the model configuration. The text completion model performs inference based on the prompt text to obtain a prediction result. When the prediction result is a non-streaming result, the predicted completed text is directly output. When the prediction result is a streaming result, an iterative generator is returned, and the completed text is generated at the word granularity based on this iterative generator.
[0150] In addition, the process of inputting the prompt text into the text completion model for streaming text prediction can also be based on the output after each round of prediction. Specifically, for the first round of prediction, the predicted word obtained by the text completion model in the first round is sent to the buffer for caching and then output. For the remaining rounds, after the text completion model obtains the predicted word of the current round and sends it to the buffer for caching, suppression processing is performed based on the caching result of the previous round, and the predicted result after suppression processing is output until the text completion model obtains the predicted word of the last round and sends it to the buffer for suppression processing to obtain the completed text of the database operation text.
[0151] Refer to Figure 11 , Figure 11 This is an optional application overall framework for the text generation method provided by the embodiments of the present disclosure. The application framework built based on the text generation method provided by the embodiments of the present disclosure includes a front end, a back end, an algorithm end, and a data end. Specifically, the front end is a text editor that supports the Language Service Protocol (LSP) and interacts with the server based on the LSP protocol. For example, it provides parameters such as database operation text and the current cursor position to the server and receives the completed text sent by the server. This application can use the Monaco editor that supports the LSP protocol as the front end. The back end is a server based on the LSP protocol, responsible for interacting with the front end and the algorithm end, and can provide the capabilities of logical processing and protocol forwarding. The algorithm end is the text generation method provided by the embodiments of the present disclosure. After receiving parameters such as database operation text and the current cursor position transmitted by the back end, it returns the completed text. The data end is responsible for coordinating relevant metadata and establishing various public interfaces and index services.
[0152] Refer to Figure 12 , Figure 12 For the Figure 11 An optional overall schematic framework on the algorithm side, and also an optional overall schematic framework of the text generation method provided by the embodiments of the present disclosure. The principle of the text generation method in the embodiments of the present disclosure is described in detail as follows:
[0153] First, using the database operation text and the current cursor position as the main inputs, and data such as the editing environment of the text editor as the secondary inputs, preprocess the input data such as the database operation text. Specifically, perform normalization conversion on the database operation text, and extract table name text based on regular expressions. Chunk the database operation text based on the text delimiter "select" to obtain multiple statement chunks, and truncate the number of lines of each statement chunk respectively to reduce irrelevant information and improve the inference speed. Parse each statement chunk based on a preset parsing library to extract field name text. Since the operation statements in the statement chunk may be incomplete, parsing exceptions may occur based on the preset parsing library. In this case, statement keywords (such as "select") can be added at the beginning of the line of the statement chunk to increase the parsing rate. Classify the extracted table name text and field name text (i.e., prefix judgment), and combine the table name text and field name text that belong to the same statement chunk and have been classified to obtain the database object identifiers in each statement chunk of the database operation text.
[0154] Next, perform multi-way parallel recall based on the database object identifiers and the business space. The recall paths corresponding to the database object identifiers include exact field recall, prefix field recall, exact table recall, and prefix table recall. Determine whether to perform field recall or table recall according to the current cursor position, and determine whether it is exact recall or prefix recall according to whether the cursor position is adjacent to the field or table. The recall paths corresponding to the business space include global popularity recall, personal popularity recall, and permission recall. Perform global popularity recall according to the table metadata in the current business space, perform personal popularity recall according to the table metadata in the data tables processed by the target account in the current business space, and perform permission recall according to the access permissions of the target account. In addition, behavior recall can also be performed according to the editing behavior of the target account, clipboard recall can be performed according to the table name text or operation text copied to the clipboard by the target account, and collaborative recall can be performed according to the data tables frequently used by other target accounts. When the database object identifier simultaneously has a complete field name, a field name prefix, and a table name prefix, cross recall can be performed based on the complete field name, the field name prefix, and the table name prefix to further narrow the recall range. In addition, when the recall data source can select the default database information, database filtering conditions can be added before the recall to narrow the recall range and improve the recall accuracy. Merge the metadata information recalled by each recall path to obtain candidate metadata information.
[0155] Next, sort the candidate metadata information, obtain the historical recall process of the historical candidate metadata information, obtain the path weight of the recall path based on the historical recall process, call the weight prediction model to obtain the feature weight of the candidate metadata information, and perform a preliminary sort on the candidate metadata information based on the weighted result of the path weight and the feature weight to obtain the preliminary sorting result. Then, eliminate the metadata information with low confidence based on the confidence level, promote the metadata information with access rights to the target account, and demote the metadata information that fails the field verification to dynamically adjust the weights of each candidate metadata information, and re-sort the preliminary sorting result according to the adjusted weights to obtain the target metadata information.
[0156] Next, construct a prompt text based on known information such as the database operation text, cursor position, and target metadata information. The prompt text includes dialect information text, metadata information text, enumeration values, enumeration text, current time, context information text, historical reference text, etc. The dialect information text is used to prompt the text completion model with the dialect used in the current part to be completed and the dialect that should be used for completion; the metadata information text is obtained by converting the target metadata information in the form of a data definition language. To avoid the converted metadata information from being too long, fields with low heat can be filtered according to the field heat information; the current time is obtained as the historical time node when the training data of the text completion model ends; retrieval enhancement is to further obtain enumeration text similar to the field name text or historical reference text similar to the database operation text through an external interface to expand the retrieval range; the context information text is the context fragment at the current cursor position, and this part needs to perceive the current editing state of the database operation text, such as the cursor position, logical association between the above and below texts, etc.
[0157] When the text completion model is trained based on the prompt text, the context fragment and the part to be completed can be regarded as three parts: prefix, part to be completed, and suffix. To conform to the text completion model's ability to continue writing during the training process, the metadata information can be arranged and organized in the order of prefix, suffix, part to be completed or suffix, prefix, part to be completed, ensuring that the part to be completed is at the end of the combination.
[0158] Next, input the prompt text and the database operation text into the text completion model for model inference to obtain the generation result (this process is in Figure 11 In the application architecture, the prompt text, the configuration of the text completion model, and the inference request can be sent to the server, and the generation result is obtained in the server. When the generation result is a streaming result, an iterative generator is returned. Based on this iterative generator, predicted words are generated round by round, and the predicted words are cached in the buffer. When the iteration is completed, the predicted result is obtained. The predicted result is suppressed line by line in the buffer (i.e., post-processing, which can be stop word truncation, special constraints, etc.). After processing each line, a processed result is output until all lines of the predicted result are processed, and the completed text of the database operation text is finally output. When the generation result is a non-streaming result, the predicted result is directly output, and the predicted result is suppressed as a whole to obtain the completed text of the database operation text.
[0159] It should also be noted that during the entire text generation process, the interfaces for the processing process, the recall process, the sorting process, the inference process, and the post-processing process need to be repeatedly accessed and queried. Therefore, the results of these processes are cached to improve the processing efficiency and response speed of the entire text generation process.
[0160] The text generation method provided by the embodiments of the present disclosure obtains the final completed text through processes such as data preprocessing, data recall, data sorting, prompt text construction, inference prediction, and post-processing of the completed text. It can mine more vertical knowledge, improve the text completion model's understanding ability of vertical knowledge, and effectively improve the accuracy when the text completion model generates the completed text of the database operation text.
[0161] In a possible implementation manner, the text generation method provided by the embodiments of the present disclosure can be applied to the code completion of SQL language. The SQL code to be completed is preprocessed to obtain the object identifier corresponding to the SQL code block. Candidate metadata information is recalled based on the object identifier. After sorting and filtering the candidate metadata information, target metadata information is obtained. A prompt text is constructed based on the target metadata information, and the prompt text is input into the text completion model for inference to obtain a predicted result. The predicted result is post-processed to obtain the part to be completed in the SQL code.
[0162] It can be understood that although the steps in each of the above flowcharts are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this embodiment, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowchart may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0163] Referring to Figure 13 , Figure 13 FIG. is a schematic structural diagram of a text generation device provided by an embodiment of the present disclosure. The text generation device 1300 includes:
[0164] An extraction module 1301, configured to obtain the currently input database operation text and the current cursor position, and extract the database object identifier in the database operation text based on the cursor position;
[0165] A recall module 1302, configured to recall candidate metadata information of multiple data tables based on the database object identifier, sort the multiple candidate metadata information and then perform screening to obtain target metadata information;
[0166] A construction module 1303, configured to obtain the context fragment adjacent to the cursor position, and construct a prompt text based on the target metadata information and the context fragment, where the prompt text is used to prompt the text completion model to complete the database operation text;
[0167] An inference module 1304, configured to input the prompt text into the text completion model for text prediction to generate a completion text of the database operation text.
[0168] Further, the extraction module 1301 is further configured to:
[0169] Perform normalization conversion on the database operation text and then extract the table name text;
[0170] Divide the database operation text into multiple statement blocks, and extract the field name text from each statement block based on a preset parsing library;
[0171] Merge the table name text and the field name text belonging to the same statement block to obtain the database object identifier in each statement block of the database operation text.
[0172] Further, the extraction module 1301 is further configured to:
[0173] Extract the field names from each statement block based on a preset parsing library. When the extraction fails based on the preset parsing library, traverse each line of the statement block;
[0174] If the database syntax keyword is missing at the beginning of the currently traversed line, add the database syntax keyword at the beginning of the currently traversed line and then extract the field name text.
[0175] Furthermore, the extraction module 1301 is further configured to:
[0176] Perform a prefix judgment on the table name text according to the positional relationship between the cursor position and the table name text to obtain a first judgment result, and classify the table name text according to the first judgment result, where the first judgment result is used to indicate that the table name text is a table name prefix or a complete table name;
[0177] Perform a prefix judgment on the field name text according to the positional relationship between the cursor position and the field name to obtain a second judgment result, and classify the field name text according to the second judgment result, where the second judgment result is used to indicate that the field name text is a field name prefix or a complete field name;
[0178] Merge the table name text and the field name text that belong to the same statement block and have been classified to obtain the database object identifiers in each statement block of the database operation text.
[0179] Furthermore, the recall module 1302 is further configured to:
[0180] Determine the business space associated with the target account of the input database operation text, and recall the metadata information of multiple recall paths based on the business space, the cursor position, and the database object identifier;
[0181] Merge the recall results of each recall path to obtain the candidate metadata information of multiple data tables.
[0182] Furthermore, the recall module 1302 is further configured to:
[0183] For the recall path corresponding to the business space, sort the data tables in the business space by popularity, and determine the first target data table from the data tables in the business space according to the sorting result, and recall the metadata information of the first target data table;
[0184] For the recall path corresponding to the cursor position and the database object identifier, determine the adjacent identifier of the cursor position in the database object identifier, and determine the data table containing the adjacent identifier as the second target data table, and recall the metadata information of the second target data table.
[0185] Furthermore, the recall module 1302 is further configured to:
[0186] Determine the second target data table that simultaneously includes adjacent complete field names, adjacent field name prefixes, and adjacent table name prefixes.
[0187] Furthermore, the recall module 1302 is also used for:
[0188] Obtain the path weights of each recall path, determine the sorting weights of the candidate metadata information based on the path weights, and perform an initial sort on multiple candidate metadata information based on the sorting weights to obtain an initial sorting result;
[0189] Based on at least one of the confidence level of the candidate metadata information, the access permission of the target account to the candidate metadata information, or the integrity verification result of the candidate metadata information, re-sort the initial sorting result to obtain a re-sorting result;
[0190] Screen and obtain the target metadata information from the re-sorting result.
[0191] Furthermore, the recall module 1302 is also used for:
[0192] For any candidate metadata information, splice the database operation text and the candidate metadata information and input them into the weight prediction model for prediction to obtain the feature weight of the candidate metadata information;
[0193] Perform a weighted sum of the feature weight and the path weight to obtain the sorting weight of the candidate metadata information.
[0194] Furthermore, the construction module 1303 is also used for:
[0195] Determine the dialect information of the context fragment, where the dialect information is used to indicate the dialect used by the context fragment;
[0196] Obtain the current time when the database operation text is input, and construct a prompt text based on the dialect information, the context fragment, the target metadata information, and the current time.
[0197] Furthermore, the construction module 1303 is also used for:
[0198] Construct a dialect information text based on the dialect information and the first identification text used to identify the dialect information;
[0199] Construct a table creation text in the form of a data definition language based on the target metadata information, and construct a metadata information text based on the table creation text and the second identification text used to identify the table creation text;
[0200] Construct a context information text based on the context fragment and the third identification text used to identify the context fragment;
[0201] Construct a prompt text based on the dialect information text, metadata information text, current time, context information text, and a fourth identification text for identifying the completed text.
[0202] Further, the construction module 1303 is further configured to:
[0203] Obtain the corresponding enumerated value according to the field name text, and add the enumerated value to the table creation text in a manner juxtaposed with the field name text;
[0204] Construct a metadata information text based on the table creation text after adding the enumerated value and a second identification text for identifying the table creation text.
[0205] Further, the construction module 1303 is further configured to:
[0206] Retrieve historical reference texts similar to the database operation text;
[0207] Construct a prompt text based on the dialect information text, metadata information text, current time, historical reference text, context information text, and a fourth identification text for identifying the completed text.
[0208] Further, the inference module 1304 is further configured to:
[0209] Input the prompt text into the text completion model for streaming text prediction, and cache the predicted words output by the text completion model in each round;
[0210] When the text completion model ends, obtain a prediction result based on the predicted words cached in multiple rounds, perform suppression processing on the prediction result, and generate a completed text of the database operation text, where the suppression processing is used to delete abnormal content in the prediction result.
[0211] In summary, the text generation device provided by the embodiments of the present disclosure can, by obtaining the currently input database operation text and the current cursor position, extract the database object identifier in the database operation text based on the cursor position, recall candidate metadata information of multiple data tables based on the database object identifier, thereby mining more vertical knowledge, sorting and screening multiple candidate metadata information, and then obtaining target metadata information with higher relevance. Subsequently, obtain the context fragment adjacent to the cursor position, construct a prompt text based on the target metadata information and the context fragment, and input the prompt text into the text completion model for text prediction, which can achieve metadata enhancement during the inference of the text completion model, improve the text completion model's understanding ability of vertical knowledge, and effectively improve the accuracy when the text completion model generates the completed text of the database operation text.
[0212] The electronic device provided by the embodiments of the present disclosure for executing the above text generation method may be a terminal. Refer to Figure 14 , Figure 14 It is a partial structural block diagram of the terminal provided by the embodiments of the present disclosure. The terminal includes components such as a camera assembly 1410, a first memory 1420, an input unit 1430, a display unit 1440, a sensor 1450, an audio circuit 1460, a wireless fidelity (WiFi) module 1470, a first processor 1480, and a first power supply 1490. Those skilled in the art can understand that Figure 14 the terminal structure shown in
[0213] does not limit the terminal, and it may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. The camera assembly 1410 can be used to collect images or videos. Optionally, the camera assembly 1410 includes a front camera and a rear camera. Generally, the front camera is arranged on the front panel of the terminal, and the rear camera is arranged on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera, so as to implement functions such as background blurring by fusing the main camera and the depth camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fusion shooting functions.
[0214] The first memory 1420 can be used to store software programs and modules. The first processor 1480 executes various functional applications and data processing of the terminal by running the software programs and modules stored in the first memory 1420.
[0215] The input unit 1430 can be used to receive input digital or character information, and generate key signal inputs related to the settings and function controls of the terminal. Specifically, the input unit 1430 may include a touch panel 1431 and other input devices 1432.
[0216] The display unit 1440 can be used to display input information or provided information and various menus of the terminal. The display unit 1440 may include a display panel 1441.
[0217] The audio circuit 1460, the speaker 1461, and the microphone 1462 can provide an audio interface.
[0218] The first power supply 1490 can be alternating current, direct current, a primary battery, or a rechargeable battery.
[0219] The number of sensors 1450 can be one or more. The one or more sensors 1450 include, but are not limited to, an acceleration sensor, a gyroscope sensor, a pressure sensor, an optical sensor, etc. Among them:
[0220] The acceleration sensor can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established by the terminal. For example, the acceleration sensor can be used to detect the components of the gravitational acceleration on the three coordinate axes. The first processor 1480 can control the display unit 1440 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor. The acceleration sensor can also be used for games or the collection of the user's motion data.
[0221] The gyroscope sensor can detect the body direction and rotation angle of the terminal. The gyroscope sensor can cooperate with the acceleration sensor to collect the user's 3D actions on the terminal. Based on the data collected by the gyroscope sensor, the first processor 1480 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.
[0222] The pressure sensor can be disposed on the side frame of the terminal and / or the lower layer of the display unit 1440. When the pressure sensor is disposed on the side frame of the terminal, it can detect the user's holding signal of the terminal, and the first processor 1480 can perform left / right hand recognition or shortcut operations according to the holding signal collected by the pressure sensor. When the pressure sensor is disposed on the lower layer of the display unit 1440, the first processor 1480 can control the operable controls on the UI interface according to the user's pressure operation on the display unit 1440. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0223] The optical sensor is used to collect the ambient light intensity. In one embodiment, the first processor 1480 can control the display brightness of the display unit 1440 according to the ambient light intensity collected by the optical sensor. Specifically, when the ambient light intensity is high, the display brightness of the display unit 1440 is increased; when the ambient light intensity is low, the display brightness of the display unit 1440 is decreased. In another embodiment, the first processor 1480 can also dynamically adjust the shooting parameters of the camera module 1410 according to the ambient light intensity collected by the optical sensor.
[0224] In this embodiment, the first processor 1480 included in the terminal can execute the text generation method of the previous embodiment.
[0225] The electronic device provided by the embodiments of the present disclosure for executing the above text generation method can also be a server. Refer to Figure 15 , Figure 15 Partial structural block diagram of the server provided by the embodiments of the present disclosure. The server may vary greatly due to configuration or performance differences, and may include one or more second processors 1510 and a second memory 1530, and one or more storage media 1540 (such as one or more mass storage devices) for storing application programs 1543 or data 1542. Among them, the second memory 1530 and the storage media 1540 may be transient storage or persistent storage. The programs stored in the storage media 1540 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the second processor 1510 may be configured to communicate with the storage media 1540 and execute a series of instruction operations in the storage media 1540 on the server.
[0226] The server may further include one or more second power supplies 1520, one or more wired or wireless network interfaces 1550, one or more input / output interfaces 1560, and / or one or more operating systems 1541, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0227] The second processor 1510 in the server may be used to execute the text generation method.
[0228] The embodiments of the present disclosure further provide a computer-readable storage medium for storing a computer program for executing the text generation method of the foregoing various embodiments.
[0229] The embodiments of the present disclosure further provide a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the text generation method described above.
[0230] In the description of the present disclosure and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so as to describe the embodiments of the present disclosure. For example, the embodiments can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0231] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist simultaneously. Here, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0232] It should be understood that in the description of the embodiments of the present disclosure, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as "greater than", "less than", "exceeding", etc. do not include the present number, and understandings such as "above", "below", "within", etc. include the present number.
[0233] In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be an indirect coupling or communication connection through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0234] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0235] In addition, the functional units in various embodiments of the present disclosure may be integrated into one processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0236] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present disclosure. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0237] It should also be understood that the various embodiments provided in the present disclosure can be combined arbitrarily to achieve different technical effects.
[0238] The above is a specific description of the preferred embodiments of the present disclosure, but the present disclosure is not limited to the above-mentioned embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present disclosure, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present disclosure.< / language> < / language>
Claims
1. A text generation method, characterized in that: include: Acquire the currently input database operation text and the current cursor position, and extract the database object identifier in the database operation text based on the cursor position; Recalling candidate metadata information of multiple data tables based on the database object identifier, sorting the multiple candidate metadata information and then screening them to obtain target metadata information; Acquire a context segment adjacent to the cursor position, and construct a prompt text based on the target metadata information and the context segment, wherein the prompt text is used to prompt a text completion model to complete the database operation text; The prompt text is input into the text completion model for text prediction to generate a completion text for the database operation text.
2. The text generation method according to claim 1, characterized in that: The step of extracting the database object identifier in the database operation text based on the cursor position includes: extracting the table name text after normalizing and converting the database operation text; Divide the database operation text into multiple statement blocks, and extract the field name text from each of the statement blocks based on a preset parsing library; The table name text and the field name text belonging to the same statement block are merged to obtain the database object identifier in each statement block of the database operation text.
3. The text generation method according to claim 2, characterized in that: The extracting the field name text from each of the sentence blocks based on a preset parsing library includes: Extracting field names from each of the statement blocks based on a preset parsing library, and traversing each line of the statement block when the extraction based on the preset parsing library fails; If the database syntax keyword is missing at the beginning of the currently traversed row, the database syntax keyword is added to the beginning of the currently traversed row and then the field name text is extracted.
4. The text generation method according to claim 2, characterized in that: The step of merging the table name text and the field name text belonging to the same statement block to obtain the database object identifier in each statement block of the database operation text includes: Performing prefix judgment on the table name text according to the positional relationship between the cursor position and the table name text to obtain a first judgment result, and classifying the table name text according to the first judgment result, wherein the first judgment result is used to indicate that the table name text is a table name prefix or a complete table name; Performing a prefix judgment on the field name text according to the positional relationship between the cursor position and the field name to obtain a second judgment result, and classifying the field name text according to the second judgment result, wherein the second judgment result is used to indicate that the field name text is a field name prefix or a complete field name; The table name text and the field name text that belong to the same statement block and are divided into types are merged to obtain the database object identifier in each statement block of the database operation text.
5. The text generation method according to claim 1, characterized in that: The step of recalling candidate metadata information of multiple data tables based on the database object identifier includes: Determine the business space associated with the target account that inputs the database operation text, and recall metadata information of multiple recall paths based on the business space, the cursor position, and the database object identifier; The recall results of each of the recall paths are merged to obtain candidate metadata information of multiple data tables.
6. The text generation method according to claim 5, characterized in that: The recalling of metadata information of multiple recall paths based on the business space, the cursor position and the database object identifier includes: For the recall path corresponding to the business space, the data tables in the business space are sorted by popularity, a first target data table is determined from the data tables in the business space according to the sorting result, and metadata information of the first target data table is recalled; For the recall path corresponding to the cursor position and the database object identifier, a neighboring identifier of the cursor position is determined in the database object identifier, a data table containing the neighboring identifier is determined as a second target data table, and metadata information of the second target data table is recalled.
7. The text generation method according to claim 6, characterized in that: The neighboring identifier includes a neighboring complete field name, a neighboring field name prefix, and a neighboring table name prefix, and determining the data table containing the neighboring identifier as the second target data table includes: A data table including the adjacent complete field name, the adjacent field name prefix and the adjacent table name prefix is determined as the second target data table.
8. The text generation method according to claim 5, characterized in that: The step of sorting and screening the plurality of candidate metadata information to obtain target metadata information includes: Acquire the path weight of each of the recall paths, determine the sorting weight of the candidate metadata information according to the path weight, and perform preliminary sorting on the plurality of candidate metadata information based on the sorting weight to obtain a preliminary sorting result; Rearranging the preliminary ranking results based on at least one of the confidence of the candidate metadata information, the access permission of the target account to the candidate metadata information, or the integrity check result of the candidate metadata information to obtain a rearranged result; Target metadata information is obtained by screening out the rearrangement results.
9. The text generation method according to claim 8, characterized in that: The determining the ranking weight of the candidate metadata information according to the path weight includes: For any of the candidate metadata information, the database operation text and the candidate metadata information are spliced and input into a weight prediction model for prediction to obtain a feature weight of the candidate metadata information; The feature weight and the path weight are weighted and summed to obtain the ranking weight of the candidate metadata information.
10. The text generation method according to claim 1, characterized in that: The constructing the prompt text based on the target metadata information and the context fragment includes: Determining dialect information of the context segment, wherein the dialect information is used to indicate the dialect used by the context segment; The current time when the database operation text is input is obtained, and a prompt text is constructed based on the dialect information, the context fragment, the target metadata information and the current time.
11. The text generation method according to claim 10, characterized in that: The constructing the prompt text based on the dialect information, the context fragment, the target metadata information and the current time includes: Constructing a dialect information text based on the dialect information and a first identification text for identifying the dialect information; Constructing a table construction text in a data definition language based on the target metadata information, and constructing a metadata information text based on the table construction text and a second identification text used to identify the table construction text; constructing a context information text based on the context segment and a third identification text for identifying the context segment; A prompt text is constructed based on the dialect information text, the metadata information text, the current time, the context information text, and a fourth identification text for identifying the completion text.
12. The text generation method according to claim 11, characterized in that: The target metadata information includes a field name text, and the step of constructing the metadata information text based on the table building text and a second identification text for identifying the table building text includes: Obtaining a corresponding enumeration value according to the field name text, and adding the enumeration value to the table creation text in parallel with the field name text; The metadata information text is constructed based on the table building text after the enumeration value is added and the second identification text used to identify the table building text.
13. The text generation method according to claim 11, characterized in that: The step of constructing the prompt text based on the dialect information text, the metadata information text, the current time, the context information text, and a fourth identification text for identifying the completion text includes: Retrieving historical reference texts similar to the database operation texts; A prompt text is constructed based on the dialect information text, the metadata information text, the current time, the historical reference text, the context information text, and a fourth identification text for identifying the completion text.
14. The text generation method according to claim 1, characterized in that: The step of inputting the prompt text into the text completion model for text prediction to generate a completion text for the database operation text includes: Inputting the prompt text into the text completion model for streaming text prediction, and caching the predicted words output by each round of the text completion model; When the text completion model ends, a prediction result is obtained based on the predicted words cached in multiple rounds, and a suppression process is performed on the prediction result to generate a completion text for the database operation text, wherein the suppression process is used to delete abnormal content in the prediction result.
15. A text generation device, characterized in that: include: An extraction module, used to obtain the currently input database operation text and the current cursor position, and extract the database object identifier in the database operation text based on the cursor position; A recall module, used to recall candidate metadata information of multiple data tables based on the database object identifier, sort the multiple candidate metadata information and then screen them to obtain target metadata information; A construction module, used to obtain a context segment adjacent to the cursor position, and to construct a prompt text based on the target metadata information and the context segment, wherein the prompt text is used to prompt a text completion model to complete the database operation text; The inference module is used to input the prompt text into the text completion model for text prediction, and generate a completion text for the database operation text.
16. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the text generation method according to any one of claims 1 to 14 is implemented.
17. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the text generation method according to any one of claims 1 to 14 is implemented.
18. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the text generation method according to any one of claims 1 to 14 is implemented.