Table building method and device based on large language model and speculation algorithm cooperation
By combining large language models with inference algorithms, the problem of inaccurate table structure definitions in existing technologies has been solved, achieving accuracy and automation in table structure design, and improving table creation efficiency and data processing performance.
Patent Information
- Application Number
- CN202511140627.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-15
AI Technical Summary
The existing "table creation during import" technology cannot accurately understand the meaning of table headers and data when faced with complex and diverse data and business scenarios. This results in inaccurate table structure definitions such as field types, constraints, and indexing strategies, which affects data storage efficiency, query performance, and data accuracy and integrity.
By combining a large language model with an inference algorithm, the target file is parsed to obtain the table header, sample data and statistical features. The semantic understanding capability of the large language model is used to correct the initial inference results of the inference algorithm, generating more accurate field inference results. Based on the corrected inference results, a visual table structure and SQL table creation statements are generated.
It achieves accuracy and automation in table structure design under complex business scenarios, reduces the cost of manual intervention, improves table creation efficiency, and enhances data storage and query performance.
Smart Images

Figure CN120725022B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of data processing and large language model technology, and more specifically, to a table building method and apparatus based on the collaboration of large language models and inference algorithms. Background Technology
[0002] With the rapid development and widespread application of emerging technologies such as big data, low-code, no-code, and data-driven approaches, software systems face a growing demand for dynamically created tables. These new technological trends require systems to be more flexible, capable of creating table structures in real time based on users' actual business needs. The traditional "create table first, then import" model clearly cannot meet this requirement. Therefore, the technology of "creating tables during import" has emerged.
[0003] Current "table creation upon import" technology struggles to understand the meaning of table headers and data when faced with complex and diverse data and business scenarios. Its inferences regarding table structure definitions, such as field types, constraints, and indexing strategies, are often inaccurate. For example, when processing data with specific business implications, inferences based solely on the data itself may fail to accurately grasp its true meaning, leading to unreasonable table structure definitions. This further impacts data storage efficiency, query performance, and data accuracy and integrity, causing numerous inconveniences for subsequent business processing. Summary of the Invention
[0004] In view of the above situation, this application provides a table building method and apparatus based on the collaboration of a large language model and a prediction algorithm, which aims to solve the above problems or at least partially solve the above problems.
[0005] In a first aspect, embodiments of this application provide a table construction method based on the collaboration of a large language model and an inference algorithm, the method comprising:
[0006] The target file is parsed to obtain the table header, sample data, and sample statistical features;
[0007] Based on the speculative algorithm, preliminary inferences are made on the table header and sample data to generate preliminary inference results for the fields. The preliminary inference results include data type, primary key identifier, field name, index identifier, and null value identifier.
[0008] The field name, sample data, sample statistical features, and preliminary inference results are input into the large language model, and the corrected inference results of the field are obtained by combining preset prompt words; the corrected inference results include data type, primary key identifier, field name, index identifier, null value identifier, and standardized naming;
[0009] The corrected inference results are used to generate a visual table structure;
[0010] In response to the user's confirmation of the visualized table structure, an SQL table creation statement is generated based on the corrected inference result, and the table is created based on the SQL table creation statement.
[0011] Secondly, embodiments of this application also provide a table-building device based on the collaboration of a large language model and an inference algorithm, the device comprising:
[0012] The parsing module is used to parse the target file and obtain the table header, sample data, and sample statistical features;
[0013] The inference module is used to make preliminary inferences on the table header and sample data based on the inference algorithm, and generate preliminary inference results for the fields. The preliminary inference results include data type, primary key identifier, field name, index identifier, and null value identifier.
[0014] The correction module is used to input field names, sample data, sample statistical features, and preliminary inference results into the large language model, and obtain the corrected inference results of the field by combining preset prompt words; the corrected inference results include data type, primary key identifier, field name, index identifier, null value identifier, and standardized naming;
[0015] The generation module is used to generate a visual table structure from the corrected inference results;
[0016] The table creation module is used to respond to the user's confirmation operation on the visualized table structure, generate an SQL table creation statement based on the corrected inference result, and create the table based on the SQL table creation statement.
[0017] Thirdly, embodiments of this application also provide an electronic device, including: a processor; and a memory arranged to store computer-executable instructions, which, when executed, cause the processor to perform the steps described in the first aspect.
[0018] Fourthly, embodiments of this application also provide a computer-readable storage medium that stores one or more programs, which, when executed by an electronic device including multiple applications, cause the electronic device to perform the steps described in the first aspect.
[0019] The at least one technical solution adopted in this application embodiment can achieve the following beneficial effects: First, the file is parsed to obtain the table header, sample data, and sample statistical features. Based on the inference algorithm, a preliminary inference is made on the table header and sample data to generate preliminary inference results for the fields. Then, the field names, sample data, sample statistical features, and preliminary inference results are input into a large language model. The preliminary inference results are corrected using preset prompts to obtain corrected inference results. By combining the inference algorithm and the large language model, and utilizing the semantic understanding and inference capabilities of the large language model, the problem of hard-coded data being unable to understand the meaning of table headers and data is solved, thus enabling more accurate inference results. Finally, a visual table structure is generated based on the corrected inference results. In response to the user's confirmation operation of the visual table structure, an SQL table creation statement is generated based on the corrected inference results, and the table is automatically created based on the SQL table creation statement. This achieves accuracy, rationality, and automation in table structure design under complex business scenarios, reduces manual intervention costs, and improves table creation efficiency. Attached Figure Description
[0020] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0021] Figure 1 This illustration shows a schematic diagram of the structure of a table building system based on the collaboration of a large language model and an inference algorithm, as provided in an embodiment of this application.
[0022] Figure 2 The flowchart of the table construction method based on the collaboration of a large language model and an inference algorithm provided in the embodiments of this application is shown.
[0023] Figure 3 A flowchart of a table-building method based on the collaboration of a large language model and an inference algorithm, according to another embodiment of this application, is shown.
[0024] Figure 4 This diagram illustrates the structure of the table-building device based on the collaboration of a large language model and an inference algorithm provided in an embodiment of this application.
[0025] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0027] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such use can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the term "comprising" and its variations should be interpreted as open-ended terms meaning "including but not limited to."
[0028] As described in the background section, the current "table creation during import" technology cannot understand the meaning of table headers and data when faced with complex and diverse data and business scenarios, and its inferences about table structure definitions such as field types, constraints, and indexing strategies are not accurate enough.
[0029] If the data type inference is inaccurate, type conversion errors may occur during data import, or the table structure may need to be manually modified later. The reasons for inaccurate data type inference are as follows: (1) Analyzing only a small amount of sample data, such as the first 100 rows, may lead to misjudgment of the type. For example, all values in the sample are integers, but the actual data contains decimals, so it is inferred to be INT instead of FLOAT; (2) The impact of null values: a certain field in the sample data is all empty, but the table header name is longitude. The semantic understanding of the table header is missing, and it may be misjudged as an empty string, so the data type is set to VARCHAR instead of the actual numeric type; (3) Missing special values: the sample does not contain extreme values or special characters, such as very long strings or scientific notation values.
[0030] The reasons for missing constraints include: (1) missing primary key / unique key: it is impossible to determine which fields constitute a unique identifier based on the sample alone; (2) non-null constraint problem: a field in the sample has no null value, but the actual data allows null values; (3) missing default value / check constraint: it is impossible to infer business rules (such as the legal value range of status fields) from the sample data.
[0031] If the field length is unreasonable, the field length may need to be modified frequently, affecting system stability and performance. The reasons affecting the field length are as follows: (1) The string field length is set based on the maximum value of the sample. For example, if the longest string in the sample is 20 characters, it is set to VARCHAR(20); (2) The actual data may contain excessively long strings, which may cause data truncation or errors.
[0032] If indexing and optimization are insufficient, fields frequently used in query conditions or join operations cannot be automatically identified.
[0033] If the mapping name is based on a preset rule, for example, if the table header field is "province", it can be directly mapped to "province". However, if the mapping rule does not include "province", it is difficult to give a reasonable field name. If the table header name is used directly, for example, if the table header field is "province", it can be used directly. However, if the table header field is "aa" or "bb", which are meaningless names, directly using "aa" or "bb" will make the field name meaningless.
[0034] Furthermore, setting larger lengths or higher precision types (such as DECIMAL(38,18)) for all fields increases storage costs, reduces read and write performance, and leads to a waste of performance and resources.
[0035] Based on this, this application proposes a table building method based on the collaboration of a large language model and a prediction algorithm. By combining the prediction algorithm and the large language model, the semantic understanding and inference capabilities of the large language model are used to correct the initial inference results of the prediction algorithm. This solves the problem that hard-coded data cannot understand the meaning of table headers and data, and can reasonably predict the most suitable table structure content, making table building more reasonable. This achieves accuracy, rationality and automation of table structure design in complex business scenarios, reduces manual intervention costs and improves table building efficiency.
[0036] Figure 1 This diagram illustrates the structure of a table-building system based on a large language model and inference algorithm according to an embodiment of this application. The table-building method based on a large language model and inference algorithm provided in this embodiment can be implemented in, for example... Figure 1 The system implementation, from Figure 1 As can be seen, the table creation system 100 based on the collaboration of a large language model and inference algorithms includes a user interaction layer, an intelligent processing layer, and a data persistence layer. Specifically, users upload files through the user interaction layer, which automatically triggers the intelligent processing layer to analyze the data and displays a real-time preview of the generated table structure on the interface. After user confirmation, the generated SQL table creation statement is sent to the database management system for execution. Once the table creation is complete, the data from the user-uploaded file is imported into the database table in the data persistence layer according to the generated table structure, achieving persistent data storage.
[0037] It should be noted that this application is not limited to Figure 1 The table-building system shown is based on the collaboration of a large language model and an inference algorithm. Any system, device, or framework that can implement the business logic of this application is acceptable. Figure 1 This is merely an illustrative example.
[0038] Figure 2 This illustration shows a flowchart of the table construction method based on the collaboration of a large language model and an inference algorithm provided in an embodiment of this application. Figure 2 It can be seen that this application includes at least steps S101-S105:
[0039] Step S101: Parse the target file to obtain the table header, sample data, and sample statistical features.
[0040] In some embodiments, a user interface is designed at the front end of the software system, providing a file upload entry that supports common file types such as CSV and Excel files. The target file is the file uploaded by the user. Based on the target file uploaded by the user, the corresponding file parser is invoked. For example, for a CSV file, a CSV parsing library is used to read the file content, extracting the header and randomly sampled data. For an Excel file, an Excel parsing tool is used to read the worksheet data, similarly extracting the header and several rows of randomly sampled data.
[0041] Furthermore, the statistical characteristics of the samples are determined based on the sample data. Specifically, the sample data is preprocessed to calculate the statistical characteristics of each field, including but not limited to the field's null value rate, the proportion of unique values, and the minimum or maximum value of numeric fields.
[0042] Step S102: Based on the inference algorithm, make preliminary inferences on the table header and sample data to generate preliminary inference results for the fields.
[0043] In some embodiments, the inference process of the inference algorithm is as follows: the default data type of the field is set to string type, the data type of the field is detected based on the priority order of data types, the field length is determined based on the maximum length of the field in the sample data and the preset length, and whether the field is a primary key is determined based on the proportion of field name keywords and unique values.
[0044] Specifically, for each field, the default data type is first set to string with a length of 255. Then, the data type is checked sequentially based on the priority order of the data types using a type detection function, for example, determining the data type of the field in the order of integer > floating-point number > date > string. During the detection process, the final data type length is determined by adding a preset length (e.g., 2 digits) to the longest length of the field in the sample data, overriding the default length. For example, if the longest length of a field in the sample data is detected to be 10 digits and it is determined to be a string type, then the data type of that field is finally set to VARCHAR (12). At the same time, the field name and the repetition rate of the data are combined to determine whether it is a primary key. If the field name contains words such as "ID" or "number" and the data repetition rate is 0 (i.e., the proportion of unique values is 100%), then it is initially inferred that the field is a primary key.
[0045] In other embodiments, for certain special fields, a pre-defined mapping between field names and data types is established. After identifying the field name based on the table header, the corresponding data type is determined based on the pre-defined mapping, and indexes are set according to the business requirements in the file. For example, if the field name is identified as "latitude" or "lat", its data type is directly set to DECIMAL (10, 6).
[0046] In some embodiments, the preliminary inference result includes data type, primary key identifier, field name, index identifier, and null value identifier. For example, the preliminary inference result is as follows:
[0047] {
[0048] "data_type":"varchar(12)",
[0049] "primary_key": false,
[0050] "field_name":"Order Date",
[0051] "nullable": false,
[0052] "index": "true"
[0053] }
[0054] Step S103: Input the field name, sample data, sample statistical features and preliminary inference results into the large language model, and obtain the corrected inference results of the field by combining the preset prompt words.
[0055] When constructing prompt words for a large language model, we fully leverage its powerful semantic understanding and inference capabilities, describing it from multiple dimensions such as field names, sample data, and statistical features. For example, field names help the model understand business meaning; "ID" usually indicates a primary key, and "amount" is likely a numeric type. Sample data provides specific data examples, helping the model recognize different formats, such as dates and email addresses. Statistical features provide quantitative indicators to assist in accurate inference. Furthermore, we force the model to output in machine-parseable JSON format, achieving a closed loop of "prompt word engineering + automatic parsing," effectively avoiding ambiguity in natural language responses and facilitating direct processing by subsequent programs.
[0056] In some embodiments, the preset prompt word includes at least one of the following:
[0057] The document describes the business scenario.
[0058] Provide indexing recommendations for the fields used in query conditions;
[0059] Identify the primary key fields;
[0060] A flag is given for fields that cannot be empty;
[0061] Standardized field names are determined based on the business scenario description, field names, and sample values in the file.
[0062] The output format is JSON.
[0063] Specifically, the system interprets the meaning of input information (e.g., business scenario description of a file, sample values, statistical features, preliminary inferences) using preset prompts; verifies preliminary inferences: requires the large language model to check whether the preliminary inference results are consistent with field names, sample values, and statistical features; corrects inconsistencies: if inconsistencies are found, requires the large language model to provide more reasonable inferences based on all input information; identifies semantic ambiguity: for semantically ambiguous field names, uses sample values and statistical features for inference; infers implicit semantics / business logic: combines field names, value meanings, and context to infer more precise meanings; identifies and corrects special formats: whether more specific formats (specific date and time formats) can be identified; refines data types / constraints: whether more precise data types can be inferred; whether other constraints should be added; infers units: whether numerical fields imply units; sets output formats: specifies how the large language model should output the corrected results in a structured manner.
[0064] In addition, in some embodiments, domain knowledge can be injected into the large language model through preset prompt words: the prompt words can contain background knowledge or terminology explanations of specific domains. For example, requiring users to describe the content of a file when importing it can serve as a description of the file's business scenario, helping the large language model make more scenario-appropriate inferences.
[0065] Specifically, after receiving data such as field names, sample data, and statistical features, the large language model preprocesses the data using built-in prompts. This stage includes text segmentation, encoding the data into a numerical form that the model can understand, and parsing parameters to configure subsequent steps. Next, the context modeling stage begins, utilizing the Transformer architecture to perform self-attention computation to capture deep relationships and semantic context within the input content. Then, the model moves to the generative inference stage, employing an autoregressive decoding mechanism to progressively generate the output sequence, while combining sampling strategies to control the diversity and quality of the generated results. Subsequently, in the post-processing stage, the generated sequence is decoded back into readable text, undergoing format adjustments (such as adding punctuation or optimizing structure) and content filtering (such as removing inappropriate or redundant information) to improve the readability and security of the final output. Finally, the corrected inference results are output in JSON format.
[0066] In some embodiments, the correction of the inference result includes data type, primary key identifier, field name, index identifier, null value identifier, and standardized naming.
[0067] For example, suppose the field name is "Order Date", the sample values are ["2023-01-01", "2023-02-15", "2023-03-20", NULL, "2023-04-10"], and the statistical features are {"null_ratio": 0.2, "unique_ratio": 0.8, "min": "2023-01-01", "max": "2023-04-10"}. These parameters are sent to the large language model, and combined with preset prompts, the model outputs a corrected inference result through the following steps:
[0068] "You are a professional data analyst. Based on the following information, infer the optimal data type and constraints for the field:"
[0069] Document Description: This document is an order form, containing information about the orders placed by customers.
[0070] Field Name: Order Date
[0071] Sample values: ["2023-01-01", "2023-02-15", "2023-03-20", NULL, "2023-04-10"]
[0072] Statistical characteristics: {“null_ratio”: 0.2, “unique_ratio”: 0.8, “min”: “2023-01-01”, “max”: “2023-04-10”}
[0073] Preliminary inference results: {
[0074] "data_type":"varchar(12)",
[0075] "primary_key": false,
[0076] "field_name":"Order Date",
[0077] "nullable": false,
[0078] "index": "true"
[0079] }
[0080] Please answer:
[0081] Recommended MySQL database data types (such as VARCHAR(255) / INT / DATETIME, etc.)
[0082] Should it be set as a primary key / unique key?
[0083] Are empty strings allowed?
[0084] Recommended indexing strategy
[0085] More appropriate field naming
[0086] Other suggested constraints
[0087] Response format (forces the model to output in a program-parseable JSON format for easier subsequent processing. This avoids ambiguity in natural language responses and achieves a closed loop of "prompt word engineering + automatic parsing"):
[0088] {
[0089] "data_type":"DATETIME",
[0090] "primary_key":false,
[0091] "nullable":true,
[0092] "index":"true",
[0093] "field_name":
[0094] "constraints":[]
[0095] }”
[0096] result:
[0097] {
[0098] "data_type":"DATE",
[0099] "primary_key":false,
[0100] "nullable":true,
[0101] "index":"true",
[0102] "field_name":"order_date",
[0103] "constraints":[]
[0104] }
[0105] In another example, suppose the field name is "weidu", the sample values are ["31.267452", "35.28061", "34.34127", "34.210476", "34.259116"] (a total of 5 values), and the statistical features are {"null_ratio": 0, "unique_ratio": 1, "min": "31.267452", "max": "35.28061"}. These parameters are sent to the large language model, and combined with preset prompts, the model outputs a corrected inference result through the following steps:
[0106] You are a professional data analyst. Based on the following information, infer the optimal data type and constraints for the field:
[0107] Document Description: This document is a store table, containing store-related information and locations.
[0108] Field name: weidu
[0109] Sample values: ["31.267452", "35.28061", "34.34127", "34.210476", "34.259116"] (5 values in total)
[0110] Statistical characteristics: {"null_ratio":0,"unique_ratio":1,"min":"31.267452","max":"35.28061"}
[0111] Preliminary inference results: {
[0112] "data_type":"double(8,6)",
[0113] "primary_key": false,
[0114] "field_name":"weidu",
[0115] "nullable": false,
[0116] index: "false"
[0117] }
[0118] Please answer:
[0119] Recommended MySQL database data types (such as VARCHAR(255) / INT / DATETIME, etc.)
[0120] Should it be set as a primary key / unique key?
[0121] Are empty strings allowed?
[0122] Recommended indexing strategy
[0123] More appropriate field naming
[0124] Other suggested constraints
[0125] Response format (forces the model to output in a program-parseable JSON format for easier subsequent processing. This avoids ambiguity in natural language responses and achieves a closed loop of "prompt word engineering + automatic parsing"):
[0126] {
[0127] "data_type": "DATETIME",
[0128] "primary_key": false,
[0129] "nullable": true,
[0130] "index": "true",
[0131] "field_name":
[0132] "constraints": []
[0133] }
[0134] result:
[0135] {
[0136] "data_type": "DECIMAL(8,6)",
[0137] "primary_key": false,
[0138] "nullable": false,
[0139] "index": "false",
[0140] "field_name": "latitude",
[0141] "constraints": [
[0142] "CHECK (latitude BETWEEN -90 AND 90)" ]
[0144] }
[0145] Step S104: Generate a visual table structure from the corrected inference results.
[0146] Step S105: In response to the user's confirmation operation on the visualized table structure, generate an SQL table creation statement based on the corrected inference results, and create the table based on the SQL table creation statement.
[0147] The generated field information and other table structure information are displayed to the user in a visual format for confirmation. If the user confirms that everything is correct, the system generates SQL table creation statements based on the results returned by the large language model.
[0148] from Figure 2 As can be seen from the method shown, this application combines the inference algorithm with a large language model, and uses the semantic understanding and inference capabilities of the large language model to correct the initial inference results of the inference algorithm. This solves the problem that hard-coded data cannot understand the meaning of table headers and data, and can reasonably infer the most suitable table structure content, making table building more reasonable. It achieves accuracy, rationality and automation of table structure design in complex business scenarios, reduces manual intervention costs and improves table building efficiency.
[0149] Figure 3 This illustration shows a flowchart of a table-building method based on a large language model and inference algorithm according to another embodiment of this application. Figure 3 As can be seen, the table construction method based on the collaboration of a large language model and an inference algorithm in this embodiment includes the following steps S201-S203:
[0150] Step S201: In response to the user's modification operation on the visual table structure, record the user's modification content.
[0151] Step S202: Adjust and correct the inference results based on the user's modifications.
[0152] Step S203: Create a table based on the adjusted and corrected inference results.
[0153] In this embodiment, the generated field information and other table structure information are displayed to the user in a visual format for confirmation. If the user finds a problem and makes adjustments, the system records the user's adjustments, such as the modified field types, lengths, primary key settings, etc., and then generates an SQL table creation statement based on the adjusted results to create the table.
[0154] Furthermore, these records will be accumulated as new business knowledge and used to optimize the rules of prompt words and inference algorithms for large language models, so as to provide more accurate table structure suggestions that better meet user needs when processing similar data in the future, forming a virtuous cycle of continuous optimization.
[0155] In the table creation method based on the collaboration of a large language model and an inference algorithm provided in this application, the large language model comprehensively understands field names, sample values, and statistical characteristics, enabling a deeper understanding of the business meaning behind the data and effectively avoiding potential biases that may occur when simply inferring data types based on sample data. For example, when facing data fields with specific business rules, the large language model can accurately determine their data type based on various information, while traditional inference algorithms may make incorrect judgments due to a lack of semantic understanding. Furthermore, the refined processing mechanism when the large language model and the inference algorithm are combined further improves the accuracy of data type inference. When the inference results from both are inconsistent, a re-examination of the sample data and multi-dimensional judgment can select the data type that best meets business requirements, thereby ensuring the accuracy and consistency of data during storage and processing.
[0156] Furthermore, based on its powerful semantic understanding capabilities, the large language model can provide precise suggestions for setting constraints according to the business meaning of fields and the characteristics of sample data. For example, for fields representing latitude and longitude, it can not only accurately infer the appropriate data type but also provide reasonable value range constraints to ensure data validity. Regarding field length settings, by combining the large language model's semantic inference and prediction algorithms with statistical analysis of sample data length, it avoids both the waste of storage space caused by excessively long field lengths and the problem of data truncation due to excessively short lengths. This precise setting not only optimizes the data storage structure but also improves data processing efficiency.
[0157] Furthermore, the large language model can intelligently determine which fields are suitable for indexing and what indexing strategy to use based on field name semantics and data distribution. For example, it automatically suggests creating indexes for frequently used fields such as "province," "city," "district," and "date," thereby significantly improving the execution speed of query operations. Compared to traditional methods that require developers to create indexes separately later, this application completes the precise index setting during the table creation stage, reducing development workload and avoiding performance issues caused by improper index settings. This ensures that the system can quickly respond to query requests when processing large amounts of data, improving the overall system performance.
[0158] Furthermore, the semantic understanding capabilities of the large language model enable it to provide more intuitive and accurate field naming suggestions based on the meaning of the table headers and sample data. This not only helps developers better understand and maintain the database structure but also facilitates data sharing and communication between different departments. Reasonable field naming reduces data usage errors caused by unclear naming, improving data readability and manageability.
[0159] Furthermore, this application effectively reduces data storage costs by rationally setting the data type, length, and precision of fields. It avoids wasted storage space caused by inappropriate data type selection or unreasonable field length settings, allowing the database to occupy less physical space when storing the same amount of data. This reduction in storage costs is particularly significant for applications with large data volumes, helping enterprises reduce hardware investment and operating costs while maintaining data processing capabilities.
[0160] In some embodiments of this application, a table building device based on the collaboration of a large language model and a prediction algorithm is provided. This table building device corresponds one-to-one with the table building methods based on the collaboration of a large language model and a prediction algorithm described in the above embodiments. Figure 4 As shown, the table building device based on the collaboration of a large language model and an inference algorithm includes a parsing module 101, an inference module 102, a correction module 103, a generation module 104, a table building module 105, and an optimization module 106.
[0161] The parsing module 101 is used to parse the target file and obtain the table header, sample data and sample statistical features;
[0162] The inference module 102 is used to make preliminary inferences on the table header and sample data based on the inference algorithm, and generate preliminary inference results for the fields. The preliminary inference results include data type, primary key identifier, field name, index identifier and null value identifier.
[0163] The correction module 103 is used to input the field name, sample data, sample statistical features and preliminary inference results into the large language model, and obtain the corrected inference results of the field by combining preset prompt words; the corrected inference results include data type, primary key identifier, field name, index identifier, null value identifier and standardized naming;
[0164] The generation module 104 is used to generate a visual table structure from the corrected inference results;
[0165] The table creation module 105 is used to respond to the user's confirmation operation on the visualized table structure, generate an SQL table creation statement based on the correction inference result, and create a table based on the SQL table creation statement.
[0166] In some embodiments of this application, the sample statistical characteristics in the above-described apparatus include at least one of the following: field null value rate, unique value ratio, minimum or maximum value of a numeric field.
[0167] In some embodiments of this application, in the above-described apparatus, the inference module 102 is specifically used to: default the data type of the field to string type; detect the data type of the field based on the priority order of data types; determine the field length based on the maximum length and preset length of the field in the sample data; and determine whether the field is a primary key based on the field name keywords and the proportion of unique values.
[0168] In some embodiments of this application, in the above-described apparatus, the inference module 102 is specifically used to determine the preliminary inference result of a field based on a preset correspondence between the field name and the data type.
[0169] In some embodiments of this application, the preset prompt words include at least one of the following: a business scenario description of the file; index recommendation rules for fields used as query conditions; an identifier for primary key fields; an identifier for fields that cannot be null; standardized field names are determined based on the business description, field name, and sample value of the fields; and the output format is JSON.
[0170] In some embodiments of this application, in the above-described apparatus, the table creation module 105 is further configured to, in response to a user's modification operation on the visual table structure, record the user's modifications; adjust and correct the inference results based on the user's modifications; and create a table based on the adjusted and corrected inference results.
[0171] In some embodiments of this application, in the above-described apparatus, the optimization module 106 is used to take the user's modified content as business rule knowledge; and to optimize the inference rules of the inference algorithm and the prompt words of the large language model based on the business rule knowledge.
[0172] It should be noted that any of the above-mentioned table building devices based on the collaboration of large language models and inference algorithms can implement the aforementioned table building method based on the collaboration of large language models and inference algorithms, which will not be elaborated here.
[0173] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Figure 5 As shown, at the hardware level, this electronic device includes a processor, and optionally also includes an internal bus, a network interface, and memory. The memory may include main memory, such as high-speed random-access memory (RAM), or it may include non-volatile memory, such as at least one disk drive. Of course, this electronic device may also include other hardware required for other business operations.
[0174] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 5 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0175] Memory is used to store programs. Specifically, programs may include program code, which includes computer operation instructions. Memory may include main memory and non-volatile memory, and provides instructions and data to the processor.
[0176] The processor reads the corresponding computer program from non-volatile memory into main memory and then runs it, forming a table-building mechanism at the logical level based on the collaboration of a large language model and inference algorithms. The processor executes the program stored in memory and specifically performs the aforementioned methods.
[0177] The processor may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0178] This electronic device can execute the table building method based on the collaboration of a large language model and an inference algorithm provided in several embodiments of this application, and is implemented as a table building device based on the collaboration of a large language model and an inference algorithm. Figure 4 The functions of the embodiments shown are not described in detail here.
[0179] This application also proposes a computer-readable storage medium that stores one or more programs, the programs including instructions that, when executed by an electronic device including multiple applications, enable the electronic device to perform the table-building method based on the collaboration of a large language model and an inference algorithm provided in several embodiments of this application.
[0180] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0181] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0182] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0183] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0184] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0185] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0186] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0187] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0188] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0189] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A table construction method based on the collaboration of a large language model and an inference algorithm, characterized in that, The method includes: The target file is parsed to obtain the table header, sample data, and sample statistical features; Based on the speculative algorithm, preliminary inferences are made on the table header and sample data to generate preliminary inference results for the fields. The preliminary inference results include data type, primary key identifier, field name, index identifier, and null value identifier. The field name, sample data, sample statistical features, and preliminary inference results are input into the large language model, and the corrected inference results of the field are obtained by combining preset prompt words; the sample statistical features include at least one of the following: field null value rate, unique value ratio, minimum value of numeric field or maximum value of numeric field; the corrected inference results include data type, primary key identifier, field name, index identifier, null value identifier, and standardized naming; The corrected inference results are used to generate a visual table structure; In response to the user's confirmation of the visualized table structure, an SQL table creation statement is generated based on the corrected inference result, and the table is created based on the SQL table creation statement.
2. The method according to claim 1, characterized in that, The preliminary inference of the table header and sample data based on the inference algorithm to generate preliminary inference results for the fields includes: The default data type of the field is set to string. The data type of a field is detected based on its priority order. The field length is determined based on the maximum and preset lengths of the fields in the sample data. Whether a field is a primary key is determined by the keyword in the field name and the proportion of unique values.
3. The method according to claim 1, characterized in that, The preliminary inference of the table header and sample data based on the inference algorithm to generate preliminary inference results for the fields includes: Based on the pre-defined correspondence between field names and data types, the preliminary inference results of the fields are determined.
4. The method according to claim 1, characterized in that, The preset prompt words include at least one of the following: The document describes the business scenario. Provide indexing recommendations for the fields used in query conditions; Identify the primary key fields; A flag is given for fields that cannot be empty; Standardized field names are determined based on the field's business description, field name, and sample values. The output format is JSON.
5. The method according to any one of claims 1-4, characterized in that, The method further includes: In response to user modifications to the visual table structure, the user's modifications are recorded. The inference results were adjusted and corrected based on the user's modifications. The table is built based on the adjusted and corrected inference results.
6. The method according to claim 5, characterized in that, The method further includes: Treat user-modified content as business rule knowledge; The inference rules of the inference algorithm and the prompt words of the large language model are optimized based on business rule knowledge.
7. A table-building device based on the collaboration of a large language model and an inference algorithm, characterized in that, The device includes: The parsing module is used to parse the target file and obtain the table header, sample data, and sample statistical features; the sample statistical features include at least one of the following: field null value rate, unique value ratio, minimum value of a numeric field, or maximum value of a numeric field; The inference module is used to make preliminary inferences on the table header and sample data based on the inference algorithm, and generate preliminary inference results for the fields. The preliminary inference results include data type, primary key identifier, field name, index identifier, and null value identifier. The correction module is used to input field names, sample data, sample statistical features, and preliminary inference results into the large language model, and obtain the corrected inference results of the field by combining preset prompt words; the corrected inference results include data type, primary key identifier, field name, index identifier, null value identifier, and standardized naming; The generation module is used to generate a visual table structure from the corrected inference results; The table creation module is used to respond to the user's confirmation operation on the visualized table structure, generate an SQL table creation statement based on the corrected inference result, and create the table based on the SQL table creation statement.
8. An electronic device, comprising: processor; as well as A memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the steps of the table-building method based on a large language model and inference algorithm as described in any one of claims 1-6.
9. A computer-readable storage medium storing one or more programs, which, when executed by an electronic device including a plurality of applications, cause the electronic device to perform the steps of the table-building method based on a large language model and inference algorithm as described in any one of claims 1-6.
Citation Information
Patent Citations
Automatic Excel template data backfilling method based on LLM semantic comprehension technology
CN119537411A
Structured query language generation method and device and nonvolatile storage medium
CN120277094A