Automatic index creation method based on large model
Through the automated indicator creation method based on large models, the full process automation from natural language requirements to SQL statements and API interfaces is realized, and the manual dependence and integration problems in traditional indicator creation methods are solved, which improves business autonomy and data interaction efficiency.
Patent Information
- Application Number
- CN202510680272.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-02
AI Technical Summary
The traditional indicator creation method has high manual participation and low efficiency, high technical threshold, lack of full-process automation capabilities, insufficient semantic understanding capabilities, and low interface standardization, resulting in low business autonomy and work efficiency, and difficult system integration.
The automated index creation method based on large models is adopted, natural language requirements are analyzed through large language models, structured parameters are generated, and SQL statements and API interfaces are automatically generated, including interactive input, data processing and interface generation modules, supporting multi-modal input and domain fine-tuning, realizing full-process automation.
It greatly improves the efficiency and accuracy of indicator creation, reduces labor costs, ensures interface normative and system compatibility, reduces integration difficulty, and improves data sharing and interaction efficiency.
Smart Images

Figure CN120578682A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a large model-based automatic indicator creation method. Background Art
[0002] In today's data-driven decision-making era, indicator creation is a key link in data analysis and business management. Traditional indicator creation methods have many drawbacks:
[0003] 1) High manual involvement and low efficiency: Business personnel must communicate requirements to data engineers, who then manually write SQL statements to calculate metrics. For example, to calculate monthly sales by region, one must write a SQL statement like SELECT region, SUM(amount) AS total_sales FROM sales_table WHERE MONTH(date_column) = [specified month] GROUP BY region; This process is time-consuming and labor-intensive, and is prone to repeated revisions due to misunderstandings of requirements.
[0004] 2) Technical barriers restrict business autonomy: Non-technical personnel find it difficult to master SQL syntax, database table structure and complex query logic (such as multi-table JOIN, subqueries, window functions, etc.), and cannot independently complete indicator creation, which limits the autonomy and work efficiency of business personnel.
[0005] 3) Complex interface development and integration: After completing the indicator calculation, the API interface needs to be manually developed, which involves interface parameter definition, response format design (such as JSON, XML), identity authentication configuration (such as OAuth2.0, API key) and security mechanism setting (such as HTTPS encryption). The process is cumbersome and the interface formats between different systems are not unified, which further increases the difficulty of integration.
[0006] Although some SQL generation tools and API gateways exist on the market, existing technologies still have obvious flaws:
[0007] 1) Lack of full-process automation capabilities: Existing tools are mostly fragmented products with single functions. They are unable to achieve full-process automation from natural language demand input to indicator calculation and interface generation. Users still need to perform a lot of manual operations and coordination work.
[0008] 2) Insufficient semantic understanding: Traditional rule engines or simple NLP technologies can only perform keyword matching and have difficulty understanding complex business semantics. For example, for a requirement to "calculate the month-over-month growth rate over the past three months," they cannot automatically derive and generate the correct calculation logic (current month's indicator minus the previous three months' indicator) or the previous three months' indicator and corresponding SQL statements.
[0009] 3) Low interface standardization: Different systems have different interface formats and specifications, which requires repeated development of adaptation code during the integration process, increasing development costs and time costs, and reducing the efficiency of data interaction between systems.
[0010] In order to solve the above problems and realize the intelligent indicator generation of demand-based interfaces, the present invention proposes an automatic indicator creation method based on a large model. Summary of the Invention
[0011] In order to overcome the shortcomings of the prior art, the present invention provides a simple and efficient method for creating automated indicators based on a large model.
[0012] The present invention is achieved through the following technical solutions:
[0013] A method for creating an automated indicator based on a large model, characterized by comprising the following steps:
[0014] Step S1: User input requirements
[0015] Users input natural language business requirement text through the interactive interface;
[0016] Step S2: Demand preprocessing
[0017] Preprocess the input text, including removing stop words, unifying time formats, and standardizing expressions;
[0018] Step S3: Semantic parsing of large language model
[0019] Use a large language model to perform named entity recognition on the pre-processed business requirement text, extract key entity information, including time entities, indicator entities, grouping entities, and data entities, and generate structured requirement parameters through semantic role labeling and logical reasoning;
[0020] Step S4: SQL statement generation
[0021] According to the structured requirement parameters, a template is matched from an SQL template library and dynamically filled to generate a preliminary SQL statement, and an executable SQL statement is obtained after syntax verification and performance optimization;
[0022] Step S5: Data connection and calculation
[0023] Connect to the data source through database connection technology, execute the SQL statement, and obtain the indicator calculation results;
[0024] Step S6: Result verification
[0025] Verify the completeness and accuracy of indicator calculation results;
[0026] Step S7: Interface definition and configuration
[0027] Define API interface parameters according to interface specifications, including interface path, request method, request parameters, response parameters, and security mechanisms;
[0028] Step S8: Generate interface document and publish
[0029] Generate an interface document containing interface functions, parameter descriptions, and calling examples, and publish the interface for external system calls.
[0030] In step S3, the extracted time entity includes time range and time granularity information;
[0031] When extracting indicator entities, first determine the indicator type to be calculated, and then associate the calculation logic corresponding to the indicator type through the built-in knowledge base;
[0032] The extracted grouping entity is the grouping dimension information;
[0033] When extracting data entities, match relevant data tables and fields from database metadata based on indicators and business scenarios;
[0034] Through semantic role labeling (SRL) and logical reasoning, the operational logic and semantic relationships in the text are analyzed to understand the complete intention of user needs.
[0035] In step S4, the SQL template library includes single-table query, multi-table association and aggregate calculation type templates;
[0036] Based on the matching of structured demand parameters, the identified grouping entity, indicator entity, and time entity information are dynamically filled into the template to generate preliminary SQL statements;
[0037] Use the SQL syntax parser to check the syntax of the generated SQL statements and correct syntax errors;
[0038] At the same time, based on database performance optimization rules, SQL statements are optimized to improve query execution efficiency.
[0039] In step S5, a connection is established with a relational database (such as MySQL, Oracle), a non-relational database (such as MongoDB) or a data warehouse (such as Snowflake, Hive) through JDBC or ODBC database connection technology, and the generated SQL statement is executed to obtain the indicator calculation results from the data source.
[0040] In step S6, the data is checked for NULL values or abnormal values, and the indicator calculation logic is verified to be correct; if the check fails, detailed error information is returned.
[0041] An automated indicator creation system based on a large model, used to implement the above method, includes an interactive input module, a large model processing module, a data processing module and an interface generation module;
[0042] The interactive input module is responsible for receiving natural language input from users, supporting text / voice multimodal input and intelligent prompts based on the requirement example library;
[0043] Large model processing module, including large language model, parsing engine and SQL generator, is used to parse demand text and generate structured parameters and SQL statements;
[0044] The parsing engine includes a named entity recognition component, a semantic role labeling component and a logical reasoning engine;
[0045] The SQL generator includes an SQL template library and a grammar optimizer;
[0046] The data processing module is responsible for connecting to relational databases, non-relational databases, or data warehouses through the data source adaptation layer, executing SQL statements, and verifying calculation results;
[0047] The data source adaptation layer supports dynamic configuration of database connection parameters, including URL, username and password, and obtains data table structure and field information through the metadata management module;
[0048] The interface generation module has built-in testing tools that support simulated request sending and preview of response results. It is responsible for defining API parameters according to interface specifications, generating interface documents and publishing them.
[0049] The large language model of the large model processing module supports domain fine-tuning and adapts customized industry terminology and database table structure metadata through incremental training.
[0050] The data processing module includes a cache mechanism that can cache high-frequency access indicator results whose access frequency reaches a custom threshold and supports timed incremental updates.
[0051] An automated indicator creation device based on a large model, characterized in that it includes a memory and a processor; the memory is used to store a computer program, and the processor is used to implement the above-mentioned method steps when executing the computer program.
[0052] A readable storage medium, characterized in that: a computer program is stored on the readable storage medium, and the computer program implements the above method steps when executed by a processor.
[0053] The beneficial effects of the present invention are: the automated indicator creation method based on the large model greatly improves work efficiency, reduces labor costs, enables business personnel to independently complete indicator creation work, and at the same time avoids indicator calculation errors caused by human understanding deviations, ensures the accuracy of indicator calculation results, reduces interface adaptation and development workload, reduces system integration difficulty, and improves data sharing and interaction efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0055] Attachment Figure 1 Schematic diagram of the large-scale model-based automated indicator creation method of the present invention.
[0056] Attachment Figure 2 Schematic diagram of the large model processing method of the present invention. DETAILED DESCRIPTION
[0057] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work should fall within the scope of protection of the present invention.
[0058] The method for creating an automated indicator based on a large model includes the following steps:
[0059] Step S1: User input requirements
[0060] Users input natural language business requirement text through the interactive interface;
[0061] For example, enter "Count the customer traffic of each store in each quarter in 2024, and generate a JSON format API interface that can be called by the front end."
[0062] Step S2: Demand preprocessing
[0063] Preprocess the input text, including removing stop words (such as "statistics" and "generate" with no real meaning), unifying the time format (converting "each quarter of 2024" to a specific time range, such as "2024-01-01 to 2024-03-31" and "2024-04-01 to 2024-06-30"), and standardizing the representation (unifying "store" with the corresponding "store" field in the database) to prepare for subsequent analysis.
[0064] Step S3: Semantic parsing of large language model
[0065] Use a large language model to perform named entity recognition on the pre-processed business requirement text, extract key entity information, including time entities, indicator entities, grouping entities, and data entities, and generate structured requirement parameters through semantic role labeling and logical reasoning;
[0066] Step S4: SQL statement generation
[0067] According to the structured requirement parameters, a template is matched from an SQL template library and dynamically filled to generate a preliminary SQL statement, and an executable SQL statement is obtained after syntax verification and performance optimization;
[0068] Step S5: Data connection and calculation
[0069] Connect to the data source through database connection technology, execute the SQL statement, and obtain the indicator calculation results;
[0070] Step S6: Result verification
[0071] Verify the completeness and accuracy of indicator calculation results;
[0072] Step S7: Interface definition and configuration
[0073] Define API interface parameters according to the interface specification, including interface path (compliant with RESTful specifications), request method (GET / POST / PUT / DELETE), request parameters (path parameters, query parameters, request body parameters), response parameters and security mechanism;
[0074] Step S8: Generate interface document and publish
[0075] Generate an interface document containing interface functions, parameter descriptions, and calling examples, and publish the interface for external system calls.
[0076] The interface documentation is the user manual for the interface, supporting online debugging and document export. It details the interface's functions, interface paths, request parameters, response parameters, call examples, error code descriptions, and other information, making it easier for developers to call and integrate the interface. After the interface is published, front-end developers can use tools to call the interface and obtain the required data according to the interface documentation.
[0077] In step S3, the extracted time entity includes time range and time granularity information, such as extracting a specific quarterly time range from "each quarter of 2024";
[0078] When extracting indicator entities, first determine the indicator type to be calculated, such as "customer flow", and then associate the calculation logic corresponding to the indicator type through the built-in knowledge base, such as COUNT(customer_id);
[0079] The extracted grouping entity is the grouping dimension information, such as "each store", which corresponds to GROUP BY store_id in the SQL statement.
[0080] When extracting data entities, match relevant data tables and fields from the database metadata based on indicators and business scenarios. For example, "customer flow" data may be stored in the "customer_count" field of the "visit_table" table.
[0081] Through semantic role labeling (SRL) and logical reasoning, the operational logic (such as aggregation, filtering, grouping) and semantic relationships in the text are analyzed to understand the complete intention of user needs.
[0082] For example, for the requirement to "calculate the month-on-month growth rate," the inferred calculation logic is (current cycle indicator - previous cycle indicator) / previous cycle indicator; for "filtering data with sales greater than 1 million by region," the corresponding SQL filter condition WHEREregion = [specified region] AND sales_amount>1000000 is generated.
[0083] In step S4, the SQL template library includes single-table query, multi-table association and aggregate calculation type templates;
[0084] For example, for group statistics requirements, select the group statistics template, dynamically fill the identified group entity, indicator entity, and time entity information into the template based on the structured requirement parameter matching, and generate preliminary SQL statements;
[0085] Use the SQL syntax parser to check the syntax of the generated SQL statements and correct syntax errors;
[0086] At the same time, based on database performance optimization rules, SQL statements are optimized, such as adding appropriate index hints, optimizing JOIN order, etc., to improve query execution efficiency.
[0087] In step S5, a connection is established with a relational database (such as MySQL, Oracle), a non-relational database (such as MongoDB) or a data warehouse (such as Snowflake, Hive) through JDBC or ODBC database connection technology, and the generated SQL statement is executed to obtain the indicator calculation results from the data source.
[0088] In step S6, the data is checked for NULL values or abnormal values to verify whether the indicator calculation logic is correct; for example, for the "gross profit margin" indicator, it is checked whether it is within a reasonable range of 0 to 100%; if the check fails, detailed error information is returned, such as "data table does not exist", "field type does not match", etc.
[0089] The large-model-based automated indicator creation system, used to implement the above method, includes an interactive input module, a large-model processing module, a data processing module, and an interface generation module;
[0090] The interactive input module is responsible for receiving natural language input from users, supporting text / voice multimodal input and intelligent prompts based on the requirement example library;
[0091] Large model processing module, including large language model, parsing engine and SQL generator, is used to parse demand text and generate structured parameters and SQL statements;
[0092] The parsing engine includes a named entity recognition component, a semantic role labeling component and a logical reasoning engine;
[0093] The SQL generator includes an SQL template library and a grammar optimizer;
[0094] The data processing module is responsible for connecting to relational databases, non-relational databases, or data warehouses through the data source adaptation layer, executing SQL statements, and verifying calculation results;
[0095] The data source adaptation layer supports dynamic configuration of database connection parameters, including URL, username and password, and obtains data table structure and field information through the metadata management module;
[0096] The interface generation module has built-in testing tools that support simulated request sending and preview of response results. It is responsible for defining API parameters according to interface specifications, generating interface documents and publishing them.
[0097] The large language model of the large model processing module supports domain fine-tuning and adapts customized industry terminology (such as finance and e-commerce indicators) and database table structure metadata through incremental training.
[0098] The data processing module includes a cache mechanism that can cache high-frequency access indicator results whose access frequency reaches a custom threshold and supports timed incremental updates.
[0099] The large-model-based automated indicator creation device includes a memory and a processor; the memory is used to store a computer program, and the processor is used to implement the above-mentioned method steps when executing the computer program.
[0100] The readable storage medium stores a computer program, which implements the above method steps when executed by a processor.
[0101] Compared with existing technologies, this large-model-based automated indicator creation method has the following characteristics:
[0102] 1) Lowering the technical threshold and improving creation efficiency: The entire process of indicator creation and interface generation is automated, eliminating the need for manual SQL statement writing and API interface development. This reduces the hours or even days required for traditional indicator creation to just a few minutes. Users do not need to master SQL programming and interface development techniques. They can quickly generate the required indicators and corresponding interfaces simply by describing their business needs in natural language. This reduces dependence on professional data engineers and developers, reduces labor costs, enables business personnel to independently complete indicator creation, and significantly improves indicator creation efficiency.
[0103] 2) Achieved precise semantic analysis and intelligent generation: Leveraging the contextual understanding and logical reasoning capabilities of a large language model, it can accurately parse complex business needs and avoid indicator calculation errors caused by human misunderstanding. Whether it is a simple statistical requirement or complex calculation logic (such as month-on-month, year-on-year, percentage calculations, multi-condition screening, etc.), it can automatically generate SQL statements and standardized API interfaces that meet the requirements, ensuring the accuracy of indicator calculations and the standardization of interfaces.
[0104] 3) Improved system compatibility and integration: Generates standardized API interfaces that comply with interface specifications, unifies interface formats and specifications, supports seamless integration with various data analysis tools and business systems, reduces interface adaptation and development workload, reduces the difficulty of system docking, and improves data sharing and interaction efficiency.
[0105] The embodiment described above is only one specific implementation of the present invention. Common changes and substitutions made by those skilled in the art within the scope of the technical solution of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for creating automated indicators based on a large model, characterized by: The following steps are involved: Step S1: User input requirements Users input natural language business requirement text through the interactive interface; Step S2: Demand preprocessing Preprocess the input text, including removing stop words, unifying time formats, and standardizing expressions; Step S3: Semantic parsing of large language model Use a large language model to perform named entity recognition on the pre-processed business requirement text, extract key entity information, including time entities, indicator entities, grouping entities, and data entities, and generate structured requirement parameters through semantic role labeling and logical reasoning; Step S4: SQL statement generation According to the structured requirement parameters, a template is matched from an SQL template library and dynamically filled to generate a preliminary SQL statement, and an executable SQL statement is obtained after syntax verification and performance optimization; Step S5: Data connection and calculation Connect to the data source through database connection technology, execute the SQL statement, and obtain the indicator calculation results; Step S6: Result verification Verify the completeness and accuracy of indicator calculation results; Step S7: Interface definition and configuration Define API interface parameters according to interface specifications, including interface path, request method, request parameters, response parameters, and security mechanisms; Step S8: Generate interface document and publish Generate an interface document containing interface functions, parameter descriptions, and calling examples, and publish the interface for external system calls.
2. The method for creating automated indicators based on a large model according to claim 1, characterized in that: In step S3, the extracted time entity includes time range and time granularity information; When extracting indicator entities, first determine the indicator type to be calculated, and then associate the calculation logic corresponding to the indicator type through the built-in knowledge base; The extracted grouping entity is the grouping dimension information; When extracting data entities, match relevant data tables and fields from database metadata based on indicators and business scenarios; Through semantic role labeling (SRL) and logical reasoning, the operational logic and semantic relationships in the text are analyzed to understand the complete intention of user needs.
3. The method for creating automated indicators based on a large model according to claim 1, characterized in that: In step S4, the SQL template library includes single-table query, multi-table association and aggregate calculation type templates; Based on the matching of structured demand parameters, the identified grouping entity, indicator entity, and time entity information are dynamically filled into the template to generate preliminary SQL statements; Use the SQL syntax parser to check the syntax of the generated SQL statements and correct syntax errors; At the same time, based on database performance optimization rules, SQL statements are optimized to improve query execution efficiency.
4. The method for creating automated indicators based on a large model according to claim 1, characterized in that: In step S5, a connection is established with a relational database, a non-relational database or a data warehouse through JDBC or ODBC database connection technology, and the generated SQL statement is executed to obtain the indicator calculation results from the data source.
5. The method for creating automated indicators based on a large model according to claim 1, characterized in that: In step S6, the data is checked for NULL values or abnormal values, and the indicator calculation logic is verified to be correct; if the check fails, detailed error information is returned.
6. An automated indicator creation system based on a large model, characterized by: Used to implement the method according to any one of claims 1 to 5, comprising an interactive input module, a large model processing module, a data processing module and an interface generation module; The interactive input module is responsible for receiving natural language input from users, supporting text / voice multimodal input and intelligent prompts based on the requirement example library; Large model processing module, including large language model, parsing engine and SQL generator, is used to parse demand text and generate structured parameters and SQL statements; The parsing engine includes a named entity recognition component, a semantic role labeling component and a logical reasoning engine; The SQL generator includes an SQL template library and a grammar optimizer; The data processing module is responsible for connecting to relational databases, non-relational databases, or data warehouses through the data source adaptation layer, executing SQL statements, and verifying calculation results; The data source adaptation layer supports dynamic configuration of database connection parameters, including URL, username and password, and obtains data table structure and field information through the metadata management module; The interface generation module has built-in testing tools that support simulated request sending and preview of response results. It is responsible for defining API parameters according to interface specifications, generating interface documents and publishing them.
7. The large model-based automated indicator creation system according to claim 6, characterized in that: The large language model of the large model processing module supports domain fine-tuning and adapts customized industry terminology and database table structure metadata through incremental training.
8. The large model-based automated indicator creation system according to claim 6, characterized in that: The data processing module includes a cache mechanism that can cache high-frequency access indicator results whose access frequency reaches a custom threshold and supports timed incremental updates.
9. An automated indicator creation device based on a large model, characterized by: The method comprises a memory and a processor; the memory is used to store a computer program, and the processor is used to implement the method steps according to any one of claims 1 to 5 when executing the computer program.
10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, which, when executed by a processor, implements the method steps according to any one of claims 1 to 5.
Citation Information
Patent Citations
Intelligent report generation method and system based on large model
CN117251455A
Resume screening method and device based on large model prompt instruction and medium
CN119807229A
Systems and Methods for Deploying a Machine-Learning Model for Performing a Specific Clinical Task
US20240404702A1
Internet of things system
WO2023030513A1
Cited By
Distributed multi-cloud-node remote sensing data synchronous transmission method, system and equipment
CN120980095A
SQL (Structured Query Language) statement structure verification system based on large model and knowledge graph fusion enhancement
CN121326958A
Query statement generation method, electronic equipment and storage medium
CN121387937A
Village folk-custom activity question and answer method and device combining large model and SQL (Structured Query Language) query
CN122309537A