A large model database self-service query data analysis system and method based on external association constraint configuration

CN122838431APending Publication Date: 2026-09-29谢云哲
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611089070.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-22
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]本发明针对现有技术依赖数据库原生外键、大模型自主推断表关联易产生查询幻觉、缺少强制约束管控机制、大批量明细数据超出模型上下文限制、成熟查询口径无法留存复用、业务数据安全管控能力不足的缺陷,提供一种基于外置关联约束配置的大模型数据库自助查询数据分析系统及方法;现有技术大多仅将外部关联规则作为参考信息,缺少前置提示词引导与后置 SQL 刚性拦截的一体化强制管控架构,本发明所述系统的核心架构包括数据存储库、大语言模型交互调用模块、外置约束存储模块、约束提示词生成模块以及查询语句校验执行模块,数据存储库内部存储若干业务原始数据表,各个业务原始数据表之间不在数据库底层设置外键关联约束,外置约束存储模块配置独立于全部业务原始数据表之外的持久化约束存储介质,约束存储介质为独立数据表,不嵌入、不依附于业务原始数据表结构,实现关联规则与业务数据分开维护,用于存储用户设定的表配对关系、字段映射关系、表间连接逻辑、筛选限定条件等关联约束规则,外置约束存储模块作为大语言模型生成 SQL 语句时唯一可引用的表间关联关系依据,约束提示词生成模块获取用户自然语言查询指令、业务数据表元数据以及外置约束存储模块内已生效的全部关联约束规则,生成携带强制约束指令的提示文本并发送至大语言模型交互调用模块,系统构建双层管控机制,第一层通过提示词限定大语言模型仅能够依据已录入约束规则生成多表连接逻辑,禁止自主推断、新增或修改表间关联关系,第二层依靠查询语句校验执行模块对大模型输出的 SQL 语句进行语法解析、SQL 注入特征拦截以及关联规则白名单比对,自动提取 SQL内部全部数据表、表间 JOIN 关系与关联字段,并逐条与外置约束规则进行匹配,提示词约束用于前置引导大模型生成逻辑,后置白名单比对作为最终刚性防线,二者配合保障不会执行约束范围以外的表关联查询,仅当所有表间连接逻辑均存在于已存储约束规则内才允许执行查询,违规关联语句直接拦截;作为优选拓展方案,系统还增设数据分级处理分析模块与查询模板固化模块,数据分级处理分析模块预设数据体量判定阈值,对查询结果数据集实施分级处理,数据体量不大于阈值则直接将明细数据集送入大模型解析,数据体量超出阈值时自动按照业务维度聚合精简数据后再开展分析,聚合过程优先采用约束内预设统计维度,无预设维度则选取查询自带分组字段完成汇总,查询模板固化模块接收人工确认指令,将核验无误的查询语句绑定对应关联约束标识存入模板库实现复用;外置约束存储介质采用独立数据表存储各类约束信息、权限标识与操作日志,关联约束规则划分为全局公共约束和个人私有约束并配置差异化操作权限,系统支持后台自定义数据体量阈值,可接入本地文档解析模块读取 Excel、Word 结构化文档作为补充元数据参与查询,默认在客户端本地存储配置信息、约束规则与查询结果,未经用户主动授权不向外网传输原始业务数据;本发明对应的实现方法包含基础执行步骤:构建底层无外键约束的业务数据表存储库并搭建独立约束存储载体,通过可视化表单点选录入或自然语言解析录入方式采集多表关联约束信息,自然语言录入模式下由大模型提取结构化约束信息并经过数据表字段合法性校验后持久化存储,接收用户自然语言查询需求后调取数据表元数据与生效约束生成强制约束提示指令,大模型仅能依托已录入关联规则生成 SQL 语句,语句经过三重校验合规后执行查询得到结果数据集;作为可选拓展步骤,系统依据预设阈值对数据集自动分流处理,超量数据聚合精简后交由大模型开展分析,大模型仅输出统计说明、异常标注、风险提示与业务优化建议,不生成经营决策类主观判定内容,同时将人工验证通过的查询语句绑定约束标识归档形成可复用查询模板;本发明脱离数据库原生外键依赖,依托独立持久化外置约束搭配提示词约束与 SQL 白名单拦截双层机制,从根源消除大模型表关联幻觉,适配分库分表生产数据库场景,大数据集自动聚合解决模型上下文超限问题,实现业务查询口径长期沉淀复用,依靠公私约束权限划分与本地数据存储机制兼顾多用户协同与数据安全,同时支持多种约束录入方式并附带合法性校验,有效降低业务人员配置多表查询规则的操作门槛

Benefits of technology

[0003]本发明针对现有技术依赖数据库原生外键、大模型自主推断表关联易产生查询幻觉、缺少强制约束管控机制、大批量明细数据超出模型上下文限制、成熟查询口径无法留存复用、业务数据安全管控能力不足的缺陷,提供一种基于外置关联约束配置的大模型数据库自助查询数据分析系统及方法;现有技术大多仅将外部关联规则作为参考信息,缺少前置提示词引导与后置 SQL 刚性拦截的一体化强制管控架构,本发明所述系统的核心架构包括数据存储库、大语言模型交互调用模块、外置约束存储模块、约束提示词生成模块以及查询语句校验执行模块,数据存储库内部存储若干业务原始数据表,各个业务原始数据表之间不在数据库底层设置外键关联约束,外置约束存储模块配置独立于全部业务原始数据表之外的持久化约束存储介质,约束存储介质为独立数据表,不嵌入、不依附于业务原始数据表结构,实现关联规则与业务数据分开维护,用于存储用户设定的表配对关系、字段映射关系、表间连接逻辑、筛选限定条件等关联约束规则,外置约束存储模块作为大语言模型生成 SQL 语句时唯一可引用的表间关联关系依据,约束提示词生成模块获取用户自然语言查询指令、业务数据表元数据以及外置约束存储模块内已生效的全部关联约束规则,生成携带强制约束指令的提示文本并发送至大语言模型交互调用模块,系统构建双层管控机制,第一层通过提示词限定大语言模型仅能够依据已录入约束规则生成多表连接逻辑,禁止自主推断、新增或修改表间关联关系,第二层依靠查询语句校验执行模块对大模型输出的 SQL 语句进行语法解析、SQL 注入特征拦截以及关联规则白名单比对,自动提取 SQL内部全部数据表、表间 JOIN 关系与关联字段,并逐条与外置约束规则进行匹配,提示词约束用于前置引导大模型生成逻辑,后置白名单比对作为最终刚性防线,二者配合保障不会执行约束范围以外的表关联查询,仅当所有表间连接逻辑均存在于已存储约束规则内才允许执行查询,违规关联语句直接拦截;作为优选拓展方案,系统还增设数据分级处理分析模块与查询模板固化模块,数据分级处理分析模块预设数据体量判定阈值,对查询结果数据集实施分级处理,数据体量不大于阈值则直接将明细数据集送入大模型解析,数据体量超出阈值时自动按照业务维度聚合精简数据后再开展分析,聚合过程优先采用约束内预设统计维度,无预设维度则选取查询自带分组字段完成汇总,查询模板固化模块接收人工确认指令,将核验无误的查询语句绑定对应关联约束标识存入模板库实现复用;外置约束存储介质采用独立数据表存储各类约束信息、权限标识与操作日志,关联约束规则划分为全局公共约束和个人私有约束并配置差异化操作权限,系统支持后台自定义数据体量阈值,可接入本地文档解析模块读取 Excel、Word 结构化文档作为补充元数据参与查询,默认在客户端本地存储配置信息、约束规则与查询结果,未经用户主动授权不向外网传输原始业务数据;本发明对应的实现方法包含基础执行步骤:构建底层无外键约束的业务数据表存储库并搭建独立约束存储载体,通过可视化表单点选录入或自然语言解析录入方式采集多表关联约束信息,自然语言录入模式下由大模型提取结构化约束信息并经过数据表字段合法性校验后持久化存储,接收用户自然语言查询需求后调取数据表元数据与生效约束生成强制约束提示指令,大模型仅能依托已录入关联规则生成 SQL 语句,语句经过三重校验合规后执行查询得到结果数据集;作为可选拓展步骤,系统依据预设阈值对数据集自动分流处理,超量数据聚合精简后交由大模型开展分析,大模型仅输出统计说明、异常标注、风险提示与业务优化建议,不生成经营决策类主观判定内容,同时将人工验证通过的查询语句绑定约束标识归档形成可复用查询模板;本发明脱离数据库原生外键依赖,依托独立持久化外置约束搭配提示词约束与 SQL 白名单拦截双层机制,从根源消除大模型表关联幻觉,适配分库分表生产数据库场景,大数据集自动聚合解决模型上下文超限问题,实现业务查询口径长期沉淀复用,依靠公私约束权限划分与本地数据存储机制兼顾多用户协同与数据安全,同时支持多种约束录入方式并附带合法性校验,有效降低业务人员配置多表查询规则的操作门槛。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
Patent Text Reader

Abstract

The application discloses a large model database self-service query data analysis system and method based on external association constraint configuration, and belongs to the technical field of large language model data query. The system comprises a data storage library, a large language model interactive calling module, an external constraint storage module, a constraint prompt word generation module and a query statement checking and executing module. A business original data table is not provided with a bottom-layer foreign key. The external constraint storage module adopts an independent persistent data table to store inter-table association constraint rules and serves as the only reference for inter-table association of a large model generated SQL statement. A double-layer management and control mechanism is constructed. The constraint prompt word is preposed to limit the large model generation logic. The query statement checking and executing module performs syntax, injection interception and association rule white list comparison on the SQL. The inter-table JOIN relationship is automatically extracted and matched with the external constraint. The illegal association statement is directly intercepted. The application is independent of the database original foreign key and relies on the external independent constraint to match the double-layer forced management and control mechanism. The large model table association inference illusion is eliminated from the root. The application is suitable for the production database scene with disabled foreign keys. The optimized scheme can realize large data automatic aggregation, query template solidification and constraint permission hierarchical management, and takes into account the self-service data analysis accuracy, business caliber reuse and data security.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of self-service query and data analysis technology of large language model database, specifically involving a natural language to SQL query and intelligent data analysis system and method that relies on external association constraints to achieve mandatory control. Background Technology

[0002] Existing self-service data retrieval solutions based on large models generally have significant drawbacks. One type of solution relies on the database's underlying native foreign keys to implement table relationships, making it difficult to adapt to production environments where foreign keys are disabled, databases are sharded, and large-scale data modifications are common. Another type relies on temporary drag-and-drop generation of relationships on the front end, and the rules cannot be persistently saved, becoming invalid after a page refresh. Yet another type of solution directly infers table relationships based on the semantics of the large model's fields, which is prone to creating illusions of master-detail table reversals and mismatched related fields, resulting in query results that do not conform to business requirements. Although some technologies introduce external business metadata to assist in generating SQL from large models, these solutions only use external relationship rules as reference information for the large model and lack a systematic and mandatory control mechanism. They cannot prevent large models from independently adding or tampering with table relationships. In addition, they generally suffer from a series of problems that urgently need to be addressed, such as massive amounts of detailed data exceeding the length of the large model's context, difficulty in preserving and reusing verified query criteria, and the risk of leakage of original business data. Summary of the Invention

[0003] This invention addresses the shortcomings of existing technologies, such as reliance on native database foreign keys, the susceptibility to query illusions due to large-scale model self-inference of table relationships, lack of mandatory constraint control mechanisms, large volumes of detailed data exceeding model context limitations, inability to retain and reuse mature query methods, and insufficient business data security control capabilities. It provides a self-service query data analysis system and method for large-scale model databases based on external association constraint configuration. Existing technologies mostly treat external association rules only as reference information, lacking pre-emptive prompts and post-SQL guidance. The rigid interception integrated mandatory control architecture of the system described in this invention includes a core architecture comprising a data repository, a large language model interaction module, an external constraint storage module, a constraint prompt word generation module, and a query statement verification and execution module. The data repository internally stores several original business data tables, and no foreign key relationships are set between these tables at the database level. The external constraint storage module is configured with a persistent constraint storage medium independent of all original business data tables. This constraint storage medium is an independent data table, not embedded in or dependent on the original business data table structure, enabling separate maintenance of association rules and business data. It stores user-defined table pairing relationships, field mapping relationships, inter-table join logic, filtering constraints, and other association constraint rules. The external constraint storage module serves as the SQL generator for the large language model. The statement is the only referenceable basis for the inter-table relationship. The constraint prompt word generation module obtains the user's natural language query command, business data table metadata, and all effective association constraint rules in the external constraint storage module. It generates a prompt text carrying the mandatory constraint command and sends it to the large language model interaction call module. The system builds a two-layer control mechanism. The first layer restricts the large language model to only generate multi-table join logic based on the entered constraint rules through prompt words, prohibiting independent inference, addition, or modification of inter-table relationships. The second layer relies on the query statement verification and execution module to perform syntax parsing, SQL injection feature interception, and association rule whitelist comparison on the SQL statement output by the large model. It automatically extracts all data tables, inter-table JOIN relationships, and related fields inside the SQL, and matches them one by one with the external constraint rules. The prompt word constraint is used to guide the large model generation logic in the front end, and the whitelist comparison in the back end serves as the final rigid defense. The two work together to ensure that table join queries outside the constraint scope will not be executed. Query execution is only allowed when all inter-table join logic exists within the stored constraint rules. Illegal join statements are directly intercepted.As a preferred expansion solution, the system also adds a data hierarchical processing and analysis module and a query template solidification module. The data hierarchical processing and analysis module presets a data volume judgment threshold and performs hierarchical processing on the query result dataset. If the data volume is less than the threshold, the detailed dataset is directly sent to the large model for parsing. If the data volume exceeds the threshold, the data is automatically aggregated and simplified according to business dimensions before analysis. The aggregation process prioritizes the preset statistical dimensions within the constraints. If there are no preset dimensions, the query's built-in grouping fields are selected to complete the summary. The query template solidification module receives manual confirmation instructions and binds the verified query statements to the corresponding related constraint identifiers and stores them in the template library for reuse. The external constraint storage medium uses an independent data table to store various constraint information, permission identifiers, and operation logs. Related constraint rules are divided into global public constraints and personal private constraints, and differentiated operation permissions are configured. The system supports custom data volume thresholds in the backend and can connect to the local document parsing module to read Excel and Word documents. Structured documents serve as supplementary metadata in queries. By default, configuration information, constraint rules, and query results are stored locally on the client side. Raw business data is not transmitted to the external network without user authorization. The implementation method of this invention includes the following basic execution steps: constructing a business data table repository without foreign key constraints and building an independent constraint storage carrier; collecting multi-table association constraint information through visual form point-and-click input or natural language parsing input; in natural language input mode, the large model extracts structured constraint information and persists it after data table field validity verification; upon receiving a user's natural language query request, it retrieves data table metadata and effective constraints to generate a mandatory constraint prompt instruction; the large model can only generate SQL based on the entered association rules. The query statement, after undergoing triple validation, is executed to obtain the result dataset. As an optional extension step, the system automatically distributes the dataset based on preset thresholds. Excess data is aggregated and streamlined before being sent to a large model for analysis. The large model only outputs statistical descriptions, anomaly annotations, risk warnings, and business optimization suggestions, without generating subjective judgments related to business decisions. Simultaneously, manually validated query statements are bound to constraint identifiers and archived to form reusable query templates. This invention eliminates the dependency on native database foreign keys, relying on a two-layer mechanism of independent persistent external constraints combined with prompt word constraints and SQL whitelist interception to fundamentally eliminate the illusion of large model table relationships. It adapts to sharded production database scenarios, automatically aggregates large datasets to resolve model context exceeding limits, and enables long-term accumulation and reuse of business query definitions. It balances multi-user collaboration and data security through public / private constraint permission division and local data storage mechanisms. It also supports multiple constraint entry methods with legality checks, effectively reducing the operational threshold for business personnel configuring multi-table query rules. Detailed Implementation

[0004] This embodiment uses MySQL to build the business data table database, independent external constraint data tables, and template repository. All original business data tables do not have underlying database foreign key constraints configured; constraint information is persistently stored in independent data tables. This differs from temporary association configurations that are temporarily stored in memory and destroyed after the query ends. Users can complete constraint configuration by selecting master and slave data tables, join types, and related fields through a front-end visual form, or by inputting natural language descriptions of table relationships. The large model parses the text to extract structured constraint information. The system automatically verifies the authenticity and validity of the data tables and related fields; only after successful verification can the information be written to the external constraint table. After a user submits a natural language query request, the system loads all effective association constraints and data table metadata to generate strong constraint warning words, forcibly restricting the large model to only generate table join logic based on registered constraints. After the large model outputs the SQL statement, the query statement verification and execution module performs syntax verification, injection attack identification, and association rule whitelist comparison. It parses and extracts all JOIN relationships and matches them against external constraints one by one. SQL statements containing unregistered association logic are directly intercepted and not executed. After a successful query returns the dataset, the system uses the default 1000... The system uses a threshold to determine the data volume. If the data volume exceeds the limit, it automatically performs aggregation and simplification before sending it to a large model for data analysis. The model only outputs objective statistical results, anomaly markers, and implementation optimization suggestions, without outputting definitive business decisions. After the user confirms that the query logic is correct, the query statement can be bound to a constraint number and stored in the template library for subsequent one-click invocation. The platform distinguishes between global public constraints that can be edited by the administrator and personal private constraints that can only be used by the creator. It also supports uploading local Excel and Word documents as temporary data tables to participate in the query process. By default, the entire system saves configuration information, query details, and original business data locally on the client. Without the user's active authorization, it will not transmit confidential original business data to external servers. Attached Figure Description

[0005] Figure 1 This is a schematic diagram of the architecture of a self-service query and data analysis system for a large model database based on external association constraint configuration, according to the present invention. Figure 2 This is a schematic diagram of the process of a self-service query and data analysis method for a large model database based on external association constraint configuration according to the present invention.

Claims

1. A self-service query and data analysis system for a large model database based on external association constraint configuration, comprising a data repository and a large language model interaction module; wherein the data repository stores several business raw data tables, and no foreign key association constraints are set between the business raw data tables at the underlying database level; characterized in that, Also includes: An external constraint storage module is configured with a persistent constraint storage medium independent of all original business data tables. This constraint storage medium is an independent data table, not embedded in or dependent on the original business data table structure. It stores user-defined multi-table association constraint rules, which at least include data table pairing relationships, field mapping relationships, inter-table join logic, and filtering conditions. The external constraint storage module serves as the sole reference for inter-table association relationships when the large language model generates SQL statements. A constraint prompt word generation module obtains the user's natural language query command, metadata of the original business data tables, and all effective association constraint rules within the external constraint storage module. It then generates prompt text carrying mandatory constraint commands and sends it to the large language model interaction call module. The query statement verification and execution module receives the database query statement output by the large language model, performs syntax parsing, SQL injection feature interception, and whitelist comparison of the SQL statement, automatically extracts all table JOIN relationships contained in the SQL and matches them one by one with external constraint rules; the system constructs a two-layer control mechanism. The first layer restricts the large language model to generate query statements only based on stored association constraint rules through prompts, prohibiting the independent inference, addition, or modification of table association relationships; the query can only be executed when all table join relationships match the registered constraint rules, and illegal association statements are directly intercepted.

2. A self-service query data analysis method for large model databases based on external association constraint configuration, characterized in that, The process includes the following steps:

1. Construct an underlying data repository and store the original business data tables. Do not configure foreign key binding constraints on any original business data tables at the underlying database level. Create an independent and persistent constraint storage carrier.

2. Obtain the multi-table association constraint information entered by the user, verify its validity, and persistently write it to the constraint storage carrier.

3. Receive the user's natural language query request, retrieve the metadata of the business data tables and the effective association constraint rules, and generate a prompt instruction with mandatory constraint effect.

4. Input the prompt instruction into the large language model to generate a database query statement. The large language model is only allowed to call the table association logic already entered in the constraint storage carrier. It is not allowed to add, modify, or infer the data table connection relationship on its own. The query statement undergoes triple verification: syntax parsing, SQL injection feature interception, and whitelist comparison of association rules. Once the verification is compliant, the query is executed to obtain the target result dataset.

3. The system according to claim 1, characterized in that, The data tables in the constraint storage medium must store at least the following fields: master data table identifier, associated slave data table identifier, connection type, master table associated fields, slave table associated fields, additional filter conditions, permission identifier, creation entity identifier, and operation modification log; the associated constraint rules in the external constraint storage module are divided into global public constraints and personal private constraints; global public constraints are only authorized accounts with editing and deletion permissions, and can be accessed by users across the platform; personal private constraints are only readable and usable by the creator account.

4. The system according to claim 1, characterized in that, It also includes a data hierarchical processing and analysis module; the data hierarchical processing and analysis module presets a data volume judgment threshold, which supports custom configuration in the background, with a default threshold of 1000 data records; it performs hierarchical processing on the result dataset: if the data volume is not greater than the threshold, the detailed dataset is directly sent to the large language model for parsing; if the data volume exceeds the threshold, the dataset is automatically aggregated and simplified to generate summary data before being sent to the large language model for data analysis; the aggregation prioritizes the preset statistical dimensions within the constraints, and if there are no preset dimensions, it selects the grouping fields provided by the query for aggregation.

5. The system according to claim 1, characterized in that, It also includes a query template solidification module, which receives manual confirmation instructions and stores the verified query statement and corresponding associated constraint identifiers into the template library, supporting direct reuse in the future.

6. The system according to claim 1, characterized in that, The system supports access to a local document parsing module to read data from Excel and Word structured documents as supplementary data table metadata and participate in the query generation process; By default, the system stores user configuration data, constraint rules, and query results locally on the client side, and does not transmit raw business data to the external network without the user's active authorization.

7. The method according to claim 2, characterized in that, The constraint input methods in step 2 include visual form point selection input and natural language parsing input. When using natural language input, the large language model parses the text to extract structured constraint information and verifies the validity of the data table fields. Only after the verification is passed can the data be persistently stored.

8. The method according to claim 2, characterized in that, It also includes step 5: the result dataset is split and processed according to a preset size threshold, and the excess dataset is aggregated and simplified before being handed over to the large model for analysis; the large language model only outputs statistical descriptions, anomaly annotations, risk warnings and business optimization suggestions, and does not generate subjective judgment content for business decision-making.

9. The method according to claim 2, characterized in that, It also includes step 6: binding the manually verified query statements with association constraint identifiers and archiving them as reusable query templates.

10. The method according to claim 2, characterized in that, It supports integration with local document parsing modules to read data from Excel and Word structured documents as supplementary data table metadata to participate in the query generation process.