Dify-based natural language data query system and method

By using a natural language data query system based on Dify, which analyzes user intent using a large language model, automatically parses parameters, and collaboratively calls multiple interfaces, the system solves the problems of high technical barriers and complexity in traditional database queries. It enables non-technical personnel to perform complex data queries and cross-data source analysis, thereby improving query efficiency and user experience.

CN120929476APending Publication Date: 2025-11-11浪潮智慧城市科技有限公司
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511014431.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Traditional database query methods have high technical barriers, are complex to operate, have difficult interface management, require frequent parameter configuration, and are difficult to coordinate with multiple interfaces. Existing natural language query systems lack support for complex multi-interface environments and do not fully utilize the semantic understanding capabilities of large language models.

Method used

Design a natural language data query system based on Dify, including a natural language understanding module, an interface metadata management module, an intelligent matching module, a parameter parsing module, and an interface calling module. The system analyzes user intent through a large language model, intelligently matches interfaces, automatically parses parameters, collaboratively calls multiple interfaces, and integrates result processing.

Benefits of technology

It lowers the technical barrier to data querying, enabling non-technical personnel to perform complex data queries, improving query efficiency and user experience, supporting comprehensive analysis across data sources, and simplifying the operation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929476A_ABST
    Figure CN120929476A_ABST
Patent Text Reader

Abstract

The invention particularly relates to a Dify-based natural language data query system and method. The Dify-based natural language data query system comprises a natural language understanding module, an interface metadata management module, an intelligent matching module, a parameter analysis module, an interface calling module and a result processing module. According to the Dify-based natural language data query system and method, a user can perform complex data query through a natural language, the use threshold is greatly reduced, the manual configuration time is shortened, the query efficiency is improved, meanwhile, the data requirement of a complex service scene can be met, cross-data-source comprehensive analysis is supported, and the user experience is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data query technology, and in particular to a natural language data query system and method based on Dify. Background Technology

[0002] In a data-driven business environment, data querying is one of the core business operations. Traditional database query methods mainly rely on SQL statements or API calls, which have the following problems:

[0003] 1) The technical threshold is high. Users need to master SQL syntax or understand API parameter specifications. Non-technical personnel will find it difficult to directly query data.

[0004] 2) The operation is complex. Complex queries require writing complex statements such as multi-table joins and conditional filtering, which are prone to errors and difficult to debug.

[0005] 3. Interface management is difficult. Enterprises usually have a large number of data interfaces, and users find it difficult to quickly find the right interface and call it correctly.

[0006] 4) Frequent parameter configuration: API calls require manual configuration of various parameters, with strict requirements on parameter format and type, which can easily lead to query failures due to configuration errors.

[0007] 5) Difficulty in multi-interface collaboration: Complex business scenarios require calling multiple interfaces and integrating the results, lacking a unified coordination mechanism.

[0008] While there are some existing natural language query systems, most are limited to single data sources or simple query scenarios, lack support for complex multi-interface environments, and do not fully utilize the semantic understanding capabilities of large language models.

[0009] To address the problems existing in the prior art, this invention proposes a natural language data query system and method based on Dify, aiming to lower the technical threshold for data query and enable non-technical personnel to perform complex data queries through natural language. Summary of the Invention

[0010] To overcome the shortcomings of existing technologies, this invention provides a simple and efficient natural language data query method based on Dify.

[0011] This invention is achieved through the following technical solution:

[0012] A natural language data query system based on Dify, characterized in that it includes a natural language understanding module, an interface metadata management module, an intelligent matching module, a parameter parsing module, an interface calling module, and a result processing module;

[0013] The Natural Language Understanding module integrates a large language model to analyze user-input natural language queries and accurately understand the user's query intent.

[0014] The interface metadata management module is used to maintain and manage the metadata information of data interfaces;

[0015] The intelligent matching module is used to match user questions with the knowledge base interface;

[0016] The parameter parsing module is used to extract and generate calling parameters that meet the interface requirements from natural language.

[0017] The API call module is used to perform collaborative calls to multiple APIs and determine the call order based on API dependencies.

[0018] The results processing module is used to integrate and format query results through data deduplication, format conversion, and exception handling to ensure the quality and consistency of the output results.

[0019] The natural language understanding module includes a large language model analysis unit, a query element extraction unit, and a standardization processing unit.

[0020] Large language model analysis unit for semantic understanding and intent recognition;

[0021] The query element extraction unit is used to extract key information, including query conditions, time range, and data type.

[0022] The standardization processing unit is used to generate standardized query requirement descriptions.

[0023] The interface metadata management module includes a metadata storage unit, a semantic indexing unit, and a version management unit;

[0024] The metadata storage unit is used to store metadata information, including interface descriptions, parameter specifications, and data formats.

[0025] The semantic indexing unit supports vector retrieval and keyword retrieval, and is used to build a semantic index of the interface description to quickly locate relevant interface metadata from the knowledge base;

[0026] The version management unit is used to manage version updates of interface metadata.

[0027] The intelligent matching module includes a semantic similarity calculation unit, a multi-interface selection unit, and a dependency analysis unit;

[0028] The semantic similarity calculation unit is used to calculate the matching degree between the user question and the interface description.

[0029] The multi-interface selection unit is used to select the interface combination with the highest matching degree;

[0030] The dependency analysis unit is used to analyze the calling order and dependencies between interfaces.

[0031] The parameter parsing module supports conditional queries, range queries, and fuzzy queries, and includes a parameter extraction unit, a format conversion unit, and a parameter verification unit.

[0032] The parameter extraction unit is used to extract parameter information from natural language, including numerical values, dates, and strings;

[0033] The format conversion unit is used to convert the parameter format and type according to the interface requirements;

[0034] The parameter verification unit is used to verify the validity and completeness of the parameters.

[0035] The interface calling module supports both serial and parallel calling modes, including:

[0036] The scheduling unit is invoked to determine the order and mode of interface calls;

[0037] Parallel processing unit, used to support parallel calls to multiple interfaces;

[0038] The exception handling unit is used to handle call exceptions and retry mechanisms.

[0039] The result processing module includes a data integration unit, a formatting unit, and a cache management unit;

[0040] The data integration unit is used to integrate the return results from multiple interfaces;

[0041] The formatting unit is used to format the output data;

[0042] The cache management unit is used to manage the query result cache.

[0043] A natural language data query method based on Dify includes the following steps:

[0044] Step S1: Natural Language Understanding

[0045] It receives natural language queries from users, analyzes the semantics of the queries using a large language model, and identifies key information, including query intent, target data type, and time range.

[0046] Extract structured information, including query conditions, sorting requirements, and aggregation methods, to generate a standardized query requirement description; Step S2: Knowledge base interface metadata matching.

[0047] Maintain a metadata knowledge base containing interface descriptions, parameter specifications, and data format information, and calculate the matching degree between user questions and interface descriptions based on semantic similarity;

[0048] It supports parallel matching of multiple interfaces and takes into account the dependencies and calling order between interfaces to select the combination of interfaces with the highest similarity.

[0049] Step S3: Intelligent Parameter Parsing and Generation

[0050] Leveraging the capabilities of large models to extract parameter information from natural language, including numerical values, dates, and strings;

[0051] Automatically convert parameter formats and types according to interface parameter specifications, handle logical relationships and constraints between parameters, and generate call parameters that meet interface requirements;

[0052] Step S4: Multi-interface collaborative call

[0053] The system determines the order and dependencies of interface calls based on the matching results, supports both serial and parallel calling modes, enables data transfer and parameter sharing between interfaces, and handles interface call exceptions and retry mechanisms.

[0054] Step S5: Result Integration and Output

[0055] Collect the return results from multiple interfaces, integrate and deduplicate the return result data according to business logic, and use a large model to organize and output user-friendly results.

[0056] A natural language data query device based on Dify, characterized in that it includes a memory and a processor; the memory is used to store a computer program, and the processor is used to execute the computer program to implement the above-described method steps.

[0057] A readable storage medium, characterized in that: a computer program is stored on the readable storage medium, and the computer program, when executed by a processor, implements the above-described method steps.

[0058] The beneficial effects of this invention are: the Dify-based natural language data query system and method allow users to perform complex data queries through natural language, greatly reducing the barrier to entry, reducing manual configuration time, improving query efficiency, meeting the data needs of complex business scenarios, supporting comprehensive analysis across data sources, and greatly enhancing the user experience. Attached Figure Description

[0059] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0060] Appendix Figure 1 This is a schematic diagram of the natural language data query method based on Dify according to the present invention. Detailed Implementation

[0061] To enable those skilled in the art to better understand the technical solutions of this invention, the technical solutions in the embodiments of this invention will be clearly and completely described below in conjunction with the embodiments of this invention. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this invention.

[0062] The Dify-based natural language data query system includes a natural language understanding module, an interface metadata management module, an intelligent matching module, a parameter parsing module, an interface call module, and a result processing module.

[0063] The Natural Language Understanding module integrates a large language model to analyze user-input natural language queries and accurately understand the user's query intent.

[0064] The interface metadata management module is used to maintain and manage the metadata information of data interfaces;

[0065] The intelligent matching module is used to match user questions with the knowledge base interface;

[0066] The parameter parsing module is used to extract and generate calling parameters that meet the interface requirements from natural language.

[0067] The API call module is used to perform collaborative calls to multiple APIs and determine the call order based on API dependencies.

[0068] The results processing module is used to integrate and format query results through data deduplication, format conversion, and exception handling to ensure the quality and consistency of the output results.

[0069] The natural language understanding module includes a large language model analysis unit, a query element extraction unit, and a standardization processing unit.

[0070] Large language model analysis unit for semantic understanding and intent recognition;

[0071] The query element extraction unit is used to extract key information, including query conditions, time range, and data type.

[0072] The standardization processing unit is used to generate standardized query requirement descriptions.

[0073] The interface metadata management module includes a metadata storage unit, a semantic indexing unit, and a version management unit;

[0074] The metadata storage unit is used to store metadata information, including interface descriptions, parameter specifications, and data formats.

[0075] The semantic indexing unit supports vector retrieval and keyword retrieval, and is used to build a semantic index of the interface description to quickly locate relevant interface metadata from the knowledge base;

[0076] The version management unit is used to manage version updates of interface metadata.

[0077] The intelligent matching module includes a semantic similarity calculation unit, a multi-interface selection unit, and a dependency analysis unit;

[0078] The semantic similarity calculation unit is used to calculate the matching degree between the user question and the interface description.

[0079] The multi-interface selection unit is used to select the interface combination with the highest matching degree;

[0080] The dependency analysis unit is used to analyze the calling order and dependencies between interfaces.

[0081] The parameter parsing module supports conditional queries, range queries, and fuzzy queries, and includes a parameter extraction unit, a format conversion unit, and a parameter verification unit.

[0082] The parameter extraction unit is used to extract parameter information from natural language, including numerical values, dates, and strings;

[0083] The format conversion unit is used to convert the parameter format and type according to the interface requirements;

[0084] The parameter verification unit is used to verify the validity and completeness of the parameters.

[0085] The interface calling module supports both serial and parallel calling modes, including:

[0086] The scheduling unit is invoked to determine the order and mode of interface calls;

[0087] Parallel processing unit, used to support parallel calls to multiple interfaces;

[0088] The exception handling unit is used to handle call exceptions and retry mechanisms.

[0089] The result processing module includes a data integration unit, a formatting unit, and a cache management unit;

[0090] The data integration unit is used to integrate the return results from multiple interfaces;

[0091] The formatting unit is used to format the output data;

[0092] The cache management unit is used to manage the query result cache.

[0093] This Dify-based natural language data query method includes the following steps:

[0094] Step S1: Natural Language Understanding

[0095] It receives natural language queries from users, analyzes the semantics of the queries using a large language model, and identifies key information, including query intent, target data type, and time range.

[0096] Extract structured information, including query conditions, sorting requirements, and aggregation methods, to generate a standardized query requirement description; Step S2: Knowledge base interface metadata matching.

[0097] Maintain a metadata knowledge base containing interface descriptions, parameter specifications, and data format information, and calculate the matching degree between user questions and interface descriptions based on semantic similarity;

[0098] It supports parallel matching of multiple interfaces and takes into account the dependencies and calling order between interfaces to select the combination of interfaces with the highest similarity.

[0099] Step S3: Intelligent Parameter Parsing and Generation

[0100] Leveraging the capabilities of large models to extract parameter information from natural language, including numerical values, dates, and strings;

[0101] Automatically convert parameter formats and types according to interface parameter specifications, handle logical relationships and constraints between parameters, and generate call parameters that meet interface requirements;

[0102] Step S4: Multi-interface collaborative call

[0103] The system determines the order and dependencies of interface calls based on the matching results, supports both serial and parallel calling modes, enables data transfer and parameter sharing between interfaces, and handles interface call exceptions and retry mechanisms.

[0104] Step S5: Result Integration and Output

[0105] Collect the return results from multiple interfaces, integrate and deduplicate the return result data according to business logic, and use a large model to organize and output user-friendly results.

[0106] Example 1: Basic Query Process

[0107] The user inputs a natural language query: "Query city GDP data for January 2024". The system, through large-scale model analysis, identifies the query target as "city GDP data" with a time range of "January 2024". The knowledge retrieval module finds the relevant GDP data interface. The parameter parsing module extracts the time parameters and converts them to the format required by the interface. The system calls the GDP data interface to obtain the query results. The result processing module formats the data and returns it to the user.

[0108] Example 2: Complex Multi-Interface Queries

[0109] The user inputs: "Comparative analysis of GDP, population, and air quality data for various cities in 2024." The system identifies the need to query data from three different dimensions. The knowledge retrieval module matches the three interfaces: GDP, population, and air quality. The parameter parsing module generates corresponding calling parameters for each interface. The system calls the three interfaces in parallel. The results processing module integrates the data and generates a comparative analysis report. The output module returns formatted analysis results.

[0110] The Dify-based natural language data query device includes a memory and a processor; the memory is used to store a computer program, and the processor is used to execute the computer program to implement the above-described method steps.

[0111] The readable storage medium stores a computer program that, when executed by a processor, implements the above-described method steps.

[0112] Compared with existing technologies, this Dify-based natural language data query method has the following characteristics:

[0113] (1) Users do not need to master SQL or API knowledge. They can perform complex data queries through natural language, which greatly reduces the threshold for use.

[0114] (2) Intelligent matching and automatic parameter parsing reduce manual configuration time and improve query efficiency.

[0115] (3) The ability to call multiple interfaces collaboratively meets the data needs of complex business scenarios and supports comprehensive analysis across data sources.

[0116] (4) Visual workflow design provides an intuitive operation interface, and error handling and fault tolerance mechanisms enhance the user experience.

[0117] (5) The modular design makes the system easy to maintain and expand, and the access to new interfaces is more convenient.

[0118] (6) Through intelligent data query methods, more users can conveniently obtain and analyze data, and give full play to the value of data.

[0119] The embodiments described above are merely one specific implementation of the present invention. Ordinary changes and substitutions made by those skilled in the art within the scope of the technical solution of the present invention should be included within the protection scope of the present invention.

Claims

1. A natural language data query system based on Dify, characterized in that: It includes a natural language understanding module, an interface metadata management module, an intelligent matching module, a parameter parsing module, an interface call module, and a result processing module; The Natural Language Understanding module integrates a large language model to analyze user-input natural language queries and understand the user's query intent. The interface metadata management module is used to maintain and manage the metadata information of data interfaces; The intelligent matching module is used to match user questions with the knowledge base interface; The parameter parsing module is used to extract and generate calling parameters that meet the interface requirements from natural language. The API call module is used to perform collaborative calls to multiple APIs and determine the call order based on API dependencies. The results processing module is used to integrate and format query results through data deduplication, format conversion, and exception handling.

2. The Dify-based natural language data query system according to claim 1, characterized in that: The natural language understanding module includes a large language model analysis unit, a query element extraction unit, and a standardization processing unit. Large language model analysis unit for semantic understanding and intent recognition; The query element extraction unit is used to extract key information, including query conditions, time range, and data type. The standardization processing unit is used to generate standardized query requirement descriptions.

3. The Dify-based natural language data query system according to claim 1, characterized in that: The interface metadata management module includes a metadata storage unit, a semantic indexing unit, and a version management unit; The metadata storage unit is used to store metadata information, including interface descriptions, parameter specifications, and data formats. The semantic indexing unit supports vector retrieval and keyword retrieval, and is used to build a semantic index of the interface description and locate relevant interface metadata from the knowledge base; The version management unit is used to manage version updates of interface metadata.

4. The Dify-based natural language data query system according to claim 1, characterized in that: The intelligent matching module includes a semantic similarity calculation unit, a multi-interface selection unit, and a dependency analysis unit; The semantic similarity calculation unit is used to calculate the matching degree between the user question and the interface description. The multi-interface selection unit is used to select the interface combination with the highest matching degree; The dependency analysis unit is used to analyze the calling order and dependencies between interfaces.

5. The Dify-based natural language data query system according to claim 1, characterized in that: The parameter parsing module supports conditional queries, range queries, and fuzzy queries, and includes a parameter extraction unit, a format conversion unit, and a parameter verification unit. The parameter extraction unit is used to extract parameter information from natural language, including numerical values, dates, and strings; The format conversion unit is used to convert the parameter format and type according to the interface requirements; The parameter verification unit is used to verify the validity and completeness of the parameters.

6. The Dify-based natural language data query system according to claim 1, characterized in that: The interface calling module supports both serial and parallel calling modes, including: The scheduling unit is invoked to determine the order and mode of interface calls; Parallel processing unit, used to support parallel calls to multiple interfaces; The exception handling unit is used to handle call exceptions and retry mechanisms.

7. The Dify-based natural language data query system according to claim 1, characterized in that: The result processing module includes a data integration unit, a formatting unit, and a cache management unit; The data integration unit is used to integrate the return results from multiple interfaces; The formatting unit is used to format the output data; The cache management unit is used to manage the query result cache.

8. A natural language data query method based on Dify, characterized in that: Includes the following steps: Step S1: Natural Language Understanding It receives natural language queries from users, analyzes the semantics of the queries using a large language model, and identifies key information, including query intent, target data type, and time range. Extract structured information, including query conditions, sorting requirements, and aggregation methods, to generate a standardized description of query requirements; Step S2: Knowledge base interface metadata matching Maintain a metadata knowledge base containing interface descriptions, parameter specifications, and data format information, and calculate the matching degree between user questions and interface descriptions based on semantic similarity; It supports parallel matching of multiple interfaces and takes into account the dependencies and calling order between interfaces to select the combination of interfaces with the highest similarity. Step S3: Intelligent Parameter Parsing and Generation Leveraging the capabilities of large models to extract parameter information from natural language, including numerical values, dates, and strings; Automatically convert parameter formats and types according to interface parameter specifications, handle logical relationships and constraints between parameters, and generate call parameters that meet interface requirements; Step S4: Multi-interface collaborative call The system determines the order and dependencies of interface calls based on the matching results, supports both serial and parallel calling modes, enables data transfer and parameter sharing between interfaces, and handles interface call exceptions and retry mechanisms. Step S5: Result Integration and Output Collect the return results from multiple interfaces, integrate and deduplicate the return result data according to business logic, and use a large model to organize and output user-friendly results.

9. A natural language data query device based on Dify, characterized in that: It includes a memory and a processor; the memory is used to store a computer program, and the processor is used to execute the computer program to implement the steps of the method as described in claim 8.

10. A readable storage medium, characterized in that: The readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method as described in claim 8.

Citation Information

Patent Citations

  • Generative artificial intelligence (AI) training and ai-assisted decisioning

    CA3254777A1

  • Interface parameter extraction method and device based on slot filling and medium

    CN120144652A

  • Interactive number asking agent system based on large language model

    CN120216656A

  • Intelligent data query method and device based on large model and medium

    CN120296126A

  • System and method for autonomously generating heterogeneous data source interoperability bridges based on semantic modeling derived from self adapting ontology

    WO2003060751A1