Domain knowledge supplement and SQL generation optimization method and device based on large model
By adopting large-model-based domain knowledge supplement and SQL generation optimization methods in the BI system, the problem of inaccurate SQL queries when handling specific knowledge of vertical fields is solved, and efficient and accurate SQL generation and optimization are achieved to meet the needs of complex business analysis.
Patent Information
- Application Number
- CN202510429243.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing question-and-answer BI systems face challenges when dealing with vertical domain-specific knowledge, especially in the absence of domain-specific private knowledge bases, the generated SQL queries are inaccurate and logical errors are frequently released, which cannot meet the complex business analysis needs.
Using a large model-based domain knowledge supplement and SQL generation optimization method, flexible, accurate and efficient SQL generation is achieved by integrating multi-source heterogeneous knowledge, intelligently analyzing unstructured documents, accurate recall of relevant knowledge, high-quality SQL generation and optimization, and dynamic update knowledge management.
It improves the efficiency and accuracy of knowledge management, reduces manual processing time and cost, enhances the user experience and system practicality. The generated SQL queries not only comply with syntax specifications, but can also be efficiently executed and accurately reflect users' needs, meeting the needs of complex business analysis.
Smart Images

Figure CN119938891A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of large language models, and specifically relates to a method and device for domain knowledge supplementation and SQL generation optimization based on a large model. Background Art
[0002] In recent years, with the rapid development of big data and artificial intelligence technology, natural language processing has been increasingly used in the field of business intelligence. Traditional BI systems usually rely on fixed report templates and predefined data queries, which limits users' flexibility and data exploration capabilities. In order to overcome these limitations, question-and-answer BI systems have emerged, allowing users to interact with the system through natural language and obtain the required analysis results.
[0003] However, existing question-and-answer BI systems face challenges in processing vertical domain-specific knowledge, especially in the absence of a domain-specific private knowledge base. The generated SQL queries are often inaccurate and prone to logical errors, and cannot meet complex business analysis needs. Summary of the invention
[0004] In order to solve at least one technical problem existing in the background technology, the present application provides a domain knowledge supplement and SQL generation optimization method based on a large model, which realizes flexible, accurate and efficient SQL generation by efficiently integrating multi-source heterogeneous knowledge, intelligently parsing unstructured documents, accurately recalling relevant knowledge, high-quality SQL generation and optimization, and dynamically updating knowledge management.
[0005] The technical solution adopted in this application is: The first aspect of the present application provides a method for domain knowledge supplementation and SQL generation optimization based on a large model, including: Integrate multi-source heterogeneous knowledge, and perform information entry and parameter configuration on the multi-source heterogeneous knowledge; Automatically parse unstructured document content and convert it into structured knowledge based on semantic segmentation model and knowledge extraction agent; Perform word segmentation based on the question input by the user, and call the retrieval algorithm to recall relevant knowledge from the knowledge base to supplement the information needed for the answer; Based on the recalled structured knowledge, the corresponding application paths are designed according to different categories of knowledge and SQL is generated.
[0006] According to an embodiment of the present application, the integration of multi-source heterogeneous knowledge and the information entry and parameter configuration of the multi-source heterogeneous knowledge are specifically as follows: Based on a visualization workbench for business experts, the multi-source heterogeneous knowledge is divided into domains, classified and semantically annotated, and the effective scope and application conditions of the knowledge are configured; or, the knowledge documents are automatically parsed and domain identification and knowledge classification are completed.
[0007] According to an embodiment of the present application, the semantic segmentation model and knowledge extraction agent are used to automatically parse the unstructured document content and convert it into structured knowledge, specifically: Use semantic segmentation models to segment unstructured document content into sentences or paragraphs, and use knowledge extraction agents to perform semantic understanding and preliminary analysis; Key elements are extracted from the segmented text and classified, and finally converted into structured knowledge entries.
[0008] According to an embodiment of the present application, the question input by the user is segmented, and a retrieval algorithm is called to recall relevant knowledge from the knowledge base to supplement the information required for the answer, specifically: Perform word segmentation on the questions input by the user and select the corresponding recall algorithm according to the characteristics of the question content; Relevant knowledge is retrieved from the knowledge base and screened, and then added to the user's query to generate answer information and SQL query information.
[0009] According to an embodiment of the present application, the structured knowledge based on the recall is used to design corresponding application paths and generate SQL according to different categories of knowledge, specifically: Combine different types of knowledge to perform syntax tuning and verification on the generated SQL.
[0010] According to an embodiment of the present application, the word segmentation process is performed on the question input by the user, and a corresponding recall algorithm is selected according to the characteristics of the question content, specifically: Based on general knowledge, semantic similarity is used for recall, and based on indicator knowledge, character similarity is used for recall.
[0011] According to an embodiment of the present application, the syntax tuning and verification of the generated SQL is performed by combining different types of knowledge, specifically: When the type is general knowledge, it supplements the background information of SQL generated by the large model; When the type is indicator knowledge, the indicators required by the large model are supplemented and the indicator generation rules of the large model are informed.
[0012] The second aspect of the present application provides a domain knowledge supplement and SQL generation optimization device based on a large model, including: A knowledge supplement module, suitable for integrating multi-source heterogeneous knowledge and performing information entry and parameter configuration on the multi-source heterogeneous knowledge; Knowledge learning module, which is suitable for automatically parsing unstructured document content and converting it into structured knowledge based on semantic segmentation models and knowledge extraction agents; The knowledge retrieval module is suitable for performing word segmentation processing based on the questions input by the user and calling the retrieval algorithm to recall relevant knowledge from the knowledge base to supplement the information required for the answer; The knowledge consumption module is suitable for recall-based structured knowledge, and designs corresponding application paths and generates SQL according to different categories of knowledge.
[0013] The third aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the domain knowledge supplementation and SQL generation optimization method based on a large model as described in any embodiment of the first aspect is implemented.
[0014] The present application also provides a non-volatile computer storage medium having computer executable instructions stored thereon, and when the computer program is executed by a processor, the method for domain knowledge supplementation and SQL generation optimization based on a large model in any embodiment of the first aspect as described above is implemented.
[0015] Due to the adoption of the above technical solution, the beneficial effects achieved by this application are as follows: According to the large model-based domain knowledge supplement and SQL generation optimization method provided in the embodiment of the present application, multi-source heterogeneous knowledge is first integrated, and information entry and parameter configuration are performed on the multi-source heterogeneous knowledge. The system can efficiently integrate knowledge resources from different sources and formats, and perform information entry and parameter configuration. This includes domain division, knowledge classification, semantic annotation, and precise configuration of the effective scope. This method ensures that business experience can be standardized and precipitated into structured knowledge, improves the efficiency and accuracy of knowledge management, and enables enterprises to better utilize their internal knowledge assets. Then, based on the semantic segmentation model and knowledge extraction agent, the unstructured document content is automatically parsed and converted into structured knowledge. This process greatly reduces the time and cost of manual processing, improves the efficiency of knowledge conversion, and ensures the quality and applicability of knowledge through expert review. Then, based on the question entered by the user, word segmentation is performed, and the retrieval algorithm is called to recall relevant knowledge from the knowledge base to supplement the information required for the answer. This method improves the accuracy of knowledge matching, enables the system to respond to user query needs more accurately, and provide high-quality answer information and SQL queries, enhancing user experience and system practicality. Finally, based on the recalled structured knowledge, the corresponding application path is designed according to different categories of knowledge and SQL is generated. This strategy ensures that the generated SQL query not only conforms to the grammatical specifications, but also can be executed efficiently and accurately reflect user needs, thereby improving the quality and performance of SQL generation and meeting the needs of complex business analysis. To summarize, the large-model-based domain knowledge supplement and SQL generation optimization method provided in the embodiment of the present application realizes flexible, accurate and efficient SQL generation by efficiently integrating multi-source heterogeneous knowledge, intelligently parsing unstructured documents, accurately recalling relevant knowledge, high-quality SQL generation and optimization, and dynamically updating knowledge management. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 A flowchart of a method for domain knowledge supplementation and SQL generation optimization based on a large model provided in an embodiment of the present application; Figure 2 A schematic diagram of the structure of a large model-based domain knowledge supplement and SQL generation optimization device provided in an embodiment of the present application; Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0017] Reference numerals: 110. Knowledge supplement module; 120. Knowledge learning module; 130. Knowledge retrieval module; 140. Knowledge consumption module; 810, processor; 820, communication interface; 830, memory; 840, communication bus. DETAILED DESCRIPTION
[0018] In order to more clearly illustrate the overall concept of the present application, a detailed description is given below in an illustrative manner in conjunction with the accompanying drawings.
[0019] In the following description, many specific details are set forth to facilitate a full understanding of the present application. However, the present application may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present application is not limited by the specific embodiments disclosed below. It should be noted that the embodiments of the present application and the features in each embodiment may be combined with each other without conflict.
[0020] In the present application, unless otherwise clearly specified and limited, a first feature "above" or "below" a second feature may be that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in an appropriate manner in any one or more embodiments or examples.
[0021] like Figure 1 As shown, the first embodiment of the present application provides a method for domain knowledge supplementation and SQL generation optimization based on a large model, including: Step 100: Integrate multi-source heterogeneous knowledge, and enter information and configure parameters for the multi-source heterogeneous knowledge.
[0022] Step 200: Based on the semantic segmentation model and knowledge extraction agent, the unstructured document content is automatically parsed and converted into structured knowledge.
[0023] Step 300: perform word segmentation based on the question input by the user, and call a retrieval algorithm to recall relevant knowledge from the knowledge base to supplement the information required for the answer.
[0024] Step 400: Based on the recalled structured knowledge and according to different categories of knowledge, design corresponding application paths and generate SQL.
[0025] In step 100, this is the first step in the domain knowledge supplement and SQL generation optimization method based on the big model. Its main purpose is to integrate knowledge from different sources and formats (i.e., "multi-source heterogeneous knowledge"). This knowledge may include but is not limited to documents, databases, web page content, etc. This step specifically includes two core activities: Integrate heterogeneous knowledge from multiple sources: This involves collecting information from different data sources and standardizing it for subsequent processing. For example, it may be necessary to parse PDF files, Word documents, Excel tables, or access relational databases and NoSQL databases.
[0026] Information entry and parameter configuration: After the knowledge is collected and standardized, the next step is to enter the information and configure the necessary parameters. This is usually done through a user-friendly interface that allows experts or administrators to enter new knowledge items and set relevant parameters to define the application scenarios of the knowledge.
[0027] In order to achieve effective knowledge integration and parameter configuration, the following technical means and technical components can be adopted: ETL tools: used to extract, transform, and load data from multiple sources into a unified data warehouse.
[0028] Natural language processing technology: used to parse unstructured text data and extract useful information from it.
[0029] Machine learning algorithms: used to identify and classify different types of knowledge and automatically suggest corresponding application paths and parameter settings.
[0030] Visual workbench: A graphical interface provided to business experts that simplifies the process of knowledge entry and parameter configuration.
[0031] Step 100 not only lays the foundation for the entire process, but also greatly enhances the organization's ability to manage and utilize its knowledge assets through efficient knowledge integration and precise parameter configuration.
[0032] In step 200, which is one of the key links in the whole method, the main purpose is to transform unstructured document content (such as PDF, Word document, web page content, etc.) into structured knowledge through advanced natural language processing technology. This process includes the following core steps: Semantic segmentation: Semantic segmentation models are used to break down unstructured text content into smaller units (such as sentences or paragraphs) for subsequent analysis.
[0033] Semantic understanding and preliminary analysis: Use knowledge extraction agents to conduct in-depth semantic understanding and preliminary classification of the segmented text to identify the key elements and information.
[0034] Feature extraction and classification: Extract key elements (such as entities, relationships, attributes, etc.) from text and classify them according to business needs.
[0035] Transformation and Review: The extracted and classified information is transformed into structured knowledge items and reviewed by experts to ensure its accuracy and applicability.
[0036] To achieve efficient and accurate knowledge transfer, the following techniques and tools can be adopted: Semantic segmentation model: It can be a deep learning based model such as BERT, RoBERTa, etc. These models can understand the context and perform effective segmentation in long texts.
[0037] Knowledge extraction agent: usually includes modules such as named entity recognition, relationship extraction, and event detection, which are used to extract useful information from text.
[0038] Natural language processing technology: In addition to semantic segmentation and knowledge extraction, it also includes part-of-speech tagging, dependency syntax analysis and other technologies to help the system better understand the text content.
[0039] Machine learning algorithms: Processes used to optimize knowledge extraction, such as training models to improve accuracy through supervised learning.
[0040] Expert review mechanism: After automated processing, an interface is provided for domain experts to review and fine-tune the extracted knowledge to ensure its quality and applicability.
[0041] In step 300, the user's query is parsed through natural language processing technology, and relevant structured knowledge is recalled from the knowledge base to generate accurate answers and SQL queries. This process includes the following main steps: Question word segmentation processing: Word segmentation: Break down the natural language questions input by users into multiple keywords or phrases (such as nouns, verbs, adjectives, etc.) for subsequent processing.
[0042] Contextual understanding: Combine contextual information to ensure that the meaning of each word accurately reflects the user's intention.
[0043] Choose an appropriate recall algorithm: Algorithm selection: The system automatically selects the most appropriate recall algorithm based on the characteristics of the question content (such as domain, knowledge category, etc.). For example, the semantic similarity algorithm is used for general knowledge, while the character similarity algorithm is used for indicator knowledge.
[0044] Dynamic adjustment: The system can dynamically adjust the algorithm based on the user's complete question or word segmentation results to ensure the best recall effect.
[0045] Knowledge base search and screening: Retrieval: Based on the selected recall algorithm, the system retrieves relevant structured knowledge from the knowledge base. The knowledge base can include relational databases, vector databases, graph databases and other storage forms.
[0046] Screening: During the retrieval process, the system will conduct a preliminary screening of the recalled knowledge items to remove irrelevant or low-quality knowledge, ensuring that only highly relevant information is ultimately returned to the user.
[0047] Additional information required for answering: Information supplement: Supplement the filtered relevant knowledge to the user's query to help the system better understand the user's query needs and generate high-quality answer information and SQL queries.
[0048] Feedback mechanism: Provides a preview interface to display the generated answer information and SQL query and its expected results for user review and confirmation.
[0049] Step 300 not only achieves efficient analysis of user queries and accurate recall of relevant knowledge, but also ensures the accuracy of knowledge matching and the practicality of the system through advanced natural language processing technology and intelligent retrieval algorithms. This step greatly improves the flexibility and responsiveness of the system, and provides enterprises with powerful data analysis and decision support capabilities.
[0050] In step 400, the structured knowledge retrieved from the knowledge base is applied to specific business scenarios and high-quality SQL queries are generated. This process includes the following main steps: Classification and path design: Knowledge classification: The system first classifies the recalled structured knowledge and identifies different types of knowledge (such as general knowledge, indicator knowledge, SQL knowledge, etc.).
[0051] Path design: The system designs different application paths based on knowledge categories. For example, general knowledge is used to provide background information, indicator knowledge is used to retrieve and calculate specific database indicators, and SQL knowledge is directly involved in generating and optimizing SQL queries.
[0052] Application of background knowledge: Additional background information: For general knowledge, the system uses it as part of the prompt word and provides it to the large model to enhance the understanding of the user's query intent and help generate more accurate SQL queries.
[0053] Application of indicator knowledge: Indicator extraction and calculation: Indicator knowledge is used to retrieve key indicators in a specific database and perform necessary calculations. The system automatically searches for relevant database fields based on the recalled knowledge entries and performs calculations based on predefined business rules.
[0054] Application and generation of SQL knowledge: SQL generation: Utilize the recalled SQL knowledge and combine it with the user's query requirements to automatically generate grammatically compliant and efficient SQL queries.
[0055] Syntax tuning and verification: During the SQL generation process, the system will perform syntax tuning and verification on the generated SQL statements to ensure their accuracy and performance optimization. For example, the system can adjust the WHERE clause in the SQL query according to business logic constraints.
[0056] Step 400 not only realizes SQL generation based on structured knowledge, but also ensures the accuracy and efficiency of SQL queries through advanced syntax parsing, optimization technology and intelligent application path design. This step greatly improves the flexibility and practicality of the system and provides enterprises with powerful data analysis and decision support capabilities.
[0057] According to the large model-based domain knowledge supplement and SQL generation optimization method provided in the embodiment of the present application, multi-source heterogeneous knowledge is first integrated, and information entry and parameter configuration are performed on the multi-source heterogeneous knowledge. The system can efficiently integrate knowledge resources from different sources and formats, and perform information entry and parameter configuration. This includes domain division, knowledge classification, semantic annotation, and precise configuration of the effective scope. This method ensures that business experience can be standardized and precipitated into structured knowledge, improves the efficiency and accuracy of knowledge management, and enables enterprises to better utilize their internal knowledge assets. Then, based on the semantic segmentation model and knowledge extraction agent, the unstructured document content is automatically parsed and converted into structured knowledge. This process greatly reduces the time and cost of manual processing, improves the efficiency of knowledge conversion, and ensures the quality and applicability of knowledge through expert review. Then, based on the question entered by the user, word segmentation is performed, and the retrieval algorithm is called to recall relevant knowledge from the knowledge base to supplement the information required for the answer. This method improves the accuracy of knowledge matching, enables the system to respond to user query needs more accurately, and provide high-quality answer information and SQL queries, enhancing user experience and system practicality. Finally, based on the recalled structured knowledge, the corresponding application path is designed according to different categories of knowledge and SQL is generated. This strategy ensures that the generated SQL query not only conforms to the grammatical specifications, but also can be executed efficiently and accurately reflect user needs, thereby improving the quality and performance of SQL generation and meeting the needs of complex business analysis. To summarize, the large-model-based domain knowledge supplement and SQL generation optimization method provided in the embodiment of the present application realizes flexible, accurate and efficient SQL generation by efficiently integrating multi-source heterogeneous knowledge, intelligently parsing unstructured documents, accurately recalling relevant knowledge, high-quality SQL generation and optimization, and dynamically updating knowledge management.
[0058] In some embodiments of the present application, multi-source heterogeneous knowledge is integrated, and information entry and parameter configuration are performed on the multi-source heterogeneous knowledge, specifically: Based on a visual workbench for business experts, multi-source heterogeneous knowledge is divided into domains, classified and semantically annotated, and the effective scope and application conditions of the knowledge are configured; or, knowledge documents are automatically parsed and domain identification and knowledge classification are completed.
[0059] Specifically, through a visual interface, business experts can define and divide knowledge in different fields. For example, knowledge in different industries or departments such as finance, healthcare, and manufacturing can be clearly distinguished. Within each field, knowledge is further classified. For example, in the financial field, knowledge can be divided into categories such as financial statement analysis, risk assessment, and market forecasting. Detailed semantic tags are added to each piece of knowledge to ensure its accuracy and interpretability in subsequent use. These tags can help the system better understand the content and application scenarios of the knowledge. According to specific business needs, configure the scope of effectiveness of the knowledge (such as time limits, geographical areas) and application conditions (such as specific user roles, data types). This step ensures that knowledge can be accurately applied in the correct context.
[0060] Supports importing knowledge documents in multiple formats (such as PDF, Word, Excel, etc.) into the system. First, preprocess these documents, including text extraction, format conversion and other operations. Use natural language processing technology (such as BERT, RoBERTa and other models) to automatically identify the field to which the document belongs. For example, by analyzing the keywords and topics in the document content, determine whether it belongs to finance, medical or other fields. After identifying the field, further classify the knowledge in the document. For example, in the financial field, the system can automatically classify the content in the document into subcategories such as financial statements, market trends, and risk management. The system automatically adds semantic tags to the identified knowledge and converts it into a structured form for subsequent storage and retrieval.
[0061] Through the visual workbench for business experts and the function of automatically parsing knowledge documents, this method achieves efficient and accurate knowledge integration and management. This not only lays the foundation for the entire process, but also ensures the quality and applicability of knowledge through advanced technology and expert review mechanisms, providing enterprises with powerful knowledge management and analysis capabilities.
[0062] In some embodiments of the present application, based on the semantic segmentation model and the knowledge extraction agent, the unstructured document content is automatically parsed and converted into structured knowledge, specifically: Use semantic segmentation models to segment unstructured document content into sentences or paragraphs, and use knowledge extraction agents to perform semantic understanding and preliminary analysis; Key elements are extracted from the segmented text and classified, and finally converted into structured knowledge entries.
[0063] Specifically, Use semantic segmentation models to segment unstructured document content into sentences or paragraphs: Document import and preprocessing: First, import unstructured documents (such as PDF, Word, web page content, etc.) into the system and perform necessary preprocessing operations, including text extraction, format conversion, and noise removal.
[0064] Semantic segmentation: Use semantic segmentation models (such as BERT, RoBERTa, etc.) to divide document content into smaller units, usually sentences or paragraphs. This segmentation not only helps with subsequent processing, but also better understands contextual information.
[0065] Semantic understanding and preliminary analysis through knowledge extraction agents: Semantic understanding: The knowledge extraction agent performs in-depth semantic understanding and preliminary classification of the segmented text. This includes identifying entities (such as names, places, dates), relationships (such as relationships between characters, causal relationships between events), and attributes (such as the specific value of a certain indicator) in the sentence.
[0066] Preliminary analysis: Based on the results of semantic understanding, the agent conducts preliminary analysis to identify key elements and important information in the text. For example, it determines which sentences are key descriptions of a specific topic and which paragraphs contain important business rules or logic.
[0067] Extract key elements from the segmented text and classify them: Key element extraction: Based on the preliminary analysis, further extract the key elements in the text. These elements can be named entities (such as company names, product names), timestamps, values, relationships, etc.
[0068] Classification: Classify the extracted key elements according to their types. For example, classify all mentioned dates into the "time" category, all amounts into the "value" category, all business rules into the "rule" category, etc.
[0069] Finally transformed into structured knowledge entries: Structured transformation: convert the classified key elements into structured knowledge entries. Each entry usually contains multiple fields, such as entity name, attribute value, relationship description, etc. For example, a structured entry may be expressed as: "The company's sales in the first quarter of 2025 were 10 million yuan."
[0070] Review and fine-tuning: The generated structured knowledge items can be displayed in the system to domain experts for review and fine-tuning to ensure their accuracy and applicability. Experts can view and modify each knowledge item through a visual interface.
[0071] By combining the semantic segmentation model with the knowledge extraction agent, this method realizes the effective transformation of unstructured document content into structured knowledge. This not only improves the flexibility and practicality of the system, but also ensures the quality and applicability of knowledge through advanced natural language processing technology and expert review mechanism, providing enterprises with powerful knowledge management and analysis capabilities.
[0072] In some embodiments of the present application, word segmentation is performed based on the question input by the user, and a retrieval algorithm is called to recall relevant knowledge from the knowledge base to supplement the information required for the answer, specifically: Perform word segmentation on the questions input by the user and select the corresponding recall algorithm according to the characteristics of the question content; Relevant knowledge is retrieved from the knowledge base and screened, and then added to the user's query to generate answer information and SQL query information.
[0073] Specifically, the natural language questions input by users are decomposed into multiple keywords or phrases (such as nouns, verbs, adjectives, etc.) for subsequent processing. For example, for the question "What is the company's sales in the first quarter of 2025?", the system will decompose it into keywords such as "2025", "first quarter", "company", and "sales". Combined with contextual information, ensure that the meaning of each word can accurately reflect the user's intention. For example, "sales" may need to be associated with a specific time period and company name.
[0074] According to the content characteristics of the question (such as domain, knowledge category, etc.), the system automatically selects the most appropriate recall algorithm. For example: for general knowledge, a recall algorithm based on semantic similarity (such as Word2Vec, BERT) is used. For indicator knowledge, a recall algorithm based on character similarity (such as Levenshtein distance, Cosine similarity) is used. The system can dynamically adjust the algorithm based on the user's complete question or word segmentation results to ensure the best recall effect. For example, if the question is a query about a specific numerical value, the system may give priority to using a character similarity algorithm to match the exact numerical value.
[0075] Based on the selected recall algorithm, the system retrieves relevant structured knowledge from the knowledge base. The knowledge base can include various storage forms such as relational databases, vector libraries, and graph databases. Example: In the above question, the system will search for records related to "first quarter of 2025", "company", and "sales" from the knowledge base. During the retrieval process, the system will perform a preliminary screening of the recalled knowledge items to remove irrelevant or low-quality knowledge to ensure that the information ultimately returned to the user is highly relevant. Example: The system may filter out sales data for other quarters that are not related to "first quarter of 2025".
[0076] Supplement the filtered relevant knowledge to the user's query to help the system better understand the user's query needs and generate high-quality answer information and SQL queries. Example: After the system finds the knowledge entry "the company's sales in the first quarter of 2025 is 10 million yuan", it uses it as part of the answer and generates the corresponding SQL query. Provide a preview interface to display the generated answer information and SQL query and its expected results for the user to review and confirm.
[0077] By segmenting the questions entered by the user and selecting the corresponding recall algorithm according to the characteristics of the question content, the system can efficiently recall relevant knowledge from the knowledge base and generate high-quality answer information and SQL queries. This not only improves the flexibility and practicality of the system, but also ensures the accuracy of knowledge matching and the practicality of the system through advanced natural language processing technology and intelligent retrieval algorithms, providing enterprises with powerful data analysis and decision support capabilities.
[0078] In some embodiments of the present application, based on the recalled structured knowledge, corresponding application paths are designed according to different categories of knowledge and SQL is generated, specifically: Combine different types of knowledge to perform syntax tuning and verification on the generated SQL.
[0079] Specifically, the structured knowledge recalled from the knowledge base is first classified. Common knowledge categories include: General Knowledge: Provides background information or context.
[0080] Indicator knowledge: involves specific numerical values, indicators or metrics.
[0081] SQL knowledge: directly used to generate and optimize SQL queries.
[0082] For general knowledge, the system uses it as part of the prompt words and provides it to the big model to enhance the understanding of the user's query intent. For example, if the user asks "What is the company's sales in the first quarter of 2025", the system can add background information such as "The company's main business scope is electronic product sales." Through background information, the system can better understand the user's query needs and take this background information into account when generating SQL.
[0083] For indicator knowledge, the system needs to retrieve key indicators from a specific database and perform necessary calculations. For example, the system can automatically find relevant database fields (such as the "sales" field) based on the recalled knowledge entries and perform calculations based on predefined business rules. The system sets specific parameters in SQL queries based on indicator knowledge, such as time range, geographic area, etc.
[0084] By using the recalled SQL knowledge and combining it with the user's query requirements, the system automatically generates SQL queries that are grammatically compliant and efficient. For example, the system can directly use the SQL templates stored in the knowledge base to generate new queries. In the process of generating SQL, the system will perform syntax tuning and verification on the generated SQL statements to ensure their accuracy and performance optimization. For example, the system can adjust the WHERE clause in the SQL query according to the business logic constraints, or optimize the JOIN operation to improve query efficiency.
[0085] Apply different types of knowledge to the SQL generation process to ensure that the generated SQL can accurately reflect the user's query intent. For example, combine background information, indicator data, and SQL templates to generate a complete SQL query. The system will perform syntax checks on the generated SQL statements to ensure that they comply with the syntax specifications of the target database. For example, check whether the keywords in the SQL are correct, whether there are spelling errors in the table name and field name, etc. The system will perform performance evaluation on the generated SQL to ensure its efficient execution. For example, optimize index usage and reduce unnecessary JOIN operations by analyzing the query plan.
[0086] After generating SQL, the system can provide a preview interface to display the generated SQL query and its expected results for users to review and confirm. Users can review the generated SQL query, confirm its correctness and make necessary modifications. The system can also provide an execution button to allow users to directly run the generated SQL query.
[0087] By combining different types of knowledge and performing syntax tuning and verification on the generated SQL, this method not only achieves efficient and accurate SQL generation, but also ensures the accuracy and efficiency of SQL queries through advanced syntax parsing, optimization technology, and intelligent application path design. This step greatly improves the flexibility and practicality of the system, providing enterprises with powerful data analysis and decision support capabilities.
[0088] In some embodiments of the present application, the question input by the user is segmented, and a corresponding recall algorithm is selected according to the characteristics of the question content, specifically: Based on general knowledge, semantic similarity is used for recall, and based on indicator knowledge, character similarity is used for recall.
[0089] Specifically, the natural language questions input by users are decomposed into multiple keywords or phrases (such as nouns, verbs, adjectives, etc.) for subsequent processing. For example, for the question "What is the company's sales in the first quarter of 2025?", the system will decompose it into keywords such as "2025", "first quarter", "company", and "sales". Combined with contextual information, ensure that the meaning of each word can accurately reflect the user's intention. For example, "sales" may need to be associated with a specific time period and company name.
[0090] For general knowledge involving background information or context, the system uses a recall algorithm based on semantic similarity. Such algorithms usually use deep learning models (such as BERT, RoBERTa) to calculate the semantic similarity between texts. For example: If a user asks "What is the company's business scope?" the system will search the knowledge base for descriptive knowledge items related to "the company's business scope" and match the most relevant answer through a semantic similarity algorithm.
[0091] For queries involving specific values or precise indicators, the system uses a recall algorithm based on character similarity. Such algorithms usually use edit distance (such as Levenshtein distance) or cosine similarity to calculate the similarity between strings. For example: If a user asks "What is the company's sales in the first quarter of 2025?" the system will search the knowledge base for specific numerical data related to "first quarter of 2025" and "sales", and match the closest answer through the character similarity algorithm.
[0092] By segmenting the questions entered by users and selecting the corresponding recall algorithm according to the characteristics of the question content (using semantic similarity recall based on general knowledge and character similarity recall based on indicator knowledge), the system can efficiently recall relevant knowledge from the knowledge base and generate high-quality answer information and SQL queries. This not only improves the flexibility and practicality of the system, but also ensures the accuracy of knowledge matching and the practicality of the system through advanced natural language processing technology and intelligent retrieval algorithms, providing enterprises with powerful data analysis and decision support capabilities.
[0093] In some embodiments of the present application, different types of knowledge are combined to perform syntax tuning and verification on the generated SQL, specifically: When the type is general knowledge, it supplements the background information of SQL generated by the large model; When the type is indicator knowledge, the indicators required by the large model are supplemented and the indicator generation rules of the large model are informed.
[0094] Specifically, general knowledge refers to providing background information or context to help the system better understand the user's query intent. Indicator knowledge refers to specific numerical values, indicators or metrics, which usually require precise matching and calculation.
[0095] When the type is general knowledge, the background information of SQL generated by the large model is supplemented. The specific implementation steps include: The system first performs word segmentation on the question input by the user, and uses natural language processing technology (such as BERT, RoBERTa) to perform semantic understanding and identify the background information requirements involved in the question. Example: For the question "What is the company's main business scope?" the system will recognize that this is a query about the company's background information. Based on the semantic similarity algorithm, the system retrieves background information related to the question from the knowledge base. This information is usually descriptive text that provides detailed descriptions on a specific topic. Example: The system finds the knowledge entry "The company is mainly engaged in the sale of electronic products". The recalled background information is provided as part of the prompt word to the big model to enhance the understanding of the user's query intent. This helps to generate more accurate SQL queries. Example: When generating SQL queries, the system can use "The company is mainly engaged in the sale of electronic products" as background information to help determine the specific tables and fields for the query. The system will perform syntax checks on the generated SQL statements to ensure that they meet the syntax specifications of the target database and perform performance optimization. Example: The generated SQL query may include selecting data related to "electronic products" from the "sales" table.
[0096] When the type is indicator knowledge, supplement the indicators required by the large model and inform the large model of the indicator generation rules. The implementation steps include: The system first performs word segmentation on the question entered by the user, and uses natural language processing technology to understand the semantics and identify the specific indicator requirements involved in the question. Example: For the question "What is the company's sales in the first quarter of 2025?" The system will recognize that this is a query about a specific sales indicator. Based on the character similarity algorithm, the system retrieves specific indicator data related to the question from the knowledge base. These data are usually numerical and provide specific values for specific indicators. Example: The system finds the knowledge entry "The company's sales in the first quarter of 2025 is 10 million yuan." The recalled indicator information is provided as part of the prompt word to the big model to enhance the understanding of the user's query intent. At the same time, the big model is informed of the indicator generation rules so that accurate SQL queries can be generated. Example: When generating SQL queries, the system can use "first quarter of 2025" and "sales" as key indicators to help determine the specific tables and fields of the query, and tell the big model how to calculate these indicators. The system will perform syntax checks on the generated SQL statements to ensure that they meet the syntax specifications of the target database and perform performance optimization. Example: The generated SQL query may include selecting sales data for the "first quarter of 2025" from the "sales" table and performing necessary aggregation operations.
[0097] By combining different types of knowledge and performing syntax tuning and verification on the generated SQL, this method not only achieves efficient and accurate SQL generation, but also ensures the accuracy and efficiency of SQL queries through advanced syntax parsing, optimization technology, and intelligent application path design. This step greatly improves the flexibility and practicality of the system, providing enterprises with powerful data analysis and decision support capabilities.
[0098] like Figure 2 As shown, the second embodiment of the present application provides a domain knowledge supplement and SQL generation optimization device based on a large model, including: The knowledge supplement module 110 is suitable for integrating multi-source heterogeneous knowledge and performing information entry and parameter configuration on the multi-source heterogeneous knowledge; A knowledge learning module 120, adapted to automatically parse unstructured document content and convert it into structured knowledge based on a semantic segmentation model and a knowledge extraction agent; The knowledge retrieval module 130 is adapted to perform word segmentation processing based on the question input by the user, and call a retrieval algorithm to recall relevant knowledge from the knowledge base to supplement the information required for the answer; The knowledge consumption module 140 is suitable for designing corresponding application paths and generating SQL based on the recalled structured knowledge and according to different categories of knowledge.
[0099] The domain knowledge supplement and SQL generation optimization device based on a big model provided in the second aspect embodiment of the present application can implement the domain knowledge supplement and SQL generation optimization method based on a big model in any embodiment of the first aspect mentioned above, and thus can achieve any technical effect in the domain knowledge supplement and SQL generation optimization method based on a big model, which will not be repeated here.
[0100] The third aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the domain knowledge supplement and SQL generation optimization method based on a large model in any embodiment of the first aspect is implemented.
[0101] Figure 3 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 3 As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830 and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the domain knowledge supplement and SQL generation optimization method based on a large model in any embodiment of the first aspect, the method comprising: Step 100: Integrate multi-source heterogeneous knowledge, and enter information and configure parameters for the multi-source heterogeneous knowledge.
[0102] Step 200: Based on the semantic segmentation model and knowledge extraction agent, the unstructured document content is automatically parsed and converted into structured knowledge.
[0103] Step 300: perform word segmentation based on the question input by the user, and call a retrieval algorithm to recall relevant knowledge from the knowledge base to supplement the information required for the answer.
[0104] Step 400: Based on the recalled structured knowledge and according to different categories of knowledge, design corresponding application paths and generate SQL.
[0105] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk or an optical disk.
[0106] On the other hand, the present invention further provides a computer program product, the computer program product includes a computer program, the computer program can be stored in a non-transitory computer-readable storage medium, when the computer program is executed by a processor, the computer can execute the domain knowledge supplement and SQL generation optimization method based on a large model provided by the above methods, the method includes: Step 100: Integrate multi-source heterogeneous knowledge, and enter information and configure parameters for the multi-source heterogeneous knowledge.
[0107] Step 200: Based on the semantic segmentation model and knowledge extraction agent, the unstructured document content is automatically parsed and converted into structured knowledge.
[0108] Step 300: perform word segmentation based on the question input by the user, and call a retrieval algorithm to recall relevant knowledge from the knowledge base to supplement the information required for the answer.
[0109] Step 400: Based on the recalled structured knowledge and according to different categories of knowledge, design corresponding application paths and generate SQL.
[0110] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to execute the domain knowledge supplementation and SQL generation optimization method based on a large model provided by the above methods, the method comprising: Step 100: Integrate multi-source heterogeneous knowledge, and enter information and configure parameters for the multi-source heterogeneous knowledge.
[0111] Step 200: Based on the semantic segmentation model and knowledge extraction agent, the unstructured document content is automatically parsed and converted into structured knowledge.
[0112] Step 300: perform word segmentation based on the question input by the user, and call a retrieval algorithm to recall relevant knowledge from the knowledge base to supplement the information required for the answer.
[0113] Step 400: Based on the recalled structured knowledge and according to different categories of knowledge, design corresponding application paths and generate SQL.
[0114] Finally, the present invention also provides a non-volatile computer storage medium having computer executable instructions stored thereon. When the computer program is executed by a processor, the method for supplementing domain knowledge and optimizing SQL generation based on a large model provided by the above methods is implemented. The method includes: Step 100: Integrate multi-source heterogeneous knowledge, and enter information and configure parameters for the multi-source heterogeneous knowledge.
[0115] Step 200: Based on the semantic segmentation model and knowledge extraction agent, the unstructured document content is automatically parsed and converted into structured knowledge.
[0116] Step 300: perform word segmentation based on the question input by the user, and call a retrieval algorithm to recall relevant knowledge from the knowledge base to supplement the information required for the answer.
[0117] Step 400: Based on the recalled structured knowledge and according to different categories of knowledge, design corresponding application paths and generate SQL.
[0118] Anything not described in this application can be achieved by adopting or drawing on existing technologies.
[0119] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.
[0120] The above is only the embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the protection scope of the present application.
Claims
1. A method for domain knowledge supplementation and SQL generation optimization based on a large model, characterized in that: include: Integrate multi-source heterogeneous knowledge, and perform information entry and parameter configuration on the multi-source heterogeneous knowledge; Automatically parse unstructured document content and convert it into structured knowledge based on semantic segmentation model and knowledge extraction agent; Perform word segmentation based on the question input by the user, and call the retrieval algorithm to recall relevant knowledge from the knowledge base to supplement the information needed for the answer; Based on the recalled structured knowledge, the corresponding application paths are designed according to different categories of knowledge and SQL is generated.
2. The method for domain knowledge supplementation and SQL generation optimization based on a large model according to claim 1 is characterized in that: The integration of multi-source heterogeneous knowledge and the input of information and parameter configuration of the multi-source heterogeneous knowledge are specifically as follows: Based on a visualization workbench for business experts, the multi-source heterogeneous knowledge is divided into domains, classified and semantically annotated, and the effective scope and application conditions of the knowledge are configured; or, the knowledge documents are automatically parsed and domain identification and knowledge classification are completed.
3. The method for domain knowledge supplementation and SQL generation optimization based on a large model according to claim 1 is characterized in that: The semantic segmentation model and knowledge extraction agent are used to automatically parse the unstructured document content and convert it into structured knowledge, specifically: Use semantic segmentation models to segment unstructured document content into sentences or paragraphs, and use knowledge extraction agents to perform semantic understanding and preliminary analysis; Key elements are extracted from the segmented text and classified, and finally converted into structured knowledge entries.
4. The method for domain knowledge supplementation and SQL generation optimization based on a large model according to claim 1 is characterized in that: The question input by the user is segmented and the retrieval algorithm is called to recall relevant knowledge from the knowledge base to supplement the information required for the answer, specifically: Perform word segmentation on the questions input by the user and select the corresponding recall algorithm according to the characteristics of the question content; Relevant knowledge is retrieved from the knowledge base and screened, and then added to the user's query to generate answer information and SQL query information.
5. The method for domain knowledge supplementation and SQL generation optimization based on a large model according to claim 1 is characterized in that: The structured knowledge based on recall is used to design corresponding application paths and generate SQL according to different categories of knowledge, specifically: Combine different types of knowledge to perform syntax tuning and verification on the generated SQL.
6. The method for domain knowledge supplementation and SQL generation optimization based on a large model according to claim 4 is characterized in that: The question input by the user is segmented and a corresponding recall algorithm is selected according to the characteristics of the question content, specifically: Based on general knowledge, semantic similarity is used for recall, and based on indicator knowledge, character similarity is used for recall.
7. The method for domain knowledge supplementation and SQL generation optimization based on a large model according to claim 5 is characterized in that: The generated SQL is grammatically tuned and verified by combining different types of knowledge, specifically: When the type is general knowledge, it supplements the background information of SQL generated by the large model; When the type is indicator knowledge, the indicators required by the large model are supplemented and the indicator generation rules of the large model are informed.
8. A domain knowledge supplement and SQL generation optimization device based on a large model, characterized in that: include: A knowledge supplement module, suitable for integrating multi-source heterogeneous knowledge and performing information entry and parameter configuration on the multi-source heterogeneous knowledge; Knowledge learning module, which is suitable for automatically parsing unstructured document content and converting it into structured knowledge based on semantic segmentation models and knowledge extraction agents; The knowledge retrieval module is suitable for performing word segmentation processing based on the questions input by the user and calling the retrieval algorithm to recall relevant knowledge from the knowledge base to supplement the information required for the answer; The knowledge consumption module is suitable for recall-based structured knowledge, and designs corresponding application paths and generates SQL according to different categories of knowledge.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method for domain knowledge supplementation and SQL generation optimization based on a large model is implemented.
10. A non-volatile computer storage medium having computer executable instructions stored thereon, characterized in that: When the computer program is executed by a processor, the method for domain knowledge supplementation and SQL generation optimization based on a large model as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Intelligent question answering system for automobile field
CN113505209A
Knowledge graph question and answer method and device, computer equipment and storage medium
CN116303923A
Open domain natural language reasoning question-answering system and method driven by large language model
CN116932708A
Knowledge question and answer method and system
CN117235211A
Question and answer method, device and equipment and readable storage medium
CN117875433A
Cited By
Water conservancy project vertical business domain knowledge base construction method and device
CN121434228A