Tobacco industry intelligent query method and system

By integrating tobacco industry metadata and converting it into natural language vectors, and using a large language model to generate standardized SQL query statements, the inefficiency and data security issues in tobacco industry database queries have been resolved, achieving efficient and accurate query results and data security assurance.

CN120832406APending Publication Date: 2025-10-24ANHUI PROVINCIAL CO OF CHINA NAT TOBACCO CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510861849.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Database queries in the tobacco industry suffer from low metadata utilization, insufficient model adaptation to the industry, biased semantic parsing, and lax access control, leading to low query efficiency and the risk of data leakage.

Method used

By integrating tobacco industry metadata and converting it into natural language vectors, a large language model is used to generate standardized SQL query statements. Combined with fine-grained access control, efficient and accurate database queries are achieved.

Benefits of technology

It improves query efficiency and accuracy, ensures data security, reduces query time, and prevents data leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832406A_ABST
    Figure CN120832406A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent query method and system for the tobacco industry, and the method comprises the steps: carrying out the keyword extraction of a natural language question in response to the natural language question inputted by a user side, and converting the natural language question into a question vector; integrating the tobacco related metadata, converting the tobacco related metadata into a natural language vector, and generating a cue word containing key metadata information based on the integrated metadata and the natural language vector; a query optimizer is used for processing the problem vectors, the keywords and the natural language vectors, normalized query information is obtained, and the query information comprises problems, entity names and associated tables and fields; processing the cue word and the normalized query information by using a large language model to generate a corresponding SQL query statement; querying a database based on the SQL query statement to obtain a corresponding query result; according to the invention, efficient and accurate query of the tobacco database is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of database intelligent query, in particular to a tobacco industry intelligent query method and system. BACKGROUND

[0002] With the deepening of the informationization construction of the tobacco industry, enterprises have accumulated massive structured data covering the whole life cycle of planting, sales, monopoly, logistics, etc. In the digital management of the tobacco industry, the traditional data query method relies on structured query language SQL or fixed report forms, and the following problems still exist:

[0003] (1) Low utilization rate of metadata: The data dictionary accumulated by the tobacco industry (such as `manager_name` indicating "customer manager name") is not effectively used for SQL generation, resulting in query logic errors;

[0004] (2) Insufficient model industry adaptation: General large language models (such as Qwen, DeepSeek, etc.) are not optimized for the business logic and data characteristics of the tobacco industry, and are inefficient in handling complex queries, and the accuracy of the generated SQL statements is not high;

[0005] (3) Deviation of industry semantic analysis: There is a lack of standardized mapping between tobacco terminology (such as "low-tar cigarette", "monopoly license", etc.) and database fields, and general natural language processing models (NLP) are difficult to accurately analyze;

[0006] (4) Coarse-grained permission control: Sensitive data (such as customer mobile phone number, ID number) is only controlled by role library table access, lacking column-level / row-level fine-grained filtering, and there is a risk of data leakage.

[0007] For example, in the related art, the patent application document with publication number CN119088927A proposes to build a large language model and a knowledge retrieval platform to realize intelligent question answering of tobacco industry professional knowledge based on large language, but this scheme focuses on the retrieval of text information of tobacco industry knowledge, and the retrieved text is fed back as the result, but does not perform standardized and normalized processing on the input information, and the accuracy of the query result is not high. The patent application document with publication number CN116010439A proposes to use a metadata maintenance system to manage and edit database tables and fields according to business requirements, but the use of metadata in this scheme needs to specify the data table used before querying, and it only manages the data. SUMMARY

[0008] The technical problem to be solved by the present application is how to efficiently and accurately query the tobacco database.

[0009] The present application solves the above technical problems through the following technical means:

[0010] A tobacco industry intelligent query method is proposed, which comprises the following steps:

[0011] In response to a natural language question input by a user, keywords of the natural language question are extracted, and the natural language question is converted into a question vector;

[0012] After integrating tobacco-related metadata, the metadata is converted into a natural language vector, and based on the integrated metadata and the natural language vector, a prompt word containing key metadata information is generated;

[0013] The question vector, the keywords and the natural language vector are processed by a query optimizer to obtain normalized query information, the query information including a question, an entity name and associated tables and fields;

[0014] The prompt word and the normalized query information are processed by a large language model to generate a corresponding SQL query statement;

[0015] The database is queried based on the SQL query statement to obtain a corresponding query result.

[0016] Further, in response to a natural language question input by a user, keywords of the natural language question are extracted, and the natural language question is converted into a question vector, which comprises:

[0017] The natural language question is subjected to keyword extraction to obtain keywords related to the question;

[0018] The natural language question is converted into a question vector by using a first Embedding model.

[0019] Further, after integrating tobacco-related metadata, the metadata is converted into a natural language vector, and based on the integrated metadata and the natural language vector, a prompt word containing key metadata information is generated, which comprises:

[0020] The tobacco-related metadata is integrated to obtain integrated metadata, the metadata including the names of analysis systems and libraries, tables and fields and their remarks;

[0021] The integrated metadata is converted into a natural language vector by using a second Embedding model;

[0022] The integrated metadata and the natural language vector are processed by a prompt word generation engine to generate a prompt word containing key metadata information.

[0023] Further, the query optimizer is used to process the problem vector, the keyword and the natural language vector to obtain normalized query information, the query information including the problem, the entity name and the associated table and field, comprising:

[0024] The natural language vector is filtered, and the table and field associated with the problem vector and the keyword are screened from the filtered natural language vector;

[0025] The problem vector and the keyword are normalized to obtain the normalized problem vector and the entity name.

[0026] Further, the large language model is used to process the prompt word and the normalized query information to generate a corresponding SQL query statement, comprising:

[0027] The large language model generates a corresponding SQL query statement for the normalized query information under the guidance of the prompt word.

[0028] Further, after the large language model is used to process the prompt word and the normalized query information to generate a corresponding SQL query statement, the method further comprises:

[0029] The SQL query statement is self-corrected to obtain a correct SQL query statement.

[0030] Further, the SQL query statement is used to query the database to obtain a corresponding query result, comprising:

[0031] The SQL query statement is sent to an AnalyticDB database by using a large form engineering, so that the SQL query statement is executed by the database to obtain a corresponding query result.

[0032] Further, after the SQL query statement is used to query the database to obtain a corresponding query result, the method further comprises:

[0033] The query result is returned to the large language model, and the query result is arranged and formatted by the large language model and then returned to the user end.

[0034] Further, the large language model adopts any one of Qwen, DeepSeek and chatGLM.

[0035] In addition, the present application also proposes a tobacco industry intelligent query system, comprising:

[0036] A problem preprocessing module is configured to extract keywords from a natural language problem input by a user end and convert the natural language problem into a problem vector.

[0037] The metadata processing module is configured to convert the integrated tobacco-related metadata into a natural language vector, and generate prompt words containing key metadata information based on the integrated metadata and the natural language vector.

[0038] The query optimization module is configured to process the question vector, the keywords and the natural language vector by using a query optimizer to obtain normalized query information, which includes the question, the entity name and the associated table and field.

[0039] The query statement generation module is configured to process the prompt words and the normalized query information by using a large language model to generate a corresponding SQL query statement.

[0040] The query module is configured to query the database based on the SQL query statement to obtain a corresponding query result.

[0041] The present application has the following advantages:

[0042] (1) The present application integrates various types of metadata in the tobacco industry, and performs vectorization operation on the integrated metadata to convert it into a natural language vector, so that the metadata can be matched and operated with the question vector representation, so as to map the natural language question with the database to obtain more industry semantic analysis query information. In addition, when generating the SQL query statement by using the large language model, the prompt words containing key metadata information generated by the metadata are combined to effectively use the tobacco industry for the generation of the SQL query statement, so as to ensure the correctness of the query logic and avoid the problems of low efficiency in processing complex queries and low accuracy of the generated SQL statement when the general large language model is not optimized for the business logic and data characteristics of the tobacco industry. Finally, the database is queried based on the accurate SQL query statement to obtain the correct query result. Moreover, the business personnel do not need to have professional SQL knowledge, but only need to use natural language to query the database, which greatly shortens the query time and improves the work efficiency.

[0043] (2) By establishing the mapping relationship between the tobacco industry professional terms and the database fields, and using the metadata for semantic calculation, the system can more accurately understand the query intention of the user and generate a SQL statement that is more in line with the demand.

[0044] (3) The fine-grained permission control mechanism ensures that only authorized users can access specific data, effectively preventing data leakage and abuse, and ensuring the security of enterprise data.

[0045] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS

[0046] Fig. 1 is a flowchart of a tobacco industry intelligent query method according to an embodiment of the present application;

[0047] Fig. 2 is a schematic diagram of a tobacco industry intelligent query principle according to an embodiment of the present application;

[0048] Fig. 3 is a structural diagram of a tobacco industry intelligent query system according to an embodiment of the present application. DETAILED DESCRIPTION

[0049] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in a clear and complete manner in conjunction with the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0050] As shown in Figs. 1-2 , the present application provides a tobacco industry intelligent query method, which comprises the following steps:

[0051] S10, in response to a natural language question input by a user terminal, extracting keywords from the natural language question, and converting the natural language question into a question vector;

[0052] It should be noted that the present embodiment converts the natural language question into a vector form, which facilitates subsequent semantic calculation and processing.

[0053] S20, after integrating tobacco-related metadata, converting the metadata into a natural language vector, and generating prompt words containing key metadata information based on the integrated metadata and the natural language vector;

[0054] It should be noted that the present embodiment integrates metadata information related to the tobacco industry, and associates a large single table covering all query information, so as to reduce the complexity of generating SQL statements subsequently.

[0055] S30, using a query optimizer to process the question vector, the keywords and the natural language vector, to obtain normalized query information, wherein the query information comprises a question, an entity name and associated tables and fields;

[0056] It should be noted that by converting the metadata into a natural language vector, the metadata can be matched and operated with the question vector representation, establishing a direct mapping of tobacco business terminology and database fields, avoiding ambiguity resolution. And by using metadata, the required data table, field, etc. information can be automatically selected according to the processed information of the question content, and the flexibility is higher.

[0057] S40, using a large language model to process the prompt word and the normalized query information, and generating a corresponding SQL query statement;

[0058] It should be noted that in the embodiment, the large language model generates query information under the guidance of the prompt word, and by converting the data dictionary and business calculation logic into executable query rules, it assists in SQL query statement generation, ensures correct query logic, avoids the problem that the general large language model is not optimized for the business logic and data characteristics of the tobacco industry, and is inefficient in processing complex queries, and the generated SQL statement is not accurate. Therefore, based on the accurate SQL query statement, the database is queried to obtain the correct query result. Therefore, the embodiment focuses more on taking industry knowledge as an example, and normalizing and standardizing the input information to improve the accuracy of SQL statement production.

[0059] S50, based on the SQL query statement, querying the database to obtain the corresponding query result.

[0060] As a further preferred technical solution, the step S10: in response to the natural language question input by the user, keyword extraction is performed on the natural language question, and the natural language question is converted into a question vector, specifically including the following steps:

[0061] S11, keyword extraction is performed on the natural language question to obtain keywords related to the question;

[0062] S12, using a first Embedding model to convert the natural language question into a question vector.

[0063] It should be noted that the natural language question input by the user is preprocessed, and existing methods such as TF-IDF, TextRank, word2vec, etc. can be used to extract keywords, entities, etc. information in the natural language question, for example, for the question "How is the cigarette sales situation in Hefei?", Extracting "Hefei", "sales situation" and other key content from it.

[0064] As a further preferred technical solution, the step S20: after integrating the tobacco-related metadata, it is converted into a natural language vector, and based on the integrated metadata and the natural language vector, a prompt word containing key metadata information is generated, specifically including the following steps:

[0065] S21, integrating tobacco-related metadata to obtain integrated metadata, the metadata including the names of analysis system and library, table, field and their remarks;

[0066] It should be noted that by comprehensively integrating the analysis system (covering indicators, dimensions and the like) and the metadata information such as the names and remarks of the library, table and field (for example, sorting out the sales indicators, regional dimensions and the like in the tobacco sales data, and the specific meanings of the fields in each library table), a large single table covering all query information is obtained.

[0067] S22, converting the integrated metadata into a natural language vector by using a second Embedding model;

[0068] S23, generating a prompt word containing key metadata information by using a prompt word generation engine on the integrated metadata and the natural language vector.

[0069] It should be noted that the embodiment specifically generates a prompt word containing key metadata information by using the PromptGenerator prompt word generation engine according to the processed metadata and natural language vector information. For example, according to the tobacco sales-related metadata, a prompt word containing key metadata such as "query the sales field data in the Hefei tobacco sales table" is generated, which is used to guide the subsequent large language model to generate an accurate SQL query statement.

[0070] As a further preferred technical solution, the step S30: processing the question vector, the keyword and the natural language vector by using a query optimizer to obtain normalized query information, the query information including the question, the entity name and the associated table and field, specifically including the following steps:

[0071] S31, filtering the natural language vector, and screening the table and field associated with the question vector and the keyword from the filtered natural language vector;

[0072] Specifically, the embodiment inputs the question in the form of a vector, the metadata in the form of a vector and the keyword and the like into the QueryOptimizer question optimizer, filters out invalid or interfering information by filtering the metadata, carries out permission judgment to determine the access permission of the user to the related database table and field according to the user role, and executes table name judgment to determine the database table involved.

[0073] The embodiment realizes precise access control of data by embedding a fine-grained permission control mechanism in the natural language processing and SQL generation process, and guarantees data security.

[0074] S32, normalize the problem vector and the keyword to obtain a normalized problem vector and an entity name.

[0075] It should be noted that the embodiment normalizes the entity name, such as standardizing "Hefei" to "Hefei City". Finally, the normalized question, entity name, and associated table and field information are output.

[0076] The embodiment integrates and manages various metadata of the tobacco industry, matches and associates the metadata and the problem vector through vectorization processing, and provides rich information support for SQL query statement generation.

[0077] As a further preferred technical solution, the step S40: processing the prompt word and the normalized query information by using a large language model to generate a corresponding SQL query statement, comprising:

[0078] Using a large language model to generate a corresponding SQL query statement for the normalized query information under the guidance of the prompt word.

[0079] It should be noted that the prompt word containing key metadata information is input into a large language model (such as Qwen, DeepSeek, etc.). The large language model generates a corresponding SQL statement according to the input prompt word, combining its language understanding and generation capabilities. For example, according to the above prompt word, a SQL statement such as "SELECT sales FROM Hefei City Tobacco Sales Table" is generated.

[0080] As a further preferred technical solution, after the step S40: processing the prompt word and the normalized query information by using a large language model to generate a corresponding SQL query statement, the method further comprises:

[0081] The SQL query statement is subjected to self-correction processing to obtain a correct SQL query statement.

[0082] It should be noted that by self-checking and correcting the generated SQL statement, possible syntax errors, semantic logic errors, and other problems are checked out and corrected accordingly, ensuring that the SQL statement is accurate and can correctly perform the query operation.

[0083] As a further preferred technical solution, the step S50: querying the database based on the SQL query statement to obtain a corresponding query result, comprising:

[0084] The SQL query statement is sent to an AnalyticDB database by using a large form engineering to execute the SQL query statement by the database to obtain a corresponding query result.

[0085] It should be understood that in actual application, the self-corrected SQL statement is sent to the AnalyticDB database, the database executes the SQL query statement, and the corresponding query result is obtained.

[0086] It should be noted that the LLM2SQL (Large Language Model to SQL) core algorithm module utilizes the powerful capabilities of the large language model DeepSeek, generates accurate SQL query statements according to the problem and metadata information, and performs self-correction to ensure the correctness and compliance of the statements.

[0087] As a further preferred technical solution, after the step S50 of querying the database based on the SQL query statement to obtain the corresponding query result, the method further comprises:

[0088] Returning the query result to the large language model for processing and formatting by the large language model and returning to the user end.

[0089] It should be noted that the large table project sends the generated SQL query statement to the database for query, and returns the result to the user, and the entire process realizes automatic conversion from natural language to SQL query, improving the query efficiency and accuracy.

[0090] In addition, as Fig. 3 shown, another embodiment of the present application also proposes a tobacco industry intelligent query system, which comprises:

[0091] A question preprocessing module 10 is configured to extract keywords from the natural language question input by the user end and convert the natural language question into a question vector.

[0092] A metadata processing module 20 is configured to integrate tobacco-related metadata and convert the integrated metadata into a natural language vector, and generate prompt words containing key metadata information based on the integrated metadata and the natural language vector.

[0093] A query optimization module 30 is configured to process the question vector, the keywords and the natural language vector using a query optimizer to obtain normalized query information, wherein the query information includes the question, the entity name and the associated table and field.

[0094] A query statement generation module 40 is configured to process the prompt words and the normalized query information using a large language model to generate a corresponding SQL query statement.

[0095] A query module 50 is configured to query the database based on the SQL query statement to obtain the corresponding query result.

[0096] As a further preferred technical solution, the problem preprocessing module 10 specifically includes:

[0097] A keyword extraction unit, configured to extract keywords from the natural language question to obtain keywords related to the question;

[0098] The first vector conversion unit is used to convert the natural language question into a question vector by using a first Embedding model.

[0099] As a further preferred technical solution, the metadata processing module 20 specifically includes:

[0100] A data integration unit, configured to integrate tobacco-related metadata to obtain integrated metadata, wherein the metadata includes the names of analysis systems and libraries, tables, and fields and their notes;

[0101] A second vector conversion unit, configured to convert the integrated metadata into a natural language vector using a second Embedding model;

[0102] The prompt sub-generation unit is used to generate prompt words containing key metadata information from the integrated metadata and natural language vectors using a prompt word generation engine.

[0103] As a further preferred technical solution, the query optimization module 30 includes:

[0104] a fine-grained processing unit, configured to filter the natural language vectors and select tables and fields associated with the question vectors and keywords from the filtered natural language vectors;

[0105] The normalization unit is used to normalize the question vector and keywords to obtain a normalized question vector and entity name.

[0106] As a further preferred technical solution, the query statement generating module 40 is specifically configured to generate corresponding SQL query statements for the normalized query information using a large language model under the guidance of prompt words.

[0107] As a further preferred technical solution, the query module 50 is specifically used to: use a large form project to send the SQL query statement to the AnalyticDB database, so that the database executes the SQL query statement and obtains corresponding query results.

[0108] As a further preferred technical solution, the system further includes an error correction module for performing self-correction processing on the SQL query statement to obtain a correct SQL query statement.

[0109] It should be noted that other embodiments or specific implementations of the tobacco industry intelligent query system described herein can refer to the above-mentioned method embodiments, and will not be repeated here.

[0110] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a list of executable instructions for implementing logic functions, which can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus or device, such as a computer-based system, a system including a processor or other system that can fetch the instructions from an instruction execution system, apparatus or device and execute the instructions, or in conjunction with these instructions execution system, apparatus or device. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transport programs for use by or in connection with an instruction execution system, apparatus or device, or in conjunction with these instruction execution systems, apparatus or devices. More specific examples (non-exhaustive list) of computer-readable medium include the following: electrical connections having one or more wires (electronic devices), portable computer diskettes (magnetic devices), random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memories), fiber optic devices, and portable compact disk read-only memories (CDROMs). In addition, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, as the program can be electronically obtained, for example, by optical scanning of the paper or other medium, followed by electronic conversion, interpretation or processing, if necessary, in other suitable ways, and then stored in a computer memory.

[0111] It should be understood that parts of the present application can be implemented in hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, and as in another embodiment, it can be implemented by any one or a combination of the following technologies known in the art: discrete logic circuit with logic gate circuit for implementing logic functions on data signals, application specific integrated circuit with suitable combination logic gate circuit, programmable gate array (PGA), field programmable gate array (FPGA) and the like.

[0112] In the description of the specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the description of the specification, the illustrative description of the above terms does not necessarily mean the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0113] In addition, the terms "first", "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the technical features indicated. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "a plurality of" is at least two, for example, two, three, etc., unless otherwise specifically limited.

[0114] Although the embodiments of the present application have been shown and described above, it is understood that the above-described embodiments are exemplary and cannot be construed as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above-described embodiments within the scope of the present application.

Claims

1. A tobacco industry intelligent query method, characterized in that, The method comprises the following steps: In response to a natural language question input by a user, keyword extraction is performed on the natural language question, and the natural language question is converted into a question vector; After integrating tobacco-related metadata, the metadata is converted into a natural language vector, and a prompt word containing key metadata information is generated based on the integrated metadata and the natural language vector; The question vector, the keyword, and the natural language vector are processed by a query optimizer to obtain normalized query information, wherein the query information includes a question, an entity name, and associated tables and fields; The prompt word and the normalized query information are processed by a large language model to generate a corresponding SQL query statement; A database is queried based on the SQL query statement to obtain a corresponding query result.

2. The tobacco industry intelligence query method of claim 1, wherein, The method comprises the following steps: The natural language question is subjected to keyword extraction to obtain keywords related to the question; The natural language question is converted into a question vector by using a first Embedding model.

3. The tobacco industry intelligence query method of claim 1, wherein, The method comprises the following steps: The tobacco-related metadata is integrated to obtain integrated metadata, wherein the metadata includes the names of analysis systems and libraries, tables, and fields as well as their remarks; The integrated metadata is converted into a natural language vector by using a second Embedding model; The integrated metadata and the natural language vector are processed by a prompt word generation engine to generate a prompt word containing key metadata information.

4. The tobacco industry intelligence query method of claim 1, wherein, The method comprises the following steps: The natural language vector is subjected to filtering processing, and tables and fields associated with the question vector and the keyword are screened from the filtered natural language vector; The question vector and the keyword are subjected to normalization processing to obtain a normalized question vector and an entity name.

5. The tobacco industry intelligence query method of claim 1, wherein, The method comprises the following steps: The large language model is used to generate a corresponding SQL query statement for the normalized query information under the guidance of the prompt word.

6. The tobacco industry intelligence query method of claim 1, wherein, After the large language model processes the prompt word and the normalized query information to generate a corresponding SQL query statement, the method further comprises the following steps: The SQL query statement is subjected to self-correction processing to obtain a correct SQL query statement.

7. The tobacco industry intelligence query method of claim 1, wherein, The method comprises the following steps: The SQL query statement is sent to an AnalyticDB database by using a large form engineering to execute the SQL query statement by the database and obtain a corresponding query result.

8. The tobacco industry intelligence query method of claim 1, wherein, After the database is queried based on the SQL query statement to obtain a corresponding query result, the method further comprises the following steps: The query result is returned to the large language model, and after the large language model processes the query result, the query result is returned to the user end.

9. The tobacco industry intelligence query method of claim 1, wherein, The large language model adopts any one of Qwen, DeepSeek and chatGLM.

10. A tobacco industry intelligent query system, characterized in that, Comprise: A question preprocessing module is configured to extract keywords from a natural language question input by a user end and convert the natural language question into a question vector; A metadata processing module is configured to integrate tobacco-related metadata into a natural language vector and generate prompt words containing key metadata information based on the integrated metadata and the natural language vector; A query optimization module is configured to process the question vector, the keywords and the natural language vector using a query optimizer to obtain normalized query information, wherein the query information includes a question, an entity name and associated tables and fields; A query statement generation module is configured to process the prompt words and the normalized query information using a large language model to generate corresponding SQL query statements; A query module is configured to query a database based on the SQL query statements to obtain corresponding query results.

Citation Information

Patent Citations

  • Visual Chinese SQL (Structured Query Language) system and query construction method

    CN116010439A

  • Intelligent question answering system and question answering method for professional knowledge of tobacco industry based on big language

    CN119088927A