Data analysis question and answer platform based on large model and knowledge vector library

Through the data analysis question and answer platform of large models and knowledge vector library, the semantic confusion and security problems of NL2SQL in enterprise-level databases are solved, and efficient, accurate and flexible data analysis question and answers are achieved, ensuring the security and user experience of data processing.

CN120448508AActive Publication Date: 2025-08-08LU ZE TECH CO LTD

Patent Information

Application Number
CN202510940236.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-08-08
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

The existing NL2SQL technology has semantic confusion, high learning costs, uncontrollable performance, data privacy and permissions in enterprise-level databases, resulting in inaccurate answers, low data processing efficiency, and inability to understand complex problems.

Method used

Using a data analysis question and answer platform based on big models and knowledge vector library, we analyze user question and answer content through big models, filter data analysis questions, identify keywords and classify them into indicators, dimensions and statistical cycles, generate SQL query statements, and display the results through the front-end user interface.

Benefits of technology

It realizes natural language data analysis and Q&A with low cost, high efficiency, high security, high accuracy and high flexibility, solves the accuracy and efficiency problems of traditional platforms and ensures the security and flexibility of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448508A_ABST
    Figure CN120448508A_ABST
Patent Text Reader

Abstract

The invention discloses a data analysis question and answer platform based on a large model and a knowledge vector library, and relates to the field of intelligent question and answer. Firstly, question and answer content of a user is obtained, the question and answer content is analyzed through a large model, and a data analysis type question is screened out; then identifying keywords in the data analysis type question; keyword classification is carried out through cooperation of a large model and a knowledge vector library, and keywords are classified into indexes, dimensions and statistical periods; the indexes, the dimensions and the statistical periods are converted into corresponding SQL fragments, and the SQL fragments are spliced to generate an SQL query statement; executing the SQL query statement to query the database, and returning a query result; and the query result is converted into structured data, and the structured data is displayed through a front-end user interface, so that natural language data analysis and question answering with low cost, high efficiency, high safety, high accuracy and high flexibility can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent question-answering technology, and in particular to a data analysis question-answering platform based on a large model and a knowledge vector library. Background Art

[0002] NL2SQL (Natural Language to SQL) is a technology that automatically converts natural language questions into SQL (Structured Query Language). Users can ask questions in everyday language, and the system translates these questions into precise SQL queries. For example, a user might enter "Show me the top five products with the highest sales in the past month." The NL2SQL system will parse this natural language and generate the corresponding SQL query.

[0003] The existing NL2SQL system directly performs lexical, syntactic, and semantic analysis on input natural language queries. It then matches keywords with the database's built-in table structure (schema) to generate SQL query statements. However, enterprise databases often contain thousands or tens of thousands of tables, and the large number of schemas can easily lead to semantic confusion. Furthermore, large models require a high learning curve for table structures. Unoptimized SQL directly written by large models can lead to a high rate of slow queries and uncontrollable performance. Furthermore, direct database operations by large models raise data privacy and permission issues. Summary of the Invention

[0004] The purpose of this application is to provide a data analysis and question-answering platform based on a large model and knowledge vector library, which can achieve low-cost, high-efficiency, high-security, high-accuracy and high-flexibility natural language data analysis and question-answering.

[0005] To achieve the above objectives, the present application provides a data analysis question-answering platform based on a large model and a knowledge vector library, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a data analysis question-answering method based on a large model and a knowledge vector library; The data analysis question answering method based on the large model and knowledge vector library includes: Obtain user Q&A content, analyze it using a large model, and filter out data analysis questions; Identify key words in data analysis questions; Use a large model with a knowledge vector library to classify keywords into indicators, dimensions, and statistical periods; Convert indicators, dimensions, and statistical periods into corresponding SQL fragments, and then concatenate the SQL fragments to generate SQL query statements; Execute SQL query statements to query the database and return the query results; Convert query results into structured data and display them through the front-end user interface.

[0006] Optionally, obtaining the user's question and answer content, and analyzing the question and answer content using a large model to screen out data analysis questions, specifically includes: Obtain user Q&A content and analyze it using a large model to determine whether the Q&A content is a data analysis question; the large model includes DeepSeek and OpenAI; data analysis questions include: questions involving datasets or data sources, inquiring about trends, patterns, predictions or statistical analysis, requesting charts, visualizations or data interpretation, and involving specific statistical methods or analytical techniques; If the question and answer content is a data analysis question, directly output the data analysis question; If the question and answer content is a non-data analysis question, the user will be prompted on the front-end user interface to rewrite the non-data analysis question into a data analysis question.

[0007] Optionally, identifying keywords in data analysis questions specifically includes: Perform text preprocessing on data analysis problems, including segmenting text into words or phrases and removing common stop words to obtain preprocessed candidate words ; Parse each candidate word The word statistical features and semantic information, and according to the formula Calculate each candidate word Total score ;in Candidate word Case feature score; Candidate word The position feature score of Candidate word The word frequency feature score of Candidate word The symbolic feature score of Candidate word Contextual relevance score of Candidate word The cross-sentence distribution feature score of The number of characters for data analysis questions; Candidate word The contextual pattern enhances the feature score; The total score Candidate words below the preset score threshold Determine as keyword .

[0008] Optionally, the candidate word The symbolic feature score of The calculation formula is: ;in, is the number of symbol types in the candidate word; For the weights of class symbols; is the symbol combination gain factor; is the position attenuation factor.

[0009] Optionally, the candidate word Contextual relevance score The calculation formula is: ;in for Left and right neighbor words; point mutual information score ; yes and Joint probability of simultaneous occurrence; and They are and The probability of a single occurrence.

[0010] Optionally, the candidate word Contextual pattern enhancement feature score The calculation formula is: ;in, Score for exact matches; Score for fuzzy matching; is the dynamic weight score; 、 and They are 、 and The weight coefficient of .

[0011] Optionally, the keyword classification is performed by using a large model in conjunction with a knowledge vector library, and the keywords are classified into indicators, dimensions, and statistical periods, specifically including: By constructing prompt words, the large model can make the first round of predictions for keywords. Perform coarse-grained classification and generate The coarse screening candidate categories are dimensions, indicators or statistical periods; Calculation keywords The cosine similarity with the standard words of each category in the knowledge vector library is used to select the category with the highest cosine similarity as the candidate category for fine screening; the candidate category for fine screening is a dimension, an indicator or a statistical period; If the fine screening candidate category is consistent with the coarse screening candidate category, determine the fine screening candidate category or the coarse screening candidate category as the keyword The final category of If the fine-screened candidate category is inconsistent with the coarse-screened candidate category, then determine whether the cosine similarity corresponding to the fine-screened candidate category exceeds the similarity threshold; If the cosine similarity of the fine-screened candidate category exceeds the similarity threshold, the fine-screened candidate category is determined to be a keyword. The final category; otherwise, the coarse screening candidate category is determined as the keyword The final category of The keywords The final category output is a JSON message in JSON format; the JSON message includes dimension keywords, indicator keywords and statistical period keywords.

[0012] Optionally, converting the indicators, dimensions, and statistical periods into corresponding SQL fragments and concatenating the SQL fragments to generate SQL query statements specifically includes: Based on the large model, the intent of the JSON message is recognized and the dimension keywords, indicator keywords, and statistical period keywords included in the JSON message are matched to the corresponding SQL fragments. Splice the SQL fragments to generate the corresponding SQL query statement.

[0013] Optionally, executing an SQL query statement to query a database and returning query results specifically includes: The large model accesses an external database tool set, executes SQL query statements, drives the SQL query tool to query in the database, and returns the data query results.

[0014] Optionally, converting the query results into structured data and displaying it through a front-end user interface specifically includes: Convert query results into structured data through semantic or structural means, and open it to the front-end user interface through the API interface; The front-end user interface renders the structured data into charts for display.

[0015] According to the specific embodiments provided in this application, this application discloses the following technical effects: The present application provides a data analysis and question-answering platform based on a large model and a knowledge vector library, which can solve problems such as inaccurate answers, low data processing efficiency, and inability to understand complex problems in traditional data analysis and question-answering platforms, meet users' needs for efficient and accurate data analysis and question-answering, and achieve low-cost, high-efficiency, high-security, high-accuracy and high-flexibility natural language data analysis and question-answering. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0017] Figure 1 This is a flowchart of a method implemented by a data analysis question-answering platform based on a large model and a knowledge vector library in this application. DETAILED DESCRIPTION

[0018] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0019] The purpose of this application is to propose a data analysis question-answering platform based on a large model and a knowledge vector library, aiming to achieve low-cost, high-efficiency, high-security, high-accuracy and high-flexibility natural language data analysis and question-answering.

[0020] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0021] This application provides a data analysis question-and-answer platform based on a large model and a knowledge vector library. The data analysis question-and-answer platform based on a large model and a knowledge vector library includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program to implement a data analysis question-and-answer method based on a large model and a knowledge vector library. The data analysis question-and-answer method based on a large model and a knowledge vector library includes the following steps 1 to 6.

[0022] Step 1: Obtain the user's question and answer content, analyze it through a large model, and filter out data analysis questions.

[0023] The data analysis Q&A platform in this application first obtains user questions and answers written in natural language, analyzes them using a trained large model, and passes data analysis questions to step 2. It then uses labeled prompts and guiding prompts to encourage users to rewrite non-data analysis questions into data analysis questions. The goal is to only allow data analysis questions to enter the entire data analysis process, while non-data analysis questions will be rejected.

[0024] The step 1 specifically includes the following steps 1.1 to 1.3.

[0025] Step 1.1: Obtain the user's Q&A content and analyze it using a large model to determine whether the Q&A content is a data analysis question.

[0026] The large language model (LLM) used in this application can be any of the various large language models currently available on the market, such as DeepSeek and OpenAI. The data analysis questions are those involving data sets or data sources, asking about trends, patterns, predictions, or statistical analysis, requesting charts, visualizations, or data interpretation, and involving specific statistical methods or analytical techniques. Taking e-commerce scenarios as an example, data analysis questions are similar to questions such as "Please tell me about the December sales of this store," "Please tell me the top 10 sales of this year's products and display them in a bar chart," and "Which product do you predict will have the worst sales this quarter?" used for operational analysis and store sales analysis.

[0027] Step 1.2: If the question and answer content is a data analysis question, directly output the data analysis question.

[0028] If the large model determines that the user's question and answer content is a data analysis question, it will directly output the data analysis question to step 2. In this way, non-data analysis questions are blocked from the entire data analysis question and answer platform, and only data analysis questions are answered.

[0029] Step 1.3: If the question and answer content is a non-data analysis question, the user is prompted on the front-end user interface to rewrite the non-data analysis question into a data analysis question.

[0030] If the large model determines that the user's question and answer content is a non-data analysis question, it will use labeled prompts and guiding prompts to allow the user to rewrite the non-data analysis question into a data analysis question.

[0031] The labeled prompts provide some question types and their definitions, including data extraction, statistical analysis, trend prediction, data transformation, and data modeling. Based on these question type definitions, the big model determines the type of question. If the question falls into one of these categories, the user's question is considered a data analysis question.

[0032] For example, the prompt template is as follows. Statistical analysis problems are defined as problems that require the collection, organization, analysis, interpretation, and presentation of data to solve or understand. For example, how much impact does a promotion have on sales? A large model can identify existing problems based on definitions and examples.

[0033] Guidance prompts provide a description that allows the model to compare the question to question types and definitions. For example, some question types and definitions, along with the user's input question, are provided in advance. The model is then trained to respond with, "The above description represents a data analysis question. Please determine whether the current input question is a data analysis question." If not, the user is directly informed.

[0034] These two types of prompts (labeled prompts and guiding prompts) are combined to form a prompt template, which can be used for large models, as shown below.

[0035] Statistical analysis questions: These are questions that require solving or understanding problems by collecting, organizing, analyzing, interpreting, and presenting data. For example, "How much impact does a promotion have on sales?" {User-entered question}. The above description represents a data analysis question. Please determine whether the question is a data analysis question. If not, directly inform the user that it is not a data analysis question.

[0036] Step 2: Identify key words in data analysis questions.

[0037] Data analysis problems require a large model to convert natural semantics into out-of-warehouse semantics, that is, to convert a piece of natural language into a formatted message that a computer can understand to transmit information. According to the data warehouse modeling theory and the construction theory based on the modern data middle platform, natural semantics are broken down into three layers of keywords: indicators, dimensions and statistical periods. Among them, indicators refer to quantifiable measurements used to measure business performance. Dimensions refer to attributes used to describe and classify indicators. Statistical period refers to the time interval for aggregating indicators. Taking the question of asking about the monthly sales of this store as an example, step 2 is used to identify the keywords in the question, here is {store / month / sales}, where the store is the dimension, the month is the statistical period, and the indicator is sales. This application identifies keywords by analyzing the statistical characteristics and semantic information of words in the text. The step 2 specifically includes the following steps 2.1 to 2.3.

[0038] Step 2.1: Perform text preprocessing on data analysis problems, including segmenting the text into words or phrases and removing common stop words, to obtain multiple candidate words after preprocessing, represented as .

[0039] For example, for the data analysis question "I want to know the store sales in the past six months", first split the sentence into words or phrases. Chinese word segmentation usually requires the help of word segmentation tools such as jieba. For this sentence, the word segmentation result may be: I / want / to know / this / past six months / of / store / sales. Stop words are usually those words that appear frequently in the text but contribute little to the semantics, such as "of", "is", "in", etc. After removing these stop words, the remaining candidate words are: past six months / store / sales. Among them, "past six months" represents the time range, which is the key time condition for the query, corresponding to the keyword category of "statistical period". "Store": represents the main object of the query, corresponding to the keyword category of "dimension". "Sales": represents the specific data indicator for the query, corresponding to the keyword category of "indicator".

[0040] Step 2.2: Analyze the word statistical features and semantic information of each candidate word and calculate the total score of each candidate word according to formula (1) The lower the total score of a candidate word, the more likely it is to be a keyword such as an indicator, dimension, or statistical period.

[0041] (1); where is the case feature score of the candidate word ; is the position feature score of the candidate word ; is the word frequency feature score of the candidate word ; is the symbol feature score of the candidate word ; is the context relevance score of the candidate word ; is the cross-sentence distribution feature score of the candidate word ; is the number of characters in the data analysis question is the context pattern enhancement feature score of the candidate word ;

[0042] For example, for the data analysis question "Please tell me the top 10 product sales this year and display them in the form of a bar chart", first split the sentence into words or phrases. The word segmentation result may be: Please / tell / me / this year / product / sales / top10 / and / in the form of / a bar chart / display. Common Chinese stop words include "Please", "tell", "me", "and", "in the form of", "display", etc. After removing these stop words, the remaining candidate words are: this year / product / sales / top10 / bar chart. Calculate the total score of each candidate word according to formula (1) ​​ For example, for the candidate words "sales" and "histogram", the total score of "histogram" is 98, and the total score of "sales" is 60. Therefore, "sales" is more likely to be an indicator, dimension or statistical period, and "sales" is identified as one of the keywords.

[0043] The following is a detailed introduction to the calculation method of the statistical features and semantic information scores of each word in formula (1).

[0044] 2.2.1 Case feature (Case): Determine whether the candidate word begins with a capital letter (such as a proper noun). Such words are more likely to be keywords. Case feature score The calculation formula is as follows: (2); The statistical methods for calculating "capitalization frequency" and "non-capitalization score" are as follows. For example, in the sentence "I want to know the store sales for the past six months," the candidate word "Sales (Revenue)" begins with a capital letter. If it appears once, its "capitalization frequency" is 1, if it appears twice, its "capitalization frequency" is 2, and so on. A "non-capitalization score" of 1 actually corresponds to a "capitalization frequency" of 0. Because log20 is negative infinity (approximately equal to 0), 1+log20≈1+0=1.

[0045] 2.2.2 Position feature: The earlier the candidate word appears in the document, the higher its importance may be. The position feature score The calculation formula is as follows: (3).

[0046] For example, for the data analysis question "Please tell me the top 10 best-selling products this year and display them in a bar chart," the corresponding word segmentation result is: Please tell me the top 10 best-selling products this year and display them in a bar chart. Therefore, the "Total Word Count in Document" is 12. After removing stop words, the remaining candidate words are: This year / product / sales / top 10 / bar chart. Therefore, for the candidate word "This year," its "First Occurrence Position" is 4.

[0047] 2.2.3 Frequency feature: The higher the frequency of a candidate word, the more important it is, but it needs to be normalized to prevent bias in long texts. The word frequency feature score The calculation formula is as follows: (4); in, Refers to the candidate word Frequency of occurrence in a document. "Maximum frequency in document" refers to the maximum frequency of all words in the entire document. For example, in a data analysis problem, "sales" appears the most frequently, ten times, so the "maximum frequency in document" is 10.

[0048] 2.2.4 Symbol feature: Accurately identify special symbols, units or keyword suffixes (such as %, $, rate, year-on-year) in indicators, reduce their scores, and improve keyword priority. The symbolic feature score of The calculation formula is as follows: (5); in, is the number of symbol types in the candidate word; For the weights of class symbols; is the symbol combination gain factor; is the position attenuation factor. This application divides the symbols into four categories and assigns weights from high to low according to priority, as shown in Table 1 below.

[0049] Table 1 Symbol types and their corresponding weights

[0050] KPI (Key Performance Indicator) refers to key performance indicators. ROI (Return on Investment) refers to return on investment. CTR (Click-Through Rate) refers to click-through rate.

[0051] For symbol combining gain factors If a candidate word contains multiple symbol combinations, its score will be further reduced and its priority will be increased. For example, in the candidate word "year-on-year growth in number of people," "year-on-year" is a time comparison word among the four types of symbols mentioned above, and "number of people" is a unit symbol. A candidate word with multiple symbol combinations is considered a symbol combination type. For such a word, its score will be multiplied by 0.4 and then 0.3, as it is more likely to be an indicator / dimension / statistical period.

[0052] Position attenuation factor It is used to suppress the interference of symbols in long-tail words (such as invalid symbols at the end of a paragraph). The calculation formula is as follows: (6).

[0053] For example, a user might enter a question incorrectly, describing it as "Let's see what the store's growth rate is." This satisfies the weighting in symbol-based hierarchical weighting. However, "rate" is placed last in the sentence. Although it is a key symbol, it is of little use. Lowering the priority of "rate" here is mainly to eliminate the influence of cluttered text. Therefore, for effective weighting, it must first meet the requirements of valid indicators / dimensions / statistical periods. Here, it is inferred by comparing the document with the relevant indicator / dimension / statistical period training data set after word segmentation. This symbol cannot appear alone, otherwise it will not meet the requirements of combined weighting. The later the symbol is in the document, the smaller the value calculated by formula (6), and the smaller the weight.

[0054] 2.2.5 Contextual Relevance (Relatedness): Measures the degree of relevance between a candidate word and surrounding words, calculated using Pointwise Mutual Information (PMI). Contextual relevance score The calculation formula is: (7); in, for The left and right neighbor words of .

[0055] PMI is calculated as follows: (8); in, yes and Joint probability of simultaneous occurrence; and They are and Probability of single occurrence; The calculated point mutual information score (PMI) effectively measures the correlation between two terms. For example, if a question frequently mentions milk powder and GMV (Gross Merchandise Volume), then the user is more likely to be asking about milk powder's GMV.

[0056] 2.2.6 Cross-sentence distribution feature (DifSentence): The more times a candidate word appears in different sentences, the higher its importance. The cross-sentence distribution feature score of The calculation formula is: (9); Among them, "total number of sentences" refers to the total number of all sentences in the document; "number of sentences that appeared" refers to the number of sentences containing the candidate word The number of sentences.

[0057] 2.2.7 Contextual Pattern Boosting Feature (PatternBoost): Identify complex contextual patterns (such as "dimension store" or "core indicator is the number of orders") and dynamically increase the weight. For example, if a question is "Please tell me what is the core indicator of the store product views this year", since "product views" is followed by "core indicator", then "product views" is more likely to be an indicator. In other words, if a candidate word If there are clearly marked keywords before and after it that are related to dimensions or indicators, then it is more likely to be a dimension / indicator. In essence, the candidate word is judged based on the characteristics of the context. Is it a dimension / indicator? Candidate word Contextual pattern enhancement feature score The calculation formula is: (10).

[0058] in, Score for exact matches; Score for fuzzy matching; is the dynamic weight score; 、 and They are 、 and The weight coefficient of . Since the total score of formula (1) is in the denominator, so The bigger the final total score The smaller.

[0059] Among them, Exact Match refers to capturing fixed sentence patterns (such as "Indicator: [candidate word]") and matching them using regular expressions. Exact Match Score The calculation formula is as follows: (11); Among them, metric is the English word for "indicator". If the question contains descriptions such as "indicator XXX", "dimension XXX", "metricXXX", etc., then these words are clearly classified into a certain classification pattern. In this case, the candidate word XXX is very likely to be an indicator or dimension. For example, for the word "indicator order rate", the corresponding pattern is " ", then the "order rate" here is the indicator, which meets the indicator model.

[0060] Among them, fuzzy matching (FuzzyMatch) refers to identifying the guiding words in the context (such as "promote", "associate", "drive", etc.) and dynamically calculating the association weight. Fuzzy matching score The calculation formula is: (12); in, Refers to the candidate word The context window is within 3 words before and after. It is a guide word weight table, such as "drive=0.6", "enhance=0.5", and "association=0.4".

[0061] Dynamic Weight refers to dynamically adjusting the weight based on the distribution of candidate words in the document. Dynamic Weight Score The calculation formula is: (13); in, Indicates candidate words The number of times it appears in the pattern, such as the number of times "User Age" appears in the pattern "Dimension: User Age" The number of times it appears in ". Indicates candidate words The total number of occurrences in the entire document.

[0062] Step 2.3: Total score Candidate words below the preset score threshold Determine as keyword .

[0063] In formula (1), is the number of characters entered by the user for the data analysis question. Is to split the text of the question into candidate words , calculate these candidate words The lower the total score, the more likely it is to be a keyword representing an indicator / dimension / statistical period. Candidate words below a certain preset score threshold Determine as keyword .

[0064] Step 2 of this application is to first identify the keywords in the data analysis question , then in the subsequent step 3, segment and classify to determine the keywords Is it an indicator, dimension or statistical period?

[0065] Step 3: Use the big model and the knowledge vector library to classify keywords into indicators, dimensions, and statistical periods.

[0066] The step 3 specifically includes the following steps 3.1 to 3.6.

[0067] Step 3.1: Let the big model make the first round of predictions by constructing prompt words. Perform coarse-grained classification and generate The coarse screening candidate categories are dimensions, indicators or statistical periods.

[0068] First, we construct a prompt word for the large model to perform a first-round prediction, perform a coarse-grained classification of keywords, and generate candidate categories (dimensions, indicators, and statistical periods). The prompt word template is shown in Table 2 below.

[0069] Table 2 Prompt word template

[0070] In Table 2, DAU refers to Daily Active Users, and Q2 refers to Quarter 2. Table 2 shows the constructed prompt word template. The large-scale model performs classification based on the template's rules. The template provides definitions of dimensions, indicators, and statistical periods. The large-scale model infers based on these rules, and classification rules can be customized based on the scenario. For example, in the e-commerce scenario, in the phrase "2024 self-operated store sales," "sales" is the indicator, the preceding non-temporal qualifier "self-operated stores" is the dimension, and the temporal qualifier "2024" is the statistical period. If the sentence is broken down in this way, for sales, the keyword is sales, the category is the indicator, and the classification rationale is that any sentence that matches the defined content (a quantifiable numerical measurement) will be assigned to the corresponding category.

[0071] Step 3.2: Calculate keywords The cosine similarity with the standard words of each category in the knowledge vector library is used to select the category with the highest cosine similarity as the candidate category for fine screening; the candidate category for fine screening is also a dimension, an indicator or a statistical period.

[0072] This application first constructs prompt words to allow the large model to make the first round of predictions for keywords. Perform coarse-grained classification to generate coarse-screened candidate categories (dimensions, indicators, or statistical periods). Then verify and disambiguate using the knowledge vector library (abbreviated as vector library) of the business domain (such as the e-commerce domain) to prevent large-scale model classification errors. Specifically, calculate the keywords to be classified The cosine similarity with the standard words of each category in the vector library is used to select the category with the highest score as the candidate category for fine screening.

[0073] The vector library used in this application is a database that stores words in the form of vector coordinates. Business and technical personnel maintain the indicators / dimensions / statistical periods and calculation calibers used by the company in the vector library to build an off-warehouse semantic layer. This vector library is then used by the large model to classify indicators / dimensions / statistical periods, and is also used when splicing SQL fragments later. In data analysis and business statistics, "calculation caliber" is a very important concept, which refers to the specific methods, rules and standards used when calculating a certain indicator. The calculation caliber ensures the consistency and accuracy of the data, allowing data from different times and different departments to be effectively compared and analyzed.

[0074] The off-warehouse semantic layer allows business teams to directly maintain metric / dimension / statistical period definitions in a low-code manner, improving the flexibility of data analysis. Business personnel directly extract core elements of metrics, dimensions, and statistical periods, including calculation scope (through direct SQL fragmentation of the data model) and definitions. Because data analysis problems are highly logical, to facilitate the analysis of such problems, this application uses independent feature vector processing to perform word segmentation and vectorization, vectorizing the quantitatively defined keywords (metrics / dimensions / statistical periods) into a vector library for use in the large model and other steps. Before performing word segmentation and vectorization, it is necessary to first define and model the metrics, dimensions, and statistical periods, including the following 3.2.1) to 3.2.3).

[0075] 3.2.1) Indicator Definition Modeling: Model indicator-related definitions, including the indicator name (in both Chinese and English), indicator definition, and calculation caliber. The calculation caliber here refers to the SQL snippet generated based on the current data warehouse table design and field definition. This is used to generate the select clause of the SQL query statement.

[0076] An example SQL query statement is as follows: select order_id from order_detail_d where shop_id = 12345 . This query extracts the "order_id" metric data for the shop ID "12345" from the data source "order_detail_d." The select clause "select order_id" is the select clause, and the where clause "shop_id = 12345" is the where clause.

[0077] The input format of indicator definition is shown in Table 3 below.

[0078] Table 3 Input format of indicator definition

[0079] Among them, name is the Chinese name of the defined indicator measurement, col is the English name corresponding to the indicator measurement, definition is the calculation caliber definition, and calculation_basis is the SQL fragment.

[0080] 3.2.2) Dimension Definition Modeling: Model dimension-related definitions. Dimensions limit the calculation scope of indicators. Here, you need to provide the dimension name (in Chinese and English), dimension definition, and calculation caliber to ultimately generate the where clause of the SQL query statement.

[0081] 3.2.3) Statistical Period Definition Modeling: This statistical period refers to the time dimension used by various businesses, such as date, year, and month. The definition should also include the business time type used by the business, such as order placement time and return time in e-commerce scenarios, to ensure that business time corresponds to business time in the data warehouse.

[0082] Traditional vectorization methods concatenate all metadata into a single text for vectorization. For example, concatenating all metadata in Table 3 into "sales_amt total amount of goods sold SUM(per_gmv)" and directly vectorizing it results in: 1) technical features (such as field names) being overwhelmed by natural language descriptions; 2) semantic conflicts between different attributes (such as field names using snake case, which differs from the grammatical structure of Chinese descriptions); and 3) inability to optimize weights for different query types (such as being unable to focus when a user explicitly searches for a field name). Therefore, to improve retrieval efficiency and accuracy, this application splits the metadata structure into four independent feature dimensions: ① indicator name / dimension name / statistical period name, ② technical field name, ③ calculation caliber (SQL fragment), and ④ Chinese description of the calculation caliber. Each feature dimension is then independently vectorized. The word segmentation vectorization process includes the following steps 3.2.4) through 3.2.6).

[0083] 3.2.4) The vectorization process of the indicator name / dimension name / statistical period name and the Chinese description of the calculation caliber.

[0084] Model selection: BAAI / bge-base-zh-v1.5 (Chinese semantic model). This Chinese semantic model takes as input the indicator name, dimension name, statistical period name, or a Chinese description of the calculation caliber, and outputs the corresponding vector coordinates (in a mathematical coordinate system). This Chinese semantic model is used to convert Chinese characters into coordinates for matching. For example, if you input the word "sales," its position in the coordinate system is converted and matched against the position of the standard indicator name "sales" in the vector library. The closer the two are, the more likely the indicator name is sales, and thus a match is achieved.

[0085] This Chinese semantic model is based on the BERT architecture, and its optimization goal is contrastive learning. Its advantage lies in narrowing the semantic distance between business terms and descriptions through contrastive learning, making it suitable for processing Chinese natural language. Its core formula is as follows: (14); in, Represents a query vector, such as the user inputs "sales amount". Represents a positive sample, such as the standard indicator name "sales". Representative negative samples, such as other irrelevant indicators. Represents the number of negative samples. sim() represents the cosine similarity. Represents the temperature coefficient, which is used to control the steepness of the distribution. Represents the loss function of the Chinese semantic model, The smaller the value, the better the model prediction result is, and the closer the word is to the standard indicator name "sales"; The larger the value, the worse the model prediction results. The value is a measure of the accuracy of the coordinates produced by vectorization.

[0086] 3.2.5) Vectorization process of technical field names.

[0087] Model selection: Microsoft / CodeBERT-base (Code Understanding Model). This code understanding model takes as input the technical field name of a metric, dimension, or statistical period. The technical field name is the English code for the metric name and is used for database queries. For example, sales corresponds to shop_pay. The output of this code understanding model is vectorized coordinates (in a mathematical coordinate system). This code understanding model is used to convert technical field names into a coordinate system.

[0088] This code comprehension model is based on BERT's multi-task pre-training and combines two objectives: Masked Language Modeling (MLM) and Replaced Token Detection (RTD). RTD is used to detect replaced tokens in the input and enhances sensitivity to naming conventions. Its advantage is that it can understand naming patterns of technical fields, such as camel case (saleAmt) and snake case (sale_amt). The objective function of MLM is as follows: (15); in, Represents the training data set D The input sequence (code or natural language) sampled in . Mis a set of randomly masked Token positions, each position has a probability p mask was selected. Represents the masked input sequence, replaced by MASK or random Token M The position in. is the position in the original input sequence The real token. Indicates the code comprehension model predicted position for In natural language processing and programming language processing, a token is the smallest unit into which text or code is segmented. MASK is a special masking token used to represent the masked portion of the input sequence.

[0089] 3.2.6) The vectorization process of computing Calibre (SQL fragments).

[0090] Model selection: Microsoft / UniSQL (SQL-specific model). This SQL-specific model takes as input SQL fragments representing indicators, dimensions, and statistical periods and outputs corresponding vector coordinates for vectorizing SQL fragments. Its advantage lies in its ability to parse SQL syntax structures, such as the aggregate function SUM(...) and JOIN logic, and accurately match calculation logic. This SQL-specific model uses a syntax-aware encoder based on Tree-LSTM: (16).

[0091] in, Represents the current syntax tree node The hidden state of . LSTM ( ) represents the syntax-aware encoder based on Tree-LSTM. Indicates the Token embedding of a node, such as SUM and COUNT. For child nodes The hidden state of . Represents the collection of child nodes of a tree. The smaller the value, the better the model prediction result is, and the closer the predicted words are to the current SQL; The larger the value of , the worse the model prediction result. Used to measure the accuracy of coordinates produced by vectorization.

[0092] Of course, the above models and algorithms can be interchanged based on their strengths and algorithmic logic; there is no single model choice. Independently vectorized content is stored in the knowledge vector library as a master document, containing vectors of all dimensions and associated metadata to ensure cross-referencing during retrieval. The master document format is shown in Table 4 below, with comments following the symbol " / / ."

[0093] Table 4 Format structure of the main document

[0094] Among them, id is the unique ID of the document; type is the document type, indicating whether the document type is an indicator, dimension, or statistical period; the various attributes in the metadata object are the defined metadata entered by the user above; the fields of the same name in the vectors object are their corresponding vectorized coordinates.

[0095] Step 3.3: If the fine screening candidate category is consistent with the coarse screening candidate category, determine the fine screening candidate category or the coarse screening candidate category as the keyword The final category.

[0096] Step 3.4: If the fine-screened candidate categories are inconsistent with the coarse-screened candidate categories, determine whether the cosine similarity corresponding to the fine-screened candidate categories exceeds the similarity threshold (e.g., 0.7).

[0097] Step 3.5: If the cosine similarity of the fine-screened candidate category exceeds the similarity threshold (such as 0.7), the fine-screened candidate category is determined to be a keyword The final category; otherwise, the coarse screening candidate category is determined as the keyword The final category.

[0098] If the highest cosine similarity exceeds a threshold (e.g., 0.7), the vector library classification result is trusted; otherwise, the large model classification result is relied upon. This allows the vector library classification result to override any LLM classification errors. For example, if the LLM misclassifies "quarter" as "dimension," but "quarter" and "statistical period" are more similar in the vector library, and their cosine similarity exceeds 0.7, the final category for the keyword "quarter" is determined to be "statistical period."

[0099] Step 3.6: Add keywords The final classification output is a JSON message in JSON format; the JSON message includes dimension keywords, indicator keywords, and statistical period keywords. After classification, the final output is a JSON message in JSON format similar to the one shown in Table 5 below.

[0100] Table 5 JSON format of JSON message

[0101] In the table, atom_metric represents the indicator keyword, dim represents the dimension keyword, and time_cycle represents the statistical cycle keyword.

[0102] Step 4: Convert the indicators, dimensions, and statistical periods into corresponding SQL fragments, and then concatenate the SQL fragments to generate SQL query statements.

[0103] Combining steps 1 through 3, a data analysis problem described in natural semantics is broken down into key elements such as indicators, dimensions, and statistical periods. These are then formatted using advanced operators into commonly used computer message transmission formats like JSON and Map. Ultimately, the natural semantic problem is transformed into a mapping of these analysis dimensions to SQL fragments. This is then combined to generate SQL queries, which are then output to the large model for complex multidimensional analysis. "Advanced operators" refer to calculations such as averages and medians. Step 4 specifically includes the following steps 4.1 through 4.2.

[0104] Step 4.1: Perform intent recognition on the JSON message based on the large model, and match the dimension keywords, indicator keywords, and statistical period keywords included in the JSON message to the corresponding SQL fragments.

[0105] First, the big model identifies the intent of the JSON message and matches keywords such as dimensions, metrics, and statistical periods contained in the JSON message to corresponding SQL fragments. Rules are used here to classify and identify JSON messages, allowing the big model to automatically determine whether the incoming message is JSON. For example, if the big model is given a prompt template stating "Only allow JSON format data to be received," it will only accept JSON messages in the format shown in Table 5.

[0106] Step 4.1 requires intent recognition to convert the corresponding indicator / dimension / statistical period keywords into SQL snippets. To achieve accurate matching, dynamic weighting of each dimension is used. The core logic is to automatically adjust the contribution weight of different feature dimensions (such as business indicator name, calculation caliber Chinese description, and technical field name) in the overall similarity calculation based on the user's query intent. The implementation principle and steps are as follows: 4.1.1) to 4.1.3).

[0107] 4.1.1) Query intent classification, query types, features and their weight distribution are shown in Table 6 below.

[0108] Table 6 Query intent classification table

[0109] 4.1.2) Intent recognition method: determine the intent type through keyword regular matching and large model text reading.

[0110] 4.1.3) Dynamic Weight Adjustment Formula. Based on the identified intent type, the weight coefficients of each dimension vector are dynamically adjusted to calculate the comprehensive similarity: (17); in, D Represents a dimension set, such as the business indicator name, Chinese description of the calculation caliber, and technical field name. Indicates the dimension of the current intention The weight of . Represents a query In dimension Vector With the vector in the library In this way, no matter whether the extracted keyword is the indicator (dimension) name, the specific calculation caliber description, or English terminology, it can be matched to the corresponding SQL fragment.

[0111] Taking the question of asking about the store's monthly sales as an example, the company's business and technical personnel maintain the company's own indicators / dimensions / statistical period definitions in a database (knowledge vector library) for use throughout the entire process. Here, employees maintain the definition of sales, the total order amount, and the corresponding SQL fragment SUM(order_amt) for querying the data source. The large model first filters whether the questions asked by users are data analysis questions, such as asking about the weather. Such questions will be removed. The keywords in the data analysis questions are extracted, here are store / month / sales, and the store is preliminarily classified as a dimension, the month is a statistical period, and the indicator is sales. Then, the knowledge vector library is connected for further confirmation. Finally, a JSON message is output: { "Indicator": "Sales", "Dimension": "Store", Statistical period: Current month }.

[0112] Obtain the JSON message and query the SQL fragments and information of the corresponding indicators / dimensions / statistical periods in the vector library respectively. Output the JSON message containing the SQL fragments shown in Table 7. The symbol "#" is followed by the corresponding comment.

[0113] Table 7 JSON message containing SQL fragment

[0114] The atom_metric, dim, and time_cycle in the olap object correspond to the SQL fragments converted from the metrics, dimensions, and statistical cycles mentioned above. dims is the name of the metadata field involved in the aggregation calculation. The table in the source object is the data source (i.e., the data table). The big model does not generate SQL throughout the entire process; it only recognizes user semantics and splices SQL fragments. This avoids the illusion of SQL generated directly by the big model and improves data analysis accuracy.

[0115] Step 4.2: Splice the SQL fragments to generate the corresponding SQL query statement.

[0116] The JSON message shown in Table 7 can be restored to an SQL query statement (SQL statement for short) using the prompts shown in Table 8 (the following prompts are for logical reference only).

[0117] Table 8 Prompt word template for generating SQL statements

[0118] The indicator clause, operator clause, dimension clause, and target table are obtained from the previous JSON message with SQL fragments and are completed into the SQL format shown in Table 8. The SQL fragment corresponding to the indicator clause represents the grouping criteria; the SQL fragment corresponding to the operator clause represents the indicator calculation logic; and the SQL fragment corresponding to the dimension clause represents the filtering criteria. The target table refers to the data source for the query data. The dataset / data source is each company's own business data. For example, an e-commerce company's data source includes order details and product information. The accessed data source contains all of the company's operational data. The completion operation involves filling the determined SQL fragment into the SQL format shown in Table 8; for example, filling the dim in the JSON message into group by will match it. The SQL query statement generated in this way has a small and regular word template context, and the number of tokens generated by the text is also small.

[0119] Step 5: Execute the SQL query statement to query the database and return the query results.

[0120] Large models connect to external database toolsets, execute SQL queries, and drive SQL query tools to query the database, returning data query results. Existing large model development frameworks have built-in access to database toolsets. These operations can be implemented through development framework configuration, with different external toolsets configured based on the databases used by different business systems.

[0121] For example, the SQL fragments in the JSON message shown in Table 7 are concatenated into the following SQL query: select sum(shop_gmv) from order_detail_d where shop_id = 100039454 and date_sub(2025-01-01) and '2025-01-01' group by shop_id. Using this SQL query to query the database will generate a data metric (for example, 100,000) for display in the front-end user interface. This metric is sales, and its output format is: { "Sales": "100000" }.

[0122] Step 6: Convert the query results into structured data and display them through the front-end user interface.

[0123] Query results are converted into structured data through semantic or structural methods. Specifically, query results are semantically or structurally transformed into specific structured data using dimensions, indicators, and statistical periods. This data is then exposed to the front-end user interface via an API for detailed data analysis and presentation. The front-end user interface can then render the structured data into charts and other formats for display. Depending on the usage scenario, data query results can be organized into different data formats (such as data reports and dashboards) for external presentation. For example, if the query result is "Sales volume is 300,000 yuan, a year-on-year increase of 30%," the front-end user interface can display different formats such as pie charts and bar charts based on the business needs.

[0124] In other words, the entire workflow executed by the data analysis question and answer platform of this application can be summarized as: screen out data analysis questions from the question and answer content described by the user in natural language → extract keywords in the data analysis questions → classify the keywords, put the dimensions together, put the indicators together, and put the statistical periods together → query the corresponding knowledge vector library separately to obtain the calculation caliber (SQL fragment) maintained by the user → splice the SQL fragments together in order to form an SQL query statement → execute the SQL query statement to query the business database → the query results are converted into structured data through the program, and opened to the front-end user interface through the interface → the front-end user interface renders the structured data into tables, pie charts and other charts for display.

[0125] The data analysis question-and-answer platform of this application, based on a large model and a knowledge vector library, is mainly used in data analysis / data query in various business fields to realize intelligent question-and-answer. For example, in the field of e-commerce, it is used for operation analysis and store sales analysis. Before using the data analysis question-and-answer platform of this application, R&D personnel need to manually write code to query relevant indicators for business personnel to analyze. After using the data analysis question-and-answer platform of this application, the user's own company does not need R&D personnel, but only needs to ask questions (input), such as "What is the sales volume of this store this month", and the relevant sales indicators will be automatically output.

[0126] The data analysis question-and-answer platform of this application implements a low-cost, highly secure, highly accurate, and highly flexible natural language data analysis method. By connecting to an external knowledge vector library, questions are converted into structured JSON messages containing common multi-dimensional analysis terms such as indicators, dimensions, and statistical periods. These JSON messages are then converted into SQL statements for querying through the knowledge vector library. Large models do not need to ingest excessive amounts of useless schema content, reducing token consumption. At the same time, pre-made indicator SQL fragment mappings improve the accuracy of query results.

[0127] This application's data analysis Q&A platform also significantly reduces the context required to read large models, lowering costs based on the number of tokens. If data analysis requirements change, it can be expanded quickly and flexibly without requiring excessive maintenance. Furthermore, after reading the database, SQL fragments are mapped to indicators and directly spliced together, enabling precise positioning and improving query accuracy and speed. Furthermore, direct authentication of the indicator semantic layer (knowledge vector library) ensures the security of data queries.

[0128] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by hardware associated with computer program instructions. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory or other media in the various embodiments provided herein may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, and the like. Volatile memory may include random access memory (RAM) or external cache memory, and the like. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0129] It should be noted that the information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0130] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0131] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A data analysis question-answering platform based on a large model and knowledge vector library, characterized by: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a data analysis question-answering method based on a large model and a knowledge vector library; The data analysis question answering method based on the large model and knowledge vector library includes: Obtain user Q&A content, analyze it using a large model, and filter out data analysis questions; Identify key words in data analysis questions; Use a large model with a knowledge vector library to classify keywords into indicators, dimensions, and statistical periods; Convert indicators, dimensions, and statistical periods into corresponding SQL fragments, and then concatenate the SQL fragments to generate SQL query statements; Execute SQL query statements to query the database and return the query results; Convert query results into structured data and display them through the front-end user interface; The identification of keywords in data analysis questions specifically includes: Perform text preprocessing on data analysis problems, including segmenting text into words or phrases and removing common stop words to obtain preprocessed candidate words ; Parse each candidate word The word statistical features and semantic information, and according to the formula Calculate each candidate word Total score ;in Candidate word Case feature score; Candidate word The position feature score of Candidate word The word frequency feature score of Candidate word The symbolic feature score of Candidate word Contextual relevance score of Candidate word The cross-sentence distribution feature score of The number of characters for data analysis questions; Candidate word The contextual pattern enhances the feature score; The total score Candidate words below the preset score threshold Determine as keyword .

2. The data analysis question-answering platform based on a large model and a knowledge vector library according to claim 1 is characterized in that: The user's question and answer content is obtained and analyzed using a large model to screen out data analysis questions, specifically including: Obtain user Q&A content and analyze it using a large model to determine whether the Q&A content is a data analysis question; the large model includes DeepSeek and OpenAI; data analysis questions include: questions involving datasets or data sources, inquiring about trends, patterns, predictions or statistical analysis, requesting charts, visualizations or data interpretation, and involving specific statistical methods or analytical techniques; If the question and answer content is a data analysis question, directly output the data analysis question; If the question and answer content is a non-data analysis question, the user will be prompted on the front-end user interface to rewrite the non-data analysis question into a data analysis question.

3. The data analysis question-answering platform based on a large model and a knowledge vector library according to claim 1 is characterized in that: The candidate word The symbolic feature score of The calculation formula is: ;in, is the number of symbol types in the candidate word; For the weights of class symbols; is the symbol combination gain factor; is the position attenuation factor.

4. The data analysis question-answering platform based on a large model and a knowledge vector library according to claim 1, characterized in that: The candidate word Contextual relevance score The calculation formula is: ;in for Left and right neighbor words; point mutual information score ; yes and Joint probability of simultaneous occurrence; and They are and The probability of a single occurrence.

5. The data analysis question-answering platform based on a large model and a knowledge vector library according to claim 1 is characterized in that: The candidate word Contextual pattern enhancement feature score The calculation formula is: ;in, Score for exact matches; Score for fuzzy matching; is the dynamic weight score; 、 and They are 、 and The weight coefficient of .

6. The data analysis question-answering platform based on a large model and a knowledge vector library according to claim 1 is characterized in that: The keyword classification is performed by using a large model in conjunction with a knowledge vector library, and the keywords are classified into indicators, dimensions, and statistical periods, specifically including: By constructing prompt words, the large model can make the first round of predictions for keywords. Perform coarse-grained classification and generate The coarse screening candidate categories are dimensions, indicators or statistical periods; Calculation keywords The cosine similarity with the standard words of each category in the knowledge vector library is used to select the category with the highest cosine similarity as the candidate category for fine screening; the candidate category for fine screening is a dimension, an indicator or a statistical period; If the fine screening candidate category is consistent with the coarse screening candidate category, determine the fine screening candidate category or the coarse screening candidate category as the keyword The final category of If the fine-screened candidate category is inconsistent with the coarse-screened candidate category, then determine whether the cosine similarity corresponding to the fine-screened candidate category exceeds the similarity threshold; If the cosine similarity of the fine-screened candidate category exceeds the similarity threshold, the fine-screened candidate category is determined to be a keyword. The final category; otherwise, the coarse screening candidate category is determined as the keyword The final category of The keywords The final category output is a JSON message in JSON format; the JSON message includes dimension keywords, indicator keywords and statistical period keywords.

7. The data analysis question-answering platform based on a large model and a knowledge vector library according to claim 6 is characterized in that: The conversion of indicators, dimensions, and statistical periods into corresponding SQL fragments and the concatenation of the SQL fragments to generate SQL query statements specifically include: Based on the large model, the intent of the JSON message is recognized and the dimension keywords, indicator keywords, and statistical period keywords included in the JSON message are matched to the corresponding SQL fragments. Splice the SQL fragments to generate the corresponding SQL query statement.

8. The data analysis question-answering platform based on a large model and a knowledge vector library according to claim 7 is characterized in that: Executing the SQL query statement to query the database and returning the query results specifically includes: The large model accesses an external database tool set, executes SQL query statements, drives the SQL query tool to query in the database, and returns the data query results.

9. The data analysis question-answering platform based on a large model and a knowledge vector library according to claim 8, characterized in that: The conversion of query results into structured data and displaying it through the front-end user interface specifically includes: Convert query results into structured data through semantic or structural means, and open it to the front-end user interface through the API interface; The front-end user interface renders the structured data into charts for display.

Citation Information

Patent Citations

  • Search engine based on parameter

    CN101122915A

  • High-precision semantic search system oriented to judicial field

    CN110674252A

  • Automatic topic annotation method based on YAKE! Keyword extraction

    CN117349591A

  • User generated content standing detection method and system based on target information identification

    CN118070774A

  • Retrieval enhancement method combining keyword extraction and semantic analysis

    CN118535682A

Cited By

  • Multi-modal retrieval method and device

    CN120910110A

  • A multimodal retrieval method and apparatus

    CN120910110B

  • Data analysis method and device based on natural language, equipment and medium

    CN121009108A

  • Financial intelligent question number precision method based on dynamic calculation optimization

    CN121071105A

  • Large model-based coal preparation plant data query method, apparatus and device, and medium

    CN121958359A