A Visualization Method for a Knowledge-Enhanced Large Model Data Analysis Agent
Through the knowledge-enhanced large-model data analysis agent visualization method, the problems of low data analysis and visualization efficiency, high configuration cost, and inaccurate intention identification in the existing technology are solved, and fast, accurate and intuitive data analysis and visualization effects are achieved.
Patent Information
- Application Number
- CN202411759158.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-12-03
AI Technical Summary
The prior art has problems in data analysis and visualization, such as data insights that are too long, data visualization configuration costs, inaccurate identification of analysis intentions, limitations of natural language input, difficult to understand visual data sets, and insufficient intuitiveness of data-driven decision-making.
A knowledge-enhanced large model data analysis agent visualization method is proposed, syntactic analysis is performed through PCFG algorithm, keyword segmentation technology is used to extract keywords, correct and complete word segmentation results through recall strategies, generate SQL statements based on large models, analyze data sets, automatically recommend chart styles and generate analysis reports.
It achieves rapid data insights, reduces latency in traditional task delivery links, reduces dependence on professionals, improves the efficiency and effectiveness of data analysis and visualization, ensures the accuracy and intuitiveness of the chart, and adapts to the knowledge level and understanding ability of different audiences.
Smart Images

Figure CN119226387B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence driven technology, and in particular to a knowledge-enhanced large model data analysis agent visualization method. Background Art
[0002] With the rapid development of large model technologies (such as deep learning models, natural language processing models, etc.), especially the emergence of large-scale pre-trained models such as the GPT series and BERT, all walks of life have ushered in unprecedented opportunities for change. The continuous updating of AI technology has made it possible to quickly retrieve, analyze and visualize data through natural language (i.e. oral or written expression), significantly improving the efficiency and intuitiveness of information processing, and meeting the growing demand for intelligent data processing in various industries. However, although existing technologies have made significant progress in data analysis and visualization, there are still many problems that need to be solved.
[0003] First, the long time required for data insight and analysis is a common problem. Many companies rely on business cockpits, which are operated and maintained by BI or business analysis teams. When making demands for data and in-depth insights, companies not only want to obtain a visual presentation of the data, but also want to understand the reasons behind the data changes and make decisions accordingly. However, the traditional task delivery link is inefficient, and the requirements need to be conveyed to different teams step by step. Missing data requires repeated communication, resulting in delayed result return time, affecting the timeliness and accuracy of real-time decision-making.
[0004] Secondly, the configuration cost of data visualization is high. Current BI products are complex and highly flexible, but this also leads to a higher threshold for use. Users who lack professional training find it difficult to master various functions, and generating data conclusions often requires manual summarization, resulting in reliance on professionals for data visualization configuration, which increases labor costs. In addition, the complex configuration process also limits the ability of enterprises to quickly respond to data needs and reduces overall work efficiency.
[0005] In addition, inaccurate analysis intent recognition is also a major challenge faced by existing technologies. Although large models perform well in natural language processing, in specific fields, when analyzing data through natural language, it is difficult for the model to accurately identify the expression intent of a specific field, resulting in query results that do not meet user expectations. This inaccuracy in intent recognition makes it impossible for charts generated based on data features to accurately reflect the actual situation of the data, affecting the effectiveness and reliability of data analysis.
[0006] In addition, the limitations of natural language input are particularly prominent in complex domain scenarios. When the input content of the user is incomplete or the expression is ambiguous, it is difficult for the current large model to accurately capture and parse the subtle differences unique to the domain, resulting in obstacles in the processes of sentence decomposition, completion, and in-depth analysis. This not only affects the accuracy of the finally analyzed sentences but also limits the model's ability to recommend appropriate charts to intuitively display data, thereby affecting the overall information transmission effect.
[0007] Furthermore, the problem of difficult-to-understand visual datasets cannot be ignored. The datasets used for visualization often contain multiple dimensions and variables, and the process of selecting and determining which dimensions and variables to visualize is a complex one. In addition, different audiences have different abilities to understand data visualization. How to select appropriate chart types to adapt to the knowledge levels and understanding abilities of different audiences is a huge challenge. This results in the sometimes difficult-to-clearly-convey analysis results and non-intuitive display effects, affecting the effectiveness of data-driven decision-making.
[0008] Finally, the lack of intuitiveness in data-driven decision-making is also a major pain point in the existing technologies. Due to the differences in each person's data interpretation ability, preferences, and background knowledge, there may be significant differences in the understanding and focus of the same set of data. This is particularly insufficient when it is necessary to quickly and efficiently intuitively display decision-making information according to personal data reading habits, resulting in the inability of data-driven decision-making to achieve the expected effects and efficiency. Summary of the Invention
[0009] To overcome the deficiencies of the existing technologies, the present invention proposes a visualization method for a knowledge-enhanced large model data analysis intelligent agent, which significantly improves the efficiency and effect of data analysis and visualization by combining knowledge enhancement and large model technologies.
[0010] To achieve the above object, a visualization method for a knowledge-enhanced large model data analysis intelligent agent of the present invention includes the following steps:
[0011] Step S1: Use the PCFG algorithm for syntactic analysis, identify the syntactic components of the sentence and their relationships, and split them into logically complete clauses;
[0012] Step S2: Use word segmentation technology combined with a business word segmentation library to extract the keyword and entity words of the sentence;
[0013] Step S3: Recommend and summarize candidate word segments through multiple recall strategies, perform rough ranking, fine ranking, and detailed ranking on them, and correct and complete the word segmentation results;
[0014] Step S4: Split and analyze natural language questions based on a fine-tuned large model, and generate SQL statements that meet business requirements;
[0015] Step S5: Parse the dataset generated by SQL, extract field information and business meanings, and identify the association relationships between fields;
[0016] Step S6: Automatically recommend chart styles based on data characteristics and user preferences, assemble the analysis report, and distribute it at a specified frequency.
[0017] Further, Step 1 is specifically as follows:
[0018] Step S11: Perform word segmentation, part-of-speech tagging, and sentence boundary detection on the input sentence, use probabilistic context-free grammar (PCFG) to perform syntactic parsing on the sentence, generate a syntactic tree, construct the syntactic tree, display the hierarchical structure and dependency relationships of each word in the sentence, determine the main phrases (such as noun phrases, verb phrases, etc.) in the sentence and their internal structures, determine the grammatical dependency relationships between words in the sentence, and distinguish relationships such as subject-predicate-object.
[0019] Step S12: Analyze each node (phrase or word) of the syntactic tree one by one, determine the literal meaning of each part, understand its basic role in the sentence, and analyze possible implicit meanings, such as rhetorical devices, context-related meanings, etc.;
[0020] Step S13: Select a top-down or bottom-up method to traverse the syntactic tree, identify potential splitting points, use relevant algorithms to reconstruct the sentence according to syntactic components and meanings, generate multiple clauses, and the split clauses should each express a clear meaning and maintain the overall information of the original sentence;
[0021] Step S14: Check the logical coherence of the split clauses to ensure there are no logical loopholes, ensure that the clauses can be naturally connected to form a complete expression, and confirm that the split clauses can independently convey a complete meaning without missing key information;
[0022] Step S15: Merge the logically related and coherent clauses, optimize the expression structure, generate multiple finally logically complete and practically meaningful clauses, and complete the intention recognition and splitting of the sentence.
[0023] Further, Step 2 is specifically as follows:
[0024] Step S21: Update information such as dimension names, organization names, metric names, business terms, and keywords commonly used in business to the word segmentation library;
[0025] Step S22: Use HanLp word segmentation technology to perform secondary word segmentation on the sentence, and use forward maximum matching, reverse maximum matching, and bidirectional maximum matching algorithms to match and segment each substring of the sentence with the words in the dictionary;
[0026] Step S23: Select the optimal word segmentation result based on the number of words and matching situation of the forward and reverse word segmentation results. Step S24: Extract business-related keywords and entity words from the final word segmentation result.
[0027] Further, step 3 is specifically as follows:
[0028] Step S31: Establish a recall algorithm pool, including various algorithms based on content, features, and semantics;
[0029] Step S32: Determine the main recall algorithm and the bypass recall algorithm, and configure the corresponding recall data volume for each strategy, where the main recall algorithm is responsible for recalling more data;
[0030] Step S33: Use the word segmentation result to find similar words in the knowledge base through full match, partial match, or fuzzy match, calculate the similarity based on the vector space model (VSM), and return candidate words according to the quantity configured by the main and bypass strategies;
[0031] Step S34: Coarse ranking: Conduct a preliminary filtering on a large number of candidate data recalled, and remove the inconsistent data with a similarity lower than 0.5 or a word count exceeding 50% of the original;
[0032] Fine ranking: Use logistic regression and the Sigmoid function to score the remaining data, eliminate approximately 50% of the data with low similarity, and retain the data with high similarity;
[0033] Step S35: Remove duplicate data and data with lower scores, retain the Top5 or Top10 results, and recommend these optimized word segmentation results to the front end for selection or directly replace the original word segmentation to achieve the correction and completion of word segmentation.
[0034] Further, step 4 is specifically as follows:
[0035] Step S41: Understand various structures and forms of SQL queries, including SELECT queries, JOIN operations, WHERE clauses, subqueries, and nested queries, etc.;
[0036] Step S42: Design natural language and SQL templates covering different query types to ensure the diversity and flexibility of the generated queries;
[0037] Step S43: Use scripts to automatically fill in the placeholders in the templates to generate complete natural language queries and corresponding SQL statements, and verify their syntax and logical correctness;
[0038] Step S44: Increase the complexity and variability of the queries, and enrich the query content by randomly adding or omitting clauses, diversifying conditional expressions, introducing nested queries, and multiple JOIN operations, etc.;
[0039] Step S45: Test and validate the generated data, evaluate its accuracy, diversity, and consistency, and execute queries in the actual database to ensure that the results meet expectations, thereby optimizing and improving the quality of the training data.
[0040] Further, step 5 is as follows:
[0041] Step S51: Use named entity recognition technology to identify the part-of-speech of each field in the dataset, including nouns, verbs, adjectives, adverbs, etc.;
[0042] Step S52: Classify the fields into business meaning categories according to their part-of-speech. For example, classify nouns as dimensions or organizations, classify verbs as metrics or measures, and obtain other attributes of the fields from the predefined library tables, such as usage scope, call times, and supported business system information;
[0043] Step S53: Parse the generated SQL statements, extract the table names, fields, and their associated relationships, and obtain the record counts of each table through the data source to establish the associated relationships between tables;
[0044] Step S54: Integrate the information from the previous steps to construct a structure tree containing dimensions, metrics, and statistical information, and annotate the characteristics of each dimension in detail so that the data can be accurately identified and used when generating reports.
[0045] Further, step 6 is as follows:
[0046] Step S61: Extract key features from the dataset, including data type, data volume, range, trend of change, supported business, data distribution (such as mean, variance, skewness, kurtosis), and relationships between variables (such as correlation, causality);
[0047] Step S62: Use the decision tree machine learning algorithm to train the model based on historical data and user behavior, and learn the relationships between different data features and chart types. The specific process includes feature selection, data splitting, recursively constructing the decision tree, and using the trained model for prediction;
[0048] Step S63: Recommend the most suitable chart type to the user according to the extracted data features and the prediction results of the trained model. For example, recommend a line chart for time series data and a bar chart for comparing categorical data, etc.;
[0049] Step S64: Based on the recommended chart type, combine the data in the dataset, automatically draw the chart and fill it into the report to generate a complete analysis report. After the report is generated, the color and style of the chart can be set according to the user's preferences, and the report can be distributed at a specified frequency to promote the sharing of insights.
[0050] Further, step 4 also includes fine-tuning the pre-trained model using a high-quality dataset, and optimizing the model training process by adjusting strategies such as hyperparameters, optimization algorithms, and learning rates, as follows:
[0051] Step 451: Conduct a strict data preprocessing and screening process. This includes removing duplicate, invalid, or incorrectly formatted data, correcting incorrect SQL statements and natural language descriptions;
[0052] Step 452: Normalize the text, such as converting to lowercase, removing punctuation marks and stop words, to reduce the model's sensitivity to text format. At the same time, convert SQL statements to a unified standard format, ensuring consistent keyword case and removing extra spaces;
[0053] Step 453: Perform data augmentation through methods such as synonym replacement and sentence restructuring to expand the training data and enhance the model's adaptability to different language expressions and query structures;
[0054] Step 454: Freeze the parameters of the original model to avoid expensive update operations on a large number of parameters;
[0055] Step 455: Add a trainable low-rank matrix on the basis of the original model. By decomposing the parameter matrix to be updated into the product of two low-rank matrices, significantly reduce the computational and storage costs;
[0056] Step 456: In the fine-tuning stage, only train the newly added matrix, keep the parameters of the original model unchanged, and adjust the newly added parameters by optimizing the objective function to improve the model's performance on specific tasks;
[0057] Step 457: During inference, add the product of the trained matrix to the parameters of the original model to obtain the final parameter matrix without adding additional inference latency;
[0058] Step 458: After completing the model fine-tuning, use the optimized TextToSQL model, combined with prompt engineering, to intelligently generate industry-specific SQL statements. By understanding natural language queries and mapping them to structured SQL queries, ensure that the generated SQL statements meet the business requirements and logical requirements of a specific industry, thereby achieving accurate operations on the database.
[0059] Compared with the prior art, the beneficial effects of the present invention are:
[0060] 1. The present invention provides a visualization method for a knowledge-enhanced large model data analysis agent, with automated natural language understanding and real-time data processing capabilities, enabling users to quickly obtain data insights, reducing latency in the traditional task assignment link, supporting real-time decision-making; simplifying the complex BI tool usage process, reducing dependence on professionals, automatically generating intuitive charts through intelligent configuration, and reducing labor and time costs.
[0061] 2. The present invention provides a visualization method for a knowledge-enhanced large model data analysis agent, which can accurately identify the analysis intentions of users in specific fields through precise natural language understanding and domain adaptation training, ensuring that the generated queries and charts highly match the user needs; even in the case of complex or ambiguous expressions, the method can still effectively decompose, complement, and parse user inputs, ensuring the accuracy of analysis statements and the relevance of chart recommendations.
[0062] 3. The present invention provides a visualization method for a knowledge-enhanced large model data analysis agent, which intelligently selects and determines the dimensions and variables for visualization, generates charts suitable for the knowledge levels of different audiences, ensures that the analysis results are clear and intuitive, and enhances the information transmission effect; automatically generates charts and reports that conform to user habits according to personal preferences and data characteristics, ensures the intuitive display of decision-making information, and promotes efficient and accurate decision-making.
[0063] 4. The present invention provides a visualization method for a knowledge-enhanced large model data analysis agent. Through the combination of knowledge enhancement and large models, the system can adapt to different fields and changing data requirements, has good scalability, and meets diverse business needs; the automated and intelligent processes optimize data processing and visualization display, and users can obtain high-quality analysis results without complex operations, improving the overall user experience and work efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0065] Figure 1 is the schematic diagram of the step flow of the present invention;
[0066] Figure 2 is the word segmentation recall and sorting diagram;
[0067] Figure 3 is the understanding and decomposition of dataset information;
[0068] Figure 4 It is a schematic diagram of syntactic tree parsing (T1);
[0069] Figure 5 is a schematic diagram of syntactic tree parsing (T2). Detailed implementation manners
[0070] The technical solution of the present invention will be described more clearly and completely below in conjunction with the accompanying drawings and through the description of the preferred implementation manners of the present invention.
[0071] As Figure 1 shown, the present invention is specifically as follows:
[0072] Step S1: Use the PCFG algorithm for syntactic analysis, identify the syntactic components of the sentence and their relationships, and split them into logically complete clauses;
[0073] Step S2: Use the word segmentation technology combined with the business word segmentation library to extract the keywords and entity words of the sentence;
[0074] Step S3: Recommend and summarize candidate word segmentations through the recall strategy, perform rough ranking, fine ranking and detailed ranking on them, and correct and supplement the word segmentation results;
[0075] Step S4: Split and analyze natural language problems based on the fine-tuned large model, and generate SQL statements that meet business requirements;
[0076] Step S5: Parse the data set generated by SQL, extract field information and business meanings, and identify the association relationships between fields;
[0077] Step S6: Automatically recommend chart styles according to data characteristics and user preferences, assemble the analysis report and distribute it at a specified frequency.
[0078] As a specific implementation manner,
[0079] 1. Clause splitting for natural sentence intention recognition
[0080] Syntactic analysis is a crucial task in NLP, which is used to parse the sentence structure in order to better understand the meaning and composition of the sentence. Based on the understanding of the sentence, the sentence is split. To achieve the splitting of the sentence, the main implementation steps are as follows:
[0081] The first step: Segment the sentence
[0082] Word segmentation is the process of splitting a continuous text string into independent morphemes (Tokens). Since Chinese characters are written continuously without obvious delimiters (such as spaces), in order to make the word segmentation closer to the words required by the business, two joining methods are provided for word segmentation:
[0083] 1). Rule - and - dictionary - based method: Use predefined dictionaries and rules for word segmentation. Maximum Forward Matching (MFM), scan from left to right and select the longest word; Maximum Backward Matching (MBM), scan from right to left and select the longest word, and the Bi - directction Matching method (BM). The implementation process is as follows:
[0084] 1.1) Forward maximum matching idea MM
[0085] Forward maximum matching segmentation is usually simply referred to as the MM method. Its basic idea is: Assume that the longest word in the segmentation dictionary has i Chinese characters. Then use the first i characters in the current string of the document to be processed as the matching field and search the dictionary. If such an i - character word exists in the dictionary, the matching is successful, and the matching field is segmented as a word.
[0086] If such an i - character word cannot be found in the dictionary, the matching fails. Remove the last character from the matching field, and re - perform the matching process on the remaining string. Proceed in this way until the matching is successful, that is, a word is segmented or the length of the remaining string is zero. In this way, one round of matching is completed, and then the next i - character string is taken for matching processing until the document is scanned completely.
[0087] Its algorithm description is as follows:
[0088] (1) Initialize the current position counter and set it to 0;
[0089] (2) Starting from the current counter, take the first i characters as the matching field until the end of the document;
[0090] (3) If the length of the matching field is not 0, search for a word of the same length in the dictionary for matching processing.
[0091] If the matching is successful, then:
[0092] a) Segment this matching field as a word and put it into the word - segmentation statistical table;
[0093] b) Add the value of the current position counter by the length of the matching field;
[0094] c) Jump to step (2);
[0095] Otherwise:
[0096] a) Remove the last character from the matching field;
[0097] b) Subtract 1 from the length of the matching field; go to step 3).
[0098] 1.2 Reverse Maximum Matching Algorithm RMM
[0099] The Reverse Maximum Matching method is usually simply referred to as the RMM method. The basic principle of the RMM method is the same as that of the MM method. The difference is that the direction of word segmentation is opposite to that of the MM method. The Reverse Maximum Matching method starts from the end of the document to be processed for matching and scanning. Each time, the last i characters are taken as the matching field. If the matching fails, the first character of the matching field is removed and the matching continues. In actual processing, the document is first inverted to generate an inverted document. Then, according to the inverted dictionary, the inverted document is processed using the Forward Maximum Matching method.
[0100] 1.3 Bi-directction Matching method, BM
[0101] Bi-directional maximum matching word segmentation combines the algorithms of the previous two. First, the document is roughly segmented according to punctuation marks, and the document is decomposed into several sentences. Then, these sentences are scanned and segmented using the Forward Maximum Matching method and the Reverse Maximum Matching method. If the matching results obtained by the two word segmentation methods are the same, it is considered that the word segmentation is correct. Otherwise, it is processed according to the minimum set. The accurate result is often the one with fewer word segmentations.
[0102] 2) Statistical-based word segmentation
[0103] The method based on rules and dictionaries is highly dependent on the integrity and quality of the word library. When segmenting words, it does not consider the context in which the words are located and does not find the optimal solution from a global perspective. Therefore, the word segmentation method based on statistical machine learning emerged as the times require.
[0104] Here, the Hidden Markov Model HMM in statistical word segmentation is adopted. The sequence of states randomly generated by the hidden Markov chain is called the state sequence; each state generates an observation, and the resulting random sequence of observations is called the observation sequence; each position in the sequence can also be regarded as a moment.
[0105] The Hidden Markov Model consists of three major elements, namely the initial state probability vector , the state transition probability matrix A, and the observation probability matrix B, where and A determine the state sequence, and B determines the observation sequence.
[0106] ;
[0107] The formula definition assumes that Q is the set of all possible states, and V is the set of all possible observations:
[0108] ;
[0109] I is defined as the state sequence of length T, and O is defined as the corresponding observation sequence:
[0110] ;
[0111] The algorithm implementation process includes:
[0112] Task 1: Define what the state sequence is
[0113] In word segmentation, the state sequence is the label corresponding to the word block, that is, B / M / E / S {B: begin, M: middle, E: end, S: single}, which is the ultimate prediction target based on HMM word segmentation. Its state is invisible and unknown and needs to be solved;
[0114] Task 2: Define what the observation sequence is
[0115] In the word segmentation scenario, the observation sequence is the set of character sequences, which is visible and countable. At this time, the word problem is transformed into one of the three major problems of HMM, that is:
[0116] (1) Probability calculation problem
[0117] Given the model input and the observation sequence , calculate the probability that the observation sequence O appears under the model . .
[0118] (2) Learning problem
[0119] Given the observation sequence , estimate the parameters of the model so that the probability of the observation sequence under this model is the largest.
[0120] (3) Prediction problem
[0121] Also known as the decoding problem. Given the model and the observation sequence , find the state sequence with the maximum conditional probability given the observation sequence . That is, given the observation sequence, solve the most likely state sequence.
[0122] After the HMM word segmenter is trained, the word segmentation problem is transformed into the prediction problem of HMM, that is, the decoding problem.
[0123] For example:
[0124] Student Xiao Q applied for a 65-yuan mobile phone package | Observation sequence input
[0125] After prediction by the HMM word segmenter, the state labels corresponding to each character position are obtained:
[0126] BEBEBMEBEBME | Predicted state sequence
[0127] Based on this state label sequence, word segmentation can be done as follows:
[0128] BE / BE / BME / BE / BME
[0129] The corresponding word segmentation result is:
[0130] Xiao Q / Student / applied for / 65 yuan / mobile phone package
[0131] Step 2: Part-of-speech tagging
[0132] After the sentence is segmented, the nature of each word needs to be labeled. Common part-of-speech tags include noun (Noun, N), verb (Verb, V), adjective (Adjective, Adj), adverb (Adverb, Adv), etc.
[0133] For example, in the sentence "I love China", the word "love" can be a verb, noun, adjective, auxiliary word, etc. Just looking at this "love" alone, assuming the probability that it is a verb in our cognition is 7, the probability of being a noun is 1, the probability of being an adjective is 1, and the probability of being an auxiliary word is 1; but in fact, this word "love" not only depends on our cognition, but also on the context. For example, an adjective often follows a verb. In this case, if "love" is preceded by a verb, then the probability that the word "love" is an adjective becomes higher, and the probability becomes 7. By calculating the maximum probability of different hidden states of each observation node and recording the path, the path with the maximum probability is finally returned. The algorithm derivation process is as follows:
[0134] 1) The input parameters of the function include the length of the observation sequence (observation_len), the length of the hidden sequence (hidden_len), the initial probability (init_p), the transition probability matrix (trans_p), and the emission probability matrix (emit_p);
[0135] 2) The function first creates two two-dimensional arrays max_probabilities and paths to store the maximum probability and path of different hidden states of each observation node;
[0136] 3) Then, the function calculates the maximum probability by traversing each hidden state of the first observation node and records the path. Next, the function traverses each subsequent observation node, calculates the cumulative probability according to the formula of the Viterbi algorithm, obtains the maximum probability of each hidden state, and updates the path;
[0137] 4) Finally, the function returns the path with the maximum probability.
[0138] Taking "I love China" as an example, the detailed description of part-of-speech tagging using the algorithm is as follows:
[0139] Observation sequence: ['I', 'love', 'China']
[0140] Hidden sequence: ['AT', 'BEZ', 'IN', 'NN', 'PERIOD']
[0141] Output: I / AT love / BEZ China / NN
[0142] Maximum probability matrix:
[0143] V0(j) = init(j) × bj(o0) (init is the initial probability, b is the emission probability matrix)
[0144] Vt(j) = max(Vt−1(i) × aij) × bj(ot) (a is the transition probability matrix, b is the emission probability matrix)
[0145] Use paths to update and save i (hidden state) when the j (observation state) path reaches the maximum probability for backtracking
[0146] According to the maximum probability matrix max_probabilities, find the hidden state corresponding to the maximum probability of the last observation state "China" as the final part-of-speech tagging result.
[0147] Step 3: Syntactic analysis
[0148] Syntactic analysis is the process of obtaining the syntactic structure from a word string. Its goal is to generate a parse tree to show the relationships between the components of the sentence and the overall syntactic structure of the sentence. The parse tree includes constituency parsing and dependency parsing. Combining the word segmentation and part-of-speech of the previous step, use the PCFG syntactic analysis algorithm to construct multiple trees for the words and calculate a probability for each word at the node. When assembling the sentence, select the tree with the maximum probability as the syntactic tree of the input string to determine the syntactic structure of the sentence or the dependency relationship between the words.
[0149] The PCFG consists of five components G = (T, N, S, R, P).
[0150] 1) T represents the terminals, and the main elements are the words in the sentence, i.e., the leaf nodes of the syntactic tree;
[0151] 2) N represents the non-terminals, including various part-of-speech and functional phrase tags;
[0152] 3) S represents a special non-terminal that starts (S ∈ N), i.e., the root node of the syntactic tree;
[0153] 4) R represents a series of rules in the format X → γ, where X ∈ N and γ ∈ (N ∪ T);
[0154] 5) P represents the probability function. The range of P is [0, 1], and the sum is 1. The formula is as follows:
[0155] ;
[0156] Based on the idea of the PCFG algorithm, combined with characteristics such as word segmentation and part-of-speech, the probability of generating a sentence is calculated as:
[0157] ;
[0158] According to the above algorithm, for syntactic construction of word segmentation, a bottom-up approach is adopted:
[0159] 1) First, each word segment is used as the leaf node of the tree. At the same time, according to the characteristics of the tree, each node has only one upper vertex, and a vertex has only two child nodes for tree construction;
[0160] 2) According to the unary reduction method, the part-of-speech of the leaf node is reduced to the upper node;
[0161] 3) According to the part-of-speech of the words, calculate pairwise combination. Through the best probability generation algorithm, calculate the probability value of the combination of two words, and select the two words with the greater probability for secondary reduction. After secondary reduction, phrase groups such as noun phrase groups and verb phrase groups will be generated;
[0162] The best probability generation algorithm for syntax is as follows:
[0163] ;
[0164] 4) For the newly generated nodes, perform secondary reduction again according to the above 3 until there is only one root node finally;
[0165] For example: The parse trees for "Big data promotes more comprehensive decision-making analysis and recommends precise solutions" are the following two trees, where S represents a sentence; NP, VP, PP, and ADJP are noun phrases, verb phrases, prepositional phrases, and adjective phrases; N, V, P, and ADJ are nouns, verbs, prepositions, and adjectives respectively. In the process of constructing the graph, a bottom-up approach is adopted, and the bottom layer is the leaf nodes of the tree.
[0166] Calculate the probabilities of grammars composed of two words and three words respectively according to the algorithm described above until the last layer. Finally, backtrack through the entire sentence to obtain the best parse tree, as Figures 4-5 shown;
[0167] Calculate the probabilities of the two trees:
[0168] P(T1) = S * NP * N * PP * P * VP * V * NP * ADJ * NP * VP * V * NP * ADJ * N
[0169] = 0.5 * 0.6 * 0.7 * 0.3 * 0.3 * 0.7 * 0.5 * 0.5 * 0.4 * 0.6 * 0.4 * 0.5 * 0.5 * 0.3 * 0.7
[0170] = 0.0000167
[0171] P(T2) = S * NP * N * PP * P * VP * V * NP * VP * V * VP * NP * ADJP * P * ADJ
[0172] = 0.5 * 0.6 * 0.7 * 0.3 * 0.7 * 0.3 * 0.6 * 0.4 * 0.4 * 0.3 * 0.7 * 0.5 * 0.5 * 0.6 * 0.4
[0173] = 0.0000160
[0174] Then compare the final probability values of the two syntactic trees and select T1 as the final syntactic tree.
[0175] Step 4: Sentence splitting and clause generation
[0176] Spell the sentence according to the syntactic tree. The spelling adopts a top-down and same-level priority assembly method, and the spelling follows the following principles:
[0177] 1) There is one and only one word (ROOT, the virtual root node, abbreviated as the virtual root) that does not depend on other words.
[0178] 2) All other words must depend on other words.
[0179] 3) Each word cannot depend on multiple words.
[0180] 4) If word A depends on B, then word C located between A and B can only depend on words between A, B or AB.
[0181] These 4 axioms respectively restrict the uniqueness of the root node, connectivity, acyclicity and projectivity of the syntactic tree, thus the dependency relationships of each word in the syntactic tree.
[0182] According to the above principles, combined with the probabilities of the nodes on the tree, spell the sentence. In the same layer, place the probability values on the left and assemble the sentence in order. When encountering a "coordinating conjunction (CC)" during the assembly process, that is, a coordination relationship, split the sentence. When splitting, the leftmost word of the tree needs to be brought into the clause. For example Figure 4 For the sentence shown, it can be split into:
[0183] Sentence 1: Big data promotes more comprehensive decision-making analysis
[0184] Sentence 2: Big data recommends precise solutions
[0185] Step 5: Check the coherence of the sentence
[0186] Check the coherence of the sentence according to the structural form of the sentence. The structural relationship between words in the sentence, that is, the component dependency relationship of a word in the sentence. This relationship is mainly represented by dependency relationship symbols. The symbols represent part of speech, and judge the coherence of the sentence from the dependency relationship of part of speech:
[0187] 1. HED (Core relationship): Represents the core of a sentence, usually the subject or predicate.
[0188] 2. SBV (Subject-Verb relationship): Represents the relationship between the subject and the predicate.
[0189] 3. VOB (Verb-Object relationship): Represents the relationship between the object and the predicate.
[0190] 4. COO (Coordinating relationship): Represents a coordinating relationship between two or more words, and they share the same dependency relationship type.
[0191] 5. Att (Modifying relationship): Represents a modifying relationship. The word being modified is called the head word, and the modifying word is called the modifier.
[0192] Based on the above relationships, use the part of speech of the words that make up the sentence as the judgment basis to judge whether the relationship between words in the sentence conforms to relationships such as SBV, VOB, Att, HED, etc., so as to judge the coherence of the sentence.
[0193] II: Word segmentation extraction of business meaning
[0194] Using the HanLp word segmentation technology, perform secondary word segmentation on the sentence. Before word segmentation, update information such as commonly used dimension names, organization names, indicator names, business terms, and keywords in the business to the word segmentation library. Perform secondary extraction of word segmentation based on the plan. Its implementation principle is to regularly maintain a dictionary (timely record new words, delete old words, etc.). When segmenting the sentence, use each substring of the sentence to match each word in the dictionary one by one for segmentation. If there is no match, it is segmented as a single character. There are three algorithm combination methods: forward maximum matching algorithm, reverse maximum matching algorithm, and bidirectional maximum matching algorithm to achieve this:
[0195] Forward maximum matching algorithm:
[0196] 1) Overlappingly take m characters of the sentence from left to right as the matching character substring, where m is the number of characters of the longest word in the machine dictionary;
[0197] 2) When the m-character substring in the original sentence is matched with all the words in the dictionary, if the match is successful, take this matching string as a word;
[0198] 3) If the match is unsuccessful, remove the last character of the m characters and use m-1 characters as the new matching field. That is
[0199] m = m - 1 (m > 1), repeat steps 1 to 3 until all words are segmented.
[0200] Reverse maximum matching algorithm:
[0201] 1) Overlappingly take m characters of the sentence from right to left as the matching character substring, where m is the number of characters of the longest word in the machine dictionary;
[0202] 2) When the m-character substring in the original sentence is matched with all the words in the dictionary, if the match is successful, take this matching string as a word;
[0203] 3) If the match is unsuccessful, remove the last character of the m characters and use m-1 characters as the new matching field. That is m = m - 1 (m > 1), repeat steps 1 to 3 until all words are segmented.
[0204] Bidirectional maximum matching algorithm
[0205] 1) Combine the forward maximum matching algorithm and the reverse maximum matching algorithm;
[0206] 2) If the number of words in the forward and reverse word segmentation results is different, take the result with the fewer number of segmented words;
[0207] 3) If the number of words in the word segmentation results is the same, but the word segmentation results are different, return the result with fewer single characters in the word segmentation results. Otherwise, return the word segmentation result of the reverse maximum matching algorithm.
[0208] Through the above algorithm, information such as keywords and business entity words with business meanings are proposed.
[0209] III: Word Segmentation Recall Completion and Correction
[0210] After splitting a sentence into multiple word segments, there are situations where the business meaning of the word segmentation is incomplete and the content of the word segmentation is incomplete, and it cannot truly express the business meaning corresponding to the word segmentation. It is necessary to complete the true business meaning of the word segmentation from the knowledge base. Therefore, it is necessary to implement multi-channel recall from the knowledge base to find the words closest to the business meaning expressed by the word segmentation. It is mainly implemented by using multi-channel recall strategies, fine ranking, and fine ranking and detailed ranking. The steps are as follows:
[0211] 1) Establish a recall algorithm pool, and the algorithms in the pool mainly include types such as content-based, feature-based, and semantic-based;
[0212] 2) Arrange the algorithms in the algorithm pool to determine the main and bypass strategies of the recall algorithm, and at the same time configure the number of data items recalled by the algorithm. The data recalled by the main recall algorithm is more than that of the bypass recall algorithm.
[0213] 3) Recall algorithm implementation:
[0214] Based on the word segmentation, perform full matching, partial matching, or fuzzy matching to search for similar words in the text knowledge base and return them. The number of returns is based on the number of items configured by the main and bypass algorithms. The algorithm for finding similar words is: taking the field as the vector dimension, the values represented by the vectorization of each dimension are not necessarily numerical values, but the form is still a vectorized form, that is, the so-called Vector Space Model (VSM for short). The similarity between two objects can be calculated in the following way. Assume that the vector representations of the two objects are respectively:
[0215] ;
[0216] ;
[0217] The similarity between these two objects can be expressed as:
[0218] ;
[0219] It represents the similarity between two components of a vector. Various methods such as Jacard similarity are used to calculate the similarity between two components. Different weight strategies are adopted for different components in the above formula. See the following formula, where is the weight of the t-th component (feature), and the specific weight value can be set manually according to the understanding of the business;
[0220] ;
[0221] 4) Coarse ranking: For the large amount of data recalled, the role of the coarse ranking stage is to preliminarily rank the candidate data and filter out the data that obviously does not meet the requirements, such as data with a similarity value less than 0.5, data with a word count exceeding 50% of the original, etc. Some obviously non-conforming data are filtered out by such rules;
[0222] 5) Fine ranking: As a key link in data recall processing, use the features of the data to score and rank. Since most of the recalled data has been pre-processed related work, the logistic regression algorithm can be used to calculate the probability of data similarity. The logistic function (also known as the sigmoid function) maps the linear combination of the input to a probability value between [0,1], which is used to predict the probability of a binary classification problem. According to the predicted probability value, it is judged which category the data belongs to. Judge whether the data after coarse ranking is more similar to the original data. If the similarity is high, it means 1 is the positive class, and if the similarity is low, it means 0 is the negative class.
[0223] The Sigmoid function produces an S-shaped curve and always returns a probability value between 0 and 1, converting any real number into a number between 0 and 1. Its mathematical formula is:
[0224] ;
[0225] Note:
[0226] F(x) represents the function
[0227] x: represents the annual feature value of the phrase
[0228] e: represents the base of the natural logarithm, approximately equal to 2.71828, which is a constant
[0229] The value calculated by the above formula is compared with a selected threshold. For example, when the threshold is 0.5, if f(x) > 0.5, then this x is classified into the category of 1, and if f(x) < 0.5, then x is classified into the category of 0. Thus, about 50% of the data is eliminated through the fine ranking algorithm, and the more similar data is retained.
[0230] 6) Fine sorting: By processing the data after fine sorting, mainly based on the characteristics of the data, duplicate data is removed, data with less scores is removed, and data within the top 10 or top 5 is retained. The data after fine sorting is recommended to the front end for selection. If exactly one condition is returned, the original data can be directly replaced, thereby realizing the correction and supplementation of the original data. For example, Figure 2 as shown.
[0231] IV: SQL Assembly of Large Models
[0232] The TextToSQL algorithm is a technology that converts natural language text into structured query language (SQL). By understanding the words of natural language and then converting them into SQL query statements that can be understood by a computer, database operations can be realized. The TextToSQL algorithm uses neural network models, such as recurrent neural networks (RNNs) and attention mechanisms, to learn the mapping relationship between natural language queries and SQL queries. When generating SQL statements for a specific industry, the algorithm is usually fine-tuned based on industry data. The fine-tuning process depends heavily on the diversity and accuracy of the data. Therefore, in order to fully utilize the advantages of the TextToSQL algorithm in accurately generating SQL, it is necessary to strengthen the collection and annotation of domain data. This includes obtaining high-quality annotated data from reliable sources as templates, ensuring the diversity and coverage of the data, and establishing a continuous data update and model optimization mechanism to improve the performance and accuracy of the TextToSQL algorithm. To obtain a large amount of useful data, a large amount of high-quality training data is generated based on the annotated data as a template. The following are some steps and considerations to achieve this goal:
[0233] Step 1: Understand the SQL structure and diversity
[0234] Understand the various structures and forms that SQL queries can have, including:
[0235] SELECT queries, including various aggregate functions (COUNT, SUM, AVG, MIN, MAX, etc.), GROUP BY, HAVING, ORDER BY, and LIMIT clauses.
[0236] JOIN operations, including INNER JOIN, LEFT JOIN, RIGHT JOIN, and FULL JOIN, and their applications on different tables.
[0237] Conditional expressions in the WHERE clause, involving comparison operators (=, !=, >, <, >=, <=), logical operators (AND, OR, NOT), and possible NULL value handling.
[0238] Subqueries and nested queries, which can appear in the SELECT, FROM, or WHERE clauses.
[0239] Step 2: Design templates
[0240] When designing templates, consider the diversity of SQL queries. The templates should be flexible enough to generate queries of different difficulties and complexities.
[0241] Template 1: SELECT query with conditions
[0242] Natural language query template: "Find records in table [table name] where [column name 1] is greater than [value] and [column name 2] is not equal to '[text value]'."
[0243] SQL template: "SELECT * FROM [table name] WHERE [column name 1]>[value] AND [column name 2]!= '[text value]';"
[0244] Template 2: Aggregation query
[0245] Natural language query template: "Calculate the average value of [column name] in table [table name] and group by [grouping column]."
[0246] SQL template: "SELECT [grouping column], AVG([column name]) FROM [table name] GROUP BY [grouping column];"
[0247] Template 3: JOIN operation
[0248] Natural language query template: "Join table [table name 1], table [table name 2], and table [table name 3], based on their associated columns, and select [column name 1] of [table name 1], [column name 2] of [table name 2], and [column name 3] of [table name 3]."
[0249] SQL template: "SELECT [table name 1].[column name 1], [table name 2].[column name 2], [table name 3].[column name 3] FROM [table name 1] JOIN [table name 2] ON [table name 1].[associated column]= [table name 2].[associated column] JOIN [table name 3] ON [table name 2].[another associated column] = [table name 3].[associated column];"
[0250] Template 4: SELECT query with CASE expression
[0251] Natural language query template: "Select [column name 1] in table [table name], if [column name 2] is greater than [value], display 'High', otherwise display 'Low'."
[0252] SQL template: "SELECT [column name 1], CASE WHEN [column name 2]>[value] THEN 'High' ELSE 'Low' END AS status FROM [table name];"
[0253] Template 5: Query with time function
[0254] Natural language query template: "Select records updated in the past 7 days from table [table name]."
[0255] SQL template: "SELECT * FROM [table name] WHERE [update date column]>= DATE_SUB(CURDATE(), INTERVAL 7 DAY);"
[0256] Template 6: SELECT query with subquery
[0257] Natural language query template: "Find all records in table [table name 1] where [column name 1] is greater than the average value of [column name 2] in table [table name 2]."
[0258] SQL template: "SELECT * FROM [table name 1] WHERE [column name 1]>(SELECT AVG([column name 2])FROM [table name 2]);"
[0259] Step 3: Data filling and variation
[0260] After having the basic templates, write scripts to automatically fill in the placeholders (such as [table name], [column name], [value], etc.) in these templates. Use randomly generated data, data from an existing database, or a combination of both to achieve this. Ensure that the generated queries are correct both syntactically and logically and can be executed in an actual database.
[0261] Step 4: Introduce complexity and variability
[0262] To increase the complexity and variability of the generated queries, you can:
[0263] 1) Randomly add or omit certain clauses (such as GROUP BY, HAVING, ORDER BY)
[0264] 2) Use different numbers and types of conditional expressions in the WHERE clause
[0265] 3) Nest queries or use subqueries
[0266] 4) Combine multiple JOIN operations to join multiple tables
[0267] 5) Use different aggregation functions and calculation expressions
[0268] Step 5: Testing and verification
[0269] Finally, verify the generated data according to the test and verification data quality evaluation formula:
[0270] (Quality\ Score = \alpha \cdot Accuracy + \beta \cdot Diversity + \gamma \cdot Consistency)
[0271] (Accuracy): The accuracy of the data, that is, whether the labeled SQL statements match the natural language descriptions.
[0272] (Diversity): The diversity of the data, covering queries in different fields and of different complexities.
[0273] (Consistency): The consistency of the data, ensuring that the same or similar natural language descriptions correspond to the same or similar SQL statements.
[0274] (\alpha, \beta, \gamma): Are weight factors, adjusted according to the specific task and dataset.
[0275] Ensure through verification that they do not error when executed in the actual database and that the returned results meet the expectations of the natural language queries.
[0276] Through these steps, utilize the capabilities of the large model to generate diverse and high-quality Text-to-SQL training data. As the model's ability to handle more complex queries improves, the data generation templates need to be continuously updated and expanded to generate Text-to-SQL data that meets your specific training requirements.
[0277] Fine-tune the pre-trained model using a high-quality dataset, and optimize the training process of the model by adjusting strategies such as hyperparameters, optimization algorithms, and learning rates, as follows:
[0278] The first step, data preprocessing and screening: To ensure the quality of the fine-tuning data, formulate a strict data preprocessing and screening process. This includes removing duplicate samples, correcting incorrect SQL statements and natural language descriptions, and ensuring the accuracy and consistency of the prepared data through manual review and data cleaning. The main methods are:
[0279] 1) Data cleaning: Remove duplicate, invalid, or incorrectly formatted data to ensure the accuracy and consistency of the training data.
[0280] 2) Text normalization: Convert the text to a standard format, such as lowercasing, removing punctuation marks, stop words, etc., to reduce the model's sensitivity to text format.
[0281] 3) SQL statement normalization: Convert SQL statements to a standard format, such as unifying keyword case, removing extra spaces, etc., to facilitate model learning and comparison.
[0282] 4) Data augmentation: Expand the training data by means of synonym replacement, sentence restructuring, etc., to increase the model's adaptability to different language expressions and query statement structures.
[0283] Second step, domain adaptation training: Considering that the data distributions in different domains may vary, a domain adaptation training strategy is adopted. By introducing a domain-specific pre-trained model (Tongyi Qianwen), combined with the fine-tuning technique LoRA and the above-mentioned domain data for relevant auxiliary task training, the pre-trained model is fine-tuned on a small-scale dataset through model fine-tuning to adapt to the tasks and data distributions in the target domain, so as to enhance the model's ability to process data in different domains. The implementation process of the LoRA fine-tuning algorithm is as follows:
[0284] 1), Freeze the parameters of the original model: During the fine-tuning process, the parameters of the original large model remain unchanged, which can avoid expensive update operations on a large number of parameters.
[0285] 2), Add trainable parameters: Based on the original model, adjust the model by introducing additional network layers or matrices. The number of newly added parameters is relatively small, usually the product of two smaller matrices obtained through low-rank factorization.
[0286] 3), Low-rank factorization: Decompose the parameter matrix W0 to be updated into the product of two low-rank matrices A and B, that is, W0 = B * A. The ranks of A and B are much smaller than the dimension of W0, thus significantly reducing the computational and storage costs.
[0287] 4), Train the newly added parameters: During the fine-tuning stage, only train the newly added matrices A and B, while the parameters of the original model remain unchanged. By optimizing the objective function (such as the loss function), adjust the values of A and B to make the model achieve better performance on specific tasks.
[0288] 5), Merging during inference: During the inference stage, add the product of the trained matrices B * A to the parameter matrix of the original model to obtain the final parameter matrix. This does not introduce additional inference latency because the merged parameter matrix can be directly used for calculation.
[0289] Among them, fine-tuning based on LoRA has many advantages and can achieve good application effects with fewer resources:
[0290] 1) Reduce computational and storage costs: Since only a small number of parameters need to be trained, LoRA fine-tuning significantly reduces computational and storage costs, making it possible to fine-tune large models with limited resources.
[0291] 2) Maintain the performance of the original model: By freezing the parameters of the original model and adding a small number of trainable parameters, LoRA fine-tuning can achieve task-specific optimization while maintaining the performance of the original model.
[0292] 3) Flexibility and scalability: LoRA fine-tuning can be applied to various types of large models, including models in the fields of natural language processing, computer vision, etc.
[0293] LoRA fine-tuning of large models is an effective model adjustment strategy. It optimizes the original model by adding a small number of trainable parameters, while significantly reducing computational and storage costs. This makes it possible to fine-tune large models with limited resources, providing more flexibility and scalability for the TextToSQL scenario.
[0294] The third step: Generate SQL: Based on the fine-tuned TextToSQL and combined with prompt engineering, intelligently generate industry-specific SQL statements.
[0295] V: Dataset understanding
[0296] After executing the SQL generated by TextToSQL, a dataset is produced. It is necessary to clarify the meaning expressed by the fields in the dataset, including the part of speech of the fields, business meaning, as well as the full table records associated with the data, feature recognition, and field relationships, etc. As Figure 3 shown, the implementation process is as follows:
[0297] The first step: Through named entity recognition technology, identify the part of speech of the fields. The parts of speech include nouns (Noun, N), verbs (Verb, V), adjectives (Adjective, Adj), adverbs (Adverb, Adv), etc.
[0298] The second step: Based on the part of speech of the fields, classify the business meaning of the fields. For example, the business meaning of nouns is classified as dimensions, organizations, etc.; verbs are expressed as metrics, measures, etc.; obtain other attributes of the fields from the agreed-upon database tables, such as usage scope, call times, information supporting the business system, etc.
[0299] The third step: Parse the sql, obtain the table names, fields, and their inter-table association relationships, obtain the record counts of each table through the data source, and establish the association relationships between the tables.
[0300] Step 4: Through the above three steps, construct a structure tree of dimensions, metrics, and statistical information, and detail the characteristics of the dimensions to facilitate data identification during report generation.
[0301] VI. Automatic Report Generation
[0302] Based on the understanding of the data set, according to the types and characteristics of dimensions, metrics, organizations, measures, etc. in the data set, and additional information such as personal preferences, generate reports automatically and intelligently. The implementation method is as follows:
[0303] Step 1: Feature extraction: Extract key features from the data. These features include data type, data volume, range, change trend, supporting business, data distribution (such as mean, variance, skewness, kurtosis, etc.), and the relationships between variables (such as correlation, causal relationship).
[0304] Step 2: Model training: Use the machine learning algorithm decision tree to train the model based on historical data and user behavior, and learn the relationships between different data features and chart types. The implementation of decision tree training is as follows
[0305] Step 1: Feature selection: Select a feature and a set of values of this feature to calculate the metric using the information gain (ID3 algorithm) of the decision tree; where the information gain represents the degree to which the uncertainty of Y is reduced after learning the information of feature X. The derivation process of the information gain (ID3 algorithm) is as follows:
[0306] (1) Information entropy of the classification system
[0307] Set a sample space (D, Y) of the classification system, where D represents the samples (with m features) and Y represents n categories. The possible values are C1, C2,..., Cn, and the probability of each category appearing is P(C1), P(C2),..., P(Cn). The entropy of this classification system is:
[0308] ;
[0309] In the data distribution, the probability P(Ci) of the category Ci appearing can be obtained by dividing the number of times this category appears by the total number of samples.
[0310] (2) Conditional entropy
[0311] According to the definition of conditional entropy, the conditional entropy in a classification system refers to the information entropy when a certain feature X of the sample is fixed. Since the possible values of this feature X can be (x1, x2,..., xn), when calculating the conditional entropy and needing to fix it, each possibility needs to be fixed, and then the statistical expectation is calculated. Therefore, the probability that the sample feature X takes the value xi is Pi, and the conditional information entropy when this feature is fixed to the value xi is H(C|X = xi). Then H(C|X) is the conditional entropy when the feature X in the classification system is fixed (X = (x1, x2,..., xn)):
[0312] ;
[0313] If this feature of the sample has only two values (x1 = 0, x2 = 1) corresponding to (appearing, not appearing), such as the appearance or non - appearance of a certain word in text classification. Then for the case of binary features, we use T to represent the feature and t to represent the occurrence of T, indicating that this feature appears. Then:
[0314] ;
[0315] Comparing with the formula of conditional entropy in the front, P(t) is the probability that T appears, and is the probability that T does not appear. Combining with the calculation formula of information entropy, we can get:
[0316] ;
[0317] ;
[0318] The probability P(t) that the feature T appears is just the number of samples in which T appears divided by the total number of samples; P(Ci|t) represents the probability that the category Ci appears when T appears, which is just the number of samples that appear T and belong to the category Ci divided by the number of samples that appear T.
[0319] (3) Information gain
[0320] According to the formula of information gain, the definition of the information gain of the feature X in the classification system:
[0321] Gain(D, X) = H(C) - H(C|X)
[0322] Information gain is for each feature. It is to see how much information there is in the system with and without a feature X. The difference between the two is the information gain brought by this feature to the system. Each time the process of selecting a feature is to calculate the information gain after dividing the data set by each feature value, and then select the feature with the highest information gain. For the case where the feature value is binary, the information gain brought by the feature T to the system is written as the difference between the original entropy of the system and the conditional entropy after fixing the feature T:
[0323] ;
[0324] (4) After obtaining a feature as the root node of the decision tree through the above round of information gain calculation, if the feature has several values, the root node will have several branches, and each branch will generate a new data subset. The remaining recursive process is to repeat the above process for each data subset until all the sub-datasets belong to the same class.
[0325] Step 2: Split the dataset into multiple subsets, with one subset for one feature and one set of values;
[0326] Step 3: Recursive construction: Repeat Steps 1 and 2 for each subset until the stopping condition is met, that is, all data points belong to the same class or there are no more features available for further splitting.
[0327] Step 4: Prediction: For a new data point, start from the root node and move down the tree according to the feature values until reaching the leaf node, and the class of the leaf node is the predicted class.
[0328] Third step, chart recommendation: Based on the features of the input data and the model prediction results, recommend the most suitable chart type for the user. For example, if the data features indicate time series data, the model may recommend a line chart; if the data features indicate a comparison of categorical data, the model may recommend a bar chart.
[0329] Fourth step, automatic report generation: Based on the recommended chart, combined with the data in the dataset, automatically draw the chart to fill the report.
[0330] As a specific embodiment, an operator wants to instantaneously analyze the business development situation through a conversational approach, and its conversational analysis statement is as follows:
[0331] "Analyze the daily data trend of the daily development volume of the electric channel in a certain place in July, and analyze the abnormal dates and reasons."
[0332] The understanding and execution process for the above natural statement is as follows:
[0333] First step, through the intention recognition of the natural statement, split the clause into the following two sentences:
[0334] Clause 1: Analyze the daily data trend of the daily development volume of the electric channel in a certain place in July
[0335] Clause 2: Analyze the abnormal dates and reasons in a certain place in July
[0336] Second step, extract the words that represent the business meaning of the above two clauses, such as: a certain place, July, daily development volume of the electric channel, data trend, abnormal dates, reasons.
[0337] Step 3: Correct and complete the extracted word segments. For example, "July" is completed to "July 2024", and "Daily development volume of the electric channel" is corrected to "Daily product development volume of the electric channel" according to the business terms stored in Rag.
[0338] Step 4: Assemble SQL through the large model after domain fine-tuning. It is necessary to combine the knowledge, sentence segmentation, and word segmentation in Rag in advance to spell the large model prompt words as follows:
[0339] { "input": "I hope you act as an SQL terminal. Facing a sample database, you only need to return SQL commands to me. The following is an instruction describing the task. Write an appropriate response to complete the request.\nThe database contains tables: app_ld_index_view_day, conf_ld_index_list, organization\nTable app_ld_index_view_day has the following fields: month: represents the accounting period, lan_id, business_id, bureau_cd: are all screening conditions for organization queries, index_id: represents the indicator code, index_value: represents the indicator value\nTable conf_ld_index_list has the following fields: index_id: represents the indicator code, index_name: represents the indicator name, data_date: represents the accounting period\nTable organization has the following fields: party_id: represents the organization code, principal: represents the organization name\nAmong them, the index_id of app_ld_index_view_day needs to be associated with the index_id of conf_ld_index_list\nAmong them, the lan_id of app_ld_index_view_day needs to be associated with the party_id of organization\nThe screening conditions for organization queries involved in the question are: app_ld_index_view_day.lan_id = 843073600000000 AND app_ld_index_view_day.business_id = -1 AND app_ld_index_view_day.bureau_cd = -1\nThe indicator name and the corresponding indicator code are as follows: Daily product development volume of the electric channel: D033004145\nTables need to be associated, and the question answered with sql. If the question involves indicators, the indicator code needs to be used as the query condition.\nAccounting period: app_ld_index_view_day.month = 202407\n
[0340] Question: Analyze the daily data trend of the daily development volume of the electricity channel in a certain place in July
[0341] The following SQL is generated through the large model:
[0342] SELECT app.day_id as billing period, org.principal as organization name, MAX(CASE WHEN app.index_id = 'D033004145' THEN app.index_value END) AS daily development volume of the electricity channel FROM app_ld_index_view_day AS app JOIN conf_ld_index_list AS conf ON app.index_id = conf.index_id JOIN organization org ON app.lan_id = org.party_id WHERE app.month = '202407' AND app.lan_id = '843073600000000' AND app.business_id = -1 AND app.bureau_cd = -1 AND app.index_id = 'D033004145' GROUP BY org.principal, app.month;
[0343] Step 5: Dataset understanding: After executing the above SQL, a dataset is generated. Each data in the dataset has a field to express others. By using the command real-time recognition technology, the dimension and metric fields are found to be: dimension field: organization name (a certain place), time dimension; metric field: daily development volume of the electricity channel.
[0344] Step 6: Automatic report generation: By calculating and recommending the historical data of data integration, this mainly shows the daily data trend and the reasons for abnormal dates. After the algorithm training based on characteristics, it will automatically recommend a fluctuation chart to display the report data, and draw warnings of different colors for the dates below the configured threshold, providing data to explain the reasons for the anomalies. Compared with the traditional report development method, this can realize the customer analysis requirements in seconds, with an efficiency improvement of more than 80%. At the same time, based on the management and intent recognition of Teling Rag, combined with efficient AI algorithms, the accuracy of converting natural sentences of the large model into industry-specific SQL is improved to 80% and above.
[0345] The above specific embodiments only describe the preferred embodiments of the present invention, rather than limiting the protection scope of the present invention. Without departing from the design concept and spirit of the present invention, various deformations, substitutions and improvements made by those of ordinary skill in the art to the technical solutions of the present invention according to the written description and drawings provided by the present invention shall fall within the protection scope of the present invention. The protection scope of the present invention is determined by the claims.
Claims
1. A knowledge-enhanced large-model data analysis agent visualization method, characterized in that: The following steps are involved: Step S1: Use the PCFG algorithm to perform syntactic analysis, identify the syntactic components of the sentence and their relationships, and split them into logically complete clauses; Step S2: extracting keywords and entity words from the sentence using word segmentation technology combined with a business word segmentation database; Step S3: recommend and summarize candidate segmented words through the recall strategy, perform rough sorting, fine sorting and detailed sorting on them, and correct and complete the segmented word results; Step S4: Split and analyze the natural language question based on the fine-tuned large model, and generate SQL statements that meet business requirements; Step S5: Parse the data set generated by SQL, extract field information and business meaning, and identify the relationship between fields; Step S6: automatically recommend chart styles based on data features and user preferences, assemble analysis reports and distribute them at a specified frequency; Step 6 is as follows: Step S61: extract key features from the data set, including data type, data volume, range, change trend, supporting business, data distribution and relationship between variables; Step S62: Using a decision tree machine learning algorithm, a model is trained based on historical data and user behavior to learn the relationship between different data features and chart types. The specific process includes feature selection, data splitting, recursive construction of a decision tree, and prediction using the trained model; Step S63: recommending chart types to the user based on the extracted data features and the prediction results of the trained model; Step S64: Based on the recommended chart type, combined with the data in the data set, the chart is automatically drawn and filled into the report to generate a complete analysis report. After the report is generated, the color and style of the chart are set according to the user's preferences, and the report is distributed at a specified frequency.
2. According to the knowledge-enhanced large-model data analysis agent visualization method of claim 1, it is characterized by: Step 1 is as follows: Step S11: perform word segmentation, part-of-speech tagging and sentence boundary detection on the input sentence, perform syntactic analysis on the sentence using a probabilistic context-free grammar, generate a syntactic tree, construct a syntactic tree, display the hierarchical structure and dependency relationship of each word in the sentence, determine the main phrases in the sentence and their internal structure, and determine the grammatical dependency relationship between words in the sentence; Step S12: Analyze each node of the syntax tree one by one, determine the literal meaning of each part, understand the basic role of the literal meaning in the sentence, and analyze the possible implicit meaning; Step S13: traverse the syntactic tree in a top-down or bottom-up manner, identify potential split points, use an algorithm to reconstruct the sentence according to syntactic components and meanings, generate multiple clauses, and each of the split clauses expresses a clear meaning while maintaining the overall information of the original sentence; Step S14: Check the logical coherence of the split clauses to form a complete expression, and confirm that the split clauses independently convey complete meanings without missing key information; Step S15: Merge logically related and coherent clauses, optimize the expression structure, generate final logically complete and meaningful clauses, and complete the sentence intent recognition and splitting.
3. The knowledge-enhanced large-model data analysis agent visualization method according to claim 1 is characterized in that: Step 2 is as follows: Step S21: updating the commonly used business information into the word segmentation library; Step S22: perform secondary word segmentation on the sentence using HanLp word segmentation technology, and use forward maximum matching, reverse maximum matching and bidirectional maximum matching algorithms to match and segment each substring of the sentence with the words in the dictionary; Step S23: Select the best word segmentation result according to the number of words and matching conditions of the forward and reverse word segmentation results; Step S24: extracting business-related keywords and entity words from the final word segmentation results.
4. The knowledge-enhanced large-model data analysis agent visualization method according to claim 1 is characterized in that: Step 3 is as follows: Step S31: Establishing a recall algorithm pool, including multiple algorithms based on content, features and semantics; Step S32: Determine the main recall algorithm and the bypass recall algorithm, and configure the recall data volume for the strategy, wherein the main recall algorithm is responsible for recalling more data; Step S33: using the word segmentation results, searching for similar words from the knowledge base through full matching, partial matching or fuzzy matching, calculating the similarity based on the vector space model, and returning the number of candidate words configured according to the main bypass strategy; Step S34: First, perform a rough sorting to preliminarily filter the large amount of candidate data recalled, and remove the non-conforming data with a similarity lower than 0.5 or a word count exceeding 50% of the original data; then perform a fine sorting to score the remaining data using logistic regression and Sigmoid function, eliminate low-similarity data, and retain high-similarity data; Step S35: Remove duplicate data and data with low scores, retain the Top 5 or Top 10 results, and recommend these optimized word segmentation results to the front end for selection or directly replace the original word segmentation to achieve word segmentation correction and completion.
5. The knowledge-enhanced large-model data analysis agent visualization method according to claim 1 is characterized in that: Step 4 is as follows: Step S41: understanding various structures and forms of SQL queries; Step S42: Designing natural language and SQL templates covering different query types; Step S43: Use the script to automatically fill in the placeholders in the template, generate a complete natural language query and corresponding SQL statement, and verify its grammatical and logical correctness; Step S44: increasing the complexity and variability of the query and enriching the query content; Step S45: Test and verify the generated data, evaluate its accuracy, diversity and consistency, and execute queries in the actual database to ensure that the results meet expectations.
6. The knowledge-enhanced large-model data analysis agent visualization method according to claim 1 is characterized in that: Step 5 is as follows: Step S51: using named entity recognition technology to identify the part of speech of each field in the data set; Step S52: classify the field into business meaning categories according to its part of speech, and obtain other attributes of the field from a predefined library table; Step S53: Parse the generated SQL statement, extract the table name, field and the relationship between them, obtain the number of records in each table through the data source, and establish the relationship between the tables; Step S54: Integrate the information from steps S51 to S53 to construct a structure tree including dimensions, indicators and statistical information, and mark the characteristics of each dimension.
7. The knowledge-enhanced large-model data analysis agent visualization method according to claim 1 is characterized in that: Step 4 also includes fine-tuning the pre-trained model using a high-quality dataset to optimize the model training process as follows: Step 451: Perform a rigorous data preprocessing and screening process, which includes removing duplicate, invalid or incorrectly formatted data, and correcting incorrect SQL statements and natural language descriptions; Step 452: normalize the text to reduce the model's sensitivity to the text format, convert the SQL statement into a unified standard format, ensure that the keywords are consistent in capitalization, and remove extra spaces; Step 453: Perform data enhancement through synonym replacement and sentence reorganization methods to expand training data and enhance the model's adaptability to different language expressions and query structures; Step 454: Freeze the parameters of the original model to avoid expensive update operations on a large number of parameters; Step 455: Add a trainable low-rank matrix based on the original model by decomposing the parameter matrix to be updated into the product of two low-rank matrices; Step 456: In the fine-tuning stage, only the newly added matrix is trained, the original model parameters are kept unchanged, and the newly added parameters are adjusted by optimizing the objective function to improve the performance of the model on a specific task; Step 457: During inference, the matrix product obtained through training is added to the original model parameters to obtain the final parameter matrix without adding additional inference delay; Step 458: After completing the model fine-tuning, use the optimized TextToSQL model in combination with the prompt word engineering to generate industry-specific SQL statements. By understanding natural language queries and mapping them to structured SQL queries, ensure that the generated SQL statements meet the business needs and logical requirements of the specific industry, thereby achieving accurate operations on the database.
Citation Information
Patent Citations
Text2SQL semantic parsing method for domain large language model
CN118377796A
Data analysis report generation method based on large language model
CN118626523A