LLM-based data analysis knowledge retrieval method and system
By adopting a hybrid recall strategy of dense vectors and sparse vectors and knowledge clustering and reorganization technology in LLM Agent, the problems of poor recall results and complex output in knowledge text retrieval are solved, and more efficient and accurate knowledge retrieval is achieved.
Patent Information
- Application Number
- CN202510592195.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-06
AI Technical Summary
When searching knowledge text, the recall effect of traditional LLM Agent is poor and cannot match the user's needs and some contents of the knowledge text, resulting in a decrease in recall rate. When sorting output, the lengthy and complex knowledge text leads to unstable performance of downstream agents.
The LLM-based data analysis knowledge retrieval method is adopted, and multiple recalls are realized through a hybrid recall strategy combining dense vectors and sparse vectors, and the recall data is roughly arranged and finely arranged. Finally, the knowledge clustering and reorganization output is output based on the LLM vector.
It significantly improves the accuracy and recall rate of search results, ensures that downstream agents can obtain accurate and complete knowledge matching results, and improve the overall application level.
Smart Images

Figure CN120104776A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data intelligent analysis, and in particular to a data analysis knowledge retrieval method, system and electronic device based on LLM. Background Art
[0002] In the research and development system of enterprise-level LLM Agent (Large Language Model Agent), knowledge retrieval has almost become an essential link. By introducing domain knowledge related to user problems and implementing RAG (Retrieval-Augmented Generation) technology, the performance of LLMAgent in various tasks can be significantly improved.
[0003] However, the traditional LLM Agent has the following technical problems when performing data retrieval and analysis: However, the traditional LLM Agent has the following technical problems when performing knowledge text retrieval: On the one hand, when performing knowledge text retrieval, traditional LLM Agents will directly process the entire knowledge text into a single dense vector, resulting in poor recall effect and failure to match user demand questions with partial content in the knowledge text, which limits the scope of recall and causes a decrease in recall rate.
[0004] On the other hand, in the sorting output of recall results, the traditional LLM Agent's RAG technology solution will directly splice the retrieved knowledge items into a text with simple line breaks and provide it to the downstream LLM Agent. However, if there are still many knowledge items and the logical complexity of their content is very high, the lengthy and difficult to understand contextual knowledge will cause the performance of the downstream Agent to be less than expected or unstable. Summary of the invention
[0005] In order to solve the above problems, the present application proposes a data analysis knowledge retrieval method, system and electronic device based on LLM.
[0006] On the one hand, this application proposes a data analysis knowledge retrieval method based on LLM, comprising the following steps: S1. Obtain the user's data analysis requirements; S2. Use LLM Agent to parse the analysis questions in the data analysis requirements and pre-extract keyword texts in the analysis questions; S3, adopting a hybrid recall strategy combining dense vectors and sparse vectors to implement multi-way recall of the analysis question and the pre-extracted keyword texts, and obtaining corresponding data knowledge recall data; S4, performing rough sorting and fine sorting on the data knowledge recall data, and screening knowledge based on the sorting results; S5. Based on the knowledge clustering and reorganization of LLM vectors, the sorted and filtered knowledge is reorganized and output.
[0007] As an optional implementation scheme of the present application, optionally, the keyword text includes at least one of the following keywords: Entity word, indicator word, or aggregation dimension.
[0008] As an optional implementation scheme of the present application, optionally, in step S3, the dense vector is: Using the LLM Embedding model in advance, all knowledge texts are vectorized and stored in the ElasticSearch search engine to obtain the dense vector; In a real-time retrieval task, constructing a corresponding retrieval task according to the analysis question and the dense vector of the pre-extracted keyword text; The retrieval task includes at least one retrieval path, and the retrieval path is configured with metadata filtering conditions corresponding to the data knowledge to be recalled.
[0009] As an optional implementation scheme of the present application, optionally, in step S3, the sparse vector is: Using Elastic Search to configure a traditional NLP word segmentation model, and according to the analysis question and the word segmentation frequency of the pre-extracted keyword text, counting the search score of the corresponding word segmentation; In a real-time retrieval task, constructing a corresponding retrieval task according to the analysis question and the sparse vector of the pre-extracted keyword text; The retrieval task includes at least one retrieval path, and the retrieval path is configured with metadata filtering conditions corresponding to the data knowledge to be recalled.
[0010] As an optional implementation scheme of the present application, optionally, in S4, the roughly sorting the data knowledge recall data includes: respectively configuring corresponding strategy weights for the recall strategy of the dense vector and the recall strategy of the sparse vector in the hybrid recall strategy; Acquire timeliness information of knowledge; Combining the timeliness information and the strategy weights configured on each recall strategy, a weight score calculation is performed on the corresponding recall knowledge in the data knowledge recall data; The weight score of each piece of recalled knowledge in the data knowledge recall data is sorted in descending order, and the recalled knowledge with a weight score not less than a preset value is retained to obtain a weighted sorted coarse-sorted knowledge data set.
[0011] As an optional implementation scheme of the present application, optionally, the weight score is calculated as follows: , Score the final weight of the recalled knowledge, where a larger value indicates a higher priority; The policy weight configured in the recall policy (preset value, range: 0~1); is the knowledge time decay factor, which is defined as the difference between the knowledge release time and the current time (unit: day); is the time attenuation coefficient (value range: 0.02~0.1), which adjusts the influence of timeliness on the weight.
[0012] As an optional implementation scheme of the present application, optionally, in S4, the fine sorting of the data knowledge recall data includes: Pre-build knowledge assessment criteria including scoring indicators for each dimension and train LLM; Using LLM, batch scoring is performed on each piece of the recalled knowledge in the rough sorted knowledge data set according to the scoring indicators of each dimension, and a comprehensive score is generated for each piece of the recalled knowledge; The comprehensive scores of the recalled knowledge are statistically sorted, and the recalled knowledge whose comprehensive scores meet a preset value is retained to obtain a refined knowledge data set.
[0013] As an optional implementation scheme of the present application, optionally, the comprehensive score is calculated as follows: , is the comprehensive score of recalled knowledge; is the weight of the i-th dimension indicator; is the original score of the i-th dimension; It is the user behavior feedback gain coefficient; The frequency of user interactions; is the time-effect attenuation factor; This is a business rule adjustment item.
[0014] On the other hand, the present application proposes a system for implementing the data analysis knowledge retrieval method based on LLM, comprising: Input module, used to obtain the user's data analysis requirements to be analyzed; A preprocessing module, used to parse the analysis questions in the data analysis requirements using LLM Agent, and pre-extract keyword texts in the analysis questions; A hybrid recall module is used to adopt a hybrid recall strategy combining dense vectors and sparse vectors to realize multi-way recall of the analysis question and the pre-extracted keyword text to obtain corresponding data knowledge recall data; A sorting module, used to perform rough sorting and fine sorting on the data knowledge recall data, and filter knowledge based on the sorting results; The output module is used to cluster and reorganize the knowledge based on the LLM vector, and to reorganize and output the sorted and filtered knowledge.
[0015] In another aspect, the present application further provides an electronic device, comprising: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement the data analysis knowledge retrieval method based on LLM when executing the executable instructions.
[0016] Technical effects of the present invention: The present invention aims at LLM Agent in the field of data analysis, and constructs a knowledge retrieval system based on LLM. According to the characteristics of data analysis knowledge, on the one hand, the natural language generation ability and text vectorization ability of LLM are fully utilized, and the problem keywords are pre-extracted by LLM. In the subsequent links, the information is used in a targeted and focused manner, which will help to significantly improve the accuracy of the retrieval results. On the other hand, a multi-source multi-modal vector parallel multi-way recall strategy is adopted, which organically combines traditional retrieval technology, so that downstream related agents can obtain an accurate and complete knowledge matching result as context in real time, thereby improving the overall level of agent application.
[0017] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the disclosure and, together with the description, serve to explain the principles of the disclosure.
[0019] Figure 1 It is a schematic diagram of the implementation process of the present invention. DETAILED DESCRIPTION
[0020] Various exemplary embodiments, features and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0021] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0022] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the following specific embodiments. It should be understood by those skilled in the art that the present disclosure can also be implemented without certain specific details. In some examples, means, components and circuits well known to those skilled in the art are not described in detail in order to highlight the main purpose of the present disclosure. Example 1
[0023] like Figure 1 As shown, on the one hand, this application proposes a data analysis knowledge retrieval method based on LLM, comprising the following steps: S1. Obtain the user's data analysis requirements; S2. Use LLM Agent to parse the analysis questions in the data analysis requirements and pre-extract keyword texts in the analysis questions; S3, adopting a hybrid recall strategy combining dense vectors and sparse vectors to implement multi-way recall of the analysis question and the pre-extracted keyword texts, and obtaining corresponding data knowledge recall data; S4, performing rough sorting and fine sorting on the data knowledge recall data, and screening knowledge based on the sorting results; S5. Based on the knowledge clustering and reorganization of LLM vectors, the sorted and filtered knowledge is reorganized and output.
[0024] The present invention aims at LLM Agent in the field of data analysis, and constructs a knowledge retrieval system based on LLM. According to the characteristics of data analysis knowledge, on the one hand, the natural language generation ability and text vectorization ability of LLM are fully utilized, and the problem keywords are pre-extracted by LLM. In the subsequent links, the information is used in a targeted and focused manner, which will help to significantly improve the accuracy of the retrieval results. On the other hand, a multi-source multi-modal vector parallel multi-way recall strategy is adopted, which organically combines traditional retrieval technology, so that downstream related agents can obtain an accurate and complete knowledge matching result as context in real time, thereby improving the overall level of agent application.
[0025] The implementation process of the present invention will be described in detail below.
[0026] The LLM Agent relied on by the present invention is a self-developed LLM Agent. Other LLM models are also applicable under the application of the present invention.
[0027] The data analysis needs to be analyzed by users can be data analysis problems in various fields.
[0028] As an optional implementation scheme of the present application, optionally, the keyword text includes at least one of the following keywords: Entity word, indicator word, or aggregation dimension.
[0029] LLM has great advantages over traditional NLP (Natural Language Processing) models in performing keyword extraction tasks. Pre-extracting keywords in user analysis needs and making targeted and focused use of this information in subsequent links will help significantly improve the accuracy of search results.
[0030] Specifically, the keywords in data analysis requirements usually include entity words, indicator words, aggregation dimensions (which can also be regarded as a type of entity word), etc. This step uses LLM Agent to parse the user's analysis questions, extract a series of keywords and return them in JSON format. The system will save these keywords for subsequent retrieval steps.
[0031] Use LLM Agent to parse the analysis questions in the data analysis requirements and pre-extract the keyword text in the analysis questions. For detailed steps and prompt words of the thinking chain, please refer to: Step 1: Initial understanding of the analysis problem Read the Question: First, carefully read the analysis question raised in the data analysis requirements to ensure a comprehensive understanding of the problem.
[0032] Identify the problem type: Determine the type of analytical problem, such as descriptive, exploratory, causal, or predictive.
[0033] Step 2: Decompose the problem structure Divide the problem into components: Break down the analysis problem into multiple components, such as problem background, analysis objectives, constraints, etc.
[0034] Identify key elements: In each component, identify the key elements that are crucial to problem analysis, such as time frame, data indicators, analysis objects, etc.
[0035] Step 3: Pre-extract keyword text Mark keywords: Based on the recognition results of step 2, mark the keywords in the question. Keywords usually include but are not limited to data indicator names, analysis object names, time range words, comparison or association words, etc.
[0036] Build a keyword list: Organize the marked keywords into a list for subsequent processing.
[0037] Step 4: Use LLM Agent for further in-depth analysis Get parsing results: LLM Agent will perform natural language processing on the input question and output the parsing results of the question, including the main idea of the question, key information, potential analysis directions, etc.
[0038] Verify keywords: Based on the analysis results of LLM Agent, verify whether the previously extracted keywords are accurate and comprehensive. If necessary, supplement or adjust the keyword list.
[0039] Step 5: Optimize the keyword extraction process Feedback Adjustment: Based on the analysis results and keyword verification of LLM Agent, optimize the keyword extraction process, such as adjusting keyword annotation rules, increasing the dimension of keyword extraction, etc.
[0040] Iterative Improvement: After dealing with multiple analysis problems, we continue to accumulate experience and iteratively improve the process and accuracy of keyword extraction.
[0041] Example The hypothetical analysis question is: "Please analyze the changes in sales of the company's products over the past year and compare the sales growth rates of different product lines." Step 1: Understand the problem as a combination of descriptive and comparative analysis.
[0042] Step 2: Question background: within the past year.
[0043] Analysis objectives: Changes in sales of the company's products and sales growth rates of different product lines.
[0044] Restrictions: No specific restrictions.
[0045] Step 3: Keyword annotations: sales, changes, product lines, growth rate, past year.
[0046] Keyword list: [sales, changes, product lines, growth rate, past year].
[0047] Get the analysis results and verify the accuracy of the keywords. The LLM Agent may emphasize that the "time range" is "the past year" and the "analysis object" is "the sales of the company's products" and "the sales growth rate of different product lines."
[0048] Step 5: According to the feedback from LLM Agent, it is confirmed that the keyword extraction is accurate and no adjustment is needed.
[0049] When dealing with subsequent issues, continue to optimize the keyword extraction process to improve accuracy and efficiency.
[0050] The present invention adopts a hybrid recall strategy combining dense vectors and sparse vectors to realize multi-way recall of the analysis question and the pre-extracted keyword text.
[0051] As an optional implementation scheme of the present application, optionally, in step S3, the dense vector is: Using the LLM Embedding model in advance, all knowledge texts are vectorized and stored in the ElasticSearch search engine to obtain the dense vector; In a real-time retrieval task, constructing a corresponding retrieval task according to the analysis question and the dense vector of the pre-extracted keyword text; The retrieval task includes at least one retrieval path, and the retrieval path is configured with metadata filtering conditions corresponding to the data knowledge to be recalled.
[0052] As an optional implementation scheme of the present application, optionally, in step S3, the sparse vector is: Using Elastic Search to configure a traditional NLP word segmentation model, and according to the analysis question and the word segmentation frequency of the pre-extracted keyword text, counting the search score of the corresponding word segmentation; In a real-time retrieval task, constructing a corresponding retrieval task according to the analysis question and the sparse vector of the pre-extracted keyword text; The retrieval task includes at least one retrieval path, and the retrieval path is configured with metadata filtering conditions corresponding to the data knowledge to be recalled.
[0053] Based on the original user questions and the pre-extracted related keyword texts, multi-way recall can be achieved. On this basis, the present invention adopts a hybrid recall strategy combining dense vectors and sparse vectors, which not only gives full play to the powerful complex semantic representation ability of LLM, can match semantically similar but differently expressed knowledge to avoid omissions, but also makes full use of the effectiveness of traditional methods, ensures retrieval quality, realizes complementary technical advantages, and improves the comprehensiveness and accuracy of query hits.
[0054] The specific method is as follows: Dense vectors refer to using the LLM Embedding model to pre-vectorize all knowledge texts and store them in the Elastic Search search engine. In real-time retrieval tasks, the above questions and keyword texts are vectorized separately as part of the query input information; sparse vectors use the BM25 retrieval model, and by using Elastic Search to configure the traditional NLP word segmentation model, retrieval scoring can be achieved based on statistical indicators such as word segmentation frequency. Therefore, user questions and keywords will be used to construct dense vectors and sparse vectors respectively to carry out retrieval tasks, forming at least 4 retrieval paths.
[0055] In addition, each pathway will also include hard conditions such as metadata filtering, such as limiting the scope of topics, limiting the scope of data tables, etc. The specific conditions depend on the upstream output of the application process and the RAG requirements of the downstream agent to ensure that irrelevant knowledge is filtered out during the initial recall.
[0056] Example of a hybrid recall strategy: combining dense and sparse vectors Dense Vector Recall Preprocessing phase: Use LLM (Language Learning Model) Embedding models, such as OpenAI's text-embedding-3-small, BAAI's BGE-M3, etc., to vectorize all knowledge texts.
[0057] The vectorized knowledge text is stored in the Elastic Search search engine, and each text corresponds to a dense vector representation.
[0058] Real-time retrieval phase: Vectorize the analysis question and pre-extracted keyword text to generate a query vector.
[0059] The query vector is input into Elastic Search, and is matched with the stored dense vector using metrics such as cosine similarity or dot product to obtain the N texts with the highest similarity as candidate results.
[0060] Sparse Vector Recall Configure Elastic Search: Configure traditional NLP word segmentation models in Elastic Search, such as IK Analyzer and Standard Analyzer, to perform word segmentation on text.
[0061] Based on the word segmentation results, the BM25 (Best Matching 25) retrieval model is used to calculate the similarity score between the text and the query.
[0062] Real-time retrieval phase: The analysis question and pre-extracted keyword text are input into Elastic Search as queries.
[0063] Elastic Search performs word segmentation and scoring on the query text and knowledge text based on the configured word segmentation model and BM25 algorithm, and obtains the M texts with the highest scores as candidate results.
[0064] Mixed recall strategy implementation Multi-way recall: Dense vector recall and sparse vector recall are performed simultaneously to obtain N and M candidate results respectively.
[0065] Merge the result sets of all strategies to obtain a candidate result set that may contain duplicates.
[0066] Result Fusion: The similarity scores of dense vectors and sparse vector related strategies are weighted and combined with the knowledge time decay factor to calculate the comprehensive scores of the candidate results for sorting.
[0067] According to the sorting results, the top K best texts are selected as the final recall results.
[0068] Optimization and Adjustment: According to the actual application effect, adjust the weights of dense vectors and sparse vectors, as well as the values of N, M, and K to optimize recall performance and accuracy.
[0069] As an optional implementation scheme of the present application, optionally, in S4, the roughly sorting the data knowledge recall data includes: respectively configuring corresponding strategy weights for the recall strategy of the dense vector and the recall strategy of the sparse vector in the hybrid recall strategy; Acquire timeliness information of knowledge; Combining the timeliness information and the strategy weights configured on each recall strategy, a weight score calculation is performed on the corresponding recall knowledge in the data knowledge recall data; The weight score of each piece of recalled knowledge in the data knowledge recall data is sorted in descending order, and the recalled knowledge with a weight score not less than a preset value is retained to obtain a weighted sorted coarse-sorted knowledge data set.
[0070] Weighted sorting, that is, based on the results of multiple recalls, different weights are assigned to each recall strategy in advance (usually higher weights are set for keyword-related strategies), and then the weight of each candidate knowledge item is automatically adjusted in real time based on the timeliness of the knowledge (based on the knowledge update date or the latest date of positive and negative feedback on related questions and answers). The knowledge with low timeliness is appropriately downgraded to form a total score for each recalled knowledge item. Finally, the knowledge items are sorted in descending order by score and the first few pieces of knowledge are retained based on the score.
[0071] See the weighting scheme below: As an optional implementation scheme of the present application, optionally, the weight score is calculated as follows: , Score the final weight of the recalled knowledge, the larger the value, the higher the priority; The policy weight configured in the recall policy (preset value, range: 0~1); is the knowledge time decay factor, which is defined as the difference between the knowledge release time and the current time (unit: day); is the time attenuation coefficient (value range: 0.02~0.1), which adjusts the influence of timeliness on the weight.
[0072] Formula structure: Weight score = strategy weight × time decay factor. Δt is the time difference from the knowledge release time to the current time, and α is the parameter for adjusting the decay speed. The formula takes into account both the importance of keywords and the timeliness of knowledge.
[0073] In an example: First, calculate the knowledge release time With current time The number of days between: , Apply an exponential decay function to limit the decay rate of timeliness: ; Secondly, the weights are calculated comprehensively: Combining strategy weights with time factors: , Carry out knowledge calculation item by item; Sort all recalled knowledge in descending order by W, and output the Top-N results (the N value is defined by the user, and the top N knowledge is output).
[0074] The various parameters can be set according to the following table 1: Table 1 parameter parameter Value suggestions K Reflects the relevance of keywords and business scenarios Static configuration (0.1~1.0) λ Control the rate at which knowledge becomes obsolete (the larger the value, the faster the decay) Finance: 0.1 Consumer: 0.02 Δt threshold Set an upper limit on the validity period of knowledge Financial sector: 30-60 days Technology sector: 180-360 days Example calculation: Assume that the parameters of a financial knowledge item are as follows: Release time: 2025-02-10 (current time: 2025-02-27, Δt=17), Strategy weight: Kw=0.8, Attenuation coefficient: λ=0.05, The weight calculation is: W=0.8×e−0.05×17≈0.8×0.427≈0.342, Explanation of results: The weight of this knowledge item has decreased by 57.3% due to reduced timeliness. The lambda value needs to be adjusted to optimize the sorting according to business needs.
[0075] Technical advantages: Dynamically balance keyword importance and knowledge freshness through the exponential decay model, and support flexible weight configuration of multiple recall strategies.
[0076] As an optional implementation scheme of the present application, optionally, in S4, the fine sorting of the data knowledge recall data includes: Pre-build knowledge assessment criteria including scoring indicators for each dimension and train LLM; Using LLM, batch scoring is performed on each piece of the recalled knowledge in the rough sorted knowledge data set according to the scoring indicators of each dimension, and a comprehensive score is generated for each piece of the recalled knowledge; The comprehensive scores of the recalled knowledge are statistically sorted, and the recalled knowledge whose comprehensive scores meet a preset value is retained to obtain a refined knowledge data set.
[0077] The refined sorting stage makes full use of the problem understanding and analysis capabilities of LLM, and introduces the agent to directly score the knowledge retained in the previous stage in batches. Based on the given fixed evaluation criteria, the LLM Agent can give each piece of knowledge a comprehensive score based on the performance of indicators in various dimensions (such as importance, completeness, relevance, logic, timeliness, etc.). Among them, for multiple knowledge items with a high degree of content overlap, it is also necessary to determine whether to eliminate redundant items.
[0078] Refer to the following refined arrangement method: As an optional implementation scheme of the present application, optionally, the comprehensive score is calculated as follows: , is the comprehensive score of recalled knowledge; is the weight of the i-th dimension indicator; is the original score of the i-th dimension; It is the user behavior feedback gain coefficient; The frequency of user interactions; is the time-effect attenuation factor; This is a business rule adjustment item.
[0079] The definitions and operations of the various formulas for calculating the comprehensive score are shown in Table 2: Table 2 symbol meaning Operational logic R Comprehensive score of recalled knowledge (for rough sorting) Dynamic adjustment items are added after weighted summation of each dimension wi The weight of the i-th dimension indicator (must satisfy ∑wi=1) Dynamic optimization through offline A / B testing or online learning (such as entropy weight method, SHAP value analysis) Si The raw score of the i-th dimension (such as relevance, quality, completeness, etc.) Normalized to the range of 0, 10, 1 (Min-Max or Z-Score normalization) λ User behavior feedback gain coefficient (default 0.2) Positive behaviors such as user clicks / collections / reposts trigger / value accumulation (decay period T=7 days) I User interaction frequency (click, favorite, etc.) Taking logarithms to avoid long-tail distribution bias (log(1+I) prevents zero values from being invalid) η Time-effect decay factor (default 0.15) Exponential decay based on the knowledge storage time t (unit: day) (τ=30 days half-life) ΔA Business rule adjustment items (such as mandatory top, blacklist filtering, etc.) Manual setting or strategy engine output (range −0.5, +0.5) Calculation example: Data input: The relevance score of a certain recalled knowledge is 0.85, the quality score is 0.7, the number of user interactions is I=15, and the storage time is t=10 days Weighted calculation: R=1.03+0.2×2.77+0.15×0.716+0.1≈1.03+0.55+0.11+0.1=1.79.
[0080] Output sorting: Sort by R value in descending order, with high-scoring segments entering the refined sorting stage first.
[0081] Through multi-level regulation of dynamic weight + time decay + business intervention, the accuracy of top knowledge is improved while ensuring sorting efficiency (computational complexity).
[0082] And finally, the knowledge clustering reorganization based on LLM vector is: The present invention designs a knowledge clustering and reorganization technology to improve this problem: First, for the knowledge items obtained after the refined sorting stage, their pre-stored Embedding vectors (generated based on the LLM Embedding model) are obtained respectively, and clustering and grouping are performed using the DBSCAN model. Secondly, based on the knowledge score in the refined sorting stage, for each group of knowledge, the score is used to sort the knowledge within the group (in descending order), and for each group, the highest score of each group is taken as the group score, and the groups are sorted (in descending order). Finally, based on the above grouping and sorting, the knowledge text is spliced and provided to the downstream LLM Agent.
[0083] Through clustering and reorganization, the knowledge text is segmented in a more reasonable and logical way, which helps to give full play to the advantages of RAG technology and improve the performance of the final LLM generation task. Obviously, those skilled in the art should understand that the implementation of all or part of the processes in the above-mentioned embodiments can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments. Those skilled in the art should understand that the implementation of all or part of the processes in the above-mentioned embodiments can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments.
[0084] Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (Flash Memory), a hard disk (Hard Disk Drive, abbreviated as: HDD) or a solid-state drive (Solid-State Drive, SSD), etc.; the storage medium can also include a combination of the above types of memory. Example 2
[0085] Based on the implementation principle of Example 1, on the other hand, the present application proposes a system for implementing the LLM-based data analysis knowledge retrieval method, including: Input module, used to obtain the user's data analysis requirements to be analyzed; A preprocessing module, used to parse the analysis questions in the data analysis requirements using LLM Agent, and pre-extract keyword texts in the analysis questions; A hybrid recall module is used to adopt a hybrid recall strategy combining dense vectors and sparse vectors to realize multi-way recall of the analysis question and the pre-extracted keyword text to obtain corresponding data knowledge recall data; A sorting module, used to perform rough sorting and fine sorting on the data knowledge recall data, and filter knowledge based on the sorting results; The output module is used to cluster and reorganize the knowledge based on the LLM vector, and to reorganize and output the sorted and filtered knowledge.
[0086] The interaction of the above modules can be understood in conjunction with the corresponding steps of the method in Example 1, and will not be described in detail in this example.
[0087] The modules or steps of the present invention described above can be implemented by a general-purpose computing system, they can be concentrated on a single computing system, or distributed on a network composed of multiple computing systems, and optionally, they can be implemented by a program code executable by a computing system, so that they can be stored in a storage system and executed by the computing system, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software. Example 3
[0088] Furthermore, in another aspect, the present application also proposes an electronic device, comprising: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement a data analysis knowledge retrieval method based on LLM as described in Example 1 when executing the executable instructions.
[0089] The electronic device of the embodiment of the present disclosure includes a processor and a memory for storing processor executable instructions. The processor is configured to implement the data analysis knowledge retrieval method based on LLM described in the above embodiment 1 when executing the executable instructions.
[0090] Here, it should be noted that the number of processors can be one or more. At the same time, the electronic device of the embodiment of the present disclosure may also include an input system and an output system. Among them, the processor, memory, input system and output system may be connected through a bus or in other ways, which are not specifically limited here.
[0091] The memory, as a computer-readable storage medium, can be used to store software programs, computer executable programs and various modules, such as the program or module corresponding to the data analysis knowledge retrieval method based on LLM in the embodiment of the present disclosure. The processor executes various functional applications and data processing of the electronic device by running the software programs or modules stored in the memory.
[0092] The input system can be used to receive input numbers or signals. The signal can be a key signal related to user settings and function control of the device / terminal / server. The output system can include display devices such as display screens.
[0093] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A data analysis knowledge retrieval method based on LLM, characterized in that: The steps include: S1. Obtain the user's data analysis requirements; S2. Use LLM Agent to parse the analysis questions in the data analysis requirements and pre-extract keyword texts in the analysis questions; S3, adopting a hybrid recall strategy combining dense vectors and sparse vectors to implement multi-way recall of the analysis question and the pre-extracted keyword texts, and obtaining corresponding data knowledge recall data; S4, performing rough sorting and fine sorting on the data knowledge recall data, and screening knowledge based on the sorting results; S5. Based on the knowledge clustering and reorganization of LLM vectors, the sorted and filtered knowledge is reorganized and output.
2. The data analysis knowledge retrieval method based on LLM according to claim 1 is characterized in that: The keyword text includes at least one of the following keywords: Entity word, indicator word, or aggregation dimension.
3. The data analysis knowledge retrieval method based on LLM according to claim 1 is characterized in that: In step S3, the dense vector is: Using the LLM Embedding model in advance, all knowledge texts are vectorized and stored in the Elastic Search search engine to obtain the dense vector; In a real-time retrieval task, constructing a corresponding retrieval task according to the analysis question and the dense vector of the pre-extracted keyword text; The retrieval task includes at least one retrieval path, and the retrieval path is configured with metadata filtering conditions corresponding to the data knowledge to be recalled.
4. The data analysis knowledge retrieval method based on LLM according to claim 1 is characterized in that: In step S3, the sparse vector is: Using Elastic Search to configure a traditional NLP word segmentation model, and according to the analysis question and the word segmentation frequency of the pre-extracted keyword text, counting the search score of the corresponding word segmentation; In a real-time retrieval task, constructing a corresponding retrieval task according to the analysis question and the sparse vector of the pre-extracted keyword text; The retrieval task includes at least one retrieval path, and the retrieval path is configured with metadata filtering conditions corresponding to the data knowledge to be recalled.
5. The data analysis knowledge retrieval method based on LLM according to claim 1 is characterized in that: In S4, the data knowledge recall data is roughly sorted, including: respectively configuring corresponding strategy weights for the recall strategy of the dense vector and the recall strategy of the sparse vector in the hybrid recall strategy; Acquire timeliness information of knowledge; Combining the timeliness information and the strategy weights configured on each recall strategy, a weight score calculation is performed on the corresponding recall knowledge in the data knowledge recall data; The weight score of each piece of recalled knowledge in the data knowledge recall data is sorted in descending order, and the recalled knowledge with a weight score not less than a preset value is retained to obtain a weighted sorted coarse-sorted knowledge data set.
6. The data analysis knowledge retrieval method based on LLM according to claim 5 is characterized in that: The weight score is calculated as follows: , Score the final weight of the recalled knowledge, where a larger value indicates a higher priority; The policy weight configured in the recall policy; is the knowledge time decay factor, which is defined as the difference between the knowledge release time and the current time; is the time decay coefficient, which adjusts the impact of timeliness on the weight.
7. The data analysis knowledge retrieval method based on LLM according to claim 5 is characterized in that: In S4, the data knowledge recall data is subjected to fine sorting processing, including: Pre-build knowledge assessment criteria including scoring indicators for each dimension and train LLM; Using LLM, batch scoring is performed on each piece of the recalled knowledge in the rough sorted knowledge data set according to the scoring indicators of each dimension, and a comprehensive score is generated for each piece of the recalled knowledge; The comprehensive scores of the recalled knowledge are statistically sorted, and the recalled knowledge whose comprehensive scores meet a preset value is retained to obtain a refined knowledge data set.
8. The data analysis knowledge retrieval method based on LLM according to claim 7 is characterized in that: The comprehensive score is calculated as follows: , is the comprehensive score of recalled knowledge; is the weight of the i-th dimension indicator; is the original score of the i-th dimension; It is the user behavior feedback gain coefficient; The frequency of user interactions; is the time-effect attenuation factor; This is a business rule adjustment item.
9. A system for implementing the LLM-based data analysis knowledge retrieval method according to any one of claims 1 to 8, characterized in that: include: Input module, used to obtain the user's data analysis requirements to be analyzed; A preprocessing module, used to parse the analysis questions in the data analysis requirements using LLM Agent, and pre-extract keyword texts in the analysis questions; A hybrid recall module is used to adopt a hybrid recall strategy combining dense vectors and sparse vectors to realize multi-way recall of the analysis question and the pre-extracted keyword text to obtain corresponding data knowledge recall data; A sorting module, used to perform rough sorting and fine sorting on the data knowledge recall data, and filter knowledge based on the sorting results; The output module is used to cluster and reorganize the knowledge based on the LLM vector, and to reorganize and output the sorted and filtered knowledge.
10. An electronic device, characterized in that include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement a data analysis knowledge retrieval method based on LLM as described in any one of claims 1 to 8 when executing the executable instructions.
Citation Information
Patent Citations
A method and apparatus for information push
CN106228386A
Commodity recommendation method and system
CN115953223A
Commodity searching method and device, equipment and medium
CN116861083A
Retrieval method and system for efficiently collecting global accurate potential customer information
CN119205277A
Legal provision retrieval method for legal data embedding optimization and retrieval effect evaluation
CN119377384A
Cited By
Intelligent analysis and text optimization method and system for external financial and economic words
CN121543600A