Language analysis method and system based on index semantic layer and large language model
By combining the index semantic layer and large language model, the traditional data analysis system is solved inequality and semantic analysis of non-technical users and the NLQ solution of large language model NLQ solution, and efficient data analysis and visualization of users through natural language is realized, improving user experience and query accuracy.
Patent Information
- Application Number
- CN202510296418.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-24
AI Technical Summary
Traditional data analysis systems are not friendly to non-technical users, and the NLQ scheme of traditional large language models has shortcomings in accuracy and semantic analysis, making it difficult to understand the implicit intentions and context of users.
By combining the metric semantic layer and the large language model, an metric semantic layer containing metadata such as indicator names, dimensions, and filters is built, and a vector representation is generated to match the user's natural language requests. The user request input to the large language model for inference, generate formatted query results, and convert them into query statements at the index semantic layer, perform data query operations, and finally visually present them through the natural language interface.
It realizes efficient data analysis and visualization by users through natural language, improves user experience and query accuracy, can understand the implicit intentions and context in user queries, and provides more intuitive and personalized data analysis results.
Smart Images

Figure CN120196645A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data visualization analysis, and particularly to a language analysis method and system based on an index semantic layer and a large language model. Background Art
[0002] Traditional data analysis systems often require users to have profound professional technical knowledge and proficient operation skills. This poses a certain entry barrier for users lacking programming backgrounds or data processing experience. These systems usually rely on complex query languages such as SQL or specific analysis tools, requiring users to have in-depth knowledge of database structures, data query syntax, and data analysis methods. Therefore, without corresponding training and support, non-technical users find it difficult to directly utilize these systems for data exploration and decision support.
[0003] In addition, the NLQ solutions of traditional large language models also have deficiencies in terms of stability and reliability, making it difficult to support the requirements of data analysis systems. Specifically, these solutions have the following two main problems:
[0004] (1) Accuracy issue: Due to the hallucination problem of large language models, their parsing accuracy still needs to be improved. Incorrect parsing will lead to incorrect data analysis results, thus affecting the quality of decisions based on the data analysis results. This is a serious challenge for enterprises that rely on accurate data analysis results.
[0005] (2) Limitations of semantic parsing: The NLQ solutions of traditional large language models often have difficulty fully understanding the implicit intentions and context in queries. Especially in specific industries or professional fields, their standard terms and expressions may be different from common usage. This makes it difficult for the system to provide accurate analysis results and unable to meet the needs of users in practical applications.
[0006] Therefore, the present invention proposes a language analysis method and system based on an index semantic layer and a large language model. Summary of the Invention
[0007] In view of this, the present invention hopes to provide a language analysis method and system based on an index semantic layer and a large language model to solve or alleviate the technical problems existing in the prior art, that is, how to support users to achieve data visualization analysis through natural language by combining the index semantic layer and the large language model, enable traditional NLQ solutions to understand the implicit intentions and context in queries, avoid parsing errors, and provide at least one beneficial option; the technical solution of the present invention is realized as follows:
[0008] In the first aspect, a language analysis method based on an index semantic layer and a large language model:
[0009] (1) Overview:
[0010] The present invention aims to achieve efficient data analysis and visualization by users through natural language by combining the indicator semantic layer with a large language model. First, an indicator semantic layer containing metadata such as indicator names, dimensions, and filters is constructed, and vector representations are generated for these metadata to match the natural language requests input by users. When a user inputs a natural language query, its text is vectorized and the Euclidean distance is calculated with the metadata vectors in the indicator semantic layer to determine the indicators involved in the user query. Subsequently, these indicators and the user request are input into the large language model for inference to generate formatted query results. Finally, these results are converted into query statements in the indicator semantic layer, data query operations are executed, and the query results are visually presented through a natural language interface, thereby providing users with convenient and intuitive data analysis and decision-making support.
[0011] (II) Technical solution:
[0012] After receiving a natural language query request input by a user through the user interface, the following steps are executed:
[0013] 2.1 Step S1, data engineering:
[0014] Read the metadata constructed by the predefined data indicator model;
[0015] Use the metadata to construct the Prompt full words of the indicators to provide a basis for subsequent natural language queries;
[0016] Generate vector representations for the same metadata in the indicator semantic layer;
[0017] These vectors will be used to match subsequent natural language requests input by users.
[0018] 2.1.1 Step S100, Prompt full word construction:
[0019] Read the metadata from the predefined data indicator model, including the names of the indicators, dimension definitions, and filter configurations.
[0020] Use the read metadata to construct the Prompt full words of the indicators.
[0021] 2.1.2 Step S101, vector representation generation:
[0022] For each metadata item in the indicator semantic layer, including the indicator name, dimension, and filter, perform text vectorization processing. Use the bag-of-words model or TF-IDF technology to convert the metadata into vector representations.
[0023] 2.2 Step S2, data processing:
[0024] Perform text vectorization on the natural language request input by the user;
[0025] Calculate the Euclidean distance between the processed vector and the metadata vector in the metric semantic layer, and determine the metric Prompt clause involved in the user query and the user's original request input to the large language model according to the threshold;
[0026] The large language model performs reasoning to generate a formatted reasoning result.
[0027] 2.2.1 Step S200, Natural language request vectorization:
[0028] Perform text preprocessing on the natural language request, including word segmentation and stop word removal.
[0029] Use the same bag-of-words model or TF-IDF technology as the metric semantic layer to convert the preprocessed text into a vector representation.
[0030] 2.2.2 Step S201, Vector matching and metric determination:
[0031] Calculate the Euclidean distance between the user request vector and the metadata vector in the metric semantic layer.
[0032] Filter out the metadata vector with the smallest distance from the user request vector according to the preset threshold, and then obtain the metric Prompt clause involved in the user query.
[0033] 2.2.3 Step S202, Large language model reasoning:
[0034] Input the metric Prompt clause and the user's original request into the large language model, perform reasoning based on the input Prompt and request to understand the user's query intention, and generate a formatted reasoning result.
[0035] 2.3 Step S3, Query request conversion:
[0036] Convert the formatted reasoning result into a query statement in the metric semantic layer according to the predefined rules, and access the database or data warehouse through the metric semantic layer to perform a data query operation.
[0037] Perform formatted analysis on the query result.
[0038] 2.3.1 Step S300, Query statement conversion:
[0039] Receive the formatted reasoning result generated by the large language model; according to the predefined rules, convert the structured information in the reasoning result into a query statement that the metric semantic layer can understand.
[0040] 2.3.2 Step S301, Data query execution:
[0041] Through the index semantic layer, the transformed query statement is sent to the database or data warehouse. Perform a data query operation to retrieve the data required by the user from the database or data warehouse.
[0042] 2.3.3 Step S302, Formatting Analysis:
[0043] S3020, Data Transformation: Transform the data in the query result into a format suitable for analysis and visualization, and perform at least one of the following operations:
[0044] Format Unification: Unify the formats of dates, times, or / and currencies, etc.;
[0045] Data Encoding: Convert categorical variables into numerical variables;
[0046] Data Aggregation: Calculate mean, sum, or / and proportion statistics.
[0047] S3021, Cleaning: Calculate descriptive statistics such as the mean, median, mode, or / and standard deviation of the data.
[0048] 2.4 Step S4, Visualization Output:
[0049] Visualize the data result after formatting analysis through the Natural Language Interface (NLI). Generate corresponding visualization elements according to the content requested by the user and return the presented result to the user.
[0050] 2.4.1 Step S400, Visualization Element Selection:
[0051] According to the specific content of the user's request and business requirements, select the visualization element type, including charts, reports, or dashboards.
[0052] 2.4.2 Step S401, Data Binding:
[0053] Bind the data result after formatting analysis to the selected visualization element;
[0054] Integrate the visualization element with the Natural Language Interface (NLI);
[0055] Present the designed visualization element in the user interface.
[0056] 2.4.3 Step S402, User Interaction and Feedback:
[0057] Receive the user's interaction operations, including clicks, swipes, or inputs, etc. Update the visualization element or provide corresponding feedback according to the user's interaction operations.
[0058] (III) Mechanism for Solving Technical Problems:
[0059] Step S1: Data Engineering
[0060] Read metadata from a predefined data metric model. These metadata details the name, dimension, and filter of the metric, and use these metadata to construct the full word of the metric's Prompt. This step provides a clear and rich context for subsequent natural language queries. For the same metadata in the metric semantic layer, we generate corresponding vector representations. These vectors will be matched with the vectors of the natural language requests input by the user in subsequent steps to achieve accurate query intent recognition.
[0061] Calculate the Euclidean distance between the processed user request vector and the metadata vector in the metric semantic layer. According to the set threshold, we can accurately determine the metric Prompt clause involved in the user query and input the original user request into the large language model. The large language model conducts in-depth reasoning on the user request and generates formatted reasoning results. These results not only contain the user's explicit needs but may also cover implicit query intents and context information.
[0062] Convert the formatted reasoning results generated by the large language model into query statements that the metric semantic layer can understand according to predefined rules. This conversion process ensures the accuracy and effectiveness of the query statements. Send the query statements to the database or data warehouse through the metric semantic layer and execute the data query operation. Subsequently, perform formatted analysis on the query results to better meet the user's analysis and visualization needs.
[0063] Second aspect, a language analysis system based on the metric semantic layer and the large language model:
[0064] This system is used to implement the above-mentioned language analysis method based on the metric semantic layer and the large language model, including:
[0065] (1) An application side for receiving user input and displaying the output results of the system, including:
[0066] (1.1) User interface module;
[0067] (1.2) Natural language processing module;
[0068] (1.3) Interaction module;
[0069] (2) A server side for receiving requests from the application side, processing data query and visualization logic, and interacting with the database side, including:
[0070] (2.1) Metric semantic layer module;
[0071] (2.2) Vector representation module;
[0072] (2.3) Query transformation module;
[0073] (2.4) Large language model module;
[0074] (2.5) Data processing and transformation module;
[0075] (3) The database side responsible for storing and managing business data, including:
[0076] (3.1) Relational database;
[0077] (3.2) Non-relational database;
[0078] (3.3) File storage system;
[0079] Compared with the prior art, the beneficial effects of the present invention are:
[0080] I. Improve user experience: In the technical solution of the present invention, users can use natural language for data query and visual analysis without learning complex query languages or operation interfaces. It can understand the implicit intentions and context in the user's query, thus providing more accurate and personalized query results.
[0081] II. Improve query accuracy: In the technical solution of the present invention, by matching the user's query with the metadata vectors in the metric semantic layer, the query intention of the user can be more accurately identified, avoiding parsing errors. The reasoning ability of the large language model further enhances the query accuracy and can handle complex query logics and implicit relationships.
[0082] III. Enhance data visualization effect: In the technical solution of the present invention, corresponding visual elements such as charts, reports or dashboards can be generated according to the specific content of the user's request, making the data results more intuitive and understandable.
[0083] IV. Improve data processing efficiency: By constructing metadata through predefined metric models, the technical solution of the present invention can quickly locate the data metrics and dimensions involved in the user's query, reducing the data processing time. The conversion and execution processes of query statements are optimized, improving the efficiency of data query and processing.
[0084] V. Reduce maintenance costs: In the technical solution of the present invention, the introduction of the metric semantic layer makes the management and maintenance of data metrics more centralized and unified, reducing the system maintenance costs. When the data metrics change, only the metadata in the metric semantic layer needs to be updated, without modifying a large number of query statements and visualization codes. Description of the Drawings
[0085] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0086] Figure 1 Schematic diagram of the method flow of the present invention;
[0087] Figure 2 Schematic diagram of the system composition of the present invention;
[0088] Figure 3 Schematic diagram of the TAML format file of the data model of the present invention;
[0089] Figure 4 Schematic diagram of the TAML format file of the indicator model of the present invention;
[0090] Figure 5 Schematic diagram of the TAML format file of the full-word indicator of the large language model constructed by the present invention;
[0091] Figure 6 Schematic diagram of the TAML format file of the first embodiment of the present invention, where the indicator metadata most relevant to the user input is the completed order volume;
[0092] Figure 7 For the first embodiment of the present invention, obtaining the formatted inference result as Figure 7 shown in the TAML format file schematic diagram;
[0093] Figure 8 Schematic diagram of the first part of the TAML format file for converting to the query statement of the indicator semantic layer in the first embodiment of the present invention;
[0094] Figure 9 Schematic diagram of the second part of the TAML format file for converting to the query statement of the indicator semantic layer in the first embodiment of the present invention;
[0095] Figure 10 Schematic diagram of the effect of visual presentation of the first embodiment of the present invention in NLI. Detailed implementation manners
[0096] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following provides a detailed description of the specific embodiments of the present invention with reference to the accompanying drawings. Many specific details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below;
[0097] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the description of the method part for relevant parts.
[0098] Glossary:
[0099] (1) Large Language Model (LLM): An artificial intelligence model designed to understand and generate human language. They are trained on large amounts of text data and can perform a wide range of tasks, including text summarization, translation, sentiment analysis, etc.
[0100] (2) Metric Semantic Layer (MSL): Through the Metric Semantic Layer, a software system can clearly describe the definition of metrics, the association of data sets, analyzable dimensions, measurable values, optional predicates, etc. Through the definition of the Metric Semantic Layer, the software system can generate clear and reusable data analysis tasks / SQL and obtain a consistent metric calculation caliber.
[0101] (3) Data Visualization (DV): A method of encoding data into graphics or visual objects, aiming to reveal patterns, trends, and insights in the data through visual communication. Data visualization converts data into a graphical or visual format, making the data easier to understand and analyze. Common types of data visualization include bar charts, line charts, pie charts, scatter plots, and heatmaps, etc. Data visualization can help people quickly identify patterns, trends, and anomalies in the data.
[0102] (4) Natural Language Query (NLQ): Allows users to query databases and other data sets in everyday language without the need to master traditional query languages.
[0103] (5) Hallucinations in large language models: The so-called "hallucinations" are actually a figurative expression, usually referring to errors or inaccuracies that occur when the model generates text or makes predictions due to reasons such as training data bias, overfitting, and anomalies in the input data.
[0104] (6) Token: "Token" refers to the basic unit used for processing and generating text. Different language models may have different definitions of tokens, but generally, a token can be a word, part of a word, or even a single character.
[0105] Example 1: As Figure 1 shown, this example will provide a language analysis method based on the metric semantic layer and large language model for the field of online car-hailing telemarketing data analysis and processing. After receiving a natural language query request input by the user through the user interface, the following steps are executed:
[0106] In this example, regarding step S1: Data engineering:
[0107] Specifically, S100 Prompt full-word construction: As Figure 3 shown, read metadata from a predefined online car-hailing telemarketing data metric model, such as the names, dimensions (such as time dimensions: day, week, month) of metrics like "call duration", "number of completed transactions", "customer rejection rate", and filter configurations (such as region, salesperson ID). Use this metadata to construct the Prompt full words of the metrics, for example: "Query the call duration and customer rejection rate of a certain salesperson within a specific time period".
[0108] Specifically, the metric model is in the form of Figure 4 shown.
[0109] Specifically, S101 vector representation generation: Perform text vectorization on the metric names "call duration", "number of completed transactions", "customer rejection rate", dimensions "day", "week", "month", and filters "region", "salesperson ID". Use the TF-IDF technique to convert these texts into vector representations for subsequent matching with user queries.
[0110] Specifically, the metric full words of the large language model constructed in step one are as Figure 5 shown, where text vectorization is only used to reduce the token length of the large language model prompt in natural language queries. Therefore, only a simple bag-of-words model can be used for text vectorization in this method.
[0111] In this example, regarding step S2: Data processing:
[0112] Specifically, S200 Natural Language Request Vectorization: Preprocess the natural language request input by the user, such as "I want to see how many calls Zhang San made this month and how many times he was rejected". Tokenize the text and remove stop words, and then use the same TF-IDF technology as the metric semantic layer to convert the text into a vector representation.
[0113] Exemplarily, if the input is "The completed order data across the country in 2023", then after text vectorization processing and distance calculation, we can obtain the metric metadata most relevant to the user input as the completed order volume as Figure 6 shown. Then input the metric prompt clause and the user request into the large language model to obtain a formatted inference result as Figure 7 shown.
[0114] Specifically, S201 Vector Matching and Metric Determination: Calculate the Euclidean distance between the user request vector and the metadata vectors in the metric semantic layer. According to the threshold, filter out the metadata vectors that best match the user request vector, and determine that the metrics involved in the user query are "call duration" and "customer rejection rate", as well as the dimension "month" and the filter "salesperson ID = Zhang San".
[0115] Specifically, S202 Large Language Model Inference: Input the determined metric Prompt clause "Query the call duration and customer rejection rate of Zhang San within a specific month" and the user's original request into the large language model.
[0116] Exemplarily, the large language model understands the user's query intention and generates a formatted inference result, such as "{Metric: Call Duration, Dimension: Month, Filter: Salesperson ID = Zhang San}, {Metric: Customer Rejection Rate, Dimension: Month, Filter: Salesperson ID = Zhang San}".
[0117] In this embodiment, regarding step S3: Query Request Conversion
[0118] Specifically, S300 Query Statement Conversion: Receive the formatted inference result generated by the large language model. According to predefined rules, convert the inference result into a query statement that the metric semantic layer can understand, such as "SELECT Call Duration, Customer Rejection Rate FROM Sales Data WHERE Salesperson ID = 'Zhang San' AND Date BETWEEN '2023-03-01' AND '2023-03-31'".
[0119] Exemplarily, the effect of converting to a query statement in the metric semantic layer is as Figures 8 - 9 shown.
[0120] Specifically, S301 data query execution: Through the metric semantic layer, send the query statement to the database. Execute the data query operation to retrieve the call duration and customer rejection rate data of Zhang San in March 2023.
[0121] Specifically, S302 formatting and analysis:
[0122] S3020 data conversion: Unify the date format in the query results to "YYYY-MM-DD", and convert categorical variables such as regions into numerical variables (such as region codes).
[0123] S3021 cleaning: Calculate descriptive statistics such as the mean, median, mode, and standard deviation of the call duration and customer rejection rate.
[0124] In this embodiment, regarding step S4: Visualization output:
[0125] Specifically, S400 visualization element selection: According to the content requested by the user, select a chart as the visualization element type. Specifically, a bar chart is used to display the call duration and customer rejection rate of Zhang San in March 2023.
[0126] S401 data binding: Bind the data results after formatting and analysis to the bar chart. Integrate the bar chart with the natural language interface (NLI). Present the designed bar chart in the user interface to display the call duration and customer rejection rate data of Zhang San in March 2023.
[0127] Exemplarily, the effect of retrieving and analyzing data through the metric semantic layer and making a visual presentation in the NLI is as Figure 10 shown.
[0128] Specifically, S402 user interaction and feedback: Receive the user's interaction operations, such as clicking on a data point in the bar chart. According to the user's interaction operations, update the visualization element or provide corresponding feedback, such as displaying the specific value of the data point or providing further analysis information.
[0129] Embodiment 2: As Figure 2 shown, on the basis of Embodiment 1, this embodiment further discloses a language analysis system for implementing the language analysis method based on the metric semantic layer and the large language model as described in Embodiment 1, including:
[0130] (1) An application end for receiving the user's input and displaying the output results of the system, including:
[0131] (1.1) A user interface module: Provide an intuitive and friendly user interface, allowing users to input query requests using natural language and display a visual representation of the query results (such as charts, reports, or dashboards).
[0132] (1.2) Natural Language Processing Module: Integrates natural language processing technology to convert the user's natural language queries into a format understandable by the system and processes the user's interaction operations (such as clicks, swipes, etc.).
[0133] (1.3) Interaction Module: Displays the data visualization results obtained from the server and provides rich interaction functions, allowing users to further explore and analyze the results.
[0134] (2) A server that is used to receive requests from the application side, process data query and visualization logic, and interact with the database side, including:
[0135] (2.1) Metric Semantic Layer Module: Stores and manages the metadata of data metrics, including metric names, dimensions, and filter information.
[0136] (2.2) Vector Representation Module: Generates vector representations for the same metadata in the metric semantic layer for matching with the natural language requests input by the user.
[0137] (2.3) Query Transformation Module: Converts user queries into query statements understandable by the metric semantic layer for subsequent execution of data query operations.
[0138] (2.4) Large Language Model Module: Receives the user requests that have undergone text vectorization processing and uses the inference ability of the large language model to generate formatted inference results. It can understand the implicit intentions and context information in the user queries, improving the accuracy of the queries.
[0139] (2.5) Data Processing and Transformation Module:
[0140] Text Vectorization: Performs text vectorization processing on the natural language requests input by the user.
[0141] Vector Matching: Matches the processed user request vectors with the metadata vectors in the metric semantic layer to determine the metric Prompt clauses involved in the user queries.
[0142] Query Execution: Accesses the database or data warehouse through the metric semantic layer, executes data query operations, and performs formatted analysis on the query results.
[0143] (3) The database side responsible for storing and managing business data, including:
[0144] (3.1) Relational Database: Such as MySQL, Oracle, etc., used to store structured data.
[0145] (3.2) Non-Relational Database: Such as MongoDB, Redis, etc., used to store unstructured or semi-structured data.
[0146] (3.3) File storage system: used to store non-database data such as attachments, multimedia files or pictures.
[0147] All of the above embodiments only express the implementation manners of the relevant practical applications of the present invention. The descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.
[0148] For those skilled in the art, it can be further realized that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0149] At the same time, those skilled in the art can understand that all or part of the processes of implementing the methods of all the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in this application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
Claims
1. A language analysis method based on an indicator semantic layer and a large language model, characterized in that: After receiving a natural language query request input by the user through the user interface, the following steps are performed: S1, reads the metadata constructed by the predefined data indicator model; uses the metadata to construct the prompt full word of the indicator; generates a vector representation for the same metadata in the indicator semantic layer; S2, after text vectorization processing of the natural language request input by the user, performs distance calculation with the metadata vector in the indicator semantic layer, and determines the indicator Prompt clause involved in the user query and the large language model of the user's original request input according to the threshold; S3, converts the formatted reasoning results into query statements of the indicator semantic layer according to predefined rules, and accesses the database or data warehouse through the indicator semantic layer to perform data query operations; S4, visualizes the formatted and analyzed data results through a natural language interface.
2. The language analysis method according to claim 1, characterized in that: The implementation method of S1 includes: S100, reading metadata from a predefined data indicator model, including the indicator name, dimension definition, and filter configuration; using the read metadata to construct the indicator prompt full word; S101, perform text vectorization on each metadata item in the indicator semantic layer, including indicator name, dimension, and filter, and convert the metadata into a vector representation using a bag-of-words model or TF-IDF technology.
3. The language analysis method according to claim 1, characterized in that: The implementation method of S2 includes: S200, preprocessing the text of the natural language request, including word segmentation and removal of stop words; using the same bag-of-words model or TF-IDF technology as the indicator semantic layer to convert the preprocessed text into a vector representation; S201, performing Euclidean distance calculation between the user request vector and the metadata vector in the indicator semantic layer; filtering out the metadata vector with the smallest distance to the user request vector according to a preset threshold, and then obtaining the indicator Prompt clause involved in the user query.
4. The language analysis method according to claim 3, characterized in that: The implementation method of S2 also includes inputting the indicator Prompt clause and the user's original request into the large language model, performing reasoning and understanding the user's query intention based on the input Prompt and request, and generating a formatted reasoning result.
5. The language analysis method according to claim 1, characterized in that: The implementation method of S3 includes: S300, receiving the formatted reasoning result generated by the large language model; converting the structured information in the reasoning result into a query statement that can be understood by the indicator semantic layer; S301, through the indicator semantic layer, the converted query statement is sent to the database or data warehouse to perform data query operations and retrieve the data required by the user.
6. The language analysis method according to claim 5, characterized in that: The S3 also includes formatting analysis of query results: S3020, converting data in the query results into a format suitable for analysis and visualization; S3021, Cleaning: Calculate the mean, median, mode and / or standard deviation descriptive statistics of the data.
7. The language analysis method according to claim 6, characterized in that: The conversion method in S3020 needs to implement at least one of the following operations: Standardized formats: Standardized formats of date, time and / or currency; Data encoding: converting categorical variables into numerical variables; Data aggregation: Calculate mean, sum, and / or ratio statistics.
8. The language analysis method according to claim 6, characterized in that: The implementation method of S4 includes: binding the data results after formatting and analysis processing with the selected visualization elements, integrating the visualization elements with the natural language interface, and presenting the designed visualization elements in the user interface.
9. A language analysis system based on an indicator semantic layer and a large language model using the language analysis method according to any one of claims 1 to 8, characterized in that: The system comprises: The application end is used to receive user input and display the system output results; The server is used to receive requests from the application side, process data query and visualization logic, and interact with the database side; The database side is responsible for storing and managing business data.
10. The language analysis system according to claim 9, characterized in that: The application end includes a user interface module, a natural language processing module and an interaction module; The server includes an indicator semantic layer module, a vector representation module, a query conversion module, a large language model module and a data processing and conversion module; The database end includes a relational database, a non-relational database and a file storage system.
Citation Information
Cited By
Data analysis method and device based on natural language, equipment and medium
CN121009108A