Intelligent question answering implementation method and device based on analysis model, and medium
By constructing an analysis model metadata vector library and a custom dictionary, and combining LangChain and business intelligence system APIs, the parsing errors and matching failures of traditional business intelligence systems when inputting natural language are solved, achieving accurate matching of indicators, dimensions and constraints, and improving the practicality of intelligent data analysis.
Patent Information
- Application Number
- CN202511164329.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-20
AI Technical Summary
Traditional business intelligence systems suffer from parsing errors, indicator matching failures, and user expression biases when processing natural language input, resulting in poor usability of intelligent query systems.
We build an analysis model metadata vector library, add a custom dictionary, perform structured parsing through a word segmenter, and use LangChain and business intelligence system APIs for accurate matching. We support multiple naming combinations and synonym records, and combine the few-shot method to improve parsing accuracy.
It improves the accuracy of natural language semantic parsing, enhances the practicality of intelligent questioning, resolves the discrepancy between user expression habits and business intelligence system APIs, and achieves precise matching of indicators, dimensions, and constraints.
Smart Images

Figure CN120723875B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of large model technology, specifically to a method, device, and medium for realizing intelligent questioning based on analytical models. Background Technology
[0002] Traditional business intelligence systems, when operated through a user interface, can restrict input content through business logic and data rules, such as allowing only existing dimensions to be selected and blocking non-standardized requests. However, when statistical requests are input in natural language (voice / text), the input information can be highly flexible due to different user understandings of the business and varying habits, leading to parsing errors or even failures to parse the data. Furthermore, traditional business intelligence systems may encounter the following problems in real-world scenarios:
[0003] ① There are strict requirements for the segmentation / word segmentation of indicator and dimension names. They cannot be simply segmented according to everyday understanding, otherwise it will lead to matching failure. For example, "adaptive training participation rate" is the name of an indicator, but it is easy to split into two words: "adaptive training" and "participation rate", which makes it difficult to match with the actual business indicator.
[0004] ② Due to memory or oral reasons, users' descriptions of indicator and dimension names may deviate from the actual situation. For example, "management service unit" is often expressed as "management unit" or "service unit".
[0005] ③ To differentiate between business areas or to follow common expression habits, users may add modifiers for business domains or processes before indicators, such as describing "settlement amount" as "education and training settlement amount" to distinguish it from "settlement amount" in other areas.
[0006] ④ Users describe the time period directly in colloquial language, but the large model only knows that it is a date dimension and cannot accurately map it to the role dimension related to the indicator. For example, "Please calculate the graduation rate for the past three years". The date dimension name associated with the "graduation rate" indicator is not specified. It may not exist or may be associated with multiple date dimensions.
[0007] ⑤ Users are used to directly inputting the physical dimension name, but the large model does not know the role-playing dimension related to the current indicator. For example, "Please count the number of training registrations in the past three years by organization", where "organization" is the physical dimension name, but the actual role-playing dimension is "management service organization".
[0008] ⑥ Users specify dimensions and constraints by using colloquial descriptions of a certain attribute value of a dimension. This is knowledge familiar to industry users, but the large model is unaware of its origin. For example, "Analyze the number of students enrolled in a certain training program under the jurisdiction of a certain education bureau in the past 3 years." Here, "a certain education bureau" indicates both its physical dimension as "organization" and its constraint as "a certain education bureau," meaning drilling up and down at the "a certain education bureau" level.
[0009] ⑦ When the dimension information provided by the user is a physical dimension, this physical dimension may have multiple role-playing dimensions in the fact table where this indicator is located. In this case, it is impossible to confirm the user's intent.
[0010] ⑧ When the user provides dimension information as an attribute, such as statistics by year, month, or quarter, the year, month, and quarter are all date-based attributes, which the large model may not be able to recognize.
[0011] Therefore, improving the accuracy of natural language semantic parsing, enhancing the practicality of intelligent querying, and overcoming the discrepancy between users' natural language expression habits and business intelligence system APIs are urgent technical problems that need to be solved. Summary of the Invention
[0012] The technical objective of this invention is to provide a method, device, and medium for implementing intelligent query analysis based on an analytical model, in order to address the problem of how to improve the accuracy of natural language semantic parsing, enhance the practicality of intelligent query analysis, and overcome the discrepancy between users' natural language expression habits and business intelligence system APIs.
[0013] The technical objective of this invention is achieved as follows: a method for implementing intelligent question counting based on an analysis model, the specific method of which is as follows:
[0014] Building an analysis model metadata vector library: Through an automated workflow, an analysis model metadata vector library is built based on the incremental changes in data in the data warehouse analysis library;
[0015] Add a custom dictionary: Load indicator names, dimension names, attribute names and attribute values from the model library as words and add them to the word segmenter. Set the word frequency of each category of words according to the parameter configuration to guide the large model in the model library to perform structured analysis on the original input content.
[0016] Semantic parsing: After the user input information is segmented by a word segmenter, the word segmentation results that meet the industry characteristics are obtained. The original text is then preliminarily parsed based on the word segmentation results to obtain structured original text parsing results. Based on the structured original text parsing results, indicator names and role-playing dimension names are matched. Then, reverse lookup dimensions are matched by physical dimension names, by dimension attribute names, and by constraint values to obtain detailed constraint information.
[0017] Calling the Business Intelligence System API: Construct a business intelligence system call request based on constraint details including metrics, dimensions, and constraint values, and receive statistical results from the business intelligence system.
[0018] As a preferred option, the analysis model metadata vector library includes a set of indicator metadata, a set of indicator dimension-related metadata, a set of indicator dimension attribute-related metadata, and a set of dimension attribute value dictionary data.
[0019] The indicator metadata set is used to query and match vector numbers by comparing the similarity of indicator names. Multiple records can be added to the same indicator within the indicator metadata set. The naming concatenation methods for the same indicator include indicator name, business domain + indicator name, business domain + business process + indicator name, and business process + indicator name. If synonyms for the indicator name exist, each synonym will also generate one of the four concatenation methods: indicator name + business domain + indicator name, business domain + business process + indicator name, and business process + indicator name. The METRIC_ID of newly added synonym records remains unchanged; only the METRIC_NAME and vector value change. In vector matching, if any of the four concatenation records (indicator name, business domain + indicator name, business domain + business process + indicator name, and business process + indicator name) is matched, the same METRIC_ID will be obtained.
[0020] The indicator dimension associated metadata set is used to query the role-playing dimension number related to the matched indicator by comparing the similarity of the dimension name. It also supports reverse lookup by physical dimension. The indicator dimension associated metadata set records the role-playing dimension and the physical dimension, which are distinguished by DIM_TYPE. Among them, the ORIGINAL_DIM_ID and ORIGINAL_DIM_NAME of the role-playing dimension row cannot be empty, and are used for reverse lookup. If there are synonyms for the dimension name, a record is added for each synonym. The DIM_ID of the newly added synonym record remains unchanged, and only the DIM_NAME and vector value change.
[0021] The indicator dimension attribute associated metadata set is used to query and match the role-playing dimension number related to the indicator by comparing the similarity of the attribute names. The indicator dimension attribute associated metadata set includes all attribute names in the indicator machine associated dimensions. When the user describes the requirements and directly describes any dimension name, the relevant physical dimension is obtained by looking up the corresponding table.
[0022] The dimension attribute value dictionary dataset is used to look up dimension information by using dimension attribute values. The dimension attribute value dictionary dataset includes key qualifiers and their field positions and levels in the dimension table, which facilitates the reverse lookup to obtain clear constraint information, providing it to the business intelligence system, greatly simplifying the search scope and improving search efficiency.
[0023] As a preferred approach, the original text is initially parsed based on the word segmentation results, and the structured original text parsing results are as follows:
[0024] Define the data structure returned by the large model using the pydanic class;
[0025] The structure to be returned is specified by the with_structured_output method of LangChain, and the word segmentation result of the previous step and the current time information are passed into the prompt word template;
[0026] The few-shot method provides multiple parsing examples, and large models can return structured parsing results including indicator names, dimension names, and constraint information. The structured parsing results include a conversational description of the date and are output according to the specified structure.
[0027] Determine whether the semantics of the structured parsing result contain drill-down markers:
[0028] If the semantics include a drill-down flag, then set the drill-down flag to True.
[0029] As a preferred option, the specific matching metric names are as follows:
[0030] Determine whether the structured original text parsing result contains at least one indicator name:
[0031] If no indicator name exists, the program will exit and prompt the user to enter one.
[0032] If an indicator name exists, perform vector matching one by one, select the result with the highest approximation among those exceeding the approximation threshold, and determine whether the match is successful.
[0033] If a match is successful, save the matching indicator number and name;
[0034] If a match fails, a failure message will be returned.
[0035] As a preferred approach, when the user inputs a role-playing dimension name that is directly related to the metric in the original text, role-playing dimension name matching is performed; specifically as follows:
[0036] Determine whether the structured original text parsing results contain dimension names:
[0037] If the structured original text parsing results contain dimension names, then after the index number is limited to "index name matching" and the results are returned, vector matching of dimension names is performed one by one.
[0038] Determine if the vector matching was successful:
[0039] If a match is successful, save the dimension number and name of the successful match;
[0040] If the match fails, continue to try "matching the physical dimension name to find the dimension in reverse".
[0041] The specific steps for reverse dimension lookup using physical dimension names are as follows:
[0042] If the structured original text parsing result contains dimension names, but "role-playing dimension name matching" fails, then physical dimension name matching will be attempted first, and the success of physical dimension name matching will be determined:
[0043] If the physical dimension name fails to match, then proceed to "reverse dimension lookup by matching dimension attribute name";
[0044] If the physical dimension name matches successfully, the role-playing dimension is further looked up using both the indicator and physical dimension constraints. The lookup results are retained as follows:
[0045] If the reverse lookup result does not exist, continue to use "match the reverse lookup dimension by dimension attribute name";
[0046] If a result exists, it means that the corresponding role-playing dimension name has been successfully matched, and the constraint details will continue to be obtained;
[0047] If there is more than one result, the user needs to specify the optional options through LangGraph HIL (Human in the loop) and return to the save point in the interrupted process to continue processing.
[0048] As a preferred method, the dimension lookup is performed by matching the dimension attribute name, as follows:
[0049] Check if the dimension attribute name matches successfully:
[0050] If the dimension attribute name fails to match, continue with "reverse dimension lookup by constraint value matching";
[0051] If the dimension attribute name matches successfully, due to the role-playing dimension, there will be multiple results with the same similarity, as follows:
[0052] If a result with the same similarity exists, it means that the corresponding dimension attribute name matches and the role-playing dimension is successfully retrieved, and the subsequent "Get Constraint Details" operation can be performed.
[0053] If there are more than one result with the same similarity, the user can specify the optional parameters through LangGraph HIL (Human in the loop), and the process will return to the save point in the interrupted flow to continue processing.
[0054] As a preferred method, the dimension lookup is performed by matching constraint values, as follows:
[0055] Determine if the constraint values match successfully:
[0056] If the constraint value fails to match, it means that all methods have been exhausted, indicating that the dimension information provided by the customer does not exist in the analysis library, and an error message will be given.
[0057] If the constraint value matches successfully, the role-playing dimension is further queried using both the indicator and physical dimensions as constraints. The results of the query are as follows:
[0058] If the reverse lookup result is not found, then attribute value matching will continue.
[0059] If a reverse lookup result exists, it means that the corresponding role-playing dimension name has been successfully matched, and subsequent operations can continue.
[0060] If there is more than one reverse lookup result, the user can specify the optional options through LangGraph HIL (Human in the loop), and the process will return to the save point in the interrupted flow to continue processing.
[0061] More preferably, the detailed constraint information is obtained as follows:
[0062] The constraint values specified by the user in natural language may have slight differences from the actual constraint content. After a successful match, the dimension attribute values will be used to replace the information entered by the user in subsequent API call requests.
[0063] Meanwhile, by analyzing the model metadata vector library, the precise constraint values in the dimension table fields (DIM_PROPERTY_NAME and DIM_PROPERTY_LEVEL) are obtained and sent to the business intelligence system API. This makes it easier for the business intelligence system to specify the constraints when constructing SQL, reduce the scope of data retrieval, and shorten the response time.
[0064] Metadata analysis, based on a consistent dimensional modeling approach, is the core architecture for data warehouse statistical analysis libraries, enabling cross-domain data integration and standardized definitions. Essentially, it ensures cross-analysis of data from different business processes through unified dimension definitions and association mechanisms. The metadata includes all indicator definitions, dimension definitions (containing detailed attribute information), and the relationships between dimensions and indicators.
[0065] Role-playing dimensions refer to the same physical dimension table being linked to multiple foreign keys in the fact table, each representing a different business role. For example, in a fact table, there might be multiple attributes linked to the date dimension, representing the business's acceptance date and completion date, respectively.
[0066] LangGraph is an open-source library developed by the LangChain team, designed specifically for building stateful, multi-step, graph-driven workflows for large language model applications. It manages task flows through a graph structure, supporting complex logic such as loops, branches, and multi-agent collaboration, making it a core tool in the LangChain ecosystem for handling dynamic scenarios. With the rapid development of large language models, obtaining analysis results through free-form questioning in natural language has become possible, with two implementation paths:
[0067] ①Chat2SQL Mode. This mode involves a large model automatically constructing SQL statements and retrieving query results based on input requirements and the object structure within the data warehouse. However, due to the vast and complex structure of the data warehouse, and the difficulty for users to clearly and accurately describe their requirements in natural language, the current intelligence level of large models is not yet capable of consistently outputting correct SQL statements. Current Chat2SQL tools, such as the Power Business Intelligence System Copilot, tend to assist developers rather than providing analytical capabilities to end users.
[0068] ②The Chat2Metrics mode primarily utilizes large models for natural language parsing and supplements the data warehouse with external knowledge bases, significantly reducing the difficulty of generating results from large models. It combines tools functions or MCP methods to call open API interfaces of business intelligence systems to obtain analysis results. However, due to the significant uncertainty in users' use of natural language expressions, it remains difficult to accurately and stably parse the results.
[0069] An electronic device includes: a memory and at least one processor;
[0070] The memory contains computer programs;
[0071] The at least one processor executes the computer program stored in the memory, causing the at least one processor to perform the intelligent question-solving method for the analysis model as described above.
[0072] A computer-readable storage medium storing a computer program that can be executed by a processor to implement the intelligent question-solving method for the analysis model as described above.
[0073] The intelligent question-answering method, device, and medium based on the analysis model of the present invention have the following advantages:
[0074] (i) This invention makes full use of the statistical analysis object definitions and inter-object relationship definitions contained in the metadata of the analysis model to maximize the conversion of user needs in natural language into API calls of business intelligence system. This greatly improves the problem of poor practicality caused by the lack of industry knowledge and illusions in large models. It makes full use of the statistical analysis object definitions and inter-object relationship definitions contained in the model metadata to greatly improve the ability to understand user input and enhance the practicality of intelligent data query products.
[0075] (ii) The present invention constructs a metadata vector library based on the analysis model, including indicators, the association between indicators-role-playing dimensions-original dimensions, the association between indicators-role-playing dimensions-dimensional attributes, key dimension attribute values, and synonyms, which facilitates the rapid matching of user request content with analysis model objects, and can also realize the association matching of indicators, indicator-related dimensions, indicator-related attributes, role-playing dimensions and physical dimensions.
[0076] (II) This invention utilizes the object names in the analysis model to construct a dedicated word segmentation dictionary, unify the terminology, resolve ambiguity, and guide the large model to perform structured parsing of the original input content. This solves the open uncertainty when describing requirements in natural language, improves the matching degree of indicator names, and adopts a method of improving compatibility by using multiple naming combinations for indicator names. It supports adding various prefixes to the original indicator names, and also supports defining synonyms for indicators and dimensions. In the process of constructing the metadata vector library set of the analysis model, it adds records of synonyms to facilitate association matching in subsequent algorithms.
[0077] (III) This invention improves the accuracy of word segmentation by extracting the names of indicators, dimensions, attributes, limited names, and synonyms from metadata as lexical units to construct a metadata word segmentation lexicon. It also allows for word segmentation according to industry / professional terms.
[0078] (iv) This invention does not require fine-tuning training of large models or importing professional knowledge. It can generate high-quality preliminary parsing results of the original text by passing in the word segmentation results through prompt words and combining them with few shot prompt words. This method can parse according to the industry / professional terminology requirements and output accurate structured data at the same time.
[0079] (v) This invention improves the success rate of indicator matching by using an indicator metadata set to match indicators with similarity.
[0080] (vi) This invention associates metadata sets with indicator dimensions, thereby improving the success rate of dimension matching by matching role-playing dimensions with similarity.
[0081] (vii) This invention associates metadata sets with indicator dimensions, enabling the reverse lookup of role-playing dimensions through physical dimensions, thereby improving the success rate of dimension matching and solving the problem that customers use physical dimension names to describe their dimension requirements due to their natural language expression habits.
[0082] (viii) This invention associates metadata sets with indicator dimension attributes, and realizes the method of first looking up the physical dimension through the physical dimension attribute, and then looking up the role-playing dimension through the physical dimension, thereby improving the dimension matching success rate. It solves the problem that customers use physical dimension attribute names to describe dimension requirements due to their natural language expression habits.
[0083] (ix) This invention uses a dictionary of dimension attribute values to enable reverse lookup of physical dimensions by constraint values, and then reverse lookup of role-playing dimensions, thereby improving the success rate of dimension matching and solving the problem that customers use physical dimension attribute names to describe dimension requirements due to their natural language expression habits.
[0084] (x) After obtaining the physical dimension name, the present invention obtains multiple role-playing dimensions associated with the indicator through reverse lookup, and then allows the user to specify the role-playing dimension through human-computer interaction, thus avoiding the misunderstanding of the user's actual needs caused by the system automatically selecting the default value.
[0085] (xi) This invention obtains complete constraint value information through a dictionary of dimension attribute values and provides it to the API interface of the business intelligence system, thereby reducing the difficulty of constructing SQL for the business intelligence system API and improving the response speed.
[0086] (xii) This invention improves parsing accuracy through a complete input parsing process, including upstream and downstream related operations, and solves the problem of discrepancies between users' natural language expression habits and business intelligence system APIs, thereby maximizing the understanding and satisfaction of user needs and enhancing the practicality of intelligent data query products.
[0087] (xiii) This invention uses LangGraph to construct a matching algorithm, and implements state management, loop logic and interactive collaboration features through graph structure. Based on the input content and upstream processing results, the following nodes are dynamically started:
[0088] ① Based on the fuzzy matching index of the vector library, it supports multi-index parsing and matching;
[0089] ②Prioritize matching of role-playing dimensions and support multiple indicator collisions to obtain common dimensions;
[0090] ③ If the role-playing dimension fails to match, the role-playing dimension will be determined by reverse lookup using the original dimension, dimension attribute, and dimension attribute value in sequence.
[0091] ④ If multiple possible results are obtained through reverse lookup, clarification can be achieved through multiple human-computer interactions;
[0092] ⑤ Supports parsing constraints, supports exact matching and range matching modes, supports reverse lookup of the attributes of the constraint value in the original dimension, and supports parsing drill-down tags from semantics. Attached Figure Description
[0093] The invention will be further described below with reference to the accompanying drawings.
[0094] Appendix Figure 1 This is a flowchart illustrating the overall process of implementing intelligent question-and-answer methods based on analytical models.
[0095] Appendix Figure 2 This is a detailed flowchart of the intelligent question-answering method based on the analysis model. Detailed Implementation
[0096] The intelligent question-answering method, device, and medium based on the analysis model of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0097] Example 1: As shown in the attached document Figure 1 As shown in the figure, this embodiment provides a method for implementing intelligent question counting based on an analysis model. The method is as follows:
[0098] S1. Construct an analysis model metadata vector library: Construct an analysis model metadata vector library based on the incremental changes in data in the data warehouse analysis library through an automated workflow;
[0099] S2. Add a custom dictionary: Load indicator names, dimension names, attribute names, and attribute values from the model library as words and add them to the word segmenter. Set the word frequency for each category according to the parameter configuration to guide the large model in the model library to perform structured analysis on the original input content. Taking the Chinese word segmentation of "stuttering" as an example, the word frequency settings for each object category are as follows:
[0100] Indicator Name: 20000 (High Frequency);
[0101] Dimension name: 20000 (High Frequency);
[0102] Attribute name: 10000 (Medium frequency);
[0103] Attribute value: 80000 (medium frequency);
[0104] This embodiment significantly improves word segmentation accuracy by adding a custom dictionary.
[0105] S3. Semantic parsing: After the user input information is segmented by the word segmenter, the word segmentation results that meet the industry characteristics are obtained. The original text is initially parsed based on the word segmentation results to obtain the structured original text parsing results. The indicator name and role-playing dimension name are matched based on the structured original text parsing results. Then, the reverse lookup dimension is matched by the physical dimension name, the reverse lookup dimension is matched by the dimension attribute name, and the reverse lookup dimension is matched by the constraint value to obtain the constraint details.
[0106] S4. Call the Business Intelligence System API: Based on the constraint details including metrics, dimensions, and constraint values, construct a business intelligence system call request and receive the statistical results from the business intelligence system.
[0107] The analysis model metadata vector library in this embodiment includes a set of indicator metadata, a set of indicator dimension-related metadata, a set of indicator dimension attribute-related metadata, and a set of dimension attribute value dictionary data.
[0108] In this embodiment, the indicator metadata set is used to query and match vector numbers by comparing the similarity of indicator names; an example of the indicator metadata set schema is shown in Table 1.
[0109] Table 1. Schema of Indicator Metadata Collection
[0110]
[0111] In this embodiment, the metadata set of the metrics allows for the addition of multiple records for the same metric. The naming concatenation methods for the same metric include metric name, business domain + metric name, business domain + business process + metric name, and business process + metric name. If there are synonyms for the metric name, each synonym will also generate four concatenation methods: metric name + business domain + metric name, business domain + business process + metric name, and business process + metric name. The METRIC_ID of the newly added synonym record remains unchanged; only the METRIC_NAME and vector value change. In vector matching, if any of the four concatenation records of metric name, business domain + metric name, business domain + business process + metric name, or business process + metric name are matched, the same METRIC_ID will be obtained.
[0112] In this embodiment, the indicator dimension associated metadata set is used to query the role-playing dimension number related to the matched indicator by comparing the similarity of the dimension name, and also supports reverse lookup by physical dimension; an example of the indicator dimension associated metadata set schema is shown in Table 2.
[0113] Table 2. Schema of Metadata Sets Associated with Indicator Dimensions
[0114]
[0115] In this embodiment, the metadata set associated with the indicator dimensions records the role-playing dimension and the physical dimension. The role-playing dimension and the physical dimension are distinguished by DIM_TYPE (role-playing dimension 1, physical dimension 9). Among them, the ORIGINAL_DIM_ID and ORIGINAL_DIM_NAME of the role-playing dimension row cannot be empty, and are used for reverse lookup. If there are synonyms for dimension names, a record is added for each synonym. The DIM_ID of the newly added synonym record remains unchanged, and only the DIM_NAME and vector value change.
[0116] In this embodiment, the indicator dimension attribute associated metadata set is used to query and match the role-playing dimension number related to the indicator by comparing the similarity of the attribute names; an example of the indicator dimension attribute associated metadata set schema is shown in Table 3.
[0117] Table 3. Schema of Metadata Sets Associated with Indicator Dimension Attributes
[0118]
[0119] In this embodiment, the metadata set associated with the indicator dimension attributes includes all attribute names in the indicator machine-associated dimension (such as year, month, and date dimension attributes). When the user describes their requirements, they can directly describe any dimension name and obtain the relevant physical dimension by looking up the corresponding table.
[0120] In this embodiment, the dimension attribute value dictionary data set is used to look up dimension information by dimension attribute value; an example of the dimension attribute value dictionary data set schema is shown in Table 4.
[0121] Table 4 Dimension Attribute Value Dictionary Data Set Schema
[0122]
[0123] The dimensional attribute value dictionary dataset includes key qualifiers and their corresponding field positions and levels in the dimension table, facilitating reverse lookups to obtain clear constraint information. This provides business intelligence systems with this information, greatly simplifying the search scope and improving search efficiency.
[0124] As attached Figure 2 As shown, the semantic parsing in step S3 of this embodiment is as follows:
[0125] S301, Word segmentation: Since the indicator names, dimension names, attribute names, and attribute values in the model library have already been segmented during the initialization phase, the user input information will be segmented by the word segmenter to obtain a word segmentation result that meets the industry characteristics.
[0126] S302, Preliminary Analysis of the Original Text: The data structure returned by the large model is defined through the pydanic class. The `with_structured_output` method of LangChain is used to specify the returned structure, and the word segmentation results of the previous step, the current time, and other information are passed into the prompt word template. The `few-shot` method provides multiple parsing examples, and the large model can then return a structured parsing result including indicator names, dimension names, and constraint information, including colloquial descriptions such as dates, and output according to the specified structure; if the semantics contain drill-down markers, the drill-down markers are set to True;
[0127] This step only involves understanding the semantics of natural language and simple structured output. Large models do not need to master complex professional knowledge and can achieve good results without fine-tuning training.
[0128] S303, Indicator Name Matching: Determine whether the original parsing result contains an indicator name. Theoretically, it should contain at least one. If it does not exist, exit and prompt the user to enter one. If an indicator name exists, perform vector matching one by one. Among the results that exceed the approximation threshold, take the result with the highest approximation. If the match is successful, save the indicator number and name of the successful match. If the match fails, return a failure message.
[0129] In this embodiment, all matching rules are compared with the approximation threshold. If there is a result greater than the threshold, the match is successful and the highest approximation result is taken; otherwise, it fails. This will not be elaborated further below.
[0130] The similarity thresholds for matching indicator names, role-playing dimension names, physical dimension names, and constraint values are all set to 0.9.
[0131] The similarity threshold for attribute name matching is set to 0.82;
[0132] S304, Role-Playing Dimension Name Matching: Role-playing dimension name matching is applicable when the user inputs role-playing dimension names directly related to the metric in the original text. If the original text parsing result contains dimension names, then the metric number is limited to "metric name matching". After returning the results, vector matching of dimension names is performed one by one. If a match is successful, the dimension number and name of the successful match are saved; if a match fails, the "reverse dimension lookup through physical dimension name matching" method is used to try again.
[0133] S305. Reverse dimension lookup by physical dimension name matching: Physical dimension matching is suitable when the user is unaware of the role-playing dimension name and directly provides the physical dimension name in the input. If the original parsed result contains the dimension name, but "role-playing dimension name matching" fails, physical dimension name matching will be attempted first. If matching fails, "reverse dimension lookup by attribute name matching" will be performed next.
[0134] If the physical dimension name matches successfully, the role-playing dimension is then retrieved using both the indicator and physical dimension constraints. The retrieved result may not exist, or it may contain one or more results. The following are the subsequent processing methods:
[0135] If it does not exist, continue to use "reverse dimension lookup by attribute name matching".
[0136] If there is one result, it means that the role-playing dimension name is successfully matched, and you can continue to "Get Constraint Details".
[0137] If there is more than one result, the user needs to specify the optional options through LangGraph HIL (Human in the loop) and return to the save point in the interrupted process to continue processing.
[0138] S306. Dimension lookup by attribute name matching: Attribute name matching is suitable for situations where the user does not specify any dimension name, but directly provides the dimension attribute name for analysis.
[0139] If the dimension attribute name fails to match, continue with "reverse dimension lookup by constraint value matching"; if a match is successful, multiple results with the same similarity may exist due to role-playing dimensions. The following are the subsequent handling methods:
[0140] If there is one result, it means that the attribute name matches and the role-playing dimension is successfully retrieved, and the subsequent "Get Constraint Details" operation can be performed.
[0141] If there are multiple results, the user needs to specify the optional options through LangGraph HIL (Human in the loop) and return to the save point in the interrupted process to continue processing.
[0142] S307. Reverse dimension lookup via constraint value matching: Constraint value matching is suitable for situations where users can indirectly select dimensions by directly providing constraint values.
[0143] If the constraint value matching fails, all methods have been exhausted, indicating that the dimension information provided by the client does not exist in the analysis library, and an error message will be given. If the matching is successful, it is still necessary to further use the dual constraints of indicators and physical dimensions to reverse look up the role-playing dimension. The reverse lookup result may not exist, or it may exist one or more. The following are the subsequent processing methods:
[0144] If it does not exist, continue using attribute value matching;
[0145] If there is one result, it means that the role-playing dimension name is successfully matched, and you can continue to the next step.
[0146] If there is more than one result, the user needs to specify the optional options through LangGraph HIL (Human in the loop) and return to the save point in the interrupted process to continue processing;
[0147] S308. Obtain detailed constraint information: The constraint values specified by the user in natural language may have slight differences from the actual constraint content. After a successful match, the dimension attribute values are used to replace the information entered by the user in subsequent interface call requests. In addition, the precise constraint values can be obtained from the vector library and sent to the business intelligence system API. This makes it easier for the business intelligence system to specify the constraint conditions when constructing SQL, reduce the scope of data retrieval, and shorten the response time.
[0148] Example 2: This example also provides an electronic device, including: a memory and a processor;
[0149] The memory stores the instructions executed by the computer.
[0150] The processor executes computer execution instructions stored in the memory, causing the processor to execute the intelligent question-and-answer implementation method based on the analysis model in any embodiment of the present invention.
[0151] The processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can be a microprocessor or any conventional processor.
[0152] Memory is used to store computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and by accessing data stored in the memory. Memory can mainly include a program storage area and a data storage area. The program storage area can store the operating system, at least one application program required for a function, etc.; the data storage area can store data created based on the use of the terminal, etc. In addition, memory can also include high-speed random access memory, and can also include non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart memory cards (SMC), secure digital cards (SD cards), flash memory cards, at least one disk storage device, flash memory devices, or other volatile solid-state storage devices.
[0153] Example 3: This example also provides a computer-readable storage medium storing multiple instructions, which are loaded by a processor to cause the processor to execute the intelligent query implementation method based on the analysis model in any embodiment of the present invention. Specifically, a system or apparatus equipped with a storage medium may be provided, on which software program code implementing the functions of any of the above embodiments is stored, and the computer (or CPU or MPU) of the system or apparatus may read and execute the program code stored in the storage medium.
[0154] In this case, the program code read from the storage medium can itself implement the function of any of the above embodiments, and therefore the program code and the storage medium storing the program code constitute part of the present invention.
[0155] Storage media embodiments for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, program code can be downloaded from a server computer via a communication network.
[0156] Furthermore, it should be clear that not only can the program code read by the computer be executed, but also the operating system or other components operating on the computer can be instructed based on the program code to perform some or all of the actual operations, thereby realizing the function of any of the embodiments described above.
[0157] Furthermore, it is understood that the program code read from the storage medium is written to the memory set in the expansion board inserted into the computer or to the memory set in the expansion unit connected to the computer. Then, based on the instructions of the program code, the CPU or other components installed on the expansion board or expansion unit execute some and all of the actual operations, thereby realizing the function of any of the embodiments described above.
[0158] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for implementing intelligent question counting based on an analytical model, characterized in that, The method is as follows: Building an analysis model metadata vector library: Through an automated workflow, an analysis model metadata vector library is built based on the incremental changes in data in the data warehouse analysis library; Add a custom dictionary: Load indicator names, dimension names, attribute names and attribute values from the model library as words and add them to the word segmenter. Set the word frequency of each category of words according to the parameter configuration to guide the large model in the model library to perform structured analysis on the original input content. Semantic parsing: After the user input information is segmented by a word segmenter, the word segmentation results that meet the industry characteristics are obtained. The original text is then preliminarily parsed based on the word segmentation results to obtain structured original text parsing results. Based on the structured original text parsing results, indicator names and role-playing dimension names are matched. Then, reverse lookup dimensions are matched by physical dimension names, by dimension attribute names, and by constraint values to obtain detailed constraint information. Calling the Business Intelligence System API: Based on the constraint details including metrics, dimensions, and constraint values, construct a business intelligence system call request and receive the statistical results from the business intelligence system; The specific matching metric names are as follows: Determine whether the structured original text parsing result contains at least one indicator name: If no indicator name exists, the program will exit and prompt the user to enter one. If an indicator name exists, perform vector matching one by one, select the result with the highest approximation among those exceeding the approximation threshold, and determine whether the match is successful. If a match is successful, save the matching indicator number and name; If the match fails, a failure message will be returned; When a user inputs a role-playing dimension name that is directly related to the metric from the original text, role-playing dimension name matching is performed; specifically as follows: Determine whether the structured original text parsing results contain dimension names: If the structured original text parsing results contain dimension names, then after the index number is limited to "index name matching" and the results are returned, vector matching of dimension names is performed one by one. Determine if the vector matching was successful: If a match is successful, save the dimension number and name of the successful match; If the match fails, continue to try "matching the reverse dimension by physical dimension name"; The specific steps for reverse dimension lookup using physical dimension names are as follows: If the structured original text parsing result contains dimension names, but "role-playing dimension name matching" fails, then physical dimension name matching will be attempted first, and the success of the physical dimension name matching will be determined: If the physical dimension name fails to match, then proceed to "reverse dimension lookup by matching dimension attribute name"; If the physical dimension name matches successfully, the role-playing dimension is further looked up using both the indicator and physical dimension constraints. The lookup results are retained as follows: If the reverse lookup result does not exist, continue to use "match the reverse lookup dimension by dimension attribute name"; If a result exists, it means that the corresponding role-playing dimension name has been successfully matched, and the constraint details will continue to be obtained; If there is more than one result, the user needs to specify the optional options through LangGraph HIL and return to the save point in the interrupted process to continue processing; The specific steps for reverse dimension lookup using dimension attribute names are as follows: Check if the dimension attribute name matches successfully: If the dimension attribute name fails to match, continue with "reverse dimension lookup by constraint value matching"; If the dimension attribute name matches successfully, due to the role-playing dimension, there will be multiple results with the same similarity, as follows: If a result with the same similarity exists, it means that the corresponding dimension attribute name matches and the role-playing dimension is successfully retrieved, and the subsequent "Get Constraint Details" operation can be performed. If there are more than one result with the same similarity, the user can specify the optional options through LangGraph HIL, and the process will return to the save point in the interrupted process to continue processing; The specific steps for reverse dimension lookup using constraint value matching are as follows: Determine if the constraint values match successfully: If the constraint value fails to match, it means that all methods have been exhausted, indicating that the dimension information provided by the customer does not exist in the analysis library, and an error message will be given. If the constraint value matches successfully, the role-playing dimension is further queried using both the indicator and physical dimensions as constraints. The results of the query are as follows: If the reverse lookup result is not found, then attribute value matching will continue. If a reverse lookup result exists, it means that the corresponding role-playing dimension name has been successfully matched, and the subsequent operations can continue. If there is more than one reverse lookup result, the user can specify the optional options through LangGraph HIL, and the process will return to the save point in the interrupted process to continue processing; The details for obtaining constraint information are as follows: The constraint values specified by the user in natural language may have slight differences from the actual constraint content. After a successful match, the dimension attribute values will be used to replace the information entered by the user in subsequent API call requests. At the same time, by analyzing the model metadata vector library, the precise constraint values in the dimension table fields are obtained and sent to the business intelligence system API, making it easier for the business intelligence system to specify the constraint conditions when constructing SQL.
2. The intelligent question-answering method based on the analysis model according to claim 1, characterized in that, The analysis model metadata vector library includes a set of indicator metadata, a set of indicator dimension-related metadata, a set of indicator dimension attribute-related metadata, and a dictionary of dimension attribute values. The indicator metadata set is used to query and match vector numbers by comparing the similarity of indicator names. Multiple records can be added to the same indicator within the indicator metadata set. The naming concatenation methods for the same indicator include indicator name, business domain + indicator name, business domain + business process + indicator name, and business process + indicator name. If synonyms for the indicator name exist, each synonym will also generate one of the four concatenation methods: indicator name + business domain + indicator name, business domain + business process + indicator name, and business process + indicator name. The METRIC_ID of newly added synonym records remains unchanged; only the METRIC_NAME and vector value change. In vector matching, if any of the four concatenation records (indicator name, business domain + indicator name, business domain + business process + indicator name, and business process + indicator name) is matched, the same METRIC_ID will be obtained. The indicator dimension associated metadata set is used to query the role-playing dimension number related to the matched indicator by comparing the similarity of the dimension name. It also supports reverse lookup by physical dimension. The indicator dimension associated metadata set records the role-playing dimension and the physical dimension, which are distinguished by DIM_TYPE. Among them, the ORIGINAL_DIM_ID and ORIGINAL_DIM_NAME of the role-playing dimension row cannot be empty, and are used for reverse lookup. If there are synonyms for the dimension name, a record is added for each synonym. The DIM_ID of the newly added synonym record remains unchanged, and only the DIM_NAME and vector value change. The indicator dimension attribute associated metadata set is used to query and match the role-playing dimension number related to the indicator by comparing the similarity of the attribute names. The indicator dimension attribute associated metadata set includes all attribute names in the indicator machine associated dimensions. When the user describes the requirements and directly describes any dimension name, the relevant physical dimension is obtained by looking up the corresponding table. The dimension attribute value dictionary data set is used to look up dimension information by dimension attribute values. The dimension attribute value dictionary data set includes key qualifiers and their field positions and levels in the dimension table, which facilitates the reverse lookup to obtain clear constraint information and provides it to the business intelligence system.
3. The intelligent question-answering method based on the analysis model according to claim 1, characterized in that, Based on the word segmentation results, the original text is initially parsed, and the structured original text parsing results are as follows: Define the data structure returned by the large model using the pydanic class; The structure to be returned is specified by the with_structured_output method of LangChain, and the word segmentation result of the previous step and the current time information are passed into the prompt word template; The few-shot method provides multiple parsing examples, and large models can return structured parsing results including indicator names, dimension names, and constraint information. The structured parsing results include a conversational description of the date and are output according to the specified structure. Determine whether the semantics of the structured parsing result contain drill-down markers: If the semantics include a drill-down flag, then set the drill-down flag to True.
4. An electronic device, characterized in that, include: Memory and at least one processor; The memory contains computer programs; The at least one processor executes the computer program stored in the memory, causing the at least one processor to perform the intelligent question-answering implementation method based on the analysis model as described in any one of claims 1 to 3.
5. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that can be executed by a processor to implement the intelligent question-answering method based on the analysis model as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Interactive number asking agent system based on large language model
CN120216656A
Method and system for realizing role playing dimension
CN120407542A
Chinese natural language query system and method
US20020052871A1