Method and device for realizing natural language number asking
By constructing the topic domain knowledge graph and using the HanLP double-array dictionary tree for text matching and intent recognition, the problems of low data processing efficiency and insufficient answer accuracy in natural language queries are solved, and efficient and accurate data query and visual display are achieved.
Patent Information
- Application Number
- CN202510267644.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art has problems such as low data processing efficiency, large models relying on, insufficient answer accuracy and instability when processing natural language queries.
By sorting out the index table, semantic modeling, building a topic domain knowledge graph, and using HanLP double-array dictionary tree for text matching and intent recognition, generating query SQL for data query, and finally performing data visualization display.
The data required can be obtained by natural language querying, which reduces the threshold for data query, improves work efficiency, reduces dependence on IT departments or data analysis teams, and ensures the accuracy and consistency of data retrieval.
Smart Images

Figure CN120123366A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and specifically provides a method and device for realizing natural language question answering about numbers. Background Art
[0002] With the advent of the big data era and the rapid progress of artificial intelligence technology, a large amount of data resources have been accumulated within enterprises. However, how to effectively extract, analyze these data and transform them into valuable business insights has become a challenge.
[0003] Although traditional business intelligence (BI) tools can provide data visualization and basic data analysis functions, they have limitations in dealing with complex and variable natural language query requirements raised by non-professionals. Users often need to have certain data analysis skills to make full use of them.
[0004] Relying on the NL2SQL ability of large models to realize intelligent question answering about numbers can indeed use the statistical data in large-scale corpora to train the model and improve the parsing accuracy. However, it often depends on a large number of training data sets, and has limited generalization ability for unseen grammars and complex contexts. Moreover, the deployment of large models consumes a large amount of hardware resources.
[0005] How to solve the problems of low data processing efficiency, dependence on the capabilities of large models, and insufficient and unstable answer accuracy in the prior art is an urgent matter for those skilled in the art. Summary of the Invention
[0006] The present invention aims at the above-mentioned deficiencies of the prior art and provides a method for realizing natural language question answering about numbers with strong practicability.
[0007] A further technical task of the present invention is to provide a device for realizing natural language question answering about numbers with reasonable design, safety and applicability.
[0008] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0009] A method for realizing natural language question answering about numbers has the following steps:
[0010] S1. Sort out the index table, perform semantic modeling, and construct a topic domain knowledge graph;
[0011] S2. Regularly write into the HanLP double-array trie;
[0012] S3. Text input recommendation;
[0013] S4. Extract key elements of the user's question;
[0014] S5. Perform date and intention recognition;
[0015] S6. Intent score calculation;
[0016] S7. Generate a query SQL for data query;
[0017] S8. Data visualization display.
[0018] Furthermore, in step S1, integrate data from multiple data sources and perform data cleaning, transformation, standardization, and normalization through a data governance center;
[0019] Perform dimensional modeling on the indicator table under the subject domain, construct a summary model, set dimensions, indicators, dates, and primary key fields, construct a data set, configure the association relationships between models, and adopt a set of constructed triples in the form of <entity, relationship, entity>.
[0020] Furthermore, in step S2, start a scheduled task to write the indicators, dimensions, and dimension values defined in the semantic model into the HanLP double-array trie at regular intervals to trade space for time.
[0021] Furthermore, in step S3, listen for the event of the user inputting a question text. First, tokenize the question, and then match phrases from the double-array trie. The matched indicators, dimensions, and dimension values are recommended for the user input.
[0022] Furthermore, in step S4, first, tokenize the user text through HanLP. If there are indicators, dimensions, and dimension values in the user question, they will be directly split out;
[0023] The remaining unmatched text is first handed over to the vector library for matching. Calculate the semantic similarity according to the vector model. If it exceeds the threshold, it is considered a successful match with this element. If the vector similarity does not exceed the threshold, continue to perform matching in the HanLP double-array trie, perform prefix and suffix searches, and retain this element if the similarity score threshold is exceeded;
[0024] Finally, the user question will be parsed into a key element group of indicators, dimensions, and dimension values.
[0025] Furthermore, in step S5, when performing date recognition, build-in time parsing regular expressions to parse the time range and time granularity. Identify the time period text in the question through regular expressions, and hand it over to the specific time granularity recognition rules to identify the overall time range one by one, and finally form a time screening range;
[0026] When performing intent recognition, build-in intent recognition rules, including rules for calculating the total amount of indicators, aggregating indicators by dimension grouping, filtering by dimension value, querying the indicator trend by date granularity, TOPN query, and detailed query intent recognition rules, and perform intent matching according to the key elements extracted in step S5 and step S6.
[0027] Further, in step S6, HanLP is used to perform part-of-speech tagging on the words in the user's question. The length of the question after removing the useless words is used as the denominator, and the lengths of the words of the metrics, dimensions, and aggregation functions matched by the intent are used as the numerator for score calculation. If the score is lower than the threshold, this intent match will be considered unable to accurately answer the user's question and will be excluded. The top 3 intent recognition results are retained according to the scores.
[0028] Further, in step S7, after the recognition and parsing in steps S4 and S5, the user's question has been comprehensively disassembled. The jsqlparser parser is used to generate the corresponding query SQL for different types of databases.
[0029] The generated SQL is handed over to the columnar database clickHouse for online analytical processing (OLAP) to implement the query and analysis of large datasets.
[0030] Further, in step S8, after the database finishes executing the query SQL, the data results are returned. The data results are automatically adapted and converted into easy-to-understand visual charts through the number of data result fields and field types, and are fed back to the user through dynamic visualization components.
[0031] When processing sensitive data, encryption technology is used to desensitize the data and implement access control policies.
[0032] An apparatus for realizing natural language querying numbers includes: at least one memory and at least one processor;
[0033] The at least one memory is used to store machine-readable programs;
[0034] The at least one processor is used to call the machine-readable program to execute a method for realizing natural language querying numbers.
[0035] Compared with the prior art, a method and an apparatus for realizing natural language querying numbers according to the present invention have the following outstanding beneficial effects:
[0036] According to the present invention, by asking questions in natural language, the required data can be obtained, which greatly reduces the data query threshold and improves the work efficiency of employees at all levels; it responds in real time to complex and changing data requirements, helps enterprises make data-based decisions quickly, and reduces the time cost of waiting for the assistance of the IT department or the data analysis team.
[0037] The effective integration and intelligent search of massive data resources enable various types of data within the enterprise to be fully mined and utilized, avoiding the idle and waste of data resources; through a series of rule parsing and recognition capabilities and precise query technologies, the accuracy of data retrieval can be improved, and the consistency of data interpretation across departments and business scenarios can be ensured.
[0038] It does not need to rely on large models and has low requirements for deployment resources. It can be deployed on ordinary virtual machines, is ready to use out of the box, and is easy to get started. Newly recruited employees or those who are not familiar with data systems can also quickly get started with query work, reducing the training time and cost specifically for data query systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0040] The appendix Figure 1 is a flowchart showing a method for implementing natural language question answering about numbers. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] In order to enable those skilled in the art of the present technology to better understand the solutions of the present invention, the following will further elaborate on the present invention in combination with specific embodiments. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0042] The following gives a best embodiment:
[0043] As Figure 1 shown, a method for implementing natural language question answering about numbers in this embodiment has the following steps:
[0044] S1. Sort out the index table;
[0045] Introduce external data sources, integrate data from multiple data sources, including databases, data warehouses, WEB services, Excel, etc., and perform data cleaning, transformation, standardization, and normalization processing through the data governance center to ensure data quality. The data quality directly determines the accuracy of question answering.
[0046] S2. Semantic modeling, construct a topic domain knowledge graph;
[0047] Perform dimensional modeling on the indicator table under the subject domain, construct a summary model, set dimensions, indicators, dates, and primary key fields, construct a data set, configure the association relationships between models, and use a graph database and a vector database for storage to enhance the semantic understanding ability of the system. The semantic modeling function supports synonym configuration to facilitate more accurate matching of elements such as indicators, dimensions, and dimension values. The formula uses a set of constructed triples: <entity, relationship, entity>, for example: (Data Catalog A, belongs to department, Department B).
[0048] S3. Write to the HanLP double-array trie regularly;
[0049] Start a scheduled task to regularly write the indicators, dimensions, and dimension values defined in the semantic model to the HanLP double-array trie. Trade space for time, and use the common prefix of strings to reduce the time overhead of word matching, and extract the key elements in the question faster and more accurately.
[0050] S4. Text input recommendation;
[0051] Listen for the event of the user inputting the question text. First, segment the question, and then match the phrases from the double-array trie. The information such as the indicators, dimensions, and dimension values that are matched is recommended for the user input.
[0052] S5. Extract the key elements of the user question;
[0053] First, segment the user text through HanLP. If there are indicators, dimensions, dimension values, etc. in the user question, they will be directly split out; the remaining text that is not fully matched is first handed over to the vector library for matching. Calculate the semantic similarity according to the vector model. If it exceeds the threshold, it is considered a successful match with this element. If the vector similarity does not exceed the threshold, continue to perform prefix and suffix searches in the HanLP double-array trie, and retain this element if the similarity score exceeds the threshold;
[0054] Finally, the user question will be parsed into a key element group such as indicators, dimensions, and dimension values.
[0055] S6. Date recognition. Build rich time parsing regular expressions to parse time intervals and time granularities. Identify the time period text in the question through regular expressions, and hand it over to the specific time granularity recognition rules to identify the overall time interval one by one, and finally form a time screening range.
[0056] S7. Intent recognition. Build intent recognition rules, including intent recognition rules such as calculating the total amount of indicators, aggregating indicators by dimension grouping, filtering by dimension value, querying the indicator trend by date granularity, TOPN query, and detailed query. Perform intent matching according to the key elements extracted in steps S5 and S6.
[0057] The built-in aggregation function recognition class recognizes aggregation functions such as maximum, minimum, sum, average, and count;
[0058] Built-in sorting recognition conditions, which can be sorted in ascending or descending order according to the metrics;
[0059] Built-in time trend granularity recognition rules, which can perform time trend queries by day, month, week, quarter, or year;
[0060] Built-in metric range filtering rules, such as "Which middle schools have a population greater than 100 and less than 10,000?", to identify the metric filtering range.
[0061] S8, Intent score calculation;
[0062] Use HanLP to perform part-of-speech tagging on the words in the user's question, identify verbs, nouns, adverbs, stop words, etc., and remove the useless words in the question.
[0063] Take the length of the question after removing the useless words as the denominator, and the lengths of the words such as the metrics, dimensions, and aggregation functions matched by the intent as the numerator to calculate the score. If it is lower than the threshold, this intent match will be considered unable to accurately answer the user's question and will be excluded. Retain the top 3 intent recognition results according to the scores.
[0064] S9, Generate query SQL;
[0065] After the recognition and parsing in steps S5, S6, and S7, the user's question has been comprehensively disassembled, forming metrics, dimensions, the data set (semantic model) involved in the metrics, dimension values, aggregation conditions, time conditions, metric filtering conditions, time granularity, sorting conditions, etc.
[0066] Use the jsqlparser parser to generate corresponding query SQL for different types of databases.
[0067] S10, Data query;
[0068] The SQL generated in step S9 is handed over to the columnar database clickHouse for online analysis (OLAP) to achieve fast query and analysis of large data sets.
[0069] S11, Data visualization display;
[0070] After the database executes the query SQL, it returns the data results, identifies the data result field types and metric fields, dimension fields, etc., and automatically adapts and converts them into easy-to-understand visual charts through the number and type of data result fields, and feeds them back to the user through dynamic visualization components, such as line charts, bar charts, tables, pie charts, and other visualization components.
[0071] S12. Security and privacy protection mechanisms;
[0072] When processing sensitive data, the system uses encryption technology to desensitize the data and implements strict access control policies to effectively ensure data security and user privacy while providing intelligent question-and-answer services.
[0073] For example:
[0074] Users can ask questions such as "What is the GDP trend of XX City in 2024?" in natural language. The conversational BI system will first extract key elements: [2024 (time condition), XX City (dimension value), GDP (indicator), trend (time dimension)]. Based on the matched key elements, it performs intent recognition and parses this question into an SQL query statement.
[0075] For example: `select MONTH(create_time) as month, city, sum(GDP) as GDP from tbl_city_gdp where city = 'Jinan City' and YEAR(create_time) = 2024 group by MONTH(create_time), city`
[0076] The query SQL is handed over to the real-time analysis data warehouse for execution and calculation of the results, which are finally presented to the user in the form of charts or tables. At the same time, the system will continuously optimize its understanding and answering ability for similar questions based on user feedback and recommend similar historical questions or generate new questions.
[0077] Based on the above method, a device for realizing natural language question answering in this embodiment includes: at least one memory and at least one processor;
[0078] The at least one memory is used to store machine-readable programs;
[0079] The at least one processor is used to call the machine-readable program to execute a method for realizing natural language question answering.
[0080] The above specific implementation manners are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above specific implementation manners. Any technical solution that conforms to the technical solutions described in the above specific implementation manners of the present invention and any appropriate changes or substitutions made by those of ordinary skill in the relevant technical field shall fall within the patent protection scope of the present invention.
[0081] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for implementing natural language number query, characterized in that: The steps are as follows: S1. Sort out the indicator table, perform semantic modeling, and build a subject domain knowledge graph; S2, write HanLP double array dictionary tree regularly; S3, text input recommendation; S4, extract key elements of user questions; S5, perform date and intention recognition; S6, intention score calculation; S7, generate query SQL and perform data query; S8. Data visualization display.
2. A method for implementing natural language numbering according to claim 1, characterized in that: In step S1, data is integrated from various data sources, and the data is cleaned, converted, standardized and normalized through the data governance center; Perform dimensional modeling on the indicator table under the subject domain, build a summary model, set dimensions, indicators, dates, and primary key fields, construct data sets, configure the association relationship between models, and use a set of constructed triples in the form of <entity, relationship, entity>.
3. A method for implementing natural language numbering according to claim 2, characterized in that: In step S2, a scheduled task is started to write the indicators, dimensions, and dimension values defined in the semantic model into the HanLP double array dictionary tree in a scheduled manner, exchanging space for time.
4. A method for implementing natural language number query according to claim 3, characterized in that: In step S3, the user input question text event is monitored, the question is first segmented, and then the phrases are matched from the double-array dictionary tree, and the matched indicators, dimensions and dimension values are recommended to the user for input.
5. A method for implementing natural language questioning according to claim 4, characterized in that: In step S4, first, the user text is segmented by HanLP. If there are indicators, dimensions and dimension values in the user question, they will be directly split out; The remaining texts that are not fully matched are first matched by the vector library. The semantic similarity is calculated according to the vector model. If it exceeds the threshold, it is considered to be matched successfully with this element. If the vector similarity does not exceed the threshold, the HanLP double array dictionary tree matching is continued to perform prefix and suffix search. If it exceeds the similarity score threshold, this element is retained. Ultimately, user questions are parsed into key element groups of metrics, dimensions, and dimension values.
6. A method for implementing natural language questioning according to claim 5, characterized in that: In step S5, when performing date recognition, the built-in time parsing regular expression is used to parse the time interval and time granularity, and the time period text in the question is recognized by the regular expression, and the unique time granularity recognition rule is used to recognize the overall time interval one by one, and finally a time screening range is formed; When performing intent recognition, built-in intent recognition rules are used, including the total amount of indicators, indicator aggregation by dimension grouping, filtering by dimension value, querying indicator trends by date granularity, TOPN query and detail query intent recognition rules, and intent matching is performed based on the key elements extracted in steps S5 and S6.
7. A method for implementing natural language questioning according to claim 6, characterized in that: In step S6, HanLP is used to perform part-of-speech tagging on the words in the user's question. The question length after removing useless words is used as the denominator, and the indicator, dimension, and aggregation function vocabulary length matched by the intent are used as the numerator to calculate the score. If it is lower than the threshold, this intent match will be considered as unable to accurately answer the user's question and will be eliminated. The top three intent recognition results will be retained according to the score.
8. A method for implementing natural language questioning according to claim 7, characterized in that: In step S7, after identification and analysis in steps S4 and S5, the user's question has been fully analyzed, and the jsqlparser parser is used to generate corresponding query SQL for different types of databases; The generated SQL is handed over to the online analytical (OLAP) column-based database clickHouse to realize the query and analysis of large data sets.
9. A method for implementing natural language questioning according to claim 8, characterized in that: In step S8, after the database executes the query SQL, it returns the data results, which are automatically converted into easy-to-understand visualization charts through the number and type of data result fields and fed back to the user through the dynamic visualization component; When processing sensitive data, encryption technology is used to desensitize the data and implement access control strategies.
10. A device for implementing natural language number asking, characterized in that: include: at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is configured to call the machine-readable program to execute the method according to any one of claims 1 to 9.
Citation Information
Cited By
Vehicle structured data query and visualization method based on large model
CN121412303A