Data query method and device, medium and equipment

Through Chinese language processing of user problems matching the double-array dictionary tree with the topic domain knowledge graph, key elements are extracted and structured query statements are generated, which solves the problem of cumbersome operations of traditional data analysis tools and realizes the convenience and accuracy of natural language data query.

CN120277092APending Publication Date: 2025-07-08SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510296340.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Traditional data analysis tools are cumbersome to operate, and users need to have professional skills, which cannot meet the needs of enterprises to quickly and accurately obtain data insights.

Method used

Chinese language is used to process user problems matching the double-array dictionary tree and topic domain knowledge graph, extract key elements, and generate structured query statements based on intent recognition rules to realize natural language data query.

Benefits of technology

It lowers the threshold for data query, improves data analysis efficiency and accuracy, and users can quickly obtain data insights without professional skills, and supports data visualization display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277092A_ABST
    Figure CN120277092A_ABST
Patent Text Reader

Abstract

The invention provides a data query method and device, a medium and equipment. The method comprises the steps of receiving a user question; carrying out matching operation on the user question and a subject domain knowledge graph by utilizing a Chinese language processing double-array dictionary tree so as to extract key elements in the user question; wherein the subject domain knowledge graph is constructed according to business data in a data source, and is accessed to the Chinese language processing double-array dictionary tree; identifying a user intention according to a pre-configured intention identification rule and the key element; and generating a corresponding structured query statement according to the identified user intention, performing query operation according to the structured query statement, and feeding back a query result to the user. The threshold of data query is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data query, and in particular to a data query method, device, medium, and equipment. Background Art

[0002] With the advent of the big data era, enterprises are faced with a vast amount of data resources and complex data analysis requirements. Traditional data analysis tools often require users to have professional data analysis skills and are cumbersome to operate, unable to meet the needs of enterprises to quickly and accurately obtain data insights. Therefore, how to lower the threshold of data analysis has become an urgent problem to be solved. Summary of the Invention

[0003] In view of the above at least one technical problem, embodiments of the present invention provide a data query method, device, medium, and equipment.

[0004] According to a first aspect, the data query method provided by the embodiments of the present invention includes:

[0005] Receiving a user question;

[0006] Using a Chinese language processing double-array trie to match the user question with a topic domain knowledge graph to extract key elements in the user question; wherein, the topic domain knowledge graph is constructed based on business data in a data source and is connected to the Chinese language processing double-array trie;

[0007] Identifying the user intention according to pre-configured intention recognition rules and the key elements;

[0008] Generating a corresponding structured query statement according to the identified user intention, performing a query operation according to the structured query statement, and feeding back the query result to the user.

[0009] In one embodiment, the operation of using a Chinese language processing double-array trie to match the user question with a topic domain knowledge graph includes:

[0010] Performing word segmentation on the user question to obtain a word segmentation result;

[0011] Using the Chinese language processing double-array trie to perform attribute element matching in the word segmentation result to extract attribute elements in the word segmentation result; wherein, the attribute information includes at least one key element such as index information, dimension information, and label object;

[0012] Calculate the similarity between the remaining word segmentation results after extracting the attribute information and the elements corresponding to the element vectors in the vector library; if the similarity is higher than the preset threshold, use the element corresponding to the element vector in the vector library as the key element; if the similarity is not higher than the preset threshold, use the Chinese language processing double-array trie to search for elements with a similarity higher than the preset threshold, and if found, use the searched element as the key element.

[0013] In one embodiment, the identifying the user intention according to the pre-configured intention recognition rule and the key element includes:

[0014] Identify a plurality of potential intentions according to the intention recognition rule and the key element;

[0015] Score each potential intention to obtain the score corresponding to the potential intention;

[0016] Among the potential intentions with scores greater than the preset score, select the preset number of potential intentions with the highest scores as the user intention.

[0017] In one embodiment, the scoring each potential intention includes:

[0018] Perform part-of-speech tagging on each word of the user question;

[0019] Remove invalid words according to the tagged part of speech;

[0020] Use the length of the user question after removing invalid words as the denominator;

[0021] Use the length of each potential intention as the numerator;

[0022] Calculate the ratio between the numerator of each potential intention and the denominator, and use the ratio as the score of the potential intention.

[0023] In one embodiment, the method further includes: if there is no potential intention with a score greater than the preset score, use a pre-trained large model to identify the user intention from the user question.

[0024] In one embodiment, the method further includes: using a time regular expression to extract a time filtering range from the user question;

[0025] Correspondingly, the identifying the user intention according to the pre-configured intention recognition rule and the key element includes: identifying the user intention according to the pre-configured intention recognition rule, the key element, and the time filtering range.

[0026] In one embodiment, generating a corresponding structured query statement according to the recognized user intention includes: generating a corresponding structured query statement according to the recognized user intention, the key elements, and the intention recognition rules.

[0027] According to a second aspect, the data query device provided by an embodiment of the present invention includes:

[0028] A problem receiving module, configured to receive a user problem;

[0029] An element matching module, configured to use a Chinese language processing double-array trie to match the user problem with a topic domain knowledge graph to extract key elements in the user problem; wherein, the topic domain knowledge graph is constructed based on business data in a data source and accesses the Chinese language processing double-array trie.

[0030] An intention recognition module, configured to recognize a user intention according to pre-configured intention recognition rules and the key elements.

[0031] A statement generation module, configured to generate a corresponding structured query statement according to the recognized user intention, perform a query operation according to the structured query statement, and feed back the query result to the user.

[0032] In one embodiment, the element matching module is specifically configured to: segment the user problem to obtain a segmentation result; use the Chinese language processing double-array trie to perform attribute element matching in the segmentation result to extract attribute elements in the segmentation result; wherein, the attribute information includes at least one key element among index information, dimension information, and label objects; calculate the similarity between the remaining segmentation result after extracting the attribute information and the elements corresponding to the element vectors in the vector library; if the similarity is higher than a preset threshold, use the element corresponding to the element vector in the vector library as the key element; if the similarity is not higher than the preset threshold, use the Chinese language processing double-array trie to search for an element with a similarity higher than the preset threshold, and if an element can be searched, use the searched element as the key element.

[0033] In one embodiment, the intention recognition module includes:

[0034] A first recognition unit, configured to recognize a plurality of potential intentions according to the intention recognition rules and the key elements;

[0035] An intention scoring unit, configured to score each potential intention to obtain a score corresponding to the potential intention;

[0036] An intention selection unit, configured to select the preset number of potential intentions with the highest scores from among the potential intentions with scores greater than a preset score as the user intention.

[0037] In one embodiment, the intention scoring unit is specifically configured to: perform part-of-speech tagging on each word of the user question; eliminate invalid words according to the tagged part of speech; use the length of the user question after eliminating invalid words as the denominator; use the length of each potential intention as the numerator; calculate the ratio between the numerator of each potential intention and the denominator, and use this ratio as the score of this potential intention.

[0038] In one embodiment, the device further includes:

[0039] A large model calling module, configured to, if there is no potential intention with a score greater than the preset score, use a pre-trained large model to identify the user intention from the user question.

[0040] In one embodiment, the device further includes:

[0041] A time filtering module, configured to extract a time filtering range from the user question by using a time regular expression;

[0042] Correspondingly, the intention recognition module is specifically configured to: recognize the user intention according to pre-configured intention recognition rules, the key elements, and the time filtering range.

[0043] In one embodiment, the statement generation module is specifically configured to: generate a corresponding structured query statement according to the recognized user intention, the key elements, and the intention recognition rules.

[0044] According to a third aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method provided in the first aspect.

[0045] According to a fourth aspect, a computing device provided by an embodiment of the present invention includes a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method provided in the first aspect is implemented.

[0046] The data query method, device, medium, and equipment provided by the embodiments of the present invention, when receiving a user question, use a Chinese language processing double-array trie to match the user question with a topic domain knowledge graph to extract key elements in the user question, and then identify the user intention according to the intention recognition rule and the key elements. Furthermore, a corresponding structured query statement is generated according to the user intention, and then a query operation can be performed according to the structured query statement, and finally the query result is fed back to the user. The method provided by the embodiments of the present invention can understand various data query requirements put forward by users in natural language form and perform accurate data analysis tasks accordingly. Users do not need to have professional data analysis skills, which reduces the threshold of data query. Moreover, the embodiments of the present invention use a Chinese language processing double-array trie to extract key elements in the user question, and then perform intention recognition according to the key elements. This method can accurately understand users and improve the accuracy of subsequent data query. Description of the Drawings

[0047] Figure 1 It is a schematic flowchart of the data query method in an embodiment of the present invention;

[0048] Figure 2 It is a structural block diagram of the data query device in an embodiment of the present invention. Detailed Embodiments

[0049] In a first aspect, the embodiments of the present invention provide a data query method. Refer to Figure 1 and the method includes the following steps S110 to S140:

[0050] S110. Receive a user question;

[0051] Among them, the user question refers to the question for the user to perform data query. For example, the user question is: Which cities have a GDP greater than 100 billion and less than 200 billion in 2023?.

[0052] S120. Use a Chinese language processing double-array trie to match the user question with a topic domain knowledge graph to extract key elements in the user question; among them, the topic domain knowledge graph is constructed based on the business data in the data source and is accessed to the Chinese language processing double-array trie;

[0053] Among them, the Chinese language processing double-array trie, namely the HanLP double-array trie, and the full spelling of HanLP is Double-Array Trie. The trie tree is a data structure used for natural language processing, mainly for tasks such as string matching, retrieval, and word frequency statistics. The trie tree is a tree-shaped data structure used to store a set of strings, where each node represents a prefix of a string. The core idea of the trie tree is to use the common prefix of strings to reduce the query time, thereby improving efficiency. The trie tree uses space for time to reduce the query time overhead by using the common prefix of strings.

[0054] Among them, a thematic domain knowledge graph is constructed based on the business data in the data source. The thematic domain knowledge graph is connected to the Chinese language processing double-array trie. The data source can be a database, a data warehouse, a WEB service, Excel, etc. After the business data in the data source is standardized and normalized through the data governance center, the thematic domain knowledge graph is constructed to ensure data quality. The data quality directly determines the accuracy of the question and answer.

[0055] Among them, the thematic domain knowledge graph can also be understood as a semantic model, and this semantic model includes the knowledge of multiple thematic domains. Each thematic domain has a corresponding index table, and the index table includes multiple fields, such as fields like dimension, date, primary key, etc. The dependency relationship between fields can also be constructed, and the dependency relationship can be expressed in the form of a triple set, that is, <entity, relationship, entity>, for example: (data catalog A, belongs to department, department B). The thematic domain knowledge graph can be stored using a graph database and a vector database to enhance the semantic understanding ability of the system. The semantic model supports synonym configuration in terms of function, which is convenient for more accurate element matching.

[0056] Specifically, the thematic domain knowledge graph can be connected to the HanLP double-array trie by starting a scheduled task. When the HanLP double-array trie performs field matching, it uses space for time to reduce the time overhead of word matching by using the common prefix of strings.

[0057] In one embodiment, the operation of matching the user question with the thematic domain knowledge graph by using the Chinese language processing double-array trie in S120 may include S121 to S123:

[0058] S121. Segment the user question to obtain a segmentation result;

[0059] S122. Use the Chinese language processing double-array trie to perform attribute element matching in the segmentation result to extract the attribute elements in the segmentation result; among them, the attribute information includes at least one key element such as index information, dimension information, and label object;

[0060] Among them, the index information, for example, GDP.

[0061] Among them, the dimension information may include the dimension and the dimension value. The dimension can be the year, and the dimension value is the specific year, for example, 2024.

[0062] Among them, the labeled object, for example, Shandong Province.

[0063] It can be seen that the user question includes key elements. The Chinese language processing double-array trie matches the tokenization results after tokenizing the user question with the topic domain knowledge graph, so as to identify the key elements in the user question that match the knowledge in the topic domain knowledge graph.

[0064] S123. Calculate the similarity between the tokenization results remaining after extracting the attribute information and the elements corresponding to the element vectors in the vector library; if the similarity is higher than the preset threshold, then use the element corresponding to the element vector in the vector library as the key element; if the similarity is not higher than the preset threshold, then use the Chinese language processing double-array trie to search for elements with a similarity higher than the preset threshold, and if found, use the searched element as the key element.

[0065] Among them, the vector library includes multiple element vectors, and each element vector corresponds to an element.

[0066] It can be understood that after the element matching is completed through the Chinese language processing double-array trie, some content in the analysis result of the user question fails to match successfully. The content that fails to match successfully is the remaining content after matching. Calculate the similarity between the remaining content after matching and the elements corresponding to the element vectors in the vector library. If the similarity is higher than the preset threshold, it means that the remaining content has a high similarity with this element. Therefore, this element is also used as a key element of the user question.

[0067] Of course, if there is still no successful match by calculating the similarity, the HanLP double-array trie matching method can be continued to perform prefix and suffix searches, and the element is retained if it exceeds the similarity score threshold.

[0068] Finally, the user question will be split into multiple key elements.

[0069] S130. Identify the user intention according to the pre-configured intention recognition rules and the key elements;

[0070] Among them, the intention recognition rules can include finding the total amount of indicators, aggregating indicators by dimension grouping, filtering by dimension value, querying indicator trends by date granularity, TOPN query, detailed query, etc. The intention recognition rules can be built-in. For example, the built-in aggregation function recognition class recognizes aggregation functions such as maximum, minimum, sum, average, and quantity; the built-in sorting recognition conditions can be arranged in ascending or descending order according to the indicators; the built-in time trend granularity recognition rules can be used to query time trends by day, month, week, quarter, and year; the built-in indicator range screening rules can be filtered by ranges such as greater than, less than, equal to, not greater than, between, and between.

[0071] In one embodiment, the step of identifying the user intent according to the preconfigured intent identification rule and the key element in S130 includes S131 to S133:

[0072] S131, identifying multiple potential intentions according to the intention identification rule and the key elements;

[0073] S132, scoring each potential intention to obtain a score corresponding to the potential intention;

[0074] S133. Select a preset number of potential intentions with the highest scores from among the potential intentions with scores greater than a preset score as the user intentions.

[0075] That is to say, based on the intention recognition rules and the key elements, multiple potential intentions of the user can be identified, each potential intention can be scored, and then the potential intentions with the highest scores are selected as the user intentions from among the potential intentions with scores higher than a certain value.

[0076] Furthermore, the scoring of each potential intention described in S132 includes the following five steps:

[0077] 1. Tag the parts of speech for each word in the user's question;

[0078] 2. Eliminate invalid words according to the marked parts of speech;

[0079] 3. The length of the user's question after removing invalid words is used as the denominator;

[0080] 4. Use the length of each potential intention as the numerator;

[0081] 5. Calculate the ratio between the numerator and the denominator of each potential intent, and use the ratio as the score of the potential intent.

[0082] For example, HanLP is used to perform part-of-speech tagging on the words in the user's question to identify verbs, nouns, adverbs, stop words, etc., and remove the useless words in the question. The length of the question after removing the useless words is used as the denominator, and the length of the potential intention is used as the numerator to calculate the score. If the score is lower than the threshold, this potential intention will be considered unable to accurately answer the user's question and will be excluded. Among the potential intentions with scores higher than the threshold, the three potential intentions with the highest scores are selected as the user's intention.

[0083] In one embodiment, the method may further include: if there is no potential intention with a score greater than the preset score, the pre-trained large model is used to identify the user intention from the user's question.

[0084] That is to say, if it is found in the above S133 that there is no user intention with a score greater than the preset score, then the preset large model can be called to process the user's question. The preset large model can adopt the advanced pre-trained model of the Transformer architecture - the Hairo model. This model is adapted to the intelligent question scenario to ensure that the model can understand and accurately answer the natural language questions input by the user, accurately identify the user's query intention, extract keywords and entities from the user's question, such as elements like metrics, time, location, dimension, limiting conditions, etc., and convert them into the user's intention.

[0085] S140. According to the identified user intention, generate a corresponding structured query statement, perform a query operation according to the structured query statement, and feedback the query result to the user.

[0086] Specifically, the jsqlparser parser can be used to generate corresponding structured query statements for different types of databases.

[0087] In one embodiment, the generating a corresponding structured query statement according to the identified user intention in S140 may include: generating a corresponding structured query statement according to the identified user intention, the key elements, and the intention recognition rules.

[0088] Since the user's question has been comprehensively disassembled before to form metrics, dimensions, semantic models, dimension values, aggregation conditions, time conditions, metric screening conditions, time granularity, sorting conditions, user intentions, etc. When generating the structured query statement, not only considering the user intention, but also considering the key elements and the intention recognition rules, etc., a more convenient structured query statement for subsequent accurate query can be generated.

[0089] In an actual scenario, to ensure efficient data query, an efficient data index system can be established to preprocess and organize the original data for fast retrieval and analysis. At the same time, the query efficiency of the data set can be improved through vector retrieval technology. After the database finishes executing the query SQL, the query result is returned, and the field types, metric fields, dimension fields, etc. in the query result are identified, and the query result is automatically adapted and transformed into an easy-to-understand visualization chart, and the visualization chart is fed back to the user through a dynamic visualization component, such as line charts, bar charts, tables, pie charts and other visualization components.

[0090] In an actual scenario, when dealing with sensitive data, encryption technology can be used to desensitize the data, and strict access control policies can be implemented to effectively ensure data security and user privacy while providing intelligent question-and-answer services.

[0091] In one embodiment, the method may further include: extracting a time filtering range from the user question by using a time regular expression;

[0092] Correspondingly, the identifying the user intention according to the pre-configured intention recognition rule and the key element includes: identifying the user intention according to the pre-configured intention recognition rule, the key element and the time filtering range.

[0093] It can be understood that rich built-in relevant time regular expressions are used to parse the time interval and time granularity. The time period text in the user question is identified through regular expressions, and the overall time interval is identified one by one by the specific time granularity recognition rule, and finally the time filtering range is formed. When identifying the user intention, the time filtering range can also be considered to improve the accuracy of user intention recognition.

[0094] For example, the user can ask a question like "The GDP situation of each city in Shandong Province in 2022" in natural language. In the embodiment of the present invention, a semantic model is constructed, and the semantic model is connected to the Chinese language processing double-array trie. The key elements in the user's question are extracted by using the Chinese language processing double-array trie: [2022 (time condition), Shandong Province (dimension, dimension value), GDP (indicator), each city (aggregation condition)]. According to the matched key elements, intention recognition is performed, and the score of each potential intention is calculated. Among the potential intentions with scores higher than the threshold, the 3 potential intentions with the highest scores are selected as the user's intention. An SQL statement is generated according to the user's intention and other information. The SQL statement is converted into a physical query statement that matches the syntax according to the data type. For example, select MONTH(create_time) as month, province, city, sum(GDP) as GDP from tbl_city_gdp where province = 'Shandong Province' and YEAR(create_time) = 2022 group by MONTH(create_time), province, city`. Then the query SQL is handed over to the real-time analysis data warehouse for execution and calculation of the results, and finally presented to the user in the form of a chart or a table.

[0095] The embodiment of the present invention has the following beneficial effects:

[0096] (1) It reduces the threshold of data query: Users do not need to have professional data analysis skills, and can interact with the system in natural language to quickly obtain the required data insights.

[0097] (2) It accelerates data-driven decision-making: It responds to complex and changing data requirements in real time, helps enterprises make data-based decisions quickly, and reduces the time cost of waiting for the assistance of the IT department or the data analysis team.

[0098] (3) It improves the efficiency and convenience of data analysis: It can quickly understand the user's query requirements and give accurate analysis results, and at the same time supports data visualization display to help users better understand the data.

[0099] (4) It realizes the unified governance of data: By constructing a unified semantic model, integrating massive data resources, forming a unified data view, and improving the management efficiency and use convenience of data.

[0100] (5) It improves the utilization rate of data resources: Through the effective integration and intelligent search of massive data resources, various types of data within the enterprise can be fully mined and utilized, avoiding the idle and waste of data resources.

[0101] (6) Have the ability of self - learning. As users use and provide feedback, the large model can continuously optimize its own understanding and generation capabilities, and continuously improve the service quality.

[0102] In the embodiments of the present invention, by constructing a complete semantic model and recognition rules, for natural - language questions, the key information in the user's question is extracted to generate a structured query statement. Among them, through HanLP trie - tree word segmentation matching, the key elements in the user's question are extracted; combined with a vector library for similarity calculation to assist semantic understanding and extract key elements with similar semantics. Abundant built - in intent recognition rules are included, including but not limited to index screening mode, detailed query mode, ratio query mode, etc.; abundant built - in time - range and time - granularity recognition rules can accurately extract the time interval in the user's question; structured query SQL statements for multiple database types are assembled and adapted covering elements such as metrics, dimensions, dimension values, entities, time ranges, and screening conditions. The embodiments of the present invention are applicable to fields such as leadership cockpits, data operation dashboards, and data analysis, greatly improving the data query efficiency, lowering the data - analysis threshold, and providing users with intuitive and conversational convenient data - insight services.

[0103] That is to say, the embodiments of the present invention aim to construct an intelligent data - query solution based on a large model and semantic modeling, which can understand various data - query requirements put forward by users in the form of natural language and perform accurate data - analysis tasks accordingly. By applying the large model to the scenarios of enterprise - internal knowledge Q&A and data analysis, not only is the convenience and intelligence of data query greatly improved, but also front - line employees without professional data - analysis skills can easily obtain the required data information, helping enterprises and government services to achieve digital and intelligent transformation.

[0104] In short, the embodiments of the present invention realize the intelligent understanding and quick response to user questions, improve the accuracy and reliability of answers, enable non - technical users to easily query and analyze data in daily language, realize the popularization, personalization, and real - time nature of complex data analysis, and provide users with a more convenient, efficient, and accurate conversational page - display service.

[0105] In the second aspect, the embodiments of the present invention provide a data - query device. Refer to Figure 2 , the device 100 includes:

[0106] A question - receiving module 110, configured to receive user questions;

[0107] An element - matching module 120, configured to use a Chinese language processing double - array trie to match the user question with a topic - domain knowledge graph to extract the key elements in the user question; wherein, the topic - domain knowledge graph is constructed based on business data in a data source and is connected to the Chinese language processing double - array trie.

[0108] An intent recognition module 130, configured to recognize a user intent according to pre-configured intent recognition rules and the key elements;

[0109] A statement generation module 140, configured to generate a corresponding structured query statement according to the recognized user intent, perform a query operation according to the structured query statement, and feed back the query result to the user.

[0110] In one embodiment, the element matching module is specifically configured to: segment the user question to obtain a segmentation result; use the Chinese language processing double-array trie to perform attribute element matching in the segmentation result to extract the attribute elements in the segmentation result; where the attribute information includes at least one key element of index information, dimension information, and label object; calculate the similarity between the remaining segmentation result after extracting the attribute information and the elements corresponding to the element vectors in the vector library; if the similarity is higher than a preset threshold, use the element corresponding to the element vector in the vector library as the key element; if the similarity is not higher than the preset threshold, use the Chinese language processing double-array trie to search for an element with a similarity higher than the preset threshold, and if it can be found, use the searched element as the key element.

[0111] In one embodiment, the intent recognition module includes:

[0112] A first recognition unit, configured to recognize a plurality of potential intents according to the intent recognition rules and the key elements;

[0113] An intent scoring unit, configured to score each potential intent to obtain a score corresponding to the potential intent;

[0114] An intent selection unit, configured to select a preset number of potential intents with the highest scores from the potential intents with scores greater than a preset score as the user intent.

[0115] In one embodiment, the intent scoring unit is specifically configured to: perform part-of-speech tagging on each word of the user question; eliminate invalid words according to the tagged part-of-speech; use the length of the user question after eliminating invalid words as the denominator; use the length of each potential intent as the numerator; calculate the ratio between the numerator and the denominator of each potential intent, and use the ratio as the score of the potential intent.

[0116] In one embodiment, the device further includes:

[0117] A large model call module, configured to, if there is no potential intent with a score greater than the preset score, use a pre-trained large model to recognize the user intent from the user question.

[0118] In one embodiment, the apparatus further includes:

[0119] A time filtering module, configured to extract a time filtering range from the user question by using a time regular expression;

[0120] Correspondingly, the intention recognition module is specifically configured to: recognize the user intention according to pre-configured intention recognition rules, the key elements, and the time filtering range.

[0121] In one embodiment, the statement generation module is specifically configured to: generate a corresponding structured query statement according to the recognized user intention, the key elements, and the intention recognition rules.

[0122] It can be understood that for the explanations, specific implementation manners, beneficial effects, examples, etc. of the relevant content in the apparatus provided by the embodiments of the present invention, reference can be made to the corresponding parts in the method provided in the first aspect, and details are not described herein again.

[0123] In a third aspect, an embodiment of the present invention provides a computer-readable medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the processor is caused to execute the method provided in the first aspect.

[0124] Specifically, a system or device equipped with a storage medium can be provided, on which software program code for implementing the functions of any one of the above embodiments is stored, and the computer (or CPU or MPU) of the system or device is caused to read and execute the program code stored in the storage medium.

[0125] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments, so the program code and the storage medium storing the program code constitute a part of the present invention.

[0126] Examples of storage media for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer via a communication network.

[0127] In addition, it should be clear that not only can the actual operations be completed in part or in whole by executing the program code read by the computer, but also by means of instructions based on the program code, the operating system operating on the computer, etc., so as to implement the functions of any one of the above embodiments.

[0128] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion module connected to the computer, and then based on the instructions of the program code, the CPU or the like installed on the expansion board or the expansion module is made to execute part or all of the actual operations, thereby implementing the functions of any one of the above embodiments.

[0129] It can be understood that the explanations, specific implementation manners, beneficial effects, examples, etc. of the relevant content in the computer-readable medium provided by the embodiments of the present invention can be referred to the corresponding parts in the method provided in the first aspect, and will not be elaborated here.

[0130] In a fourth aspect, an embodiment of this specification provides a computing device, including a memory and a processor, where an executable code is stored in the memory, and when the processor executes the executable code, the method in any one of the embodiments in the specification is implemented.

[0131] It can be understood that the explanations, specific implementation manners, beneficial effects, examples, etc. of the relevant content in the computing device provided by the embodiments of the present invention can be referred to the corresponding parts in the method provided in the first aspect, and will not be elaborated here.

[0132] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0133] Those skilled in the art should be able to realize that in the above one or more examples, the functions described in the present invention can be implemented by hardware, software, a plug-in, or any combination thereof. When implemented by software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0134] The specific implementation manners described above further elaborate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above is only the specific implementation manners of the present invention, and is not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made on the basis of the technical solution of the present invention should be included in the protection scope of the present invention.

Claims

1. A data query method, characterized in that, Including: Receiving a user question; Using a Chinese language processing double-array trie to match the user question with a topic domain knowledge graph to extract key elements in the user question; wherein, the topic domain knowledge graph is constructed based on business data in a data source and accesses the Chinese language processing double-array trie; Identifying the user intent according to pre-configured intent recognition rules and the key elements; Generating a corresponding structured query statement according to the identified user intent, performing a query operation according to the structured query statement, and feeding back the query result to the user.

2. The method according to claim 1, wherein The operation of using a Chinese language processing double-array trie to match the user question with a topic domain knowledge graph includes: Segmenting the user question to obtain a segmentation result; Using the Chinese language processing double-array trie to perform attribute element matching in the segmentation result to extract attribute elements in the segmentation result; wherein, the attribute information includes at least one key element among index information, dimension information, and label objects; Calculating the similarity between the remaining segmentation result after extracting the attribute information and the elements corresponding to the element vectors in the vector library; if the similarity is higher than a preset threshold, taking the element corresponding to the element vector in the vector library as the key element; if the similarity is not higher than the preset threshold, using the Chinese language processing double-array trie to search for elements with a similarity higher than the preset threshold, and if found, taking the searched element as the key element.

3. The method according to claim 1, characterized in that, The operation of identifying the user intent according to pre-configured intent recognition rules and the key elements includes: Identifying multiple potential intents according to the intent recognition rules and the key elements; Scoring each potential intent to obtain the score corresponding to the potential intent; Selecting the preset number of potential intents with the highest scores among the potential intents with scores greater than the preset score as the user intent.

4. The method according to claim 3, wherein The operation of scoring each potential intent includes: Performing part-of-speech tagging on each word of the user question; Removing invalid words according to the tagged part-of-speech; Taking the length of the user question after removing invalid words as the denominator; Taking the length of each potential intent as the numerator; Calculating the ratio between the numerator and the denominator of each potential intent, and taking the ratio as the score of the potential intent.

5. The method according to claim 3, wherein It also includes: If there are no potential intents with scores greater than the preset score, using a pre-trained large model to identify the user intent from the user question.

6. The method according to claim 1, characterized in that, It also includes: Using a time regular expression to extract a time filtering range from the user question; Correspondingly, the operation of identifying the user intent according to pre-configured intent recognition rules and the key elements includes: identifying the user intent according to pre-configured intent recognition rules, the key elements, and the time filtering range.

7. The method according to claim 1, wherein The operation of generating a corresponding structured query statement according to the identified user intent includes: Generating a corresponding structured query statement according to the identified user intent, the key elements, and the intent recognition rules.

8. A data query device, characterized in that, Including: A question receiving module for receiving a user question; An element matching module, which is used to match the user question with the topic domain knowledge graph by using a Chinese language processing double-array trie to extract key elements in the user question; wherein, the topic domain knowledge graph is constructed based on business data in a data source and accesses the Chinese language processing double-array trie. An intention recognition module, which is used to recognize the user intention according to pre-configured intention recognition rules and the key elements. A statement generation module, which is used to generate a corresponding structured query statement according to the recognized user intention, perform a query operation according to the structured query statement, and feedback the query result to the user.

9. A computer-readable storage medium, characterized in that, It stores a computer program, and when the computer program is executed on a computer, it causes the computer to execute the method described in any one of claims 1 to 7.

10. A computing device, characterized in that, It includes a memory and a processor. An executable code is stored in the memory, and when the processor executes the executable code, the method described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Visual knowledge graph query template construction method, device and system and storage medium

    CN112507135A

  • Expert domain knowledge graph query method based on natural language questions

    CN112597272A

  • Question and answer method and system based on tin smelting knowledge graph

    CN118779424A

  • Index query method and device, electronic equipment and storage medium

    CN119311819A