Table-based query method and device, computer equipment, readable storage medium and program product
The method improves query accuracy in table data processing by using a pattern extraction model to match user queries with standard table data, addressing ambiguity and vagueness issues in NLP models.
Patent Information
- Application Number
- CN202510789636.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-15
AI Technical Summary
The existing tabular data query methods face ambiguity and ambiguity when dealing with user problems, resulting in the model-generated answers that do not meet the actual needs of users and lack accuracy.
By obtaining query information, the information is extracted to obtain the to-processed column names and data, the pattern extraction model is used to train and identify the column names and column values, and match the standard column names and data in the table in combination with semantics, syntax and context matching methods. The target column names are determined using a multi-way verification mechanism, and processed through the big model to improve accuracy.
Improve the accuracy of table data query, ensure the accuracy of large-scale model processing, reduce user interaction, and improve system response speed and user experience.
Smart Images

Figure CN120316147A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a table-based query method, device, computer device, computer-readable storage medium, and computer program product. Background Art
[0002] With the rapid development of information technology, the generation and accumulation of data show an explosive growth trend. This phenomenon is particularly prominent in the field of structured data, especially the generation of tabular data. Tabular data, with its intuitive and structured characteristics, widely exists in various fields such as finance and commerce. These data not only have a large volume but also a complex variety. With the sharp increase in the amount of data, how to efficiently extract useful information and knowledge from these huge tabular data has become an urgent challenge to be solved. Traditional manual data analysis and query methods are no longer able to meet the needs of modern data processing.
[0003] In this context, natural language processing (NLP) technology has developed rapidly. In recent years, large language models (such as GPT-3, GPT-4, etc.) have performed well in various NLP tasks. Especially in dealing with the combination problem of natural language and structured data, the progress of language models provides new solutions for table question-answering systems. Existing solutions usually involve converting tabular data into descriptive language, such as by designing complex prompts (methods like chain of thought) to assist large models in generating table question answers, or extracting header information, the format of table content, and the first few rows of table data, etc. These descriptive information are provided as input to the large model, and the model will generate the code or answer required to solve the user's problem based on its understanding of language and data. This method has improved the processing efficiency of tabular data to a certain extent, but still faces many challenges in practical applications.
[0004] Specifically, the large model faces the following difficulties when dealing with tabular data problems:
[0005] Ambiguity and vagueness of user questions: The questions raised by users may be vague or ambiguous, affecting the accuracy of the model. For example, the question "What is the salary of employees" may lack a clear time point, or the question "profit" may refer to "gross profit" or "net profit", etc. These may all cause the answers generated by the model not to meet the actual needs of users. Summary of the Invention
[0006] Based on this, it is necessary to provide a table-based query method, device, computer device, computer-readable storage medium, and computer program product that can improve the accuracy of query results for the above technical problems.
[0007] In a first aspect, this application provides a table-based query method, the method includes:
[0008] Obtain query information;
[0009] Perform information extraction on the query information to obtain the column names to be processed and the data to be processed;
[0010] Match the column names to be processed and the data to be processed with the standard column names and standard data in the table respectively to obtain the target column names;
[0011] Process the target column names through a large model to obtain a query result corresponding to the query information.
[0012] In one embodiment, the performing information extraction on the query information to obtain the column names to be processed and the data to be processed includes:
[0013] Process the query information through a pre-trained pattern extraction model to obtain the column names to be processed and the data to be processed; the pattern extraction model is trained based on sample natural language and the corresponding table column names and column values of the sample natural language.
[0014] In one embodiment, the matching the column names to be processed and the data to be processed with the standard column names and standard data in the table respectively to obtain the target column names includes:
[0015] Match the column names to be processed and the data to be processed with the standard column names and standard data in the table respectively through different matching methods to obtain the initial column names corresponding to each matching method;
[0016] Determine the verification logic;
[0017] Select the target column names from the initial column names based on the verification logic.
[0018] In one embodiment, the matching the column names to be processed and the data to be processed with the standard column names and standard data in the table respectively through different matching methods to obtain the initial column names corresponding to each matching method includes at least one of the following:
[0019] Match the column names to be processed and the data to be processed with the standard column names and standard data in the table respectively based on semantic matching to obtain the first initial column names, where the semantic matching is performed based on a pre-trained language model;
[0020] Match the column names to be processed and the data to be processed with the standard column names and standard data in the table respectively based on syntactic matching to obtain the second initial column names, where the syntactic matching is performed based on character distance or regular expressions;
[0021] Match the to-be-processed column name and the to-be-processed data with the standard column names and standard data in the table respectively in a context matching manner to obtain a third initial column name. The context matching manner is to determine the scenario classification corresponding to the to-be-processed column name and the to-be-processed data based on the query information, screen the standard column names and standard data in the table based on the scenario classification, and match the screened standard column names and standard data with the to-be-processed column name and the to-be-processed data respectively.
[0022] In one embodiment, the selecting the target column name from the initial column names based on the verification logic includes at least one of the following:
[0023] Obtain the priorities of different matching methods set based on the scenario, and select the initial column name corresponding to the target priority as the target column name;
[0024] Match the initial column names corresponding to each matching method, and determine the number of matching methods with the same initial column name; select the initial column names corresponding to each matching method whose quantity meets the requirements as the target column names;
[0025] Cross-check the initial column name obtained based on the to-be-processed column name and the initial column name obtained based on the to-be-processed data, and determine the target column name based on the result of the cross-check.
[0026] In one embodiment, the method further includes:
[0027] When there are at least two target column names determined based on the verification logic, output each target column name;
[0028] Receive a selection instruction for the target column name, and determine the final target column name based on the selection instruction.
[0029] In a second aspect, the present application further provides a table-based query device, and the device includes:
[0030] A query information acquisition module, configured to acquire query information;
[0031] An extraction module, configured to perform information extraction on the query information to obtain a to-be-processed column name and a to-be-processed data;
[0032] A matching module, configured to match the to-be-processed column name and the to-be-processed data with the standard column names and standard data in the table respectively to obtain a target column name;
[0033] A query module, configured to process the target column name through a large model to obtain a query result corresponding to the query information.
[0034] In a third aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the method in any one of the above embodiments are implemented.
[0035] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method in any one of the above embodiments are implemented.
[0036] In a fifth aspect, the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method in any one of the above embodiments are implemented.
[0037] The above table-based query method, device, computer device, computer-readable storage medium, and computer program product obtain query information; perform information extraction on the query information to obtain the column names to be processed and the data to be processed; match the column names to be processed and the data to be processed with the standard column names and standard data in the table respectively to obtain the target column names; process the target column names through a large model to obtain a query result corresponding to the query information. In this way, by pre-extracting the query information to obtain the column names to be processed and the data to be processed, and then matching them with the standard column names and standard data, the target column names can be accurately determined, improving the accuracy of subsequent large model processing, and thus improving the accuracy of the query result. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained without creative efforts based on these drawings.
[0039] Figure 1 It is an application environment diagram of the table-based query method in an embodiment;
[0040] Figure 2 It is a flowchart of the table-based query method in an embodiment;
[0041] Figure 3 It is a flowchart of the matching step in an embodiment;
[0042] Figure 4 It is a structural block diagram of the table-based query device in an embodiment;
[0043] Figure 5Internal structure diagram of a computer device in an embodiment. Detailed implementation
[0044] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0045] The query method based on a table provided by the embodiments of the present application can be applied to, for example Figure 1 The application environment shown. Among them, the terminal 102 communicates with the server 104 through a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed in the cloud or other network servers.
[0046] Among them, the terminal 102 can receive the input query information, which can be input through at least one of text, voice, and actions. The present application does not specifically limit the input method of the query information. The terminal 102 sends the query information to the server 104. The server 104 extracts the information from the query information to obtain the column names to be processed and the data to be processed; the column names to be processed and the data to be processed are respectively matched with the standard column names and standard data in the table to obtain the target column names; the target column names are processed through a large model to obtain a query result corresponding to the query information. Subsequently, the query result is returned to the terminal 102, and the terminal 102 displays or outputs the query result through other response methods, such as voice or actions. In this way, by pre-extracting the query information to obtain the column names to be processed and the data to be processed, and then matching them with the standard column names and standard data, the target column names can be accurately determined, improving the accuracy of subsequent large model processing, thereby improving the accuracy of the query result.
[0047] Among them, in the above embodiment, the query method based on a table is applied to a system including a terminal and a server as an example, and is implemented through the interaction between the terminal and the server. It can be understood that the query method based on a table can be applied to the terminal alone or to the server alone.
[0048] Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server 104 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0049] In an exemplary embodiment, as Figure 2 shown, a table-based query method is provided. Taking the method applied to the Figure 1 server in it as an example for illustration, it includes the following steps S202 to step S208. Among them:
[0050] S202: Obtain query information.
[0051] The query information is the data to be queried input. The query information can include questions, keywords, etc. The query information can be input in ways such as text, voice, or actions. Among them, the query information input in the form of voice or actions can be first converted into text information for subsequent processing.
[0052] S204: Perform information extraction on the query information to obtain the column names to be processed and the data to be processed.
[0053] Among them, the column names to be processed and the data to be processed are obtained by performing information extraction on the query data. The column names to be processed and the data to be processed may not be the column names or data in the table, and they are only obtained according to the query information.
[0054] In some optional embodiments, performing information extraction on the query information to obtain the column names to be processed and the data to be processed includes: processing the query information through a pre-trained pattern extraction model to obtain the column names to be processed and the data to be processed; the pattern extraction model is trained based on sample natural language and the corresponding table column names and column values of the sample natural language.
[0055] Among them, the pattern extraction model can be a relatively small model (such as Deberta and other models), which reduces the burden on large models and reduces overhead. Identify and extract possible "fake column names" (i.e., the column names to be processed above) and "fake values" (i.e., the data to be processed above) from the user's input question (i.e., the query information above), and calculate the logical relationship between the two.
[0056] Among them, the task of the pattern extraction model is to extract possible "fake column names", "fake values" from the user's natural language input (i.e., query information), as well as the corresponding calculation logic relationships. The training of this pattern extraction model depends on a large amount of labeled data. Through the method of supervised learning, the pattern extraction model gradually learns to correctly identify the column names and column values to be processed mentioned by the user in various contexts, even if they do not directly correspond to the real column names and column values in the table one by one. The result output by the pattern extraction model is a list of possible "fake column names" and "fake values", and these lists will be used as the input for the subsequent matching steps.
[0057] S206: Match the column name to be processed and the data to be processed with the standard column names and standard data in the table respectively to obtain the target column name.
[0058] Among them, the query result is based on the table. The table includes standard column names and standard data. The column name to be processed and the data to be processed obtained in the extraction step are respectively matched with the standard column names and standard data in the table to obtain the target column name.
[0059] Specifically, the server performs a first match between the column name to be processed and the standard column names in the table, and a second match between the data to be processed and the standard data in the table. The first match and the second match can be performed serially or in parallel, and no specific limitation is made here. However, for the purpose of improving processing efficiency, in this application, the first match and the second match are performed in parallel.
[0060] In some alternative embodiments, the first match may obtain a first target column name, and the second match may obtain a second target column name. In this application, the final target column name is obtained by verifying the first target column name and the second target column name. The verification method can be at least one of comparing each other and outputting for the user to select.
[0061] In some alternative embodiments, the first match and the second match may include at least one matching method. For example, these matching methods may include semantic matching methods, syntactic matching methods, and context matching methods. In other embodiments, the matching methods may also include others, and no specific limitation is made here. Among them, in each match, if multiple matching methods are used, multiple matching results may occur. In this application, the matching results corresponding to each matching method in each match (such as the first match) are processed to obtain the target matching result corresponding to each match (such as the first matching result). The processing of each matching method may include at least one of selecting the result corresponding to the optimal matching method based on priority, selecting the result supported by the most based on consistency, and outputting the result of unsuccessful matching through result cross-verification for the user to confirm.
[0062] In the above way, the corresponding column names can be determined more accurately, making the user's questions clear. In addition, the user may not be familiar enough with the table, resulting in the column names used in the question being semantically similar but not exactly the same as those in the table. For example, the user asks "salary" while "wage" is used in the table. This difference may cause the large model to fail to accurately match the correct column name. In this application, the above semantic matching method, grammar matching method, and context matching method are used to avoid this situation and improve the processing efficiency.
[0063] In addition, users sometimes ask questions about specific content in the table, and this content may not be directly reflected as column names. Due to the input length limit of the large model, it is impossible to use all the content of the table as input, which makes the model face difficulties in processing complex questions involving specific data content. Therefore, in this application, the matching of the data to be processed and the standard data is also introduced, so that the content of the table can also be introduced into the processing process to improve the processing accuracy.
[0064] S208: Process the target column name through the large model to obtain a query result corresponding to the query information.
[0065] After determining the target column name, the large model processes based on the target column name, so that the user's question is targeted at the corresponding target column name, which can improve the processing accuracy of the large model, that is, the accuracy of the query result corresponding to the query information is improved.
[0066] In other embodiments, the target column name and the query information can be input into the large model, so that the large model processes the query information with reference to the target column name to improve the accuracy of the query result.
[0067] The above table-based query method includes: obtaining query information; performing information extraction on the query information to obtain the column name to be processed and the data to be processed; matching the column name to be processed and the data to be processed with the standard column name and the standard data in the table respectively to obtain the target column name; processing the target column name through the large model to obtain a query result corresponding to the query information. By pre-extracting the query information to obtain the column name to be processed and the data to be processed, and then matching them with the standard column name and the standard data, the target column name can be accurately determined, improving the processing accuracy of the subsequent large model, and thus improving the accuracy of the query result.
[0068] In some alternative embodiments, matching the column name to be processed and the data to be processed with the standard column names and standard data in a table respectively to obtain a target column name includes: matching the column name to be processed and the data to be processed with the standard column names and standard data in the table respectively through different matching methods to obtain initial column names corresponding to each matching method; determining a verification logic; and selecting a target column name from the initial column names based on the verification logic.
[0069] As shown in combination with Figure 3 FIG. Figure 3 is a schematic flowchart of the matching step in an embodiment. In this embodiment, the column name to be processed is subjected to a first match with the standard column names in the table, and the data to be processed is subjected to a second match with the standard data in the table, and the first match and the second match each include multiple matching methods.
[0070] Among them, each matching method in the first match and the second match respectively obtains corresponding initial column names, and then a target column name is selected from the initial column names based on the verification logic.
[0071] The verification logic includes two layers. The first layer is to process the initial column names of each matching method in the first match and the second match respectively to obtain a first target column name corresponding to the first match and a second target column name corresponding to the second match. The second layer is to obtain a target column name based on the first target column name of the first match and the second target column name of the second match.
[0072] The first layer may include at least one of selecting the result corresponding to the optimal matching method based on the priority, selecting the result with the most support based on the consistency, and outputting the result of unsuccessful matching through result cross-verification for the user to confirm.
[0073] The second layer may also include at least one of selecting the result corresponding to the optimal matching method based on the priority, selecting the result with the most support based on the consistency, and outputting the result of unsuccessful matching through result cross-verification for the user to confirm.
[0074] In the above embodiments, the target column name is determined through multiple matching mechanisms and verification logic, improving the accuracy of the target column name.
[0075] In one of the optional embodiments, the to-be-processed column names and to-be-processed data are respectively matched with the standard column names and standard data in the table through different matching methods to obtain the initial column names corresponding to each matching method, including at least one of the following: matching the to-be-processed column names and to-be-processed data with the standard column names and standard data in the table respectively through a semantic matching method to obtain the first initial column names, where the semantic matching is performed based on a pre-trained language model; matching the to-be-processed column names and to-be-processed data with the standard column names and standard data in the table respectively through a syntactic matching method to obtain the second initial column names, where the syntactic matching is performed based on character distance or regular expressions; matching the to-be-processed column names and to-be-processed data with the standard column names and standard data in the table respectively through a context matching method to obtain the third initial column names. The context matching method is to determine the scenario classification corresponding to the to-be-processed column names and to-be-processed data based on the query information, screen the standard column names and standard data in the table based on the scenario assignment, and match the screened standard column names and standard data with the to-be-processed column names and to-be-processed data respectively.
[0076] Among them, in order to further improve the accuracy of matching, the present application designs a three-way matching mechanism to match the "fake column names" and "fake values" from different perspectives. This three-way matching mechanism includes:
[0077] The semantic matching method: Based on word vectors or pre-trained language models, compare the "fake column names" with the column names in the table in the semantic space to find the most similar column names.
[0078] The syntactic matching method: Use text pattern recognition technology to match the "fake column names" with the table column names at the character level to capture possible spelling mistakes or other syntactic differences.
[0079] The context matching method: Based on the context information, match the "fake column names" with the table column names in combination with the overall context of the question, so as to ensure the consistency of the matching result with the question context. First, classify the extracted "fake column names" and "fake values". In the present application, the scenario classification corresponding to the to-be-processed column names and to-be-processed data can be determined based on the query information. For example, it can be classified into time-related, location-related, event-related, etc. Then, filter in the corresponding real column names, that is, the standard column names and standard data are classified in advance or in real time based on the corresponding scenarios, and then only the standard column names and standard data corresponding to the scenario classification are obtained. Match the obtained standard column names and standard data corresponding to the scenario classification with the to-be-processed column names and to-be-processed data to ensure that the answer does not deviate from the user's question. For example, when the user asks a time-related question, it will not match the location-related column names.
[0080] In one alternative embodiment, selecting a target column name from the initial column names based on the verification logic includes at least one of the following: obtaining the priorities of different matching methods set based on scenarios, and selecting the initial column name corresponding to the target priority as the target column name; matching the initial column names corresponding to each matching method, and determining the number of matching methods with the same initial column names; selecting the initial column names corresponding to each matching method whose quantity meets the requirements as the target column names; performing cross-verification on the initial column names obtained based on the column names to be processed and the initial column names obtained based on the data to be processed, and determining the target column name based on the result of the cross-verification.
[0081] In one alternative embodiment, the method further includes: when there are at least two target column names determined based on the verification logic, outputting each target column name; receiving a selection instruction for the target column name, and determining the final target column name based on the selection instruction.
[0082] After the three-way matching is completed, the system will obtain two or several candidate column name lists according to the matching of "fake column name" and "fake value". To ensure the accuracy of the matching result, a multi-way verification mechanism is introduced in this article. The verification and determination of the final result are carried out through the following steps:
[0083] The specific ways of priority setting include: setting the priorities of different matching methods according to scenarios. For example, in some application scenarios, for example, the accuracy of semantic matching is higher, so the priority of semantic matching is the highest. Therefore, the result of semantic matching is preferentially adopted based on the priority.
[0084] The specific ways of consistency verification include: if there are conflicts among the three-way matching results, then through the consistency verification mechanism, preferentially select the column names supported in multiple matching methods as the final result. For example, the results of semantic matching and syntactic matching are consistent, and the result of context matching is inconsistent. Therefore, the results of semantic matching and syntactic matching are selected as the final result.
[0085] The specific ways of result cross-verification include: when there is a conflict between the two column names obtained by matching "fake column name" and "fake value", if the "fake value" can be 100% matched, then the column name corresponding to the "fake value" is adopted; if not, the user confirmation mechanism is adopted to further clarify the final result, and the result after user confirmation will be submitted to the large model as the final column name for question answering.
[0086] For the convenience of understanding, a specific embodiment is given. The table-based query method of the present application specifically includes the following steps:
[0087] 1. Data Preparation: Construct a training set containing a large amount of labeled data, which includes natural language questions and their corresponding table column names and column values. This data covers different types of tables and various user Q&A scenarios to ensure the generalization ability of the model.
[0088] 2. Model Training: Use a pre-trained language model based on the Transformer architecture (such as BERT or GPT) and fine-tune it on the above dataset. The model is trained to identify key phrases in the user's question and map them to possible "fake column names" and "fake values".
[0089] 3. Model Application: In actual use, when the user enters a question, the system calls the trained pattern extraction model to extract "fake column names" and "fake values" from the user's question. For example, when the user enters "What is the growth rate of sales in Q1 of 2023?", the model may extract "sales" (fake column name) and "Q1 of 2023" (fake value).
[0090] 4. Semantic Matching: Adopt a pre-trained word vector model (such as Word2Vec, GloVe) or a deep learning-based language model (such as BERT, GPT) to calculate the semantic similarity between the "fake column name" and the column names in the table. For example, if the "fake column name" is "sales", the system may find column names with similar semantics such as "revenue", "turnover", etc.
[0091] 5. Syntactic Matching: Use edit distance (such as Levenshtein distance) or regular expression matching algorithms to perform character-level comparison between the "fake column name" and the table column names. For example, when the "fake column name" is "Research Institute" and the actual column name is "Research Fellow", the system can correct spelling mistakes through syntactic matching.
[0092] 6. Context Matching: The system analyzes the context information of the user's question (such as time, location, business type, etc.) and combines it with the structure information of the table (such as table title, column name description) for matching. For example, in questions related to quarterly reports, context matching may preferentially select column names related to quarterly data.
[0093] 7. Priority Setting: According to the application scenario, the system sets different priorities for semantic matching, syntactic matching, and context matching. For example, in the financial data scenario, the system may preferentially adopt the results of semantic matching.
[0094] 8. Consistency Check: The system cross - validates the results of the three - way matching. If the results obtained by multiple matching methods are consistent, the column name and column value are directly selected as the final result. For example, if both semantic matching and context matching point to the same column name "Income", the system can directly confirm this column name.
[0095] 9. User Confirmation: When there are multiple candidate column names in the verification result or the verification is inconsistent, the system presents these candidate column names to the user and requests the user to select the most appropriate column name. For example, if the system identifies three possible column names "Income", "Revenue", and "Sales Amount", the system will request the user to confirm the column name that best matches their intention.
[0096] In the above - mentioned embodiment, a dedicated pattern extraction model is trained and used to identify and extract possible "fake column names" and "fake values" from the natural - language questions input by the user. Based on large - scale labeled data, this model can accurately extract information related to the user's intention in various contexts.
[0097] Second, three different matching methods - semantic matching, syntactic matching, and context matching - are adopted to find the real column names and column values in the table that are most similar to the "fake column names" and "fake values":
[0098] Among them, the semantic - matching method uses word vectors or pre - trained language models to compare the "fake column names" and table column names in the semantic space. The syntactic - matching method captures possible spelling mistakes or syntactic differences through character - level pattern recognition. The context - matching method combines the overall context of the question to ensure that the matching result is consistent with the context.
[0099] Third, the multi - path verification mechanism determines the finally - matched column names and column values by setting priorities and consistency checks.
[0100] Through the multi - path verification mechanism and step - by - step decomposition, the accuracy of the large model in the table question - answering task is effectively improved, especially in terms of column - name and column - value recognition, reducing the occurrence of incorrect matches and improving the accuracy of question - answering.
[0101] Through the mechanism combining automation and user confirmation, unnecessary user interactions are reduced, and while ensuring accuracy, the system's response speed and user experience are improved, enhancing the user experience.
[0102] The three - way matching mechanism enables the system to accurately identify the user's intention in various complex application scenarios, and is applicable to table question - answering tasks of different types and structures, adapting to complex scenarios.
[0103] It should be understood that although the steps in the flowcharts involved in the above embodiments are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0104] Based on the same inventive concept, an embodiment of the present application also provides a table-based query device for implementing the above-mentioned table-based query method. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the table-based query device provided below can refer to the limitations on the table-based query method in the above text, and will not be repeated here.
[0105] In an exemplary embodiment, as Figure 4 shown, a table-based query device is provided, including: a query information acquisition module 401, an extraction module 402, a matching module 403, and a query module 404, where:
[0106] The query information acquisition module 401 is used to acquire query information;
[0107] The extraction module 402 is used to perform information extraction on the query information to obtain the to-be-processed column names and to-be-processed data;
[0108] The matching module 403 is used to match the to-be-processed column names and the to-be-processed data with the standard column names and standard data in the table respectively to obtain the target column names;
[0109] The query module 404 is used to process the target column names through a large model to obtain a query result corresponding to the query information.
[0110] In one of the optional embodiments, the above extraction module 402 is specifically used to process the query information through a pre-trained pattern extraction model to obtain the to-be-processed column names and to-be-processed data; the pattern extraction model is trained based on sample natural language and the corresponding table column names and column values of the sample natural language.
[0111] In one optional embodiment, the above-mentioned matching module 403 is specifically configured to match the to-be-processed column name and the to-be-processed data with the standard column name and the standard data in the table respectively through different matching methods to obtain initial column names corresponding to each matching method; determine the verification logic; and select a target column name from the initial column names based on the verification logic.
[0112] In one optional embodiment, the above-mentioned matching module 403 is specifically configured to match the to-be-processed column name and the to-be-processed data with the standard column name and the standard data in the table respectively through at least one of the following matching methods to obtain initial column names corresponding to each matching method: matching the to-be-processed column name and the to-be-processed data with the standard column name and the standard data in the table respectively based on semantic matching to obtain a first initial column name, where the semantic matching is performed based on a pre-trained language model; matching the to-be-processed column name and the to-be-processed data with the standard column name and the standard data in the table respectively based on syntactic matching to obtain a second initial column name, where the syntactic matching is performed based on character distance or regular expressions; matching the to-be-processed column name and the to-be-processed data with the standard column name and the standard data in the table respectively based on context matching to obtain a third initial column name, and the context matching method is to determine the scenario classification corresponding to the to-be-processed column name and the to-be-processed data based on the query information, screen the standard column name and the standard data in the table based on the scenario classification, and match the screened standard column name and standard data with the to-be-processed column name and the to-be-processed data respectively.
[0113] In one optional embodiment, the above-mentioned matching module 403 is specifically configured to select a target column name from the initial column names based on the verification logic based on at least one of the following methods: obtain the priorities of different matching methods set based on the scenario, and select the initial column name corresponding to the target priority as the target column name; match the initial column names corresponding to each matching method to determine the number of matching methods with the same initial column name; select the initial column names corresponding to each matching method whose quantity meets the requirements as the target column names; perform cross-verification on the initial column name obtained based on the to-be-processed column name and the initial column name obtained based on the to-be-processed data, and determine the target column name based on the result of the cross-verification.
[0114] In one optional embodiment, the above-mentioned matching module 403 is specifically configured to output each target column name when there are at least two target column names determined based on the verification logic; receive a selection instruction for the target column name, and determine the final target column name based on the selection instruction.
[0115] Each module in the above table-based query device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0116] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 5 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store table data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a table-based query method.
[0117] Those skilled in the art can understand that Figure 5 the structure shown in
[0118] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0119] In an embodiment, a computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps in each of the above method embodiments.
[0120] In an embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, it implements the steps in each of the above method embodiments.
[0121] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0122] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in this application can include at least one of a relational database and a non-relational database. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0123] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.
[0124] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.
Claims
1. A table-based query method, characterized in that, The method includes: Obtain query information; Perform information extraction on the query information to obtain the column names to be processed and the data to be processed; Match the column names to be processed and the data to be processed with the standard column names and standard data in the table respectively to obtain the target column names; Process the target column names through a large model to obtain a query result corresponding to the query information.
2. The method according to claim 1, characterized in that The performing information extraction on the query information to obtain the column names to be processed and the data to be processed includes: Process the query information through a pre-trained pattern extraction model to obtain the column names to be processed and the data to be processed; the pattern extraction model is trained based on sample natural language and the corresponding table column names and column values of the sample natural language.
3. The method according to claim 1, characterized in that, The matching the column names to be processed and the data to be processed with the standard column names and standard data in the table respectively to obtain the target column names includes: Match the column names to be processed and the data to be processed with the standard column names and standard data in the table respectively through different matching methods to obtain the initial column names corresponding to each matching method; Determine the verification logic; Select the target column names from the initial column names based on the verification logic.
4. The method according to claim 3, characterized in that The matching the column names to be processed and the data to be processed with the standard column names and standard data in the table respectively through different matching methods to obtain the initial column names corresponding to each matching method includes at least one of the following: Match the column names to be processed and the data to be processed with the standard column names and standard data in the table respectively based on semantic matching to obtain the first initial column names, where the semantic matching is performed based on a pre-trained language model; Match the column names to be processed and the data to be processed with the standard column names and standard data in the table respectively based on syntactic matching to obtain the second initial column names, where the syntactic matching is performed based on character distance or regular expressions; Match the column names to be processed and the data to be processed with the standard column names and standard data in the table respectively based on context matching to obtain the third initial column names. The context matching method determines the scenario classification corresponding to the column names to be processed and the data to be processed based on the query information, filters the standard column names and standard data in the table based on the scenario assignment, and matches the filtered standard column names and standard data with the column names to be processed and the data to be processed respectively.
5. The method according to claim 3, wherein The selecting the target column names from the initial column names based on the verification logic includes at least one of the following: Obtain the priorities of different matching methods set based on scenarios, and select the initial column names corresponding to the target priority as the target column names; Match the initial column names corresponding to each matching method, and determine the number of matching methods with the same initial column names; select the initial column names corresponding to each matching method that meet the requirements as the target column names; Perform cross-verification on the initial column names obtained based on the column names to be processed and the initial column names obtained based on the data to be processed, and determine the target column names based on the cross-verification results.
6. The method according to claim 3, characterized in that, The method further includes: When there are at least two determined target column names based on the verification logic, output each target column name; Receive a selection instruction for the target column name and determine the final target column name based on the selection instruction.
7. A table-based query device, characterized in that, The device includes: A query information acquisition module, configured to acquire query information; An extraction module, configured to perform information extraction on the query information to obtain to-be-processed column names and to-be-processed data; A matching module, configured to match the to-be-processed column names and the to-be-processed data with standard column names and standard data in a table respectively to obtain target column names; A query module, configured to process the target column names through a large model to obtain a query result corresponding to the query information.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method, system and device for converting natural language query into SQL and storage medium
CN114547072A
User intention recognition method and device, electronic equipment and medium
CN115860012A
Method and device for determining target chart, electronic equipment and program product
CN119149714A
Structured query statement generation method and device, electronic equipment and storage medium
CN119576965A
Data processing method and apparatus, and device and medium
WO2024179581A1