Natural language driven table search method and device based on differentiable search index
By using differentiable search index and dual-view hierarchical clustering algorithm in table search and combining large language models to generate synthetic queries, the problems of insufficient error accumulation and interaction in existing table search methods are solved, and more accurate and efficient table search is achieved.
Patent Information
- Application Number
- CN202510400669.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-04-01
AI Technical Summary
Existing table search methods cannot be optimized end-to-end, error accumulation affects the accuracy of search results, and the interaction between the query and the table is insufficient, so semantic relationships cannot be fully captured.
A natural language-driven table search method based on differentiable search index is adopted, and a unique identifier is assigned to each table through a dual-view hierarchical clustering algorithm, and a synthetic query is created through a query generator based on a large language model. The encoder-decoder language model is trained to achieve deep interaction between the query and the table.
It effectively solves the problems of error accumulation and insufficient interaction, improves the accuracy and efficiency of table search, meets user query needs and improves user experience.
Smart Images

Figure CN119917646A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data discovery, and in particular, relates to a natural language driven table search method and device based on a differentiable search index. Background Art
[0002] In recent years, with the continuous growth of data volume, tabular data has been increasingly used in various institutions, enterprises and networks. These tabular data contain a large amount of information and can support decision-making. However, due to the large size of tabular data, it becomes difficult for users to locate relevant tables in large table repositories or data lakes. Existing table search methods are mainly based on keywords, basic tables and natural language queries, among which natural language queries have attracted much attention due to their user-friendliness and accuracy.
[0003] Natural language driven table search provides an intuitive way for non-technical users to express their needs, allowing them to obtain the required information more accurately. This approach improves the efficiency and accuracy of information retrieval by understanding the user's natural language query and automatically finding the most relevant table in the table repository.
[0004] In the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art:
[0005] First, existing table search methods cannot be optimized end-to-end, and the accumulated errors in the representation, indexing, and search stages will affect the final search results. Existing methods usually adopt the traditional dense vector search process, that is, first encode the table into an embedding vector of fixed dimension, then build the index, and finally search. In this process, the errors in each stage will accumulate to the next stage, resulting in a decrease in the accuracy of the final search results. Secondly, the interaction between the query and the table is insufficient, resulting in the inability to fully capture the semantic relationship between the query and the table. Existing methods usually encode the query and the table into independent embedding vectors respectively, and then search by calculating the similarity between the two. However, this method ignores the deep semantic interaction between the query and the table, fails to fully utilize the contextual information between the query and the table, and cannot accurately understand the intent of the query when processing complex natural language queries. Summary of the invention
[0006] In view of the deficiencies in the prior art, the purpose of the embodiments of the present application is to provide a natural language driven table search method and device based on a differentiable search index, which achieves more accurate and efficient table search by unifying the index and search process.
[0007] According to a first aspect of an embodiment of the present application, a natural language driven table search method based on a differentiable search index is provided, comprising:
[0008] Acquire a table repository, wherein each table in the table repository includes a table title, attributes, and cell values;
[0009] Based on the title, attributes, and cell values of each table, a text encoder is used to extract the semantic features of the table, and the embedding vectors corresponding to the metadata view and instance data view of each table are obtained.
[0010] According to the embedding vector of each table, a unique identifier is assigned to each table through a two-view hierarchical clustering algorithm;
[0011] Generate a number of synthetic queries for each table by a query generator, wherein the query generator is built based on a large language model;
[0012] The synthetic query and the corresponding table and its identifier are constructed as training samples to train an encoder-decoder language model;
[0013] Get the query given by the user, use the trained encoder-decoder language model to generate table identifiers, and thus get the table related to the query.
[0014] Furthermore, for the table The metadata view corresponds to a text sequence consisting of the table title and attributes. }, the instance data view corresponds to the text sequence of table attributes and cell values interlaced and concatenated ,in is the title of the table, is the attribute of the table, is the cell value of the table, Represents the cell value of the k-th row and m-th column in the i-th table. is the number of attributes in the table, that is, the total number of columns in the table cells, is the total number of rows of cells in the table; the metadata view and the instance data view are encoded into corresponding embedding vectors respectively using a text encoder.
[0015] Furthermore, according to the embedding vector of each table, a prefix-aware unique identifier is assigned to each table through a dual-view hierarchical clustering algorithm, including:
[0016] Based on the embedding vectors corresponding to the metadata view and the instance data view, respectively, a clustering tree is constructed by a dual-view hierarchical clustering algorithm;
[0017] Each table is assigned a unique identifier based on the clustering tree.
[0018] Furthermore, the dual-view hierarchical clustering algorithm is as follows: clustering is performed using the embedding vector corresponding to the metadata view until the current level of the number of clusters reaches a predetermined maximum depth of the first view; clustering is continued using the embedding vector corresponding to the instance data view; wherein the root node contains all tables, and for each child node obtained by clustering, the condition for continuing to cluster downward is that the number of tables contained must be greater than the predetermined minimum size of the leaf node, otherwise, for this child node, downward clustering is stopped.
[0019] Furthermore, the identifier corresponding to each table includes a path from the root node to the leaf node where the table is located and a number of the table in the leaf node.
[0020] Furthermore, for the table , its training samples are divided into two categories, the first type of training samples takes the synthetic query corresponding to the table as input and the corresponding identifier as output; the second type of training samples takes the serialized table is the input, and the corresponding identifier is the output, where is the title of the table, is the attribute of the table, is the cell value of the table, Represents the cell value of the k-th row and m-th column in the i-th table. is the number of attributes in the table, that is, the total number of columns in the table cells, The total number of rows of cells in the table.
[0021] According to a second aspect of an embodiment of the present application, a natural language driven table search device based on a differentiable search index is provided, comprising:
[0022] An acquisition module, used for acquiring a table repository, wherein each table in the table repository includes a table title, attributes, and cell values;
[0023] An extraction module is used to extract the semantic features of each table based on the title, attributes, and cell values of each table using a text encoder to obtain the embedding vectors corresponding to the metadata view and instance data view of each table;
[0024] an assignment module, for assigning a unique identifier to each table through a two-view hierarchical clustering algorithm according to the embedding vector of each table;
[0025] A generation module, configured to generate a plurality of synthetic queries for each table through a query generator, wherein the query generator is constructed based on a large language model;
[0026] A training module, configured to construct the synthetic query and the corresponding table and its identifier as a training sample, and train an encoder-decoder language model;
[0027] The search module is used to obtain a query given by a user and generate a table identifier using the trained encoder-decoder language model to obtain a table related to the query.
[0028] According to a third aspect of an embodiment of the present application, a computer program product is provided, comprising a computer program / instruction, which implements the method described in the first aspect when executed by a processor.
[0029] According to a fourth aspect of an embodiment of the present application, there is provided an electronic device, including:
[0030] one or more processors;
[0031] A memory for storing one or more programs;
[0032] When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in the first aspect.
[0033] According to a fifth aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method described in the first aspect are implemented.
[0034] The technical solution provided by the embodiments of the present application may have the following beneficial effects:
[0035] It can be seen from the above embodiments that the present application realizes natural language driven table search by constructing a differentiable search index. The method assigns a prefix-aware unique identifier to each table through a dual-view hierarchical clustering algorithm, and creates a synthetic query through a query generator based on a large language model, and then stores the mapping between the query and the table and its identifier in the encoder-decoder language model, thereby realizing deep interaction between the query and the table. This method effectively solves the problems of error accumulation and insufficient query-table interaction in existing table search methods, improves the accuracy and efficiency of table search, and improves the user experience while meeting user query needs.
[0036] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0038] Figure 1 The invention is a flowchart showing a natural language driven table search method based on a differentiable search index according to an exemplary embodiment.
[0039] Figure 2 The invention is a block diagram showing a natural language driven table search device based on a differentiable search index according to an exemplary embodiment.
[0040] Figure 3 The diagram is a schematic diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0041] Here, exemplary embodiments are described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application.
[0042] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms of "a", "said" and "the" used in this application and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0043] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0044] Figure 1 FIG. 1 is a flowchart of a natural language driven table search method based on a differentiable search index according to an exemplary embodiment. Figure 1 As shown, the method may include the following steps:
[0045] S11: Acquire a table repository, wherein each table in the table repository includes a table title, attributes, and cell values;
[0046] S12: Based on the title, attributes, and cell values of each table, a text encoder is used to extract the semantic features of the table, and the embedding vectors corresponding to the metadata view and instance data view of each table are obtained;
[0047] S13: According to the embedding vector of each table, a unique identifier is assigned to each table through a two-view hierarchical clustering algorithm;
[0048] S14: generating a plurality of synthetic queries for each table by a query generator, wherein the query generator is constructed based on a large language model;
[0049] S15: constructing the synthetic query and the corresponding table and its identifier as a training sample, and training the encoder-decoder language model;
[0050] S16: Obtain a query given by the user, and use the trained encoder-decoder language model to generate a table identifier, thereby obtaining a table related to the query.
[0051] It can be seen from the above embodiments that the present application provides a natural language driven table search method based on a differentiable search index. The method converts table information into an embedded vector using a text encoder model to extract the semantic features of the table. Then, a prefix-aware unique identifier is assigned to each table through a dual-view hierarchical clustering algorithm. In addition, a synthetic query is created for each table by a query generator based on a large language model. The mapping between the synthetic query and the table and its corresponding identifier is stored in the parameters of the encoder-decoder language model through training to achieve deep interaction between the query and the table. Finally, the trained model can directly generate a table identifier based on a user-given query to obtain a table related to the query. This method effectively solves the problems of error accumulation and insufficient query-table interaction in existing table search methods, improves the accuracy and efficiency of table search, and improves the user experience while meeting user query needs.
[0052] In a specific implementation of S11, a table repository is obtained, wherein each table in the table repository includes a table title, attributes, and cell values;
[0053] Specifically, the present invention can be applied to multiple practical fields such as enterprise data management, institutional data openness, and enterprise data lake table retrieval engine. Without loss of generality, the present invention represents the table repository as , where each table Contains structured data such as headers, attributes, and cell values. Indicates the number of tables in the table repository. For example, in enterprise data management, the table repository It can include employee information table, sales record table and other tables. Each table It can be described by its title (such as "employee information", "sales record", etc.), attributes (such as "name", "age", "sales", etc.) and cell values (such as specific employee name, age, sales, etc.). These tabular data can be used to support corporate decision-making, data analysis and other tasks.
[0054] In a specific implementation of S12, based on the title, attributes, and cell values of each table, a text encoder is used to extract semantic features of the table, and embedding vectors corresponding to the metadata view and the instance data view of each table are obtained;
[0055] Specifically, each table Include Title ,property and cell value ,in is the number of attributes in the table (total number of columns), The total number of rows of cells in the table, for example: Represents the cell value of the first row and first column of the i-th table. Using the pre-trained text encoder, the two perspectives of each table, namely the metadata view representation and the instance data view representation, are encoded, where the metadata view is represented as a table A text sequence consisting of the title and attributes of }, the instance data view is represented as a table A text sequence consisting of attributes and cell values interlaced and concatenated The metadata representation and instance data representation of the table are input into the text encoder to obtain the corresponding embedding vector These embedding vectors capture the semantic features of the table. The metadata view representation provides high-level semantic information of the table, helping the model understand the subject and structure of the table, while the instance data view representation provides fine-grained information of the table, reflecting the actual data content and distribution in the table. Both views serve to generate more accurate table identifiers later, helping the model understand the table at both the macro and micro levels.
[0056] In a specific implementation of S13, a unique identifier is assigned to each table through a dual-view hierarchical clustering algorithm according to the embedding vector of each table;
[0057] Specifically, Figure 2 As shown, this step may include the following sub-steps:
[0058] S31: constructing a clustering tree by a dual-view hierarchical clustering algorithm based on the embedding vectors corresponding to the metadata view and the instance data view respectively;
[0059] The dual-view hierarchical clustering algorithm is as follows: clustering is performed using the embedding vector corresponding to the metadata view until the current level of the number of clusters reaches the predetermined maximum depth of the first view; clustering is continued using the embedding vector corresponding to the instance data view; wherein the root node contains all tables, and for each child node obtained by clustering, the condition for continuing to cluster downward is that the number of tables contained must be greater than the predetermined minimum size of the leaf node, otherwise, for this child node, downward clustering is stopped. Using the metadata view first and then the instance data view ensures that each generated table identifier has coarse-to-fine semantics.
[0060] Specifically, for the process of building a clustering tree, the input of the algorithm is a set of tables The metadata view embedding vector collection and instance data view embedding vector set , the number of clusters per layer K, the minimum size of leaf nodes L and the maximum depth of the first view d; the output is the root node of the clustering tree ; The steps of the algorithm are:
[0061] Initialize the current level to 0 and set the root node The view is set to the metadata view embedding vector collection ;
[0062] Cluster all tables. Similar tables calculated based on metadata view embedding vectors are clustered into a cluster. New child nodes are created in the clustering tree based on the clusters. , while the level of the child nodes is +1, each child node contains a set of tables belonging to the same cluster and a corresponding set of embedding vectors;
[0063] Repeat the above clustering process for each child node until the current level reaches d, then switch the view and use the instance data view to embed the vector replace The vectors are clustered, that is, the nodes in the dth layer and beyond are clustered based on the instance data view embedding vectors;
[0064] For each child node If the number of tables it contains is less than the predetermined minimum size L of leaf nodes, clustering is stopped, otherwise clustering continues until the number of tables contained in each node is less than L.
[0065] S32: assigning a unique identifier to each table according to the clustering tree;
[0066] Specifically, the process of assigning a unique identifier to each table according to the above-obtained clustering tree is as follows: except for the root node of the clustering tree, starting from the second layer, each node is assigned a unique number according to its level and position in the clustering tree. For example, each child node (cluster) under the root node is assigned a number ranging from 0 to K-1 (K is the number of clusters at this level). Ultimately, the nodes of each layer will have a number ranging from 0 to K-1. For each table in the leaf node, according to its position in the leaf node, a number ranging from 0 to L-1 (L is the maximum number of tables in the leaf node) is assigned.
[0067] Finally, the identifier of each table is a sequence, including the path from the root node to the leaf node where the table is located and the number of the table in the leaf node, for example ,in represents the node number in the jth layer, and Indicates the number of the table in the leaf node, ranging from 0 to L-1. For example, if the number of levels of the clustering tree is 3, one of the tables may be assigned a table identifier of 00120301, where 00 represents the node number of the first layer, 12 represents the node number of the second layer, 03 represents the number of the third layer, i.e. the leaf node at this layer, and the last 01 is the number of the table in the leaf node.
[0068] The dual-view hierarchical clustering algorithm is used to assign a unique identifier to each table. The main reason is that the table data has a clear structure, including titles, attributes, and cell values. The identifier obtained by the dual-view hierarchical clustering algorithm has coarse-to-fine semantics, better captures the structural characteristics of the table, conforms to the characteristics of the model autoregressive decoding during inference, and helps improve accuracy. For example, for a table about the sales record of XX product, the table identifier assigned is 00120301, since the dual-view hierarchical clustering algorithm is used to assign IDs, tables with table identifier prefix 00 are all related to sales, tables with prefix 0012 are all related to XX product sales, and so on. The semantic information they represent is from coarse to fine.
[0069] In a specific implementation of S14, a plurality of synthetic queries are generated for each table by a query generator, wherein the query generator is constructed based on a large language model;
[0070] Specifically, a large language model is fine-tuned using training data containing table-related tasks using a low-rank adaptation (LoRA) technique, wherein the table-related tasks include table size identification, table cell extraction, table row / column extraction, table cell retrieval, table-to-text generation, table fact verification, and table question-answering, so as to enhance the model's ability to understand tabular data; the fine-tuned large language model is used as a query generator for tabular data, and the query generator is capable of generating natural language queries related to the table content; the number of synthetic queries required to be generated for each table is determined, denoted as N, sub-tables are sampled in the table, each sub-table contains part of the rows and columns, the above-mentioned query generator is called for the sub-table, and a set of M synthetic queries is generated, where M is much smaller than N, and is generally 3; the generated synthetic queries are quality checked, and after removing queries with format errors or duplicates of existing queries, the number of remaining synthetic queries is greater than or equal to 0 and less than or equal to N; the synthetic queries that pass the quality check are added to the query set of the table, and the above steps are repeated to iteratively sample tables and generate queries until the number of queries in the query set reaches N; the above process is repeated for all tables to generate a synthetic query set for each table. For example, for a table named "2024 XX Product Sales Records", the synthetic queries that may be generated include "How many pieces of XX product were sold in January 2024?", "What was the sales volume of XX product in March 2024?", and so on.
[0071] In a specific implementation of S15, the synthetic query and the corresponding table and its identifier are constructed as training samples to train the encoder-decoder language model;
[0072] Specifically, a pre-trained encoder-decoder language model is selected and initialized using its pre-trained parameters to fully utilize its pre-trained knowledge on large-scale text. For each table, its training samples are divided into two categories: (i) each synthetic query is taken as input and the corresponding identifier is taken as output to construct the first category of training samples; (ii) the serialized table is As input, the corresponding identifier is used as output to construct the second type of training samples. Finally, there are N+1 training samples for each table, and the model is finally trained through sequence-to-sequence training, that is, the input of the model is a text sequence such as a query or table, and the output is a table identifier, which is also a text sequence. In this way, the model learns the mapping relationship between the query and the table, and realizes the deep interaction between the query and the table.
[0073] This step unifies the indexing and search processes into a differentiable model parameter, and memorizes the mapping between the table and related synthetic queries and the table identifier (i.e., index) in the model parameter through training; in the subsequent search query process, the user query is input into the trained model to directly output the table identifier and locate the output of the relevant table.
[0074] In a specific implementation of S16, a query given by a user is obtained, and a table identifier is generated using a trained encoder-decoder language model, thereby obtaining a table related to the query.
[0075] Specifically, the natural language query entered by the user is input into the trained encoder-decoder language model. Based on the semantic information of the query, the model directly autoregressively generates a table identifier related to the query, and directly finds the table related to the query from the table repository based on the generated table identifier. For example, the query entered by the user is: "What is the sales performance of XX product in February 2024?" The model will directly return the table identifier 01041103 based on the user input, and directly find the table based on the 01041103 identifier.
[0076] Corresponding to the aforementioned embodiment of the natural language driven table search method based on a differentiable search index, the present application also provides an embodiment of a natural language driven table search device based on a differentiable search index.
[0077] Figure 2 is a block diagram of a natural language driven table search device based on a differentiable search index according to an exemplary embodiment. Figure 2 , the device may include:
[0078] An acquisition module 21 is used to acquire a table repository, wherein each table in the table repository includes a table title, attributes, and cell values;
[0079] An extraction module 22, configured to extract semantic features of each table using a text encoder based on the title, attributes, and cell values of each table, and obtain embedding vectors corresponding to the metadata view and the instance data view of each table;
[0080] an assignment module 23, for assigning a unique identifier to each table through a dual-view hierarchical clustering algorithm according to the embedding vector of each table;
[0081] A generating module 24, configured to generate a plurality of synthetic queries for each table by using a query generator, wherein the query generator is constructed based on a large language model;
[0082] A training module 25, configured to construct the synthetic query and the corresponding table and its identifier as a training sample to train an encoder-decoder language model;
[0083] The search module 26 is used to obtain a query given by a user and generate a table identifier using the trained encoder-decoder language model, thereby obtaining a table related to the query.
[0084] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0085] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present application scheme. A person of ordinary skill in the art can understand and implement it without paying any creative work.
[0086] Accordingly, the present application also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the natural language driven table search method based on a differentiable search index as described above.
[0087] Accordingly, the present application also provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned natural language driven table search method based on differentiable search index. Figure 3 As shown in FIG. 1 , a hardware structure diagram of a natural language driven table search device based on a differentiable search index provided by an embodiment of the present invention is provided for any device with data processing capability, except Figure 3 In addition to the processor, memory and network interface shown, any device with data processing capability in which the apparatus in the embodiment is located may also include other hardware, generally based on the actual functions of the device with data processing capability, which will not be described in detail.
[0088] Accordingly, the present application also provides a computer-readable storage medium on which computer instructions are stored, and when the instructions are executed by the processor, the natural language-driven table search method based on the differentiable search index as described above is implemented. The computer-readable storage medium can be an internal storage unit of any device with data processing capabilities described in any of the aforementioned embodiments, such as a hard disk or a memory. The computer-readable storage medium can also be an external storage device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), an SD card, a flash card (Flash Card), etc. equipped on the device. Furthermore, the computer-readable storage medium can also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and can also be used to temporarily store data that has been output or is to be output.
[0089] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the contents disclosed herein. The present application is intended to cover any variations, uses or adaptations of the present application, which follow the general principles of the present application and include common knowledge or customary technical means in the art that are not disclosed in the present application.
Claims
1. A natural language driven table search method based on a differentiable search index, characterized in that: include: Acquire a table repository, wherein each table in the table repository includes a table title, attributes, and cell values; Based on the title, attributes, and cell values of each table, a text encoder is used to extract the semantic features of the table, and the embedding vectors corresponding to the metadata view and instance data view of each table are obtained. According to the embedding vector of each table, a unique identifier is assigned to each table through a two-view hierarchical clustering algorithm; Generate a number of synthetic queries for each table by a query generator, wherein the query generator is built based on a large language model; The synthetic query and the corresponding table and its identifier are constructed as training samples to train an encoder-decoder language model; Get the query given by the user, use the trained encoder-decoder language model to generate table identifiers, and thus get the table related to the query.
2. The method according to claim 1, characterized in that For the table The embedding vector corresponding to the metadata view is a text sequence concatenated by the table title and attributes. The embedding vector corresponding to the instance data view is a text sequence composed of table attributes and cell values interlaced and concatenated ,in is the title of the table, is the attribute of the table, is the cell value of the table, Represents the cell value of the k-th row and m-th column in the i-th table. is the number of attributes in the table, that is, the total number of columns in the table cells, is the total number of rows of cells in the table; A text encoder is used to encode the metadata view and instance data view into corresponding embedding vectors respectively.
3. The method according to claim 1, characterized in that: Based on the embedding vector of each table, a prefix-aware unique identifier is assigned to each table through a two-view hierarchical clustering algorithm, including: Based on the embedding vectors corresponding to the metadata view and the instance data view, respectively, a clustering tree is constructed by a dual-view hierarchical clustering algorithm; Each table is assigned a unique identifier based on the clustering tree.
4. The method according to claim 3, characterized in that The dual-view hierarchical clustering algorithm is: clustering using the embedding vectors corresponding to the metadata views until the current level of the number of clusters reaches a predetermined maximum depth of the first view; Clustering is continued using the embedding vector corresponding to the instance data view, where the root node contains all the tables. For each child node obtained by clustering, the condition for continuing to cluster downward is that the number of tables contained must be greater than the predetermined minimum size of the leaf node, otherwise, for this child node, downward clustering is stopped.
5. The method according to claim 1, characterized in that The identifier corresponding to each table includes a path from the root node to the leaf node where the table is located and the number of the table in the leaf node.
6. The method according to claim 1, characterized in that For the table ,The training samples are divided into two categories, where the first training sample takes the synthetic query corresponding to the table as input and the corresponding identifier as output; The second training sample is in the form of serialization is the input, and the corresponding identifier is the output, where is the title of the table, is the attribute of the table, is the cell value of the table, Represents the cell value of the k-th row and m-th column in the i-th table. is the number of attributes in the table, that is, the total number of columns in the table cells, The total number of rows of cells in the table.
7. A natural language driven table search device based on a differentiable search index, characterized in that: include: An acquisition module, used for acquiring a table repository, wherein each table in the table repository includes a table title, attributes, and cell values; An extraction module is used to extract the semantic features of each table based on the title, attributes, and cell values of each table using a text encoder to obtain the embedding vectors corresponding to the metadata view and instance data view of each table; an assignment module, for assigning a unique identifier to each table through a two-view hierarchical clustering algorithm according to the embedding vector of each table; A generation module, configured to generate a plurality of synthetic queries for each table through a query generator, wherein the query generator is constructed based on a large language model; A training module, configured to construct the synthetic query and the corresponding table and its identifier as a training sample, and train an encoder-decoder language model; The search module is used to obtain a query given by a user and generate a table identifier using the trained encoder-decoder language model to obtain a table related to the query.
8. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
9. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instruction is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Video search method and system
CN105653700A
Real-time recommendation method based on intem2vec and vector clustering
CN114610960A
Large-model-assisted self-lifting multi-modal industrial equipment knowledge graph construction method
CN119577159A
Creating cognitive intelligence queries from multiple data corpuses
US20180267976A1