A method for encrypted index query of tabular data

CN122884979APending Publication Date: 2026-10-09BEIJING SHUI XIAOYI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611358343.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-09-03
Publication Date
2026-10-09

AI Technical Summary

Technical Problem

[0006]本发明的目的在于提供一种表格数据密态索引查询方法,用于解决现有密态表格查询方案中存在的异构列类型适配能力不足、多列组合查询效率较低、动态更新成本较高、查询过程泄露风险较大以及查询结果完整性难以验证的问题

Benefits of technology

1、本发明通过为不同可查询列生成密态列标识,并针对数值列、日期列、枚举列和文本列分别构建行组级密态摘要索引,使密态数据服务端能够在不直接获知明文列名、明文列值和明文查询条件的情况下,对表格数据进行候选过滤,从而提高密态表格查询场景下对异构列类型的适配能力;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122884979A_ABST
    Figure CN122884979A_ABST
Patent Text Reader

Abstract

The application discloses a table data encrypted state index query method, and belongs to the technical field of table data query. The method obtains a table to be protected and analyzes table structure information, generates an encrypted state column identifier and a column-level index key; divides the table into a plurality of row groups, constructs row group-level encrypted state abstract index for numerical columns, date columns, enumeration columns and text columns, and unifies encrypted state hit results into a candidate row token set or a candidate bit map through row-level encrypted state position index; and stores the encrypted table data in association with the encrypted state index and integrity commitment information. When querying, an encrypted state query token is generated, the server performs row group-level filtering and row-level filtering, and returns ciphertext data and integrity proof after combined operation according to a query predicate graph. The method can improve the encrypted state query efficiency of multiple column combinations, reduce the information leakage risk and dynamic update cost, and support query result integrity verification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of tabular data query technology, and specifically to a method for querying tabular data using dense indexes. Background Technology

[0002] As the scale of structured data in business systems, data warehouses, and spreadsheets continues to grow, users typically need to query and filter table data by multiple fields such as department, date, amount, status, and keywords. For ordinary plaintext tables, query efficiency can be improved through range indexes, inverted indexes, bitmap indexes, or database query optimization mechanisms. However, in scenarios where table data is encrypted and stored on servers, cloud platforms, or third-party data servers, the server cannot directly obtain plaintext column names, plaintext column values, and plaintext query conditions, making traditional plaintext indexing and query methods difficult to apply directly.

[0003] Existing encrypted table query solutions typically build indexes for a single data type or a single query method. For example, range encoding or bucketing is often used for numeric or date columns; equality labels or inverted indexes are often used for enumeration columns; and keyword labels or Bloom filters are often used for text columns. While these methods can achieve encrypted queries to some extent, when the same table contains numeric, date, enumeration, and text columns, and multi-column combined queries are required, the candidate result formats of different index types are inconsistent. The server often needs to expand the scan range or return more candidate data, which is then decrypted and filtered by the client, resulting in low query efficiency and a large amount of encrypted data transmission.

[0004] Furthermore, in scenarios where table data is frequently added, deleted, or modified, existing solutions are prone to problems such as excessively large index update ranges, complex version maintenance, or the need to rebuild large-scale indexes, making it difficult to adapt to the usage requirements of dynamic table data. At the same time, information such as query access patterns, candidate result size, and number of returned results may still be exposed during the encrypted query process. If the server fails to return, tampers with, or uses an old version of the index to return query results, the client will also find it difficult to verify the completeness and consistency of the query results in a timely manner.

[0005] Therefore, it is necessary to provide a table data dense-state index query method that can adapt to the query requirements of different types of table columns, unify the dense-state hit results of different column types into a composable candidate row representation, improve the efficiency of multi-column combination queries, reduce dynamic update overhead, reduce the risk of leakage during the query process, and support the integrity verification of query results. Summary of the Invention

[0006] The purpose of this invention is to provide a dense index query method for tabular data, which solves the problems of insufficient heterogeneous column type adaptation capability, low efficiency of multi-column combination query, high dynamic update cost, high risk of leakage during the query process, and difficulty in verifying the integrity of query results in existing dense table query schemes.

[0007] To address the aforementioned technical problems, this invention provides a method for querying dense indexes in tabular data, comprising the following steps: Obtain the table to be protected and parse it to obtain the table structure information including the set of column queryable operators. The table to be protected includes multiple data records. Columns that are not empty in the set of column queryable operators are identified as queryable columns. A secret column identifier is generated and a column-level index key is derived. The table to be protected is divided into row groups, and a secret row group identifier and row token are generated. A row group-level secret summary index containing secret index labels is generated according to the data type of the queryable columns. The secret index labels are mapped to the row-level secret position index, so that the hit results of each queryable column are represented as a set of row tokens or a candidate bitmap within the same row group. The data records are encrypted to form ciphertext data, and integrity commitment information is generated from the row group-level ciphertext digest index, the row-level ciphertext location index, and the ciphertext data and then stored in the ciphertext data server. Generate a query predicate graph and generate a dense query token containing dense column identifiers, row group-level query labels, row-level query labels, predicate graph representation, and leakage control parameters; Candidate row groups are obtained based on row group-level query tags, and comprehensive candidate row groups are obtained by combining them according to the query predicate graph. Within the comprehensive candidate row groups, row-level dense position indexes are queried based on row-level query tags, and a combination operation is performed on the row token set or candidate bitmap to obtain the candidate row identifier set. Based on the leakage control parameters, encapsulate the candidate row identifier set, return the corresponding ciphertext data and the integrity certificate generated based on the integrity commitment information, verify the integrity certificate, decrypt the ciphertext data and perform residual verification, and then output the query results.

[0008] Furthermore, the table structure information includes table identifier, column identifier, column data type, row identifier, and column queryable operator set; when generating the encrypted column identifier, the table identifier, column identifier, column data type, column security level, and index version parameter are used as input; when deriving the column-level index key, the table identifier, encrypted column identifier, index version parameter, and column security level are used as input, and subkeys for row group-level encrypted digest index, row-level encrypted position index, row token, pseudo-candidate, and integrity commitment are derived.

[0009] Furthermore, the row group identifier is processed to obtain a dense row group identifier, and a row token is generated based on the table identifier, the dense row group identifier, the row identifier, and the index version parameter. The row-level dense position index stores the set of row tokens or candidate bitmaps associated with the dense index label under the corresponding dense row group identifier. The candidate bitmap represents the candidate row according to the pseudo-random row position order within the row group.

[0010] Furthermore, when generating the row group-level dense state summary index, the cell values ​​of the numeric and date columns are mapped to bucket numbers according to the bucketing rules, and the row group-level bucket summary labels are generated from the bucket numbers; the enumeration values ​​of the enumeration column are standardized to generate equivalent dense state labels; and the text cells of the text column are standardized and fragmented to generate blinded text labels.

[0011] Furthermore, when encrypting data records, encrypted data is formed in units of rows, row groups, or data blocks, and the table identifier, encrypted row group identifier, row token, and index version parameter are used as associated data; the integrity commitment information is generated by the index tree, data tree, and deletion mark tree when there are deleted records.

[0012] Furthermore, when generating the query predicate graph, the query columns, query operators, query values, and logical relationships are identified from the table query request. Column predicates are set as predicate nodes, and intersections, unions, or differences are set as logical operation relationships. When generating dense query tokens, range queries are converted into bucket query tags, enumerated equality queries are converted into equality query tags, and text fragment queries are converted into text fragment query tags.

[0013] Furthermore, when generating the candidate row identifier set, the corresponding dense index partition is first selected based on the dense column identifier, and the candidate row group corresponding to the same query predicate is recorded as the predicate candidate result; when the query predicate graph contains multiple query predicates, the candidate results of each predicate are merged according to the combination relationship indicated by the query predicate graph to obtain a comprehensive candidate row group, and then the row-level dense position index is queried only within the comprehensive candidate row group to generate the candidate row identifier set.

[0014] Furthermore, the leakage control parameters are used to determine the number of pseudo-query tags, the number of pseudo-candidates, and the fixed return level; when the size of the candidate row identifier set is less than the fixed return level, the encrypted data server fills in the pseudo-candidates from the committed pseudo-candidate space and returns them together; after integrity verification and decryption, the client or trusted proxy identifies and removes records marked as fill type.

[0015] Furthermore, the integrity proof includes an existence or non-existence proof corresponding to the row group-level secret state digest index, an existence or non-existence proof corresponding to the row-level secret state position index, a data tree proof corresponding to the returned ciphertext data, and a deletion tag tree proof corresponding to the deleted record. These are used to verify that the row group candidate bitmap, the row-level candidate bitmap or row token set, the ciphertext data, the deletion tag, and the index version belong to the same integrity commitment information.

[0016] Furthermore, when a table to be protected is updated, deleted, or modified, the affected row groups and affected queryable columns are identified; newly added records are written to the incremental index segment of the corresponding row group, deleted records are written to the deletion flag, and modified records are updated only for the changed queryable columns; during queries, the effective baseline index segment, incremental index segment, and deletion flag are processed, and the corresponding proof is returned.

[0017] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are: 1. This invention generates encrypted column identifiers for different queryable columns and constructs row group-level encrypted summary indexes for numerical columns, date columns, enumeration columns and text columns respectively. This enables the encrypted data server to perform candidate filtering on table data without directly knowing the plaintext column names, plaintext column values ​​and plaintext query conditions, thereby improving the adaptability of encrypted table query scenarios to heterogeneous column types. 2. This invention maps the ciphertext hit results corresponding to different column types to the row-level ciphertext position index in a unified manner, so that numerical range conditions, date range conditions, enumerated equality conditions and text fragment conditions can all be converted into candidate row token sets or candidate bitmaps, and supports combination operations such as intersection, union or difference, thereby reducing the scanning range and the size of intermediate candidate sets when querying multiple columns, and reducing the amount of ciphertext data returned. 3. This invention limits the observable candidate size and return result size during the query process by using leakage control methods such as fixed return tiers, pseudo query tags, pseudo candidate row groups, or pseudo candidate row tokens. In addition to the preset leakage profile, it makes it difficult for the server to directly determine the actual number of hit rows, the actual number of query tags, and the actual result size, thereby helping to reduce the risk of information leakage during the cryptic query process. 4. This invention generates integrity commitments for encrypted indexes, encrypted data, and deletion markers, and provides existence proofs, non-existence proofs, data tree proofs, and deletion marker tree proofs when queries are returned. This enables clients or trusted agents to verify whether query results have been tampered with, missed, incorrectly combined, deleted records rolled back, or have inconsistent versions. When table data is added, deleted, or modified, only the affected row groups and affected queryable columns are locally indexed, thereby reducing the index maintenance cost in dynamic table scenarios. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the overall process of a table data dense state index query method according to the present invention; Figure 2 This is a schematic diagram of the table structure information processing and column-level index key derivation process in this invention; Figure 3 This is a schematic diagram of the structure of the row group-level dense state summary index and the row-level dense state position index in this invention; Figure 4 This is a schematic diagram of the encrypted query token generation and server-side two-layer filtering process in this invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] This embodiment provides a method for querying encrypted indexes of tabular data. This method can be applied to data processing scenarios including a client or trusted proxy and an encrypted data server. The client or trusted proxy is used to process plaintext tables, generate encrypted indexes, generate encrypted query tokens, verify integrity proofs, and decrypt query results. The encrypted data server is used to store the encrypted tabular data and encrypted indexes, and performs encrypted candidate filtering based on the encrypted query tokens. It should be noted that this invention protects the method flow; the client, trusted proxy, and encrypted data server are only used to illustrate the execution environment of the method.

[0021] In this embodiment, the encrypted data server is configured to not obtain plaintext column names, plaintext column values, plaintext query conditions, and data decryption keys. The encrypted data server can store encrypted data, encrypted indexes, and integrity commitment information, and can perform matching and combination operations based on encrypted query tokens. Its observable information includes table size classification, number of encrypted column identifiers, number of query tags, candidate row group classification, return block classification, index version parameters, and authorized access modes, but it cannot directly obtain plaintext table content, real row numbers, real number of hit rows, and real query result size from these parameters.

[0022] Example 1: like Figure 1 As shown, this embodiment provides a method for querying table data using a dense index, including the following steps: Step S1: Obtain table structure information, generate secret column identifiers and derive column-level index keys. like Figure 2As shown, the client or trusted proxy obtains the table to be protected; the table to be protected can be a spreadsheet, a CSV table, a relational database table, a data warehouse wide table, or a structured data table exported by the business system.

[0023] The client or trusted proxy performs structural parsing on the table to be protected to obtain table structure information; the table structure information includes table identifier, column identifier, column data type, row identifier, and set of column queryable operators.

[0024] In one specific implementation, the table structure information includes: table identifier. ; Plain text listing ; Column name; Column data type; Set of queryable operators for the column; Row identifier Or primary key; cell coordinates; column security level Current index version parameter .

[0025] The column data types include numeric, date, enumeration, text, boolean, monetary, or formatted identifier types; the set of column query operators includes equality query, range query, prefix query, containment query, text fragment query, and multi-column combination query.

[0026] To prevent the encrypted data server from directly knowing the plaintext column names and field business semantics, the client or trusted proxy generates an encrypted column identifier for each queryable column. The dense-state column identifier can be generated using a pseudo-random function: ; in, Represents a pseudo-random function. This represents the metadata protection key. Indicates the table identifier, Indicates plaintext column identifier, Indicates the column data type. Indicates the security level of the column. This represents the index version parameter, and the symbol || indicates the concatenation operation.

[0027] After generating the encrypted column identifier, the client or trusted agent derives the column-level index key based on the table identifier, the encrypted column identifier, the version parameter, and the column security level. It can be generated using a key derivation function: ; in, This represents the key derivation function. Indicates the root key. This represents the column-level index key.

[0028] Furthermore, to isolate different index uses, multiple subkeys can be derived from the column-level index key, including: the first bucket key used for generating the bucket digest tags of the row group level for numeric or date columns. ; The second bucket key used to generate row-level bucket position labels for numeric or date columns. ; Equivalent key used to generate equivalent labels for enumerating columns ; The text key used for generating text column fragment labels ; Row token key used for row-level position index generation ; Padding key used for pseudo-candidate expansion The commitment key used to generate the integrity commitment. .

[0029] In one implementation, the client or trusted proxy determines the index parameters based on the total number of rows in the table, column types, column value ranges, update frequency, query precision, and column security level; row group size. The number of rows can be determined based on the table size and update frequency, such as 512, 1024, 4096, or 8192. Smaller row groups are used when the update frequency is high, while larger row groups are used for query analysis tables.

[0030] For numerical columns, bucket width The column values ​​are determined based on the query precision and the range of column values. The corresponding bucket number is: ; in, The lower bound of the column values ​​for this numeric column; for date columns, the bucket width can be determined by day, week, month, or quarter; for text columns, the Bloom filter bit length... and the number of hash functions Based on the target false alarm rate and the number of text fragments Confirmed, among which , .

[0031] The fixed set of returned tiers can include 32, 64, 128, 256, 512 or larger tiers; if the candidate set size is... Then choose not less than The minimum tier is used as the returned block size; the number of pseudo-tags, pseudo-candidate row groups, and version rotation threshold are determined based on column security level, query frequency, and allowed storage overhead.

[0032] Through the above processing, keys for different columns, different versions, and different index purposes are isolated from each other; when a key is rotated for a certain column or version, it is not necessary to rebuild the entire table index, thereby reducing key maintenance costs and security risks.

[0033] Step S2: Divide the rows into groups and generate row group-level dense state summary index and row-level dense state position index. like Figure 3 As shown, the client or trusted proxy divides the table to be protected into multiple row groups according to the row grouping rules. A row group is a data unit obtained by grouping several rows in the table. Each row group includes several data records and corresponds to a row group identifier. .

[0034] The row group division rules include division by fixed number of rows, division by date window, division by business partition field, division by primary key hash result, division by data distribution, and division by update frequency.

[0035] For example, in scenarios with large table sizes and low update frequency, row groups can be formed for every thousand or ten thousand rows; in time series business tables, row groups can be divided according to months, quarters, or date windows; in multi-tenant business scenarios, row groups can be divided according to tenant identifiers or business partitions; in high-frequency update tables, records with high update frequency can be assigned to separate row groups to reduce the scope of index maintenance during subsequent updates.

[0036] To prevent the server from directly knowing the correspondence between line groups and plaintext line ranges, the client or trusted proxy can encrypt the line group identifier: ; in, Indicates the identifier of the dense row group. This indicates the row group identifier protection key.

[0037] After the row groups are divided, the client or trusted agent generates a row group-level dense state summary index for each queryable column in each row group, based on the data type of the queryable column and the set of queryable operators.

[0038] To accommodate the need for secret-state queries on different types of queryable columns within the same table, the client or trusted proxy generates distinct row-group-level secret-state summary index structures based on the column data type for different queryable columns within the same row group. Numeric or date columns correspond to row-group-level bucket summary labels, enumerated columns correspond to sets of equal-value secret-state labels, and text columns correspond to blinded text fragment filtering structures. All of these different types of row-group-level secret-state summary indexes are associated with row-level secret-state position indexes within the same row group, enabling candidate results from predicate hits on different columns to be converted into a unified set of candidate row tokens or candidate bitmaps.

[0039] The compressed candidate bitmap is encoded in ascending order of row token or row number within a row group. It can use block bitmap encoding, run-length encoding, or a combination of both to represent the set of candidate rows associated with a certain dense index label within a row group.

[0040] 1. Row-group level dense summary index for numeric or date columns. For numeric or date columns, the client or trusted agent does not directly use plaintext range indexes, nor does it expose the order of plaintext values ​​to the encrypted data server. Instead, it maps the column values ​​to bucket spaces and generates encrypted interval summaries based on the bucketing results.

[0041] Specifically, the client or trusted proxy performs the following processing: determining bucketing rules based on column value distribution, query granularity, row group data range, or column security level; mapping the numerical or date value of each cell within the row group to one or more bucket numbers. Based on the first bucket key Generate row group-level bucket summary tags for each bucket number. And based on the row group-level bucket summary labels that actually appear in the row group, a row group-level dense state summary index is formed for that row group and that column.

[0042] The row group-level bucket summary tags are generated in the following way: ;in, This is used by the server to determine whether a certain value bucket or date bucket is likely to hit the query range in the row group-level dense summary index.

[0043] Furthermore, the client or trusted agent uses the second bucket key. Generate row-level bucket location labels for bucket numbers within candidate row groups. : ; in, Used to associate the corresponding set of row tokens or candidate bitmaps within a candidate row group; the row-level bucket location label It is not included in the first secret query token, but is generated by the client or trusted agent based on the secret row group identifier of the comprehensive candidate row group after the comprehensive candidate row group is obtained by row group-level filtering, and then encapsulated into the second secret query token.

[0044] When performing subsequent range queries, the client or trusted proxy will convert the query range into one or more bucket numbers, and based on the first bucket key. Generate row group-level bucket query tags; the secret data server matches the row group-level bucket query tags with the row group-level secret summary index to determine candidate row groups that may hit the range conditions; for candidate row groups, the secret data server then queries the corresponding row token set or candidate bitmap based on the row-level bucket position tags to obtain the candidate row identifier set.

[0045] For candidate data falling within the boundary bucket, residual verification can be performed by the client or trusted agent in step S6 after decryption to eliminate data that does not meet the true range conditions.

[0046] 2. Enumerated column-level dense state summary index For enumerated columns, the client or trusted proxy performs normalization processing on the enumerated values ​​appearing in the row group and generates an equivalent dense state label set.

[0047] The standardization process includes removing leading and trailing spaces, unifying capitalization, unifying encoding, and mapping to standard enumeration numbers.

[0048] For enumeration values Equivalent dense state label Generate in the following way: ; in, This represents the standardized result of the enumerated values; For the list of salts, according to Generate row-level summary tags that do not change with row groups; row-level location tags also use row group scope salts. ,according to The server only stores and matches the encrypted data generated by the salt, and does not obtain the plaintext of the salt value.

[0049] The client or trusted proxy uses the set of equivalent dense state tags that have appeared in a row group as the row group-level dense state summary index corresponding to that row group and that enumeration column; when the query condition is the equality condition of a certain enumeration value, the dense state data server can determine the possible candidate row groups that may be hit based on the corresponding query tags.

[0050] 3. Row-group level dense-state summary index of text columns For text columns, the client or trusted proxy uses a blind index of text fragments instead of using a character-by-character ciphertext concatenation method for text queries.

[0051] Specifically, the client or trusted proxy performs the following processing: normalizes the text cells; generates text fragments based on the text content; generates blinded text labels for the text fragments; and writes the blinded text labels into the row group-level text summary index.

[0052] The text standardization process includes encoding standardization, case standardization, full-width / half-width character conversion, stop character processing, and punctuation normalization; the text fragments are n-gram fragments, word segmentation fragments, keyword fragments, prefix fragments, or formatted substrings.

[0053] For text fragments Blind text labels Generate in the following way: ; In a preferred embodiment, the row group-level text digest index uses a blinded Bloom filter; the client or trusted proxy uses multiple keyed hash functions to map the blinded text tags to several positions in the Bloom filter; the encrypted data server can determine whether a row group may contain the corresponding text fragment based on the query token, and the server cannot obtain the plaintext column name, plaintext query value and plaintext record content except for the preset leakage profile.

[0054] 4. Row-level dense state location index While generating the row group-level secret state summary index, the client or trusted agent also generates the row-level secret state location index; the row-level secret state location index is used to further locate candidate rows within the candidate row group.

[0055] For each data record in the row group, the client or trusted agent generates a row token. : ; in, Indicates plaintext line identifier or primary key, This indicates a cryptographic row token that is visible to the server but whose plaintext row identifier cannot be deduced from it.

[0056] For each encrypted index tag, the client or trusted proxy establishes an association between the tag and the corresponding row token set; the row-level encrypted position index adopts at least one of the following: pseudo-random row token set, salted inverted list, compressed candidate bitmap, and block candidate bitmap.

[0057] In one implementation, the client or trusted agent generates a pseudo-random row position for each row group and maps each record in the row group to a pseudo-random bitmap position. When an index label corresponds to several rows, the corresponding position is set to valid in the compressed candidate bitmap. When the dense data server performs intersection, union, or difference operations on the candidate bitmap, it can only observe the pseudo-random position and cannot know the real row number or primary key.

[0058] With the above settings, the row group-level secret state summary index is used to coarsely filter candidate row groups according to column type, and the row-level secret state position index is used to uniformly represent the secret state hit results corresponding to different column types as a set of candidate row tokens or a candidate bitmap. The secret state data server can perform intersection, union, or difference operations on the set of candidate row tokens or candidate bitmaps corresponding to numeric columns, date columns, enumeration columns, and text columns without recognizing plaintext column names, plaintext column type semantics, and plaintext cell values, thereby completing multi-column combination queries. Therefore, this invention does not simply superimpose bucket indexes, inverted indexes, or bitmap indexes, but uniformly maps the secret state summary results of heterogeneous column types to the row-level secret state position space, reducing the index scan range, the size of intermediate candidate sets, and the amount of secret text data transmission when performing multi-type column combination queries.

[0059] Step S3: Encrypt the table data and associate it with the stored encrypted index and integrity commitment information. After completing the construction of the row group-level encrypted digest index and the row-level encrypted location index, the client or trusted agent encrypts the data records in the table to be protected.

[0060] In one implementation, the client or trusted agent can encrypt data on a row-by-row, row-by-row, or data-by-block basis. During encryption, table identifiers, encrypted row group identifiers, row tokens, and version parameters can be used as associated data to prevent ciphertext data from being incorrectly replaced between different tables, different versions, or different row groups.

[0061] For a certain data record It can generate encrypted data. : ; in, This represents a symmetric encryption algorithm. This represents the data encryption key. This indicates related data. include , , and .

[0062] Simultaneously, the client or trusted agent generates integrity commitment information for the row group-level encrypted digest index, the row-level encrypted location index, and the encrypted data; the integrity commitment information can be generated using Merkle trees, hash chains, authentication accumulators, or signature commitment methods.

[0063] In one specific implementation, the client or trusted agent calculates the hash value for the index item and the ciphertext data block within the same row group, and then generates the row group commitment root GroupRoot based on these hash values; further, multiple row group commitment roots are combined to generate a global commitment root TableRoot; the client or trusted agent saves the global commitment root, or writes the global commitment root to a trusted storage location.

[0064] In one implementation, the integrity commitment information includes an index tree, a data tree, and a deletion marker tree; the leaf nodes of the index tree are: ; in, For dense state index labels, This is the set of row tokens or compressed candidate bitmaps associated with the dense index label.

[0065] The leaf nodes of the data tree are: ; in, This is the ciphertext data block for the corresponding row.

[0066] Deleting a leaf node in the marker tree is as follows: ; in, Indicates deletion status. This indicates the deletion version or deletion time parameter. The root nodes of the index tree, data tree, and deletion marker tree are used to generate the row group commitment root or global commitment root.

[0067] The leaf nodes in the index tree are used to bind the encrypted index label and its corresponding row token set or candidate bitmap; the leaf nodes in the data tree are used to bind the row token and its corresponding encrypted data block; the leaf nodes in the deletion marker tree are used to bind the row token and its deletion status; the index tree, data tree and deletion marker tree are all associated with the same index version parameter to support subsequent queries to verify index hit, data return, deletion status and version consistency.

[0068] Subsequently, the client or trusted proxy will associate and store the encrypted table data, row group-level encrypted digest index, row-level encrypted location index, and auxiliary information required for integrity proof to the encrypted data server.

[0069] The encrypted data server can store and read the above data, but it cannot deduce the plaintext table content, plaintext column names, plaintext cell values, or plaintext row identifiers based on the above data.

[0070] Step S4: Convert the table query request into a query predicate graph and generate a dense query token. like Figure 4As shown, when a user initiates a table query request, the client or trusted proxy parses the table query request and identifies the query columns, query operators, query values, logical relationships, and security level requirements contained therein.

[0071] For example, a user initiates the following query: to find employee records whose department is R&D, whose salary is within a preset range, and whose remarks contain a certain keyword; the client or trusted proxy can parse the query into three column predicates: department column equivalence predicate, salary column range predicate, and remarks column text fragment predicate.

[0072] Subsequently, the client or trusted proxy generates a query predicate graph based on the column predicates and logical relationships. The query predicate graph can be a directed acyclic graph, in which column predicates serve as predicate nodes, and AND, OR, NOT, intersection, union, or difference serve as logical operation nodes or edge relationships.

[0073] When generating the query predicate graph, the client or trusted proxy can also determine the execution order of different column predicates based on the secret selection rate summary. The secret selection rate summary can be a statistical summary obtained during the index building phase, used to characterize the potential hit candidate size of a certain column predicate. The secret selection rate summary is generated based on the ratio of the number of candidate row tokens associated with the secret index label to the row group size, and is encoded and secret-encapsulated according to a preset tier. It is used to estimate the candidate size and select a fixed return tier without exposing the exact number of hit rows. This statistical summary does not contain plaintext column values ​​and is only used to estimate the execution cost. For example, equivalent predicates with smaller candidate sizes can be executed first. Text fragment predicates with larger candidate sizes can be executed after filtering other predicates. In this way, the size of the intermediate candidate set and the server-side set operation overhead can be reduced.

[0074] After generating the query predicate graph, the client or trusted agent generates a secret query token based on the query predicate graph, column-level index key, and version parameters.

[0075] For range queries on numeric or date columns, the client or trusted proxy converts the plaintext range conditions into one or more bucket numbers and generates bucket query labels based on the bucket key.

[0076] For equality queries on enumerated columns, the client or trusted proxy performs standardization processing on the query enumerated values ​​and generates equality query tags based on the equality key.

[0077] For text column fragment queries, the client or trusted proxy performs standardization and fragmentation processing on the query text, and generates text fragment query tags based on the text key.

[0078] The first secret query token includes a secret column identifier, a set of row group-level query tags, a version parameter, a secret representation of the predicate graph structure, a query random number, a set of pseudo query tags, and a leakage control parameter; the second secret query token includes a secret row group identifier of the comprehensive candidate row group, a set of corresponding row-level query tags, a version parameter, and a query random number; the row-level query tags in the second secret query token are generated with the corresponding secret row group identifier as the scope.

[0079] For queries with high security levels, the client or trusted proxy adds several pseudo-query tags in addition to the real core query tags, and combines fixed return tier parameters, batch query encapsulation, and version rotation to reduce the leakage of query access patterns, query frequency, and result scale. The query random number, query batch number, or one-time encapsulation parameters are used to protect the freshness and integrity of the encrypted query token, rather than to change the deterministic matching relationship of the core query tags.

[0080] Step S5: The server performs row group-level filtering and row-level filtering based on the encrypted query token. After receiving the encrypted query token, the encrypted data server first determines the corresponding encrypted index partition based on the encrypted column identifier. Since the encrypted column identifier is generated by the client or trusted proxy through a pseudo-random function, the encrypted data server cannot know its corresponding plaintext column name or field meaning.

[0081] Subsequently, the secret data server performs row group-level filtering on the row group-level secret summary index based on the row group-level query tag in the secret query token.

[0082] For numeric or date columns, the secret data server matches the row group-level secret interval summary based on the row group-level bucket query label to obtain candidate row groups that may contain the target range value; and within the candidate row group, it queries the corresponding row token set or candidate bitmap based on the row-level bucket position label to obtain the candidate row identifier set corresponding to the predicate of the numeric or date column.

[0083] For enumerated columns, the dense data server matches the row group-level dense label set with the equivalent query label to obtain candidate row groups that may contain the target enumerated value.

[0084] For text columns, the dense data server matches blinded Bloom filters, salted inverted labels, or other text summary structures based on the text fragment query tags to obtain candidate row groups that may contain the target text fragment.

[0085] When the query predicate graph contains multiple column predicates, the secret data server performs combination operations on the candidate row group sets corresponding to different column predicates according to the logical relationship represented by the query predicate graph. For example, for AND relations, the intersection operation is performed; for OR relations, the union operation is performed; for NOT relations or exclusion conditions, the difference operation is performed. After row group-level combination operations, a comprehensive candidate row group is obtained.

[0086] After obtaining the comprehensive candidate row group, the encrypted data server returns the encrypted row group identifier and row group-level integrity certificate of the comprehensive candidate row group to the client or trusted agent. After the client or trusted agent verifies the information, it generates a second encrypted query token and sends it to the encrypted data server.

[0087] Then, the dense data server performs row-level filtering based on the row-level dense position index within the comprehensive candidate row group; specifically, the dense data server queries the corresponding pseudo-random row token set, salted inverted list, or compressed candidate bitmap according to the row-level query label in the dense query token to obtain the candidate row identifier set corresponding to each column predicate.

[0088] When the candidate row identifier set is represented by a compressed candidate bitmap, the encrypted data server can directly perform bitwise AND, bitwise OR, or bitwise difference operations on the bitmap; when the candidate row identifier set is represented by a pseudo-random row token set, the encrypted data server can perform intersection, union, or difference operations.

[0089] In the case of dynamic updates, the secret data server reads the currently valid baseline index segment and incremental index segment simultaneously, merges the candidate row token set or candidate bitmap obtained from the two, and performs exclusion processing on the merged candidate row identifier set according to the deletion flag.

[0090] Through the above row group-level filtering and row-level filtering, the encrypted data server obtains a set of candidate row identifiers that satisfy the query predicate graph; the encrypted data server cannot deduce the real row number, primary key or plaintext query conditions from this set of candidate row tokens or candidate bitmap.

[0091] Step S6: Perform leakage control, return encrypted data, and complete integrity verification. After obtaining the candidate row identifier set, the encrypted data server performs pseudo-candidate expansion or fixed-size encapsulation on the candidate row identifier set according to preset leakage control parameters.

[0092] The leakage control parameters are determined by the client or trusted proxy based on column security level, number of query predicates, estimated candidate size, query frequency, and business security policy, and are encapsulated in a secure query token; the leakage control parameters include security level. Fixed return set Number of pseudo-query tags (p) and number of pseudo-candidate row groups And the version rotation threshold R. Among them, the fixed return tier is used to determine the number of candidate ciphertext data blocks returned by the ciphertext data server each time, the number of pseudo query labels is used to determine the number of pseudo labels mixed in with the real core query labels, and the version rotation threshold is used to determine whether to re-derive the index labels and query labels of the affected columns.

[0093] The leakage control parameters are used to limit the scale information that the encrypted data server can observe during the query process. The client or trusted proxy determines the number of pseudo-query tags, the number of pseudo-candidates, and the fixed return tier based on the column security level, the number of query predicates, the estimated candidate size, and the allowed storage overhead. The encrypted data server returns candidate encrypted data blocks according to the fixed return tier. When the number of real candidates is less than the return tier, pseudo-candidates are added from the committed pseudo-candidate space. The pseudo-candidates are generated during the index building phase and marked as padding types in the integrity commitment information. The client or trusted proxy identifies and removes padding type records after integrity verification and decryption.

[0094] Leakage control methods include: adding rows outside the true candidate row group... A specified number of pseudo-candidate row groups are added; a specified number of pseudo-query tags are added in addition to the real core query tags; the returned results are encapsulated into data blocks of a fixed size according to the fixed return set; pseudo-candidate row tokens are added in addition to the real candidate row tokens; index version rotation is triggered when the version rotation threshold R is reached; query random numbers, query batch numbers, or one-time encapsulation parameters are used for freshness verification, replay attack protection, and integrity binding of the encrypted query tokens, without changing the core query tags used for index matching.

[0095] Clients or trusted agents can evaluate the effectiveness of parameter configurations based on index storage amplification factor, average number of candidate row groups, text query false positive rate, fixed tier return overhead, and dynamic update merging overhead, and adjust row group size, bucket width, Bloom filter parameters, fixed return tier, and version rotation threshold based on the evaluation results.

[0096] In one specific implementation, the fixed return tier set and tier selection rules are encapsulated in a cryptographic query token by the client or trusted proxy; the cryptographic data server obtains the size of the real candidate row identifier set. Then, according to the aforementioned grading selection rules, return from the fixed grading set. Select to satisfy The minimum value is used as the return value for this time. And a tiered selection proof is added to the integrity proof. When At that time, the dense data server selects from the committed pseudo-candidate space. The pseudo-candidate row identifiers are used to make the final returned candidate set size . .

[0097] Through the above leakage control measures, the information observable on the encrypted data server is limited to table size classification, candidate row group classification, return block classification, and index version parameters, rather than directly exposing the actual number of hit rows, the actual number of query tags, and the actual size of the results.

[0098] When the encrypted data server returns query results, it also returns an integrity proof corresponding to the query results. For a matched encrypted index tag, it returns an index tree existence proof to prove that the encrypted index tag and its associated row token set or candidate bitmap belong to the current version of the index tree. For a queried but unmatched encrypted index tag, it returns a non-existence proof, which can be a sparse Merkle tree proof or an adjacent leaf node proof when the encrypted index tags are stored in an ordered manner. For the returned encrypted data block, it returns a data tree proof. For deleted records, it returns a deletion tag tree proof. The client or trusted proxy verifies whether the query tag, candidate row identifier set, encrypted data, deletion tag, and index version parameter belong to the same integrity commitment information based on the saved global commitment root or trusted commitment root. If the verification fails, the query results are rejected.

[0099] The integrity proof also includes a predicate graph evaluation proof, which includes the query label corresponding to each predicate node, the hash value of the candidate row group set, the hash value of the row-level candidate bitmap or row token set, the input hash value and output hash value of each logical operation node, and the hash value of the final candidate row identifier set. The client or trusted proxy recalculates the hash values ​​and candidate sets according to the query predicate graph to verify that the encrypted data server has not missed any candidate row groups, has not missed any hit labels, and has not incorrectly performed intersection, union, or difference operations.

[0100] The encrypted data server returns the encrypted data block and integrity certificate to the client or trusted agent.

[0101] After receiving the encrypted data block and integrity certificate, the client or trusted proxy first verifies the integrity certificate based on the pre-saved global commitment root or trusted commitment root. If the verification fails, it determines that the result returned by the server may be tampered with, omitted, inconsistent in version, or inconsistent in index, and refuses to output the query result or outputs an error message.

[0102] If the integrity verification passes, the client or trusted agent decrypts the returned ciphertext data. For candidate errors caused by numerical bucketing, date bucketing, text fragment indexing, or Bloom filters, the client or trusted agent further performs residual verification to eliminate candidate records that do not meet the true query conditions.

[0103] Meanwhile, the client or trusted proxy identifies and removes false candidate records according to the false candidate generation rules, and finally outputs the true query results.

[0104] Through the above steps, this invention enables indexed queries of tabular data under dense storage conditions, while also taking into account query efficiency, leakage control, and result integrity verification.

[0105] Step S7: Dynamically update the dense state index When a table to be protected is added, deleted, or modified, the client or trusted agent determines the affected row groups and affected queryable columns based on the row identifier, the secret row group identifier, and the changed column identifier; only the secret index label, row token association, deletion flag, and corresponding integrity commitment information are updated for the affected row groups and affected queryable columns, while the unaffected row groups and queryable columns retain their original index status.

[0106] For a new record, the client or trusted agent determines the target row group to which the new record belongs, generates the row token and encrypted data block corresponding to the new record, generates the corresponding encrypted index label based on the column values ​​of each queryable column in the new record, writes the association between the encrypted index label and the row token into the incremental index segment of the target row group, and updates the corresponding index tree node and data tree node.

[0107] For deleted records, the client or trusted agent does not physically delete the corresponding index entry from the baseline index segment. Instead, it generates a deletion tag containing a row token, deletion version parameter, and deletion commitment value, and writes it into the deletion tag tree. In subsequent queries, the secret data server excludes the corresponding row token from the candidate row identifier set based on the deletion tag and returns the deletion tag tree proof.

[0108] For modified records, the client or trusted agent determines the queryable columns that have changed before and after the modification; for the queryable columns that have changed, the old encrypted index labels are invalidated, and new encrypted index labels are generated based on the modified column values. The association between the new encrypted index labels and the original row tokens is written into the incremental index segment; the queryable columns that have not changed are not re-indexed.

[0109] When the number of incremental index segments in a row group reaches a preset segment threshold, the number of deletion markers reaches a preset deletion threshold, or the number of queries for that row group reaches a preset query threshold, the client or trusted agent performs a row group-level merge of the baseline index segments, incremental index segments, and deletion markers for that row group. It then regenerates the row group-level dense summary index, row-level dense position index, and row group commitment root for that row group, and derives the index keys for the affected columns based on the new index version parameters. After the merge is completed, the old version index items will no longer participate in subsequent queries.

[0110] During the dynamically updated query process, the secret data server generates an integrity certificate based on the currently valid baseline index segment, incremental index segment, and deletion marker, and binds the certificate to the current query token through the index version parameter to avoid using old version indexes, expired index tags, or deleted records to generate query results.

[0111] When executing a query, the secret data server merges the currently valid baseline index segment, incremental index segment, and deletion marker, and returns the corresponding index tree proof, data tree proof, and deletion marker tree proof. The client or trusted agent confirms that the proof belongs to the same version as the current query token based on the index version parameter, thus avoiding the use of old version indexes or deleted records to generate query results.

[0112] Example 2: The following uses an employee information table as an example to illustrate the execution process of this invention; the employee information table includes columns such as employee number, department, date of employment, salary, and remarks; wherein, the department column is an enumerated column, the date of employment column and the salary column are date columns or numerical columns, and the remarks column is a text column.

[0113] First, the client or trusted agent obtains the employee information table and extracts the table identifier, column identifier, column data type, row identifier, and column queryable operator set. The client or trusted agent generates encrypted column identifiers for the department column, the date of employment column, the salary column, and the remarks column, and derives column-level index keys for different columns.

[0114] Then, the client or trusted proxy divides the employee information table into multiple row groups according to a fixed number of rows or department distribution; for the salary column, the client or trusted proxy generates dense bucket labels according to salary range and forms a row group-level dense range summary; for the department column, the client or trusted proxy standardizes the department name and generates an equivalent dense label set; for the remarks column, the client or trusted proxy generates n-gram fragments or keyword fragments from the remarks text and generates blinded text labels.

[0115] Simultaneously, the client or trusted agent generates a row token for each employee record and establishes an association between the encrypted index label and the row token set to obtain the row-level encrypted location index. Subsequently, the client or trusted agent encrypts the employee record and associates and stores the encrypted data, row group-level encrypted digest index, row-level encrypted location index, and integrity commitment information to the encrypted data server.

[0116] When a user queries "employee records whose department is R&D, whose salary is within a preset range, and whose remarks contain a certain keyword", the client or trusted proxy will parse the query request into a department equivalence predicate, a salary range predicate, and a remarks text fragment predicate, and construct a query predicate graph. The client or trusted proxy will generate equivalence query labels, range query labels, and text fragment query labels according to the index keys corresponding to each column, and add query random numbers, pseudo-query labels, and fixed return block parameters to form a secret query token.

[0117] After receiving the secret query token, the secret data server first determines candidate row groups based on the row group-level secret summary indexes corresponding to the department column, salary column, and remarks column, and then performs an intersection operation on the candidate row group sets according to the AND relationship in the query predicate graph. Subsequently, the secret data server queries the row-level secret position indexes within the candidate row groups to obtain the candidate row token sets corresponding to the department predicate, salary predicate, and remarks predicate, and performs an intersection operation on these candidate row token sets to obtain a comprehensive candidate row token set.

[0118] The encrypted data server adds pseudo-candidate row tokens to the comprehensive candidate row token set according to the fixed return block parameters, and returns the corresponding encrypted data block and integrity proof.

[0119] The client or trusted proxy verifies the integrity certificate; if the verification passes, the returned ciphertext data is decrypted, and residual verification is performed on the wage boundary range and remarks text conditions to remove data records and false candidate records that do not meet the true query conditions, and finally outputs the true query results; if the verification fails, the query results are rejected or a query error message is output.

[0120] When an employee information table is updated with new employees, deleting employees upon leaving, or modifications to fields such as salary, department, and remarks, the client or trusted agent only updates the row group to which the corresponding employee record belongs and the queryable columns that have changed. Through a local update mechanism consisting of incremental index segments, deletion markers, row group index merging, and version rotation, the scope of index maintenance is limited to local row groups and local columns, thereby reducing the overhead of rebuilding the entire table index in dynamic table scenarios.

[0121] As can be seen from the above examples, the present invention can achieve efficient dense combination query of table data when the plaintext table, plaintext column names and plaintext query conditions are not visible on the server side.

[0122] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for querying dense indexes in tabular data, characterized in that, Includes the following steps: Obtain the table to be protected and parse it to obtain the table structure information including the set of column queryable operators. The table to be protected includes multiple data records. Columns that are not empty in the set of column queryable operators are identified as queryable columns. A secret column identifier is generated and a column-level index key is derived. The table to be protected is divided into row groups, and a secret row group identifier and row token are generated. A row group-level secret summary index containing secret index labels is generated according to the data type of the queryable columns. The secret index labels are mapped to the row-level secret position index, so that the hit results of each queryable column are represented as a set of row tokens or a candidate bitmap within the same row group. The data records are encrypted to form ciphertext data, and integrity commitment information is generated from the row group-level ciphertext digest index, the row-level ciphertext location index, and the ciphertext data and then stored in the ciphertext data server. Generate a query predicate graph and generate a dense query token containing dense column identifiers, row group-level query labels, row-level query labels, predicate graph representation, and leakage control parameters; Candidate row groups are obtained based on row group-level query tags, and comprehensive candidate row groups are obtained by combining them according to the query predicate graph. Within the comprehensive candidate row groups, row-level dense position indexes are queried based on row-level query tags, and a combination operation is performed on the row token set or candidate bitmap to obtain the candidate row identifier set. Based on the leakage control parameters, encapsulate the candidate row identifier set, return the corresponding ciphertext data and the integrity certificate generated based on the integrity commitment information, verify the integrity certificate, decrypt the ciphertext data and perform residual verification, and then output the query results.

2. The method for querying dense indexes of tabular data according to claim 1, characterized in that, The table structure information includes table identifier, column identifier, column data type, row identifier, and set of column queryable operators; when generating the encrypted column identifier, the table identifier, column identifier, column data type, column security level, and index version parameter are used as inputs; when deriving the column-level index key, the table identifier, encrypted column identifier, index version parameter, and column security level are used as inputs, and subkeys for row group-level encrypted digest index, row-level encrypted position index, row token, pseudo-candidate, and integrity commitment are derived.

3. The method for querying dense indexes of tabular data according to claim 1, characterized in that, The row group identifier is processed to obtain the encrypted row group identifier, and a row token is generated based on the table identifier, the encrypted row group identifier, the row identifier and the index version parameter. The row-level encrypted position index stores the set of row tokens or candidate bitmaps associated with the encrypted index label under the corresponding encrypted row group identifier. The candidate bitmap represents the candidate row according to the pseudo-random row position order within the row group.

4. The method for querying dense indexes of tabular data according to claim 1, characterized in that, When generating the row group-level dense state summary index, the cell values ​​of the numeric and date columns are mapped to bucket numbers according to the bucketing rules, and row group-level bucket summary labels are generated from the bucket numbers; the enumeration values ​​of the enumeration column are standardized to generate equivalent dense state labels; the text cells of the text column are standardized and fragmented to generate blinded text labels.

5. The method for querying dense indexes of tabular data according to claim 1, characterized in that, When encrypting data records, encrypted data is formed in units of rows, row groups, or data blocks, and the table identifier, encrypted row group identifier, row token, and index version parameter are used as associated data; the integrity commitment information is generated by the index tree, data tree, and deletion mark tree when there are deleted records.

6. The method for querying dense indexes of tabular data according to claim 1, characterized in that, When generating the query predicate graph, the query columns, query operators, query values ​​and logical relationships are identified from the table query request. The column predicates are set as predicate nodes, and the intersection, union or difference are set as logical operation relationships. When generating the dense query token, the range query is converted into a bucket query label, the enumerated equality query is converted into an equality query label, and the text fragment query is converted into a text fragment query label.

7. The method for querying dense indexes of tabular data according to claim 1, characterized in that, When generating the candidate row identifier set, the corresponding dense index partition is first selected based on the dense column identifier, and the candidate row group corresponding to the same query predicate is recorded as the predicate candidate result. When the query predicate graph contains multiple query predicates, the candidate results of each predicate are merged according to the combination relationship indicated by the query predicate graph to obtain the comprehensive candidate row group. Then, the row-level dense position index is queried only within the comprehensive candidate row group to generate the candidate row identifier set.

8. The method for querying dense indexes of tabular data according to claim 1, characterized in that, The leakage control parameters are used to determine the number of pseudo-query tags, the number of pseudo-candidates, and the fixed return level; when the size of the candidate row identifier set is less than the fixed return level, the encrypted data server fills in the pseudo-candidates from the committed pseudo-candidate space and returns them together; after integrity verification and decryption, records marked as fill type are identified and removed.

9. The method for querying dense indexes of tabular data according to claim 1, characterized in that, The integrity proof includes the existence or non-existence proof corresponding to the row group-level secret state digest index, the existence or non-existence proof corresponding to the row-level secret state position index, the data tree proof corresponding to the returned ciphertext data, and the deletion tag tree proof corresponding to the deleted record. It is used to verify that the row group candidate bitmap, the row-level candidate bitmap or row token set, the ciphertext data, the deletion tag and the index version belong to the same integrity commitment information.

10. The method for querying dense indexes of tabular data according to claim 1, characterized in that, When a table to be protected is updated, deleted, or modified, the affected row groups and affected queryable columns are identified. New records are written to the incremental index segment of the corresponding row group, deleted records are written to the deletion flag, and modified records are updated only for the changed queryable columns. During a query, the effective baseline index segment, incremental index segment, and deletion flag are processed, and the corresponding proof is returned.