A customer privacy data storage method of a customer information management platform
Patent Information
- Application Number
- CN202611106115.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-24
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-07-24
AI Technical Summary
[0003]在相关技术中,通用的加密算法会破坏数据的局部有序性,使得数据库无法针对加密后的数据直接建立有效的索引结构,导致查询操作通常退化为全表扫描,严重影响检索效率
在本发明提供的一种客户信息管理平台的客户隐私数据存储方法中,通过将识别数据与业务数据分离存储,并在写入时基于识别数据生成截断索引键值用于构建数据库索引,同时将全量哈希序列存储于备份列作为后续重构的熵源,有效解决了加密数据无法直接建立索引以及全量哈希索引键过长导致写入性能低下的技术问题。在此基础上,本发明通过在内存中维护哈希碰撞计数矩阵,实时统计截断索引键值的出现频率并确定哈希分布均匀度,实现了对索引碰撞程度的零输入/输出(input/output,I/O)开销感知,能够在业务无感知的情况下精准识别索引健康度;当哈希分布均匀度低于预设偏斜阈值且当前索引截取长度尚未达到上限时,客户信息管理平台自动调整索引截取长度并基于备份列中存储的全量哈希序列对索引列的截断索引键值进行重构更新,整个重构过程无需读取和解密识别数据列中的密文,完全规避了中央处理器(Central Processing Unit,CPU)密集型的解密操作,极大降低了索引调整的资源消耗和时间成本,同时通过调整状态标记的开关控制实现了写入和查询模块的协同切换,确保在索引重构期间业务不中断、数据不丢失。由此,本发明在保障客户隐私数据加密存储安全性的前提下,实现了高并发写入吞吐量与精确检索效率的动态平衡,并提供了低成本的索引自适应演进能力,显著提升了客户信息管理平台在数据规模激增或分布变化时的稳定性和运维效率。
Smart Images

Figure CN122614843B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and more specifically to a method for storing customer privacy data in a customer information management platform. Background Technology
[0002] In the context of digital development, customer information management platforms have become crucial carriers for enterprises to manage customer data. The customer identification data and related business data stored on these platforms contain a significant amount of private information. With increasingly stringent data security regulations, enterprises are demanding higher levels of encrypted storage for customer privacy data. Furthermore, under multi-tenant, multi-business-line operating models, enterprises need to group and manage customer data according to different business dimensions. The high-concurrency data storage and retrieval demands also place higher industry requirements on the security, efficiency, and flexibility of customer privacy data storage.
[0003] In related technologies, common encryption algorithms can disrupt the local order of data, making it impossible for databases to directly build effective index structures for encrypted data. This often leads to query operations degenerating into full table scans, severely impacting retrieval efficiency. To address this issue, some solutions use storing the full hash sequence and building an index. However, an excessively long full hash sequence significantly increases the maintenance overhead of the database index tree in high-concurrency write scenarios, reducing write throughput. If a shorter prefix of the full hash sequence is used as the index to improve write performance, hash collisions can reduce retrieval accuracy. Furthermore, when collisions are severe, it is usually necessary to decrypt the entire encrypted data and rebuild the index, a process that consumes significant resources and causes prolonged business interruption. Summary of the Invention
[0004] To address the technical challenge of balancing write performance and retrieval efficiency in encrypted data storage under high-concurrency write scenarios, this invention aims to provide a method for storing customer privacy data in a customer information management platform. The specific technical solution adopted is as follows: Firstly, a method for storing customer privacy data in a customer information management platform is provided. This method includes: acquiring a record to be stored, the record including identification data and business data associated with the identification data, and acquiring the business group identifier to which the identification data belongs; generating encrypted data and a full hash sequence corresponding to the identification data, and truncating the full hash sequence according to the current index truncation length corresponding to the business group identifier to generate a truncated index key value for constructing a database index; storing the truncated index key value, the full hash sequence, the encrypted data, and the business data respectively in an index column, a backup column, an identification data column, and a... The business data column maintains a hash collision count matrix corresponding to the business group identifier in memory. The hash collision count matrix is updated according to the truncated index key value, and the hash distribution uniformity used to characterize the degree of collision of the truncated index key value is determined according to the hash collision count matrix. The hash collision count matrix is used to count the frequency of occurrence of different truncated index key values. If the hash distribution uniformity is less than the preset skew threshold and the current index truncation length is less than the preset maximum truncation length, the current index truncation length is adjusted, and the truncated index key value of the index column is reconstructed and updated based on the adjusted index truncation length and the full hash sequence stored in the backup column.
[0005] In one possible design, encrypted data and a full hash sequence corresponding to the identification data are generated. The full hash sequence is then truncated according to the current index truncation length corresponding to the business group identifier to generate a truncated index key value for building a database index. This includes: encrypting the identification data using a preset symmetric encryption algorithm based on the encryption key corresponding to the business group identifier to generate encrypted data; performing hash calculation on the identification data using a preset hash algorithm based on the hash key corresponding to the business group identifier to obtain a full hash sequence with a preset fixed length; accessing the configuration center to determine the index truncation length associated with the business group identifier as the current index truncation length; and, if no record corresponding to the business group identifier is found in the configuration center, determining the preset baseline length as the current index truncation length and storing the current index truncation length associated with the business group identifier in the configuration center; and truncating the data from the beginning of the full hash sequence by the number of bytes corresponding to the current index truncation length to generate a truncated index key value.
[0006] In one possible design, maintaining a hash collision counting matrix corresponding to the business group identifier in memory includes: initializing the hash collision counting matrix corresponding to the business group identifier in memory, configuring a preset hash function with independent seed parameters for each row of the hash collision counting matrix, setting the initial count value of all counting units in the hash collision counting matrix to zero, the hash collision counting matrix being a two-dimensional counting matrix, the number of rows of the two-dimensional counting matrix corresponding to the number of preset hash functions, and the number of columns corresponding to the number of hash buckets; asynchronously retrieving truncated index key values from the message queue, and mapping the truncated index key values to the corresponding counting units in the hash collision counting matrix through multiple preset hash functions, and accumulating and updating the count value of the counting units.
[0007] In one possible design, the hash distribution uniformity, used to characterize the degree of collision of truncated index key values, is determined based on the hash collision count matrix. This includes: constructing a set of count values based on the count values of each column in any row of the hash collision count matrix at a preset period, or constructing a set of count values based on the average of the count values of each column in any number of rows of the hash collision count matrix; extracting non-zero values from the set of count values to obtain a set of non-zero count values; determining the sum of all count values in the set of non-zero count values; and determining the hash distribution uniformity based on the sum of the squares of the proportions of each count value in the set of non-zero count values to the sum of count values.
[0008] In one possible design, the above method further includes: after determining the uniformity of hash distribution, attenuating the count values of all counting units in the hash collision counting matrix; resetting the hash collision counting matrix to zero if the index truncation length adjustment is not triggered for a preset duration; and resetting the hash collision counting matrix to zero after the index truncation length adjustment is triggered.
[0009] In one possible design, when the hash distribution uniformity is less than a preset skew threshold and the current index truncation length is less than a preset maximum truncation length, the current index truncation length is adjusted. This includes: reading the preset skew threshold and the preset maximum truncation length stored in the configuration center. The preset skew threshold is used to determine whether the hash collision level has reached the critical state where the index length needs to be adjusted, and the preset maximum truncation length is used to limit the upper limit of the index truncation length. When the hash distribution uniformity is less than the preset skew threshold and the current index truncation length is less than the preset maximum truncation length, the adjustment status flag corresponding to the business group identifier is set to the enabled state, and the current index truncation length is increased by a preset step size to obtain the adjusted index truncation length.
[0010] In one possible design, based on the adjusted index truncation length and the full hash sequence stored in the backup column, the truncated index key value of the index column is reconstructed and updated. This includes: reading the full hash sequence stored in the backup column in the database, and extracting data from the full hash sequence corresponding to the adjusted index truncation length in bytes to generate a new truncated index key value; writing the new truncated index key value into the index column of the database, replacing the original truncated index key value; and after updating the index column of all existing data corresponding to the business group identifier, setting the adjustment status flag to the off state.
[0011] In one possible design, the above method further includes: when the adjustment state is marked as enabled, in response to a new data storage request, generating a first truncated index key value corresponding to the current index truncation length, and generating a second truncated index key value corresponding to the adjusted index truncation length; and writing the first truncated index key value and the second truncated index key value together into the index column of the database.
[0012] In one possible design, the above method further includes: in response to a data query request, obtaining the data to be queried; when the adjustment state is marked as enabled, generating a first query key value corresponding to the current index truncation length based on the data to be queried, and generating a second query key value corresponding to the adjusted index truncation length; when the adjustment state is marked as disabled, generating a third query key value corresponding to the current index truncation length based on the data to be queried; performing a union search in the database containing the first and second query key values, or performing an equality search in the database on the third query key value, to obtain a candidate record set; decrypting the encrypted data in the candidate record set and comparing it with the data to be queried to determine the target query result.
[0013] In one possible design, the encrypted data in the candidate record set is decrypted and compared with the data to be queried to determine the target query result. This includes: traversing each record in the candidate record set, decrypting the encrypted data in the record according to the encryption key corresponding to the business group identifier to obtain candidate identification data; comparing the candidate identification data with the data to be queried, and determining the business data in the record where the comparison result is consistent as the target query result.
[0014] The present invention has the following beneficial effects: In the customer privacy data storage method of the customer information management platform provided by the present invention, by storing identification data and business data separately, and generating truncated index key values based on the identification data to build a database index during writing, and storing the full hash sequence in a backup column as the entropy source for subsequent reconstruction, the technical problems of encrypted data not being able to be directly indexed and the low write performance caused by the full hash index key being too long are effectively solved. Based on this, the present invention maintains a hash collision counting matrix in memory, counts the frequency of occurrence of truncated index key values in real time, and determines the hash distribution uniformity. This achieves zero input / output (I / O) overhead awareness of the degree of index collision, enabling accurate identification of index health without business awareness. When the hash distribution uniformity is lower than a preset skew threshold and the current index truncation length has not yet reached its upper limit, the customer information management platform automatically adjusts the index truncation length and reconstructs and updates the truncated index key values of the index column based on the full hash sequence stored in the backup column. The entire reconstruction process does not require reading and decrypting the ciphertext in the identification data column, completely avoiding the CPU-intensive decryption operation, greatly reducing the resource consumption and time cost of index adjustment. At the same time, the coordinated switching of the write and query modules is achieved by adjusting the status flag switch, ensuring that business is not interrupted and data is not lost during index reconstruction. Therefore, this invention achieves a dynamic balance between high-concurrency write throughput and accurate retrieval efficiency while ensuring the security of encrypted storage of customer privacy data. It also provides low-cost index adaptive evolution capabilities, significantly improving the stability and operational efficiency of the customer information management platform when data scale surges or distribution changes. Attached Figure Description
[0015] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a schematic diagram of the structure of a customer information management platform provided in one embodiment of the present invention; Figure 2 This is a flowchart illustrating a customer privacy data storage method for a customer information management platform, as provided in one embodiment of the present invention. Detailed Implementation
[0017] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a customer privacy data storage method for a customer information management platform proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0018] In embodiments of the present invention, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" or "for example" in embodiments of the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0019] In the description of this invention, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" and "more than one" refer to two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0021] The following description, in conjunction with the accompanying drawings, details a specific scheme for a customer privacy data storage method for a customer information management platform provided by the present invention.
[0022] Please see Figure 1 It shows a schematic diagram of the structure of a customer information management platform provided in one embodiment of the present invention, such as... Figure 1 As shown, the customer information management platform 10 includes an acquisition unit 11, a generation unit 12, a storage unit 13, a monitoring unit 14, and an adjustment unit 15.
[0023] The acquisition unit 11 is used to acquire the record to be stored, wherein the record to be stored includes identification data and business data associated with the identification data, and to acquire the business group identifier to which the identification data belongs.
[0024] In some embodiments, the acquisition unit 11 receives a customer data storage request initiated by the business terminal, parses and extracts the request data to obtain a record to be stored. The record to be stored includes identification data (such as ID card number, mobile phone number, etc.) used to identify customers and business data associated with the identification data (such as order number, amount, etc.). At the same time, the business group identifier to which the identification data belongs is determined according to the business dimensions such as the tenant and business line to which the identification data belongs.
[0025] The generation unit 12 is used to generate encrypted data and a full hash sequence corresponding to the identification data, and to truncate the full hash sequence according to the current index truncation length corresponding to the business group identifier, so as to generate a truncated index key value for building a database index.
[0026] In some embodiments, the generation unit 12 encrypts the identification data using a preset symmetric encryption algorithm based on the encryption key corresponding to the service group identifier, generating encrypted data; it then performs hash calculation on the identification data using a preset hash algorithm based on the hash key corresponding to the service group identifier, obtaining a full hash sequence with a preset fixed length; simultaneously, the generation unit 12 accesses the configuration center, determines the index truncation length associated with the service group identifier as the current index truncation length, and if no record corresponding to the service group identifier is found in the configuration center, it determines the preset baseline length as the current index truncation length and stores the current index truncation length associated with the service group identifier in the configuration center; finally, it extracts the number of bytes corresponding to the current index truncation length from the beginning position of the full hash sequence to generate a truncated index key value.
[0027] Storage unit 13 is used to store truncated index key values, full hash sequences, encrypted data, and business data in the index column, backup column, identification data column, and business data column of the database, respectively.
[0028] In some embodiments, the storage unit 13 stores the truncated index key value in the index column of the database according to the column field partitioning rules of the database. This index column will build a database index to support subsequent fast data retrieval; stores the full hash sequence in the backup column of the database. This backup column does not build a database index and only serves as a cold backup data source for subsequent index reconstruction; stores encrypted data in the identification data column of the database to ensure the privacy and security of customer identification data; and stores business data in the business data column of the database to complete the direct persistence of business data.
[0029] The monitoring unit 14 is used to maintain a hash collision count matrix corresponding to the service group identifier in memory, update the hash collision count matrix according to the truncated index key value, and determine the hash distribution uniformity used to characterize the degree of collision of the truncated index key value according to the hash collision count matrix. The hash collision count matrix is used to count the frequency of occurrence of different truncated index key values.
[0030] In some embodiments, the monitoring unit 14 initializes a hash collision counting matrix corresponding to the service group identifier in memory, configures a preset hash function with independent seed parameters for each row of the hash collision counting matrix, sets the initial count value of all counting units in the hash collision counting matrix to zero, and the hash collision counting matrix is a two-dimensional counting matrix, the number of its rows corresponding to the number of preset hash functions, and the number of its columns corresponding to the number of hash buckets. The monitoring unit 14 asynchronously obtains the truncated index key value transmitted by the storage unit 13 from the message queue, and maps the truncated index key value to the corresponding counting unit in the hash collision counting matrix through multiple preset hash functions, and accumulates and updates the count value of the counting unit. The monitoring unit 14 constructs a count value set based on the count value of each column in any row of the hash collision counting matrix according to a preset period, or constructs a count value set based on the average of the count values of each column in any multiple rows of the hash collision counting matrix; extracts non-zero values from the count value set to obtain a non-zero count value set; determines the total count of all count values in the non-zero count value set; and determines the hash distribution uniformity based on the sum of the squares of the proportion of each count value in the non-zero count value set to the total count.
[0031] The adjustment unit 15 is used to adjust the current index truncation length when the hash distribution uniformity is less than the preset skew threshold and the current index truncation length is less than the preset maximum truncation length, and to reconstruct and update the truncated index key value of the index column based on the adjusted index truncation length and the full hash sequence stored in the backup column.
[0032] In some embodiments, the adjustment unit 15 reads a preset skew threshold and a preset maximum truncation length stored in the configuration center. The preset skew threshold is used to determine whether the hash collision level has reached a critical state requiring index length adjustment, and the preset maximum truncation length is used to limit the upper limit of the index truncation length. When the hash distribution uniformity is less than the preset skew threshold and the current index truncation length is less than the preset maximum truncation length, the adjustment unit 15 sets the adjustment status flag corresponding to the service group identifier to the enabled state and increases the current index truncation length by a preset step to obtain the adjusted index truncation length. When the current index truncation length reaches the preset maximum truncation length, the adjustment unit 15 generates and sends an alarm message, which includes the service group identifier. After updating the index columns of all existing data corresponding to the service group identifier, the adjustment unit 15 sets the adjustment status flag to the disabled state.
[0033] Please see Figure 2 The diagram illustrates a flowchart of a customer privacy data storage method for a customer information management platform according to an embodiment of the present invention, including the following steps S201-S205.
[0034] S201. Obtain the record to be stored and obtain the business group identifier to which the identification data belongs.
[0035] As one possible implementation, a customer data storage request initiated by the business side is received through a data access interface, the request message is parsed and processed, and the complete record to be stored is extracted.
[0036] The records to be stored include identification data and business data associated with the identification data.
[0037] It should be noted that identification data refers to information used to uniquely identify or locate a customer, such as sensitive information like ID card numbers, passport numbers, mobile phone numbers, or membership accounts. This identification data will be used to build indexes and as query conditions. Business data refers to information associated with the identification data that needs to be returned during queries, such as order numbers, transaction amounts, travel dates, and ticketing records—non-sensitive or, while sensitive, business information not used as the basis for index construction. By storing identification data and business data together in the same record, the customer information management platform establishes a mapping relationship from identification data to business data, enabling subsequent retrieval of the corresponding business data based on the identification data.
[0038] Business group identifiers are used to distinguish different tenants or different business lines, such as enterprise identifiers or department identifiers. Business group identifiers can be obtained in two ways:
[0039] First, the business group identifier is explicitly carried in the data storage request by the business unit that initiates the customer data storage request. For example, when the customer account opening system initiates an account opening data storage request, it attaches the business group identifier "BIZ_OPEN_ACCOUNT" in the request header or request body. The customer information management platform parses and extracts this business group identifier from the data storage request.
[0040] Second, when a customer data storage request does not include a business group identifier, the customer information management platform automatically determines the business group identifier to which the identified data belongs by performing mapping and matching based on contextual information such as the source system identifier of the data storage request, the request interface path, or the field type of the identified data, using preset routing rules. For example, when the identified data is a bank card number, the identified data is mapped to the business group identifier "BIZ_BANK_CARD" according to the preset field type and business group mapping table.
[0041] It should be noted that different business group identifiers correspond to independent encryption keys and hash keys. These keys are used for encryption and hash calculation of the identification data, respectively, thus isolating the privacy data encryption systems between different business groups. Furthermore, different business group identifiers correspond to independent current index truncation length configurations. The current index truncation length controls the granularity of the full hash sequence truncation, allowing each business group to independently adjust its indexing strategy based on its own data distribution characteristics, avoiding interference between different business groups due to differences in data distribution. In addition, the business group identifiers are also used to associate with independent hash collision counting matrices, enabling the customer information management platform to independently monitor the collision status of truncated index key values for each business group, providing accurate data support for subsequent dynamic adjustment of the index truncation length.
[0042] S202. Generate encrypted data and a full hash sequence corresponding to the identification data, and truncate the full hash sequence according to the current index truncation length corresponding to the business group identifier to generate truncated index key values for building database indexes.
[0043] As one possible implementation, firstly, based on the encryption key corresponding to the business group identifier, the identification data is encrypted using a preset symmetric encryption algorithm to generate encrypted data.
[0044] In some embodiments, the encryption key is uniformly generated and managed by the key management module of the customer information management platform. Different business group identifiers correspond to different encryption keys, thereby achieving isolation of the encryption systems between different business groups. The preset symmetric encryption algorithm is AES-256 (Advanced Encryption Standard, 256-bit key length). After obtaining the corresponding encryption key by initiating a key query request to the key management module through the business group identifier, the original byte sequence of the identified data is used as plaintext input, and the AES-256 algorithm is called to perform encryption operations, outputting a ciphertext byte sequence of fixed length, i.e., the encrypted data.
[0045] It should be noted that the encryption key is stored securely in the key management module. The encryption key is dynamically obtained during each encryption process, rather than being persistently stored locally, in order to reduce the risk of key leakage.
[0046] Furthermore, based on the hash key corresponding to the business group identifier, the identification data is hashed using a preset hash algorithm to obtain a full hash sequence with a preset fixed length.
[0047] In some embodiments, the hash key is also uniformly generated and managed by the key management module. Different business group identifiers correspond to different hash keys. The preset hash algorithm is HMAC-SHA256 (Hash-based Message Authentication Code, based on the SHA-256 hash function). This algorithm introduces a key mechanism, namely a hash key, on the basis of the standard SHA-256 hash function, so that the same identification data produces different hash outputs under different hash keys, thereby preventing attackers from reverse engineering the original identification data by pre-compiling rainbow tables. After obtaining the corresponding hash key from the key management module through the business group identifier, the original byte sequence of the identification data is used as the message input, and the hash key is used as the key input. The HMAC-SHA256 algorithm is called to perform hash operation, and the output is a fixed-length hash byte sequence of 256 bits (i.e., 32 bytes), which is the full hash sequence.
[0048] It should be noted that the preset fixed length is the inherent output length of the hash algorithm, which is 32 bytes in this embodiment; if other hash algorithms (such as HMAC-SHA512) are used, the preset fixed length will be adjusted accordingly to the inherent output length of the algorithm (such as 64 bytes).
[0049] Simultaneously, the configuration center is accessed to query the index truncation length associated with the service group identifier, and the queried index truncation length is determined as the current index truncation length. If the configuration center does not find the record corresponding to the service group identifier (i.e., the service group is being accessed for the first time), the preset baseline length is determined as the current index truncation length, and the current index truncation length is associated with the service group identifier and stored in the configuration center. For example, the preset baseline length can be set to 4 bytes (32 bits), which ensures a certain degree of distinguishability while minimizing index storage overhead.
[0050] Then, the customer information management platform truncates the full hash sequence according to the current index truncation length.
[0051] In some embodiments, starting from the beginning of the full hash sequence (i.e., the first byte), a number of bytes corresponding to the current index truncation length are extracted to generate a truncated index key value. For example, when the current index truncation length is 4 bytes, the first 4 bytes are truncated from the beginning of the full hash sequence as the truncated index key value; when the current index truncation length is adjusted to 6 bytes, the first 6 bytes are truncated from the beginning of the full hash sequence as the truncated index key value.
[0052] It should be noted that the truncation length of index key values directly affects the distinguishability and storage efficiency of database indexes. The shorter the truncation length, the smaller the value space of the index key, and the higher the probability that different identified data will produce the same truncated index key value (i.e., hash collision), leading to a decrease in index distinguishability and deterioration in query performance. The longer the truncation length, the larger the value space of the index key, and the lower the probability of hash collision, but the storage space occupied by the index column increases accordingly.
[0053] S203. Store the truncated index key value, the full hash sequence, the encrypted data, and the business data in the index column, backup column, identification data column, and business data column of the database, respectively.
[0054] As one possible implementation, after obtaining the truncated index key value, full hash sequence, encrypted data, and business data based on the above steps, these four data items are written into four predefined columns in the database for persistent storage.
[0055] It should be noted that the database is a relational database (such as MySQL, PostgreSQL, etc.). The database predefines a four-column storage structure for each business group identifier: an index column, a backup column, an identification data column, and a business data column. Each column performs a different storage function, and the four data items are linked through the primary key of the same row record. This allows for unified management of index building, data backup, privacy protection, and business information delivery within a single record.
[0056] The truncated index key value is stored in the index column of the database. The index column is the core index field in the database used to support equality retrieval. The database builds a B+ tree index or hash index on the index column, enabling the application to perform efficient equality matching retrieval based on the truncated index key value when initiating a data query request. The length of the truncated index key value is determined by the current index truncation length, and its value space depends on the number of bytes truncated. In this embodiment, when the current index truncation length is 4 bytes, the value space of the truncated index key value is... There are several possible values; when the current index truncation length is adjusted to 6 bytes, the value space of the truncated index key value expands to... There are several possible values. The truncated index key value stored in the index column is not the original identification data itself, but a derived value after hash calculation and truncation. Therefore, even if the index column data in the database is illegally obtained, the attacker cannot directly restore the original identification data, thus achieving a certain degree of privacy protection at the index level.
[0057] The full hash sequence is stored in a backup column in the database. The backup column is a dedicated field in the database used to store the complete hash output. Its content is a complete hash byte sequence with a preset fixed length, calculated using a preset hash algorithm. In this embodiment, the length of the full hash sequence is 32 bytes. The existence of the backup column is the key data foundation for this solution to dynamically adjust the index truncation length. When it is determined that the current index truncation length needs to be increased, there is no need to re-acquire the original identification data and perform hash calculations. Instead, the stored full hash sequence is directly read from the backup column, and data corresponding to the adjusted index truncation length is extracted from the beginning position of the full hash sequence. A new truncated index key value is generated and written to the index column for replacement. This design avoids repeated access to and repeated hash calculations of the original identification data during index reconstruction, significantly reducing the computational overhead and time cost of index adjustment operations. Especially in business scenarios with large amounts of existing data, it can effectively shorten the completion time of index reconstruction and reduce the impact window on normal business queries.
[0058] Encrypted data is stored in the identification data column of the database. The identification data column is a dedicated field in the database used to store the encrypted identification data; its content is the ciphertext byte sequence obtained by encrypting the original identification data. There is no reversible derivation relationship between the encrypted data in the identification data column and the truncated index key value in the index column. The truncated index key value comes from the truncation result of a hash operation, while the encrypted data comes from the output of a symmetric encryption operation. They are generated independently based on different cryptographic principles. This design ensures that even if an attacker simultaneously obtains the truncated index key value in the index column and the encrypted data in the identification data column, they cannot deduce the original identification data through the relationship between the two, thus achieving defense-in-depth at the storage level.
[0059] Business data is stored in business data columns within the database. These columns are fields used to store business attribute information, containing business attribute information extracted from data storage requests and associated with the identification data. Examples of such information include customer name, account opening date, account balance, membership level, and transaction record summary. Business data is stored directly in plaintext within these columns without encryption, allowing the business end to directly access and use it after authentication, avoiding the performance overhead of frequent encryption / decryption operations. It's important to note that the business data itself does not contain sensitive identification information directly linked to a specific individual. Its security relies on the protection mechanism of the encrypted data in the aforementioned identification data columns. Only after a query is performed through a dual verification process—namely, a truncated index key value retrieval and a comparison of the decrypted encrypted data—can the querying party obtain the business data associated with the specific identification data, thus indirectly achieving access control over the business data.
[0060] In some embodiments, a transaction mechanism is employed to ensure the atomicity of the write operation when performing the write operations of the above four data items. The truncated index key value, full hash sequence, encrypted data, and business data are treated as four components of the same logical record and written within the same database transaction. If the write operation fails for any column, the entire transaction is rolled back to ensure the consistency of the four columns and avoid inconsistencies where some columns are written while others are not. For example, if the index column and backup column have been successfully written, but the write to the identification data column fails due to insufficient disk space, the entire transaction will be rolled back, and the data already written to the index column and backup column will be undone, thus ensuring the integrity and consistency of the four columns in the same row record.
[0061] S204. Maintain a hash collision count matrix corresponding to the business group identifier in memory, update the hash collision count matrix according to the truncated index key value, and determine the hash distribution uniformity used to characterize the degree of collision of the truncated index key value according to the hash collision count matrix.
[0062] The hash collision count matrix is used to count the frequency of occurrence of different truncated index key values.
[0063] As one possible implementation, the hash collision counting matrix corresponding to the business group identifier is first initialized in the application server's memory. The hash collision counting matrix is a two-dimensional counting matrix, where the number of rows corresponds to the number of preset hash functions, and the number of columns corresponds to the number of hash buckets.
[0064] In some embodiments, the preset number of hash functions is 1. The number of hash buckets is Therefore, the dimension of the hash collision counting matrix is OK The column contains a total of Counting units. Preset number of hash functions. and the number of hash buckets During the initialization phase, the operations and maintenance personnel configure the system based on the volume and accuracy requirements of the business data. For example, in this embodiment, we take... , The hash collision counting matrix contains 4 rows and 1024 columns, totaling 4096 counting units.
[0065] It should be noted that upon first receiving the truncated index key value corresponding to a certain service group identifier, a hash collision count matrix corresponding to that service group identifier is initialized in memory. Each row of the hash collision count matrix is configured with a preset hash function with an independent seed parameter, i.e., the [number]th row of the hash collision count matrix. OK( The corresponding preset hash function It has a unique seed parameter that is different from other rows. This ensures that the preset hash functions for different rows map the same truncated index key value to different hash bucket locations, thereby improving the robustness and accuracy of hash collision statistics through independent mapping across multiple rows. After configuring the preset hash functions, the initial count value of all counting units within the hash collision counting matrix is set to zero.
[0066] The lifecycle of the hash collision counting matrix is consistent with the operational cycle of the customer information management platform. During the normal operation of the customer information management platform, the hash collision counting matrix always resides in memory to support real-time monitoring of the distribution of truncated index key values. The hash collision counting matrix is isolated at the business group identifier level. Different business group identifiers correspond to different hash collision counting matrices, which do not affect each other, thus ensuring that the hash collision statistics of each business group are calculated and judged independently.
[0067] Furthermore, in this embodiment, after the truncated index key value is generated and written to the database, it is simultaneously sent to a message queue for asynchronous transmission. The customer information management platform asynchronously retrieves the truncated index key value from the message queue to avoid synchronous blocking of the main data storage process and ensure that the response latency of the data write operation is not affected by hash collision statistics processing.
[0068] It should be noted that the message queue can be a distributed message queue such as Kafka or RabbitMQ, or it can be a concurrent linked list queue in memory (ConcurrentLinkedQueue). This is used to decouple the writing module and the monitoring module of the customer information management platform, so as to avoid the monitoring process blocking the main business writing process.
[0069] Furthermore, the truncated index key value is mapped to the corresponding counting unit in the hash collision counting matrix through multiple preset hash functions, and the count value of the counting unit is accumulated and updated.
[0070] In some embodiments, for the hash collision count matrix of the first OK( The truncated index key value is taken as input, and the corresponding preset hash function is called. Perform mapping calculations to obtain the mapping position, denoted as... Its formula is expressed as In the formula, To truncate the index key value at the 1st... The hash bucket index mapped to in the row has a value range of 1. , For the first The preset hash function corresponding to the row To truncate the index key value, Indicates based on the first The preset hash function corresponding to the row performs calculations on the truncated index key value. The modulo operator, The number of hash buckets, whose value is not zero, is used to determine the mapping position. Then, the corresponding number in the hash collision count matrix... Line number The count value of the column's counting cell is incremented by 1, that is... , The number of collisions in the hash collision count matrix Line number The count value of the column's counting unit. This is performed once for each row (i.e., for each preset hash function), so for each truncated index key value obtained, the hash collision count matrix will contain... The count value of each counting unit is accumulated and updated.
[0071] Understandably, based on the aforementioned update mechanism, the count value of each counting unit in the hash collision counting matrix reflects the frequency of occurrence of the truncated index key value under the corresponding preset hash function and the corresponding hash bucket combination. When multiple different identification data produce the same truncated index key value after hash calculation and truncation (i.e., a hash collision occurs), these truncated index key values will be mapped to the same counting unit, making the count value of that counting unit significantly higher than other counting units. Therefore, the distribution of the count values of each counting unit in the hash collision counting matrix can intuitively reflect the degree of collision and the uniformity of distribution of truncated index key values in the hash bucket space.
[0072] Furthermore, the customer information management platform constructs a set of count values based on the count values of each column in any row of the hash collision counting matrix, or based on the average of the count values of each column in any multiple rows of the hash collision counting matrix, according to a preset period.
[0073] For example, the customer information management platform triggers a snapshot analysis at a preset period (such as every 1000 truncated index key values processed or every minute). When using a single-row snapshot, the customer information management platform directly extracts the count values of each column in that row to construct a count value set; when using a multi-row snapshot, the customer information management platform calculates the average of the count values at the same column position in multiple rows to construct a count value set.
[0074] Furthermore, the customer information management platform extracts the non-zero values from the set of count values to obtain a set of non-zero count values, and determines the total count of all count values in the set of non-zero count values.
[0075] In some embodiments, the set of non-zero count values is denoted as ,in, The number of non-zero count values. ( ) is the first There are 10 non-zero count values. The sum of all count values in the set of non-zero count values is denoted as . , .
[0076] Furthermore, the customer information management platform determines the hash distribution uniformity based on the sum of the squares of the proportions of each count value in the set of non-zero count values to the total count.
[0077] In some embodiments, the formula for determining the uniformity of hash distribution is as follows: In the formula, The hash distribution uniformity has a value range of 1. , The first in the set of non-zero count values One non-zero count value, It is the sum of all counts in the set of non-zero counts. The number of non-zero count values. For the first The proportion of non-zero count values to the total count sum It is the sum of the squares of the proportions of all non-zero counts in the set of non-zero counts.
[0078] It needs to be explained that, in When the hash value equals zero, it means there is no set of non-zero count values, so it does not participate in the above calculations, and the hash is directly assigned to the uniformity of the hash values. Set to 1.
[0079] Specifically, when the truncated index key values are completely uniformly distributed in the hash bucket space, that is, the count values in each mapped hash bucket are equal ( ),at this time Reaching its maximum value indicates that the distribution of truncated index key values is most uniform, and the degree of hash collision is lowest. When truncated index key values are extremely concentrated in the hash bucket space, that is, all truncated index key values map to the same hash bucket (…), the hash value is considered to be at its maximum. ),at this time The minimum value indicates that the distribution of truncated index key values is most concentrated, and the degree of hash collision is highest.
[0080] Therefore, hash distribution uniformity The closer the value is This indicates that the more uniform the distribution of truncated index key values, the higher the index's distinguishability, and the better the performance of database equality retrieval; hash distribution uniformity The closer the value is to 0, the more concentrated the distribution of the truncated index key values is, the lower the index's distinguishability, the more candidate records need to be compared during database equality retrieval, and the worse the query performance.
[0081] In some embodiments, after performing the step of determining the hash distribution uniformity, the customer information management platform performs the state management step of the hash collision counting matrix.
[0082] The customer information management platform manages the state of the hash collision counting matrix, including: after determining the uniformity of hash distribution, decaying the count values of all counting units in the hash collision counting matrix; resetting the hash collision counting matrix to zero if the index truncation length adjustment is not triggered for a preset duration; and resetting the hash collision counting matrix to zero after the index truncation length adjustment is triggered.
[0083] First, after determining the uniformity of the hash distribution, the customer information management platform performs a decay process on the count values of all counting units in the hash collision counting matrix. This decay process prevents counter overflow and ensures that the monitored metrics reflect recent rather than historical data distribution. Optionally, the customer information management platform divides the count value of each counting unit in the hash collision counting matrix by 2 and then rounds it down. This operation preserves the relative distribution characteristics of the data while reducing the absolute value, effectively preventing the counter from overflowing after prolonged operation. After the decay process, the hash collision counting matrix continues to be used for statistical updates in the next cycle.
[0084] Secondly, if the customer information management platform does not trigger an index truncation length adjustment within a preset time period, it will reset the hash collision count matrix to zero. The preset time period can be configured according to business scenarios, for example, set to 1 hour. If the customer information management platform never triggers an index truncation length adjustment operation within the preset time period, it indicates that the current index configuration has remained stable for a relatively long period, and the count values of all counting units in the hash collision count matrix will be reset to zero. This operation eliminates interference from historical accumulated data, ensuring that the statistical data accurately reflects the index key value distribution during the current period.
[0085] Secondly, after triggering the index truncation length adjustment, the customer information management platform resets the hash collision count matrix to zero. When the customer information management platform determines that the adjustment conditions are met and executes the index truncation length adjustment operation (e.g., increasing the current index truncation length by a preset step), the count values in the hash collision count matrix based on the old index truncation length are no longer applicable to distribution monitoring under the new index truncation length. Therefore, after completing the index truncation length adjustment, the customer information management platform immediately resets the count values of all counting units in the hash collision count matrix to zero. After the reset, the hash collision count matrix begins to re-count the truncated index key values generated based on the new index truncation length, providing an accurate data foundation for subsequent index health monitoring.
[0086] Understandably, by managing the state of the hash collision count matrix, periodic decay and selective reset of the hash collision count matrix are achieved. This prevents counter overflow and ensures that the monitoring indicators always reflect the recent distribution of index key values, providing accurate and real-time decision-making basis for adaptive index adjustments. The state-managed hash collision count matrix will continue to be used for the next period's truncated index key value statistical update, forming a complete monitoring closed loop.
[0087] S205. When the hash distribution uniformity is less than the preset skew threshold and the current index truncation length is less than the preset maximum truncation length, adjust the current index truncation length and reconstruct and update the truncated index key value of the index column based on the adjusted index truncation length and the full hash sequence stored in the backup column.
[0088] As one possible implementation, the customer information management platform first reads the preset skew threshold and preset maximum truncation length stored in the configuration center.
[0089] It should be noted that the preset skew threshold is used to determine whether the degree of hash collision has reached the critical state requiring adjustment of the index length; for example, it can be set to 0.3. When the hash distribution uniformity is lower than this threshold, it indicates that the current index truncation length has caused severe hash collisions, the index key value distribution is excessively concentrated, the candidate record set size is too large during retrieval, and the retrieval efficiency is affected. The preset maximum truncation length is used to limit the upper limit of the index truncation length; for example, it can be set to 16 bytes to prevent the index from growing indefinitely, leading to excessive storage overhead or algorithm deadlock.
[0090] Furthermore, the hash distribution uniformity is compared with a preset skew threshold, and the current index truncation length is compared with a preset maximum truncation length. If the hash distribution uniformity is less than the preset skew threshold and the current index truncation length is less than the preset maximum truncation length, the customer information management platform determines that the adjustment conditions are met and initiates the index reconstruction process.
[0091] In some embodiments, the adjustment status flag corresponding to the business group identifier is first set to the enabled state. Then, the customer information management platform increases the current index truncation length by a preset step size to obtain the adjusted index truncation length, and stores the adjusted index truncation length in the configuration center in association with the business group identifier for subsequent write and query operations. The adjustment status flag indicates whether the system is currently in the index reconstruction phase, and the preset step size can be set to 1 byte, meaning that the index truncation length is increased by 1 byte with each adjustment.
[0092] It should be noted that after the customer information management platform adjusts the status flag to the enabled state, the write module will execute dual-write logic (i.e., generate and store both old and new index key values simultaneously) when it receives a new data storage request. The query module of the customer information management platform will execute dual-mode retrieval (i.e., retrieve records corresponding to both old and new index key values simultaneously) when executing a data query request, ensuring that business is not interrupted and data is not lost during the reconstruction.
[0093] In some embodiments, when the adjustment state is marked as enabled, in response to a new data storage request, a first truncated index key value (old index key) corresponding to the current index truncation length is generated, and a second truncated index key value (new index key) corresponding to the adjusted index truncation length is generated; then the first truncated index key value and the second truncated index key value are written together into the index column of the database.
[0094] Furthermore, the customer information management platform initiates a background reconstruction task to re-index and reconstruct the existing data. The background reconstruction task runs with low priority and employs a batch submission mechanism to minimize the impact on online business writes.
[0095] In some embodiments, the full hash sequence stored in the backup column of the database is read, and data corresponding to the adjusted index truncation length is extracted from the full hash sequence to generate a new truncated index key value. Since the backup column stores the full hash sequence, the customer information management platform does not need to read the encrypted data in the identification data column, nor does it need to decrypt the encrypted data, completely avoiding CPU-intensive decryption operations. The generation of the new index key value can be completed simply by truncating bytes in memory, greatly reducing the reconstruction cost.
[0096] Furthermore, the customer information management platform writes the newly generated truncated index key value into the index column of the database, replacing the original truncated index key value. For each record in the existing data, the customer information management platform performs the above operations of reading the backup column, truncating and generating the new index, and replacing the old index. Optionally, the customer information management platform adopts a batch commit method, for example, committing a transaction every 1000 records processed, to reduce the number of database interactions and minimize the impact on online business.
[0097] After updating the index columns of all existing data corresponding to the business group identifier, the customer information management platform sets the adjustment status flag to the off state. With the adjustment status flag off, the write module stops dual-write logic and reverts to generating only a single index key value corresponding to the new index truncation length; the query module stops dual-modal retrieval and reverts to retrieving only a single query key value corresponding to the new index truncation length. Additionally, the customer information management platform performs index cleanup, deleting all truncated index key values in the database's index columns that correspond to the current index truncation length, to free up storage space and ensure retrieval accuracy in a stable state.
[0098] If, during the adjustment process, the current index truncation length has reached the preset maximum truncation length, but the hash distribution uniformity is still less than the preset skew threshold, the circuit breaker protection mechanism will be triggered. The customer information management platform will generate and send an alarm message, which includes a business group identifier. This alarm message is used to notify operations and maintenance personnel that there may be a high degree of repetition in the data source itself (such as a large number of null values or default values), requiring manual intervention to prevent the algorithm from looping infinitely.
[0099] In one design, the customer information management platform responds to a data query request and performs a retrieval operation, including: obtaining the data to be queried; when the adjustment status is marked as enabled, generating a first query key value corresponding to the current index truncation length based on the data to be queried, and generating a second query key value corresponding to the adjusted index truncation length; when the adjustment status is marked as disabled, generating a third query key value corresponding to the current index truncation length based on the data to be queried; performing a union search in the database containing the first and second query key values, or performing an equality search in the database on the third query key value, to obtain a candidate record set; decrypting the encrypted data in the candidate record set and comparing it with the data to be queried to determine the target query result.
[0100] In some embodiments, the customer information management platform first receives a data query request from the business end, which carries identification data to be queried. This identification data is information of the same type as the identification data in the records to be stored, such as an ID card number, mobile phone number, or membership account. The customer information management platform parses and extracts the identification data to be queried from the data query request, using it as input for subsequent searches.
[0101] Secondly, the customer information management platform accesses the configuration center to read the adjustment status flag associated with the business group identifier and the current index truncation length. The adjustment status flag is used to indicate whether the current stage is index reconstruction. Based on the value of the adjustment status flag (e.g., a value of 1 for the enabled state indicates that the index is being reconstructed and truncated, and a value of 0 for the disabled state indicates that the index is not being reconstructed), the branch processing logic is executed.
[0102] If the adjustment status is marked as enabled, it indicates that index reconstruction is currently underway. The database contains data with two index lengths: one part is existing data that has not yet been processed by the background reconstruction task, whose index key values are generated based on the current index truncation length; the other part is existing data that has already been reconstructed, as well as data newly written during the transition period, whose index key values are generated based on the adjusted index truncation length. To ensure no data is missed, the customer information management platform generates two query key values based on the data to be queried: the first query key value is generated by truncating the data from the beginning of the full hash sequence of the data to be queried, corresponding to the current index truncation length; the second query key value is generated by truncating the data from the beginning of the full hash sequence, corresponding to the adjusted index truncation length. The customer information management platform then performs a union retrieval in the database, containing both the first and second query key values, i.e., an OR query in Structured Query Language (SQL), to obtain a candidate record set.
[0103] If the adjustment status is set to "off," it indicates that index reconstruction has been completed, and all index key values in the database are generated based on the current index truncation length. The customer information management platform generates a query key value based on the data to be queried: it extracts a number of bytes corresponding to the current index truncation length from the beginning of the full hash sequence of the data to be queried, generating a third query key value. The customer information management platform then performs an equality search on the third query key value in the database to obtain a candidate record set.
[0104] Through the above retrieval operations, the customer information management platform utilizes the database's B+ tree index to quickly locate records, narrowing the search scope from the entire table to a candidate record set. The candidate record set contains records with matching index key values; however, due to potential collisions caused by hash truncation, the candidate record set may contain false positives (i.e., records with the same index key value but different identification data).
[0105] Finally, the customer information management platform decrypts the encrypted data in the candidate record set and compares it with the data to be queried to determine the target query result.
[0106] In some embodiments, the customer information management platform traverses each record in the candidate record set, and decrypts the encrypted data in the record using the same symmetric encryption algorithm as when it was stored, based on the encryption key corresponding to the business group identifier, to obtain candidate identification data. The customer information management platform performs a byte-by-byte comparison between the candidate identification data and the identification data to be queried. If the two are completely identical, the record is determined to be the target data, and the business data in the corresponding record is identified as the target query result; if the two are inconsistent, the record is determined to be false positive data generated by hash truncation and is excluded in this query. This decryption and comparison process is executed in the secure memory area of the application server, using the efficiency of in-memory computing to compensate for the uncertainty caused by hash truncation in the storage layer, thus achieving accurate retrieval of encrypted data.
[0107] Based on this, the customer information management platform enables data consistency querying during the dynamic evolution of the index. Whether in a stable state or a transitional state of index reconstruction, it can accurately return business data that matches the data to be queried, ensuring the integrity and reliability of the retrieval function.
[0108] It is understandable that in the customer privacy data storage method of the customer information management platform provided in the embodiments of the present invention, by storing identification data and business data separately, and generating truncated index key values based on the identification data to build a database index during writing, while storing the full hash sequence in a backup column as the entropy source for subsequent reconstruction, the technical problems of encrypted data not being able to be directly indexed and the low write performance caused by the full hash index key being too long are effectively solved. Building upon this foundation, this invention maintains a hash collision counting matrix in memory, continuously monitors the frequency of truncated index key values, and determines the uniformity of hash distribution. This achieves zero I / O overhead awareness of index collision levels, enabling accurate identification of index health without business oversight. When the hash distribution uniformity falls below a preset skew threshold and the current index truncation length has not yet reached its upper limit, the customer information management platform automatically adjusts the index truncation length and reconstructs and updates the truncated index key values of the index column based on the full hash sequence stored in the backup column. The entire reconstruction process eliminates the need to read and decrypt the ciphertext in the identification data column, completely avoiding CPU-intensive decryption operations and significantly reducing the resource consumption and time cost of index adjustment. Furthermore, by adjusting the status flag switch, the collaborative switching of the write and query modules is achieved, ensuring uninterrupted business operations and no data loss during index reconstruction. Therefore, this invention achieves a dynamic balance between high-concurrency write throughput and accurate retrieval efficiency while ensuring the security of encrypted storage of customer privacy data. It also provides low-cost index adaptive evolution capabilities, significantly improving the stability and operational efficiency of the customer information management platform when data volume surges or distribution changes.
[0109] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0110] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
Claims
1. A method for storing customer privacy data in a customer information management platform, characterized in that, The method includes: Obtain the record to be stored, the record to be stored including identification data and business data associated with the identification data, and obtain the business group identifier to which the identification data belongs; Generate encrypted data and a full hash sequence corresponding to the identification data, and truncate the full hash sequence according to the current index truncation length corresponding to the business group identifier to generate a truncated index key value for building a database index; The truncated index key value, the full hash sequence, the encrypted data, and the business data are respectively stored in the index column, backup column, identification data column, and business data column of the database; Maintain a hash collision count matrix corresponding to the service group identifier in memory. Update the hash collision count matrix according to the truncated index key value. Determine the hash distribution uniformity to characterize the degree of collision of the truncated index key value based on the hash collision count matrix. The hash collision count matrix is used to count the occurrence frequency of different truncated index key values, including: constructing a count value set based on the count values of each column in any row of the hash collision count matrix according to a preset period, or constructing the count value set based on the average of the count values of each column in any multiple rows of the hash collision count matrix; extract... The non-zero values in the count value set are used to obtain a non-zero count value set; the total count of all count values in the non-zero count value set is determined; the hash distribution uniformity is determined based on the sum of the squares of the proportions of each count value in the non-zero count value set to the total count; after determining the hash distribution uniformity, the count values of all counting units in the hash collision counting matrix are decayed; if the index truncation length adjustment is not triggered for a preset duration, the hash collision counting matrix is reset to zero; if the index truncation length adjustment is triggered, the hash collision counting matrix is reset to zero. When the hash distribution uniformity is less than a preset skew threshold and the current index truncation length is less than a preset maximum truncation length, the current index truncation length is adjusted, and based on the adjusted index truncation length and the full hash sequence stored in the backup column, the truncated index key value of the index column is reconstructed and updated. This includes: reading the preset skew threshold and the preset maximum truncation length stored in the configuration center. The preset skew threshold is used to determine whether the hash collision degree has reached a critical state requiring index length adjustment, and the preset maximum truncation length is used to limit the upper limit of the index truncation length; when the hash distribution uniformity is less than the preset skew threshold and the current index truncation length is less than a preset maximum truncation length, the current index truncation length is adjusted, and the current index truncation length is less than a preset maximum truncation length. If the length is less than the preset maximum truncation length, set the adjustment status flag corresponding to the business group identifier to the enabled state, and increase the current index truncation length by a preset step to obtain the adjusted index truncation length; read the full hash sequence stored in the backup column of the database, and extract the number of bytes corresponding to the adjusted index truncation length from the full hash sequence to generate a new truncated index key value; write the new truncated index key value into the index column of the database, replacing the original truncated index key value; after updating the index column of all existing data corresponding to the business group identifier, set the adjustment status flag to the disabled state.
2. The customer privacy data storage method of the customer information management platform according to claim 1, characterized in that, Generate encrypted data and a full hash sequence corresponding to the identification data, and truncate the full hash sequence according to the current index truncation length corresponding to the business group identifier to generate truncated index key values for constructing database indexes, including: Based on the encryption key corresponding to the service group identifier, the identification data is encrypted using a preset symmetric encryption algorithm to generate the encrypted data; Based on the hash key corresponding to the business group identifier, the identification data is hashed using a preset hash algorithm to obtain the full hash sequence with a preset fixed length; Access the configuration center, determine the index truncation length associated with the service group identifier as the current index truncation length; and if no record corresponding to the service group identifier is found in the configuration center, determine the preset base length as the current index truncation length, and associate the current index truncation length with the service group identifier and store it in the configuration center; The truncated index key value is generated by extracting the number of bytes corresponding to the current index truncation length from the beginning position of the full hash sequence.
3. The customer privacy data storage method of the customer information management platform according to claim 1, characterized in that, Maintaining a hash collision count matrix corresponding to the service group identifier in memory, including: Initialize the hash collision counting matrix corresponding to the business group identifier in memory, and configure a preset hash function with independent seed parameters for each row of the hash collision counting matrix. Set the initial count value of all counting units in the hash collision counting matrix to zero. The hash collision counting matrix is a two-dimensional counting matrix. The number of rows of the two-dimensional counting matrix corresponds to the number of preset hash functions, and the number of columns corresponds to the number of hash buckets. The truncated index key value is asynchronously obtained from the message queue, and the truncated index key value is mapped to the corresponding counting unit in the hash collision counting matrix through multiple preset hash functions. The count value of the counting unit is then accumulated and updated.
4. The customer privacy data storage method of the customer information management platform according to claim 1, characterized in that, The method further includes: When the adjustment state is marked as enabled, in response to a new data storage request, a first truncated index key value corresponding to the current index truncation length is generated, and a second truncated index key value corresponding to the adjusted index truncation length is generated. The first truncated index key value and the second truncated index key value are written together into the index column of the database.
5. The customer privacy data storage method of the customer information management platform according to claim 1, characterized in that, The method further includes: In response to a data query request, retrieve the identification data to be queried; When the adjustment state is marked as enabled, a first query key value corresponding to the current index truncation length is generated based on the data to be queried, and a second query key value corresponding to the adjusted index truncation length is generated. When the adjustment status is marked as closed, a third query key value corresponding to the current index truncation length is generated based on the identification data to be queried; Perform a union search on the database containing the first query key value and the second query key value, or perform an equality search on the database for the third query key value, to obtain a candidate record set; The encrypted data in the candidate record set is decrypted and compared with the data to be queried to determine the target query result.
6. The customer privacy data storage method of the customer information management platform according to claim 5, characterized in that, The encrypted data in the candidate record set is decrypted and compared with the data to be queried to determine the target query result, including: Traverse each record in the candidate record set, and decrypt the encrypted data in the record according to the encryption key corresponding to the service group identifier to obtain the candidate identification data; The candidate identification data is compared with the identification data to be queried, and the business data in the records where the comparison results are consistent are determined as the target query result.
Citation Information
Patent Citations
Data structure for hash operation and hash table storage and query method based on data structure
CN111625534A
File deduplication storage method based on Hash algorithm
CN122240576A