Data processing method, data query method, equipment and storage medium

By splicing and compressing the high-frequency query fields of the data to be written in the distributed database and generating target shard keys, the problem of uneven data distribution is solved, efficient data storage and query is realized, and it is suitable for distributed database systems.

CN120277072APending Publication Date: 2025-07-08BEIJING CHENGSHI WANGLIN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510405447.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Traditional shard key design leads to uneven distribution of data in distributed databases, resulting in unbalanced system load and reduced query efficiency, especially in high-frequency query scenarios, problems are more prominent.

Method used

By splicing and compressing multiple high-frequency query fields to be written to the data, logical shard identification is generated, target shard keys are dynamically generated, and target physical tables are determined based on target shard keys, so as to achieve uniform storage and efficient query of data.

Benefits of technology

It effectively avoids the problem of data tilt, improves data query efficiency, and meets the needs of multi-dimensional high-frequency query scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277072A_ABST
    Figure CN120277072A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method, a data query method, equipment and a storage medium. The method comprises the following steps: identifying a plurality of high-frequency query fields contained in to-be-written data, compressing data corresponding to the plurality of high-frequency query fields to obtain first compressed data, intercepting second compressed data serving as a logic fragment identifier from the first compressed data, and writing the second compressed data into the to-be-written data; and embedding the data corresponding to the plurality of compressed high-frequency query fields and the logic fragment identifier into a preset target fragment key structure to dynamically generate a target fragment key corresponding to the to-be-written data, and determining a logic fragment corresponding to the to-be-written data and a target physical table based on the logic fragment identifier in the target fragment key. According to the method, the to-be-written data and the target fragmentation key are written into the corresponding target physical table, so that the data are uniformly fragmented and stored into the distributed database, the problem of data skew is effectively avoided, cross-fragmentation query is avoided, and the data query efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data storage, and in particular, to a data processing method, a data query method, a device, and a storage medium. Background Art

[0002] With the continuous development of application services, the amount of data in databases has shown exponential growth. However, the limited resources of physical servers have gradually made databases face bottlenecks in terms of data carrying capacity and processing capacity. To address this challenge, the sharding and partitioning technology has been widely applied as an effective solution.

[0003] In the sharding and partitioning technology, the selection of the sharding key will directly affect the data distribution uniformity during data sharding storage and the efficiency of data sharding query. The sharding key is a field used to determine the storage location of data in a distributed database. In traditional solutions, the sharding key usually consists of a business field or a system field (such as a unique identifier, a timestamp, etc.). This design of the sharding key based on a single field can achieve distributed storage of data to a certain extent, but there are some significant problems in practical applications.

[0004] For example, during the data storage process, if the distribution characteristics of the field selected as the sharding key are uneven, it will lead to uneven data distribution in the distributed system finally stored, resulting in uneven system load and decreased query efficiency. In addition, during the data query process, if the query condition uses not the single field that generates the sharding key, but queries based on other business fields, it is necessary to scan all shards to obtain the required data, resulting in an increase in the overhead of cross-shard query and a significant reduction in query efficiency. Especially in high-frequency query scenarios, this problem will be more prominent. Summary of the Invention

[0005] Multiple aspects of this application provide a data processing method, a data query method, a device, and a storage medium, which are used to evenly shard and store the data to be written into the database, effectively avoid the data skew problem, avoid cross-shard query, and at the same time can also improve the data query efficiency to meet the requirements of multi-dimensional high-frequency query scenarios.

[0006] An embodiment of this application provides a data processing method, which is applied to a distributed database. The distributed database includes multiple physical tables, and the physical tables are obtained by logically partitioning the distributed database. The method includes:

[0007] In response to a data writing request triggered by a client for the data to be written, determine multiple high-frequency query fields corresponding to the data to be written, and screen out the data corresponding to the multiple high-frequency query fields from the data to be written;

[0008] Concatenate the data corresponding to the multiple high-frequency query fields, and perform compression processing on the concatenation result to obtain the first compressed data;

[0009] Extract the low-order field data with a preset bit width bit by bit from the first compressed data to generate the second compressed data, and use the second compressed data as the logical shard identifier;

[0010] Perform a merging process on the first compressed data and the second compressed data to determine the target shard key corresponding to the data to be written;

[0011] Query a preset mapping table to determine the target physical table corresponding to the second compressed data, where the preset mapping table stores the mapping relationship between the logical shard identifier and the physical table identifier;

[0012] Write the data to be written and the target shard key into the target physical table;

[0013] Among them, the first compressed data is used to determine the shard key containing the query data; the second compressed data is used to determine the target physical table corresponding to the query data.

[0014] An embodiment of the present application provides a data processing device located in a distributed database. The distributed database includes multiple physical tables, and the physical tables are obtained by logically sharding the distributed database. The device includes:

[0015] A response module, configured to respond to a data writing request triggered by a client for data to be written, determine multiple high-frequency query fields corresponding to the data to be written, and filter out the data corresponding to the multiple high-frequency query fields from the data to be written;

[0016] A concatenation module, configured to concatenate the data corresponding to the multiple high-frequency query fields, and perform compression processing on the concatenation result to obtain the first compressed data;

[0017] A generation module, configured to extract the low-order field data with a preset bit width bit by bit from the first compressed data to generate the second compressed data, and use the second compressed data as the logical shard identifier;

[0018] A merging module, configured to perform a merging process on the first compressed data and the second compressed data to determine the target shard key corresponding to the data to be written;

[0019] A query module, configured to query a preset mapping table to determine the target physical table corresponding to the second compressed data, where the preset mapping table stores the mapping relationship between the logical shard identifier and the physical table identifier;

[0020] A writing module, configured to write the data to be written and the target sharding key into the target physical table;

[0021] Wherein, the first compressed data is used to determine the sharding key containing the query data; the second compressed data is used to determine the target physical table corresponding to the query data.

[0022] An embodiment of the present application further provides an electronic device, including: a memory and a processor; the memory is configured to store a computer program; the processor is coupled to the memory and configured to execute the computer program to implement the steps in the data processing method provided by the embodiment of the present application.

[0023] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to implement the steps in the data processing method provided by the embodiment of the present application.

[0024] An embodiment of the present application further provides a computer program product, including computer programs / instructions, which, when executed by a processor, cause the processor to implement the steps in the data processing method provided by the embodiment of the present application.

[0025] In the data processing solution provided by the present application, by compressing the concatenation result of the data corresponding to multiple high-frequency query fields in the data to be written, the first compressed data is obtained, and the low-bit field data with a preset bit width is intercepted bit by bit from the first compressed data, and the second compressed data is obtained. The first compressed data and the second compressed data are combined to obtain the target sharding key corresponding to the data to be written. In this way, the compressed data of the data corresponding to multiple high-frequency query fields is embedded into the target sharding key corresponding to the data to be written as a combined query fragment. Subsequently, the target sharding key corresponding to the query data can be quickly located based on the data corresponding to the high-frequency query fields, so as to quickly locate the target physical table corresponding to the query data based on the target sharding key, which improves the query efficiency and can meet the requirements of multi-dimensional high-frequency query scenarios. After determining the target sharding key corresponding to the data to be written, the logical sharding identifier can be determined based on the second compressed data in the target sharding key, and then the target physical table corresponding to the logical sharding identifier can be determined by querying the preset mapping table. The data to be written and the target sharding key are written into the target physical table. Since the target sharding key is generated based on the combined result of the first compressed data and the second compressed data, it is possible to avoid data distribution unevenness caused by the uniform distribution of high-frequency query field data, so as to achieve uniform sharding storage of the data to be written into the database and effectively avoid the data skew problem.

[0026] An embodiment of the present application provides a data query method, which is applied to a distributed database. The distributed database includes multiple physical tables, and the physical tables are obtained by logically sharding the distributed database. The method includes:

[0027] Receiving a query request including query data sent by a client, and identifying a target field corresponding to the query data in the distributed database;

[0028] In response to the determination result that the target field is not a sharding key field, performing compression processing on the query data to obtain first compressed sub-data;

[0029] Querying, from a reverse index dictionary, first compressed data corresponding to the first compressed sub-data. The reverse index field stores compressed sub-data corresponding to query data and compressed data corresponding to a concatenation result including the query data. The first compressed data is obtained by performing compression processing on the concatenation result, and the concatenation result is obtained by concatenating data corresponding to multiple high-frequency query fields. The data corresponding to the multiple high-frequency query fields includes the query data;

[0030] Based on the first compressed data, determining a target sharding key including the query data. The target sharding key is obtained by merging the first compressed data and second compressed data. The second compressed data is obtained by bitwise intercepting low-order field data with a preset bit width from the first compressed data;

[0031] Based on the logical sharding identifier in the target sharding key, determining a first target physical table corresponding to the query data;

[0032] Reading target data corresponding to the query data from the first target physical table.

[0033] An embodiment of the present application provides a data query device, which is located in a distributed database. The distributed database includes multiple physical tables, and the physical tables are obtained by logically sharding the distributed database. The device includes:

[0034] An identification module, configured to receive a query request including query data sent by a client, and identify a target field corresponding to the query data in the distributed database;

[0035] A compression module, configured to, in response to the determination result that the target field is not a sharding key field, perform compression processing on the query data to obtain first compressed sub-data;

[0036] A query module, configured to query, from a reverse index dictionary, first compressed data corresponding to the first compressed sub-data, where the reverse index field stores compressed sub-data corresponding to query data and compressed data corresponding to a concatenation result including the query data, the first compressed data is obtained by compressing the concatenation result, the concatenation result is obtained by concatenating data corresponding to multiple high-frequency query fields, and the data corresponding to the multiple high-frequency query fields includes the query data;

[0037] A first determination module, configured to determine, based on the first compressed data, a target shard key including the query data, where the target shard key is obtained by merging the first compressed data and second compressed data, and the second compressed data is obtained by bitwise intercepting low-order field data with a preset bit width from the first compressed data;

[0038] A second determination module, configured to determine, based on a logical shard identifier in the target shard key, a first target physical table corresponding to the query data;

[0039] A reading module, configured to read target data corresponding to the query data from the first target physical table.

[0040] An embodiment of the present application further provides an electronic device, including: a memory and a processor; the memory is configured to store a computer program; the processor is coupled to the memory and is configured to execute the computer program to implement the steps in the data query method provided by the embodiment of the present application.

[0041] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to implement the steps in the data query method provided by the embodiment of the present application.

[0042] An embodiment of the present application further provides a computer program product, including a computer program / instructions, which, when executed by a processor, causes the processor to implement the steps in the data query method provided by the embodiment of the present application.

[0043] In the data query solution provided by the embodiments of the present application, it is applied to a distributed database, which includes multiple physical tables. The physical tables are obtained by logically partitioning the distributed database. Moreover, the multiple physical tables include a first target physical table. In the scenario of database table partitioning, after receiving a query request containing query data sent by a client, the target field corresponding to the query data in the distributed database can be identified. And it is determined whether the target field is a sharding key field. If the target field corresponding to the query data is not a sharding key field, in response to the determination result that the target field is not a sharding key field, the query data is compressed to obtain a first compressed sub-data. Then, the first compressed data corresponding to the first compressed sub-data is queried from the reverse index dictionary. Wherein, the reverse index field stores the compressed sub-data corresponding to the query data and the compressed data corresponding to the concatenation result containing the query data. The first compressed data is obtained by compressing the concatenation result. The concatenation result is obtained by concatenating the data corresponding to multiple high-frequency query fields, and the data corresponding to the multiple high-frequency query fields includes the query data. Then, based on the first compressed data, the target sharding key containing the query data is determined. Wherein, the target sharding key is obtained by merging the first compressed data and the second compressed data. The second compressed data is obtained by bitwise intercepting the low-order field data with a preset bit width from the first compressed data. Then, based on the logical sharding identifier in the target sharding key, the first target physical table corresponding to the query data is determined, and the target data corresponding to the query data is read from the first target physical table.

[0044] In the above solution, when it is determined that the query data is not a sharding key value, the query data is compressed to obtain a first compressed sub-data, and the first compressed data corresponding to the first compressed sub-data is found from the reverse index dictionary to determine the first compressed data corresponding to the combined query fragment containing the query data, and the target sharding key to which the first compressed data corresponding to the combined query fragment belongs is found. Based on the logical sharding identifier in the target sharding key, the target physical table corresponding to the content to be queried of the query data can be located. In this way, only the target physical table needs to be traversed to read the target data to be queried, avoiding cross-sharding queries, thereby improving the data query efficiency. At the same time, it also supports querying using the data corresponding to each high-frequency query field in the physical table, and the target sharding keys corresponding to the data of each high-frequency query field can be quickly located to perform data query based on the target sharding key, which can meet the requirements of multi-dimensional high-frequency query scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0046] Figure 1 Flow diagram of a data processing method provided by an exemplary embodiment of the present application;

[0047] Figure 2 Flow diagram of determining first compressed data based on data corresponding to multiple high-frequency query fields provided by an embodiment of the present application;

[0048] Figure 3 Schematic diagram of a sharding key structure provided by an exemplary embodiment of the present application;

[0049] Figure 4 Flow diagram of a data query method provided by an exemplary embodiment of the present application;

[0050] Figure 5 Flow diagram of another data query method provided by an embodiment of the present application;

[0051] Figure 6 A data processing device provided by an exemplary embodiment of the present application;

[0052] Figure 7 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application;

[0053] Figure 8 A data query device provided by an exemplary embodiment of the present application;

[0054] Figure 9 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0055] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0056] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0057] The various models involved in this application (including but not limited to language models or large models) comply with relevant laws and regulations and standards.

[0058] In modern distributed database systems, the sharding technique is a commonly used data partitioning strategy. By dispersing data storage across different tables or databases, it improves system scalability and performance. In the sharding technique, the selection of the sharding key is one of the core issues. The sharding key determines how data is distributed among different physical tables.

[0059] However, traditional sharding keys usually consist of a business field or a system field (such as a unique identifier, timestamp, etc.). However, if the distribution characteristics of the field selected as the sharding key are uneven, it will lead to uneven data distribution in the distributed system, resulting in unbalanced system load and decreased query efficiency. Additionally, during the data query process, if the query condition uses not the single field that generates the sharding key but other business fields for query, all shards need to be scanned to obtain the required data, increasing the cost of cross-shard queries and significantly reducing query efficiency. Especially in high-frequency query scenarios, this problem becomes more prominent.

[0060] To address the above technical problems, the embodiments of this application propose a new data processing method and data query method. In this technical solution, dynamic sharding storage is performed on each piece of data to be written based on multiple high-frequency query fields in each piece of data to be written. Specifically, by analyzing the data to be written, multiple high-frequency query fields in the data to be written are identified, the data corresponding to the multiple high-frequency query fields is compressed to obtain first compressed data, and second compressed data used as a logical sharding identifier is intercepted from the first compressed data. Then, the compressed data of the multiple high-frequency query fields and the logical sharding identifier are embedded into a preset target sharding key structure to dynamically generate the target sharding key corresponding to the data to be written. Based on the logical sharding identifier in the target sharding key, the logical shard and the target physical table corresponding to the data to be written are determined, and the data to be written and the target sharding key are written into the corresponding target physical table to achieve uniform sharding storage of data in the distributed database and effectively avoid the data skew problem.

[0061] In addition, when performing data queries, the sharding key corresponding to each query data can be directly determined, and the specific target physical table can be located through the sharding key, avoiding cross-shard queries. This can not only improve data query efficiency but also meet the requirements of multi-dimensional high-frequency query scenarios.

[0062] The following will detail the data processing method and data query method provided by the embodiments of this application in conjunction with the accompanying drawings.

[0063] Figure 1A flowchart of a data processing method provided for an exemplary embodiment of the present application. As Figure 1 shown, this method can be applied to a distributed database, which includes multiple physical tables obtained by logically partitioning the distributed database. The execution subject of this method can be a data processing device. Specifically, this method can include the following steps:

[0064] 101. In response to a data writing request triggered by a client for data to be written, determine multiple high-frequency query fields corresponding to the data to be written, and screen out the data corresponding to the multiple high-frequency query fields from the data to be written.

[0065] 102. Concatenate the data corresponding to the multiple high-frequency query fields, and perform compression processing on the concatenation result to obtain first compressed data.

[0066] 103. Intercept low-order field data with a preset bit width from the first compressed data to generate second compressed data, and the second compressed data is used as a logical shard identifier.

[0067] 104. Perform a merging process on the first compressed data and the second compressed data to determine a target shard key corresponding to the data to be written.

[0068] 105. Query a preset mapping table to determine a target physical table corresponding to the second compressed data. The preset mapping table stores the mapping relationship between the logical shard identifier and the physical table identifier.

[0069] 106. Write the data to be written and the target shard key into the target physical table.

[0070] In the scenario of database table partitioning, when a client wants to write data to be written into a distributed database, the client can send a data writing request to the data processing device. After receiving the data writing request, in response to the data writing request triggered by the client for the data to be written, determine multiple high-frequency query fields corresponding to the data to be written.

[0071] Among them, the high-frequency query field refers to a field that is frequently accessed or used as a query condition in the service scenario. For example, in the e-commerce scenario, the high-frequency query fields can be user ID, order number, product name, inventory status, etc. In practical applications, a large number of user historical query records can be collected, and according to the historical query records, determine multiple high-frequency query fields included in the data to be written. Or determine multiple high-frequency query fields included in the data to be written according to the actual service requirements.

[0072] After determining multiple high-frequency query fields corresponding to the data to be written, filter out the data corresponding to the multiple high-frequency query fields from the data to be written. Furthermore, splice the data corresponding to the multiple high-frequency query fields to obtain a splicing result. And perform compression processing on the splicing result to obtain the first compressed data.

[0073] Among them, the data corresponding to the multiple high-frequency query fields can be directly spliced together to obtain the splicing result. For example, the data to be written is user data, and this user data includes field information such as user ID, user name, user mobile phone number, order number, etc. Among them, the user ID is user-101#bj, the user name is Li San, the user mobile phone number is 888888, and the order number is 123456. And the multiple high-frequency query fields corresponding to the data to be written determined are the user ID, user name, and order number respectively. Then the user ID, user name, and order number in the user data can be spliced to obtain the spliced string as user-101#bjLi San123456.

[0074] However, in practical applications, the data corresponding to each field in the user data may contain special characters, etc. Then when splicing the data corresponding to the multiple high-frequency query fields, the data corresponding to each high-frequency query field can be normalized first to convert it into a unified encoding format. For example, convert the above user ID: user-101#bj to user101bj.

[0075] After obtaining the splicing result of the data corresponding to the multiple high-frequency query fields, perform compression processing on the splicing result to embed the compressed splicing result as a combined query fragment into a preset shard key structure. Among them, a hash algorithm can be used to perform compression processing on the splicing result.

[0076] In order to reduce the problem of uneven data distribution caused by the data corresponding to the selected high-frequency query fields being concentrated in a certain specific range, when performing compression processing on the splicing result, a hierarchical hash compression method can be used to compress the splicing result to obtain the first compressed data. In this way, not only can the data corresponding to the multiple high-frequency query fields be compressed, but also the collision rate can be reduced.

[0077] Next, intercept the low-order field data of a preset bit width from the first compressed data to generate the second compressed data. Among them, the second compressed data can be used as the logical shard identifier corresponding to the logical shard where the data to be written is written.

[0078] In an alternative embodiment, the low-order field data of a preset bit width can be intercepted from the least significant bit of the first compressed data to the left to generate the second compressed data.

[0079] Subsequently, the first compressed data and the second compressed data are combined to determine the target shard key corresponding to the data to be written. In the embodiments of the present application, a preset shard key structure is newly designed. The shard key structure is composed of a combined query segment and a shard identification field. In this way, after obtaining the first compressed data and the second compressed data, the first compressed data and the second compressed data can be directly concatenated to obtain the target shard key corresponding to the data to be written. In addition, the preset shard key structure may further include a timestamp field, a data identification field, etc., which can be set according to actual requirements.

[0080] After determining the target shard key corresponding to the data to be written, based on the second compressed data in the target shard key (i.e., the logical shard identification in the target shard key), the logical shard corresponding to the data to be written can be determined, and by querying a preset mapping table, the target physical table corresponding to the second compressed data can be determined. The preset mapping table stores the mapping relationship between the logical shard identification and the physical table identification.

[0081] It should be noted that in the embodiments of the present application, the decoupling of the logical shard identification and the physical table position can be achieved. The relationship between the logical shard and the physical table can be dynamically managed through a preset mapping table. The physical table corresponding to the logical shard identification can be dynamically managed by directly modifying the information stored in the preset mapping table to achieve dynamic adjustment of the shard position without modifying the business code. Moreover, in this way, online expansion can be achieved. When a new physical table needs to be added, only new records need to be added to the preset mapping table and part of the logical shard data needs to be migrated.

[0082] As can be seen from the above description, the target shard key corresponding to the data to be written generated in the above manner not only embeds the combined query segment but also embeds the logical shard identification. In this way, during subsequent queries, the target shard key can be parsed to obtain the logical shard identification in the target shard key, and then based on the logical shard identification, the specific physical table can be located. Specifically, the first compressed data is used to determine the shard key containing the query data, and the second compressed data is used to determine the target physical table corresponding to the query data.

[0083] Finally, the data to be written and the target shard key are written into the target physical table, so that subsequently, it is convenient to locate each piece of data to be queried based on the shard keys in each physical table.

[0084] In summary, in the embodiment of the present application, by compressing the splicing result of the data corresponding to multiple high-frequency query fields in the data to be written, the first compressed data is obtained. The low-bit field data with a preset bit width is intercepted bit by bit from the first compressed data to obtain the second compressed data, and the first compressed data and the second compressed data are merged to obtain the target shard key corresponding to the data to be written. In this way, the compressed data corresponding to multiple high-frequency query fields is embedded into the target shard key corresponding to the data to be written as a combined query segment. Subsequently, the target shard key corresponding to the query data can be quickly located based on the data corresponding to multiple high-frequency query fields, so as to quickly locate the target physical table corresponding to the query data based on the target shard key, which improves the query efficiency and can meet the requirements of multi-dimensional high-frequency query scenarios. After determining the target shard key corresponding to the data to be written, the logical shard identifier can be determined based on the second compressed data in the target shard key, and then the target physical table corresponding to the logical shard identifier can be determined by querying the preset mapping table. The data to be written and the target shard key are written into the target physical table. Since the target shard key is generated based on the combined result of the first compressed data and the second compressed data, it is possible to avoid uneven data distribution caused by the uniform distribution of high-frequency query field data, so as to achieve uniform sharding storage of the data to be written into the database and effectively avoid the data skew problem.

[0085] For ease of description, in combination with the following embodiments, the specific implementation process of splicing the data corresponding to multiple high-frequency query fields and compressing the splicing result to obtain the first compressed data is exemplarily described.

[0086] Figure 2 FIG. is a schematic flowchart of determining the first compressed data based on the data corresponding to multiple high-frequency query fields provided by the embodiment of the present application; as Figure 2 shown, the multiple high-frequency query fields include the first high-frequency query field and the second high-frequency query field. The present method provides an implementation manner for compressing the splicing result of the data corresponding to multiple high-frequency query fields. Specifically, it may include the following:

[0087] 201. Normalize the data corresponding to the first high-frequency query field and the data corresponding to the second high-frequency query field respectively to obtain the first normalized data and the second normalized data.

[0088] 202. Perform a splicing process on the first normalized data and the second normalized data to obtain a spliced string.

[0089] 203. Perform a non-encrypted hash operation on the spliced string to obtain the first hash operation result.

[0090] 204. Perform a modulo operation on the first hash operation result to obtain a second hash operation result with a first preset length.

[0091] 205. Determine the second hash operation result as the first compressed data, and the first compressed data serves as a combined query fragment.

[0092] Among them, before splicing the data corresponding to the first high-frequency query field and the data corresponding to the second high-frequency query field, the data corresponding to the first high-frequency query field and the data corresponding to the second high-frequency query field can be normalized respectively to obtain normalized data in the same encoding format. Here, the normalization process may refer to deleting special symbols, special marking information, etc. in the data corresponding to each high-frequency query field.

[0093] After obtaining the first normalized data corresponding to the first high-frequency query field and the second normalized data corresponding to the second reassignment query field, splice the first normalized data and the second normalized data to obtain a spliced string. For example, the first normalized data is: user101bj, the second normalized data is: 123456, and the spliced string is: user101bj123456.

[0094] Next, perform a non-cryptographic hash operation on the spliced string to obtain a first hash operation result. Among them, performing a non-cryptographic hash operation on the spliced string can make the data corresponding to the obtained first hash operation result more uniform to avoid data skew problems. Specifically, in an optional embodiment, the FNV-1a hash algorithm can be used to perform a non-cryptographic hash operation on the spliced string to generate a first hash operation result with a 32-bit value.

[0095] Specifically, in an optional embodiment, the specific implementation method of using the FNV-1a hash algorithm to perform a non-cryptographic hash operation on the spliced string to generate a first hash operation result with a 32-bit value can be: obtain an initial hash value with a second preset length, and convert the spliced string into a byte stream. Perform an exclusive OR operation on the current byte in the byte stream and the least significant byte of the initial hash value to obtain an exclusive OR operation result. Shift the exclusive OR operation result to the left to obtain a left shift result, and multiply the left shift result by the FNV prime number to obtain a product value. Perform a modulo operation on the product value to obtain the hash value corresponding to the current byte. Determine the first hash operation result based on the hash values corresponding to each byte in the byte stream.

[0096] Among them, the second preset length can be 32 bit, and the corresponding preset length can be set according to actual needs. Determine the hash values corresponding to each byte in the byte stream byte by byte, and determine the first hash operation result based on the hash values corresponding to each byte in the byte stream.

[0097] After obtaining the first hash operation result, a modulo operation can also be performed on the first hash operation result to obtain a second hash operation result of a first preset length. Wherein, the first preset length is less than the second preset length. For example, the first preset length can be 12bit, 8bit, etc. The corresponding first preset length can be set by balancing the storage size occupied and the amount of data information corresponding to multiple query fields to be retained.

[0098] Finally, the second hash operation result is determined as the first compressed data. Wherein, the first compressed data serves as a combined query segment. It can be seen from the above description that the first compressed data, that is, the combined query segment, is obtained by performing two hash operations in stages. This can not only dynamically adjust the number of bits corresponding to the first compressed data, but also reduce the collision probability.

[0099] In summary, in the embodiment of the present application, by respectively performing normalization processing on the data corresponding to the first high-frequency query field and the data corresponding to the second high-frequency query field, the first normalized data and the second normalized data are obtained. The first normalized data and the second normalized data are concatenated to obtain a concatenated string. Two-stage hash operations are performed on the concatenated string to obtain the first compressed data. This can compress the data corresponding to multiple high-frequency query fields, embed the compressed data into the target shard key, reduce the storage space occupied by the target shard key, and at the same time, can also dynamically adjust the number of bits corresponding to the first compressed data to reduce the collision probability.

[0100] Wherein, after obtaining the first compressed data, the low-bit field data with a preset bit width can be intercepted from the least significant bit of the first compressed data to the left to generate the second compressed data. And the first compressed data and the second compressed data are merged to determine the target shard key corresponding to the data to be written. The following specifically describes the implementation process of generating the target shard key corresponding to the data to be written in combination with the following embodiments.

[0101] In the embodiment of the present application, the preset shard key structure is composed of four fields: a timestamp, a combined query segment, a data identifier (unique sequence), and a logical shard identifier. For example Figure 3 The shard key structure shown. And, the timestamp in the shard key can be data of 40bit size, the combined query segment is data of 12bit size, the data identifier is data of 8bit size, and the logical shard identifier is data of 4bit size. Then when generating the target shard key corresponding to the data to be written, the target shard key is composed of the numerical values corresponding to these four fields. That is, the combined query segment here is the first compressed data, and the logical shard identifier is the second compressed data.

[0102] In this way, when generating the target shard key, the reception timestamp corresponding to the query request received for the data to be written can be obtained. Then, using the open-source tool class, a unique sequence corresponding to the data to be written is generated. In the order from left to right, the reception timestamp, the first compressed data, the unique sequence, and the second compressed data are concatenated in sequence to obtain a concatenated data of a third preset length. The concatenated data is determined as the target shard key corresponding to the data to be written.

[0103] After generating the target shard key corresponding to the data to be written, the second compressed data in the target shard key is determined as the logical shard identifier of the target logical shard corresponding to the data to be written, and the target physical table identifier corresponding to the logical shard identifier is searched from the preset mapping table. Based on the target physical table identifier, the target physical table corresponding to the data to be written is determined. Furthermore, the data to be written and the target shard key corresponding to the data to be written are stored in the target physical table.

[0104] The above embodiments introduce the specific implementation process of shard-storing a data to be written into the corresponding target physical table in the distributed database in the scenario of database table sharding. The data query process will be introduced in detail below in combination with the following embodiments.

[0105] Figure 4 It is a schematic flowchart of a data query method provided by an exemplary embodiment of the present application. As Figure 4 shown, the execution subject of this method can be a data query device. Specifically, on the basis of the above embodiments, this method may include the following steps:

[0106] 401. Receive a query request containing query data sent by the client, and identify the target field corresponding to the query data in the distributed database.

[0107] 402. In response to the determination result that the target field is not a shard key field, perform compression processing on the query data to obtain the first compressed sub-data.

[0108] 403. Query the first compressed data corresponding to the first compressed sub-data from the reverse index dictionary. The reverse index dictionary stores the compressed sub-data corresponding to the query data and the compressed data corresponding to the concatenated result containing the query data.

[0109] 404. Based on the first compressed data, determine the target shard key containing the query data.

[0110] 405. Based on the logical shard identifier in the target shard key, determine the first target physical table corresponding to the query data.

[0111] 406. Read the target data corresponding to the query data from the first target physical table.

[0112] In practical applications, when a client wants to read target data from a distributed database, it can send a data query request to a data query device using the query data. After the data query device receives the query request containing the query data sent by the client, it identifies the target field corresponding to the query data in the distributed database. For example, the query data is a specific data corresponding to the user ID column field in a physical table in the distributed database.

[0113] Determine whether the target field is a shard key field. If the target field is not a shard key field, in response to the determination result that the target field is not a shard key field, the query data can be compressed to obtain the first compressed sub-data. Among them, when compressing the query data, the same compression method used when generating the combined query fragments in the shard key can be adopted. For example, a phased hash operation is used to compress the query data.

[0114] Specifically, in an optional embodiment, the implementation process of compressing the query data to obtain the first compressed sub-data may include: performing a non-encrypted hash operation on the query data to obtain a first hash value, performing a modulo operation on the first hash value to obtain a second hash value, and determining the second hash value as the first compressed sub-data.

[0115] Among them, an optional implementation manner of performing a non-encrypted hash operation on the query data to obtain the first hash value may be: obtaining an initial hash value of a second preset length, converting the query data into a byte stream. Performing an exclusive OR operation on the current byte in the byte stream and the least significant byte of the initial hash value to obtain an exclusive OR operation result, and shifting the exclusive OR operation result to the left to obtain a left shift result. Multiplying the left shift result by the FNV prime number to obtain a product value. Performing a modulo operation on the product value to obtain the hash value corresponding to the current byte. Based on the hash values corresponding to each byte in the byte stream, the first hash value is determined.

[0116] Among them, in an optional embodiment, the specific implementation manner of performing a modulo operation on the first hash value to obtain the second hash value may be: starting from the least significant bit of the first hash value, intercepting the low-order field data with a preset bit width to the left to generate the second hash value.

[0117] After obtaining the first compressed sub-data corresponding to the query data, query the reverse index dictionary to query the first compressed data corresponding to the first compressed sub-data from the reverse index dictionary. Among them, the compressed data embedded in the shard key is the concatenation result of the data corresponding to multiple high-frequency query fields, that is, it is equivalent to writing the compressed multiple query data in each row data record in the physical table into the shard key. Then, after obtaining the first compressed sub-data corresponding to a query data, it is impossible to directly determine the target shard key corresponding to the query data based on the first compressed sub-data.

[0118] Therefore, in the embodiments of the present application, a reverse index dictionary is pre-created to find the compressed data containing the query data through the reverse index dictionary, so as to realize the reverse mapping of the compressed sub-data index corresponding to the query data to the compressed data containing the query data. Among them, the reverse index dictionary stores the compressed sub-data corresponding to the query data and the compressed data corresponding to the concatenation result containing the query data. The first compressed data is obtained by compressing the concatenation result, and the concatenation result is obtained by concatenating the data corresponding to multiple high-frequency query fields. The data corresponding to multiple high-frequency query fields includes the query data.

[0119] Next, based on the first compressed data, the target shard key containing the query data is determined. That is, the target shard key corresponding to the first compressed data is searched for. The target shard key is the shard key containing the query data. Among them, the target shard key is obtained by merging the first compressed data and the second compressed data. The second compressed data is obtained by bitwise intercepting the low-order field data with a preset bit width from the first compressed data.

[0120] For example, it is assumed that each data in the reverse index dictionary is stored in the form of a key-value pair. Among them, the key key is the compressed sub-data (hash value) corresponding to a single query data, and the key value value is the common compressed data (hash operation result) containing the single query data. For example, the key is the hash value of a, and the value is the common hash value of a + b + c. In actual applications, if the query data carried in the query request is a, then the hash value of a can be calculated and the reverse index dictionary can be searched to find the common hash value of a + b + c. The common hash value of a + b + c is the combined query field in the shard key, and then the target shard key where the hash value of a + b + c is located can be searched for.

[0121] After the target shard key corresponding to the query data is determined, next, based on the logical shard identifier in the target shard key, the first target physical table corresponding to the query data is determined, and the target data corresponding to the query data is read from the first target physical table.

[0122] Specifically, in an optional embodiment, the implementation process of determining the first target physical table corresponding to the query data based on the logical shard identifier in the target shard key may be: parsing the target shard key to identify the value of the identifier bit corresponding to the logical shard identifier in the target shard key; searching for the first target physical table identifier corresponding to the value of the identifier bit from a preset mapping table. The preset mapping table stores the mapping relationship between the logical shard identifier and the physical table identifier. Based on the first target physical table identifier, the first target physical table to which the query data belongs is determined.

[0123] Among them, in addition to storing the mapping relationship between the logical shard identifier and the physical table identifier, the preset mapping table can also store the table status of the physical table corresponding to the physical table identifier. In practical applications, physical table migration may occur. Then, in order to more accurately read the target data corresponding to the query data, after determining the first target physical table to which the query data belongs, the table status corresponding to the first target physical table can be obtained, and based on the table status, it can be determined whether the first target physical table is in a migration state.

[0124] Specifically, obtain the status information corresponding to the current first target physical table. If the status information is not the migration state, directly read the target data from the first target physical table. If the status information is the migration state, in response to the determination result that the status information is the migration state, based on the migration direction, determine the second target physical table corresponding to the first target physical table, read the first query content corresponding to the query data from the first target physical table, and read the second query content corresponding to the query data from the second target physical table. Merge the first query content and the second query content to obtain the target data.

[0125] In the embodiment of the present application, when it is determined that the query data is not the shard key value, the query data is compressed to obtain the first compressed sub-data, and by searching the reverse index dictionary for the first compressed data corresponding to the first compressed sub-data, the first compressed data corresponding to the combined query fragment containing the query data is determined, and the target shard key to which the first compressed data corresponding to the combined query fragment belongs is searched. Based on the logical shard identifier in the target shard key, the target physical table corresponding to the content to be queried for the query data can be located. In this way, only by traversing the target physical table can the target data to be queried be read, avoiding cross-shard queries, thereby improving the data query efficiency. At the same time, it also supports querying using the data corresponding to each high-frequency query field in the physical table. The target shard keys corresponding to the data of each high-frequency query field can be quickly located, so as to perform data query based on the target shard keys, which can meet the requirements of multi-dimensional high-frequency query scenarios.

[0126] The above embodiments introduce the specific implementation process of data query when the query data carried in the query request sent by the client is not the shard key. In addition, in practical applications, there is also the case where the query data carried in the query request sent by the client is the shard key. Combining Figure 5 A detailed description of this specific implementation process is given.

[0127] Figure 5 It is a schematic flow chart of another data query method provided by the embodiment of the present application; as Figure 5 shown, the method may include the following steps:

[0128] 501. In response to the determination result that the target field is a shard key field, parse the query data to determine the logical shard identifier corresponding to the query data.

[0129] 502. Search for the second target physical table identifier corresponding to the logical shard identifier in the preset mapping table.

[0130] 503. Based on the second target physical table identifier, determine the third target physical table to which the query data belongs.

[0131] 504. Read the target data corresponding to the query data from the third target physical table.

[0132] Among them, when it is recognized that the target field corresponding to the query data in the distributed database is a shard key field, that is, the query data is a shard key, then the value of the identification bit corresponding to the logical shard identifier in the query data can be parsed to determine the logical shard identifier corresponding to the query data.

[0133] Next, search for the second target physical table identifier corresponding to the logical shard identifier in the preset mapping table, and based on the second target physical table identifier, determine the third target physical table to which the query data belongs. And obtain the location information of the third target physical table, and read the target data corresponding to the query data based on this location information.

[0134] In the embodiment of the present application, the logical shard identifier corresponding to the query data can be directly obtained by parsing the query data, and furthermore, based on the logical shard identifier, the corresponding third target physical table can be quickly located to read the target data corresponding to the query data from the third target physical table, without the need to perform hash modulo operation processing on the query data, simplifying the data query process and improving the data query efficiency.

[0135] The data processing device and data query device of one or more embodiments of the present application will be described in detail below. Those skilled in the art can understand that these devices can all be configured by using commercially available hardware components through the steps taught by this solution.

[0136] Figure 6 A data processing device provided for an exemplary embodiment of the present application, as Figure 6 shown, the device includes: a response module 11, a splicing module 12, a generation module 13, a merging module 14, a query module 15, and a writing module 16.

[0137] The response module 11 is configured to, in response to a data writing request triggered by a client for data to be written, determine a plurality of high-frequency query fields corresponding to the data to be written, and filter out the data corresponding to the plurality of high-frequency query fields from the data to be written.

[0138] The splicing module 12 is used to splice the data corresponding to the multiple high-frequency query fields, and perform compression processing on the splicing result to obtain the first compressed data.

[0139] The generation module 13 is used to intercept low-order field data with a preset bit width from the first compressed data to generate second compressed data, and the second compressed data is used as a logical shard identifier.

[0140] The merging module 14 is used to perform merging processing on the first compressed data and the second compressed data to determine the target shard key corresponding to the data to be written. Among them, the first compressed data is used to determine the shard key containing the query data; the second compressed data is used to determine the target physical table corresponding to the query data.

[0141] The query module 15 is used to query a preset mapping table to determine the target physical table corresponding to the second compressed data, and the preset mapping table stores the mapping relationship between the logical shard identifier and the physical table identifier.

[0142] The writing module 16 is used to write the data to be written and the target shard key into the target physical table.

[0143] Optionally, the multiple high-frequency query fields include a first high-frequency query field and a second high-frequency query field; specifically, the splicing module 12 is configured to: perform normalization processing on the data corresponding to the first high-frequency query field and the data corresponding to the second high-frequency query field respectively to obtain the first normalized data and the second normalized data; perform splicing processing on the first normalized data and the second normalized data to obtain a spliced string; perform non-encrypted hashing operation on the spliced string to obtain a first hashing operation result; perform modulo operation on the first hashing operation result to obtain a second hashing operation result with a first preset length; determine the second hashing operation result as the first compressed data, and the first compressed data is used as a combined query segment.

[0144] Optionally, the splicing module 12 is specifically configured to: obtain an initial hash value with a second preset length; convert the spliced string into a byte stream; perform exclusive OR operation on the current byte in the byte stream and the least significant byte of the initial hash value to obtain an exclusive OR operation result; perform left shift on the exclusive OR operation result to obtain a left shift result; multiply the left shift result by an FNV prime number to obtain a product value; perform modulo operation on the product value to obtain the hash value corresponding to the current byte; determine the first hashing operation result based on the hash values corresponding to the respective bytes in the byte stream.

[0145] Optionally, the generating module 13 is specifically configured to: start from the least significant bit of the first compressed data, and intercept the low-order field data with a preset bit width to the left to generate the second compressed data.

[0146] Optionally, the merging module 14 is specifically configured to: obtain the reception timestamp corresponding to the query request; use an open-source tool class to generate a unique sequence corresponding to the data to be written; sequentially splice the reception timestamp, the first compressed data, the unique sequence, and the second compressed data in order from left to right to obtain a spliced data with a third preset length; determine the spliced data as the target shard key corresponding to the data to be written. Optionally, the query module 15 is specifically configured to: determine the second compressed data as the logical shard identifier of the target logical shard corresponding to the data to be written; look up the target physical table identifier corresponding to the logical shard identifier in a preset mapping table; based on the target physical table identifier, determine the target physical table corresponding to the data to be written.

[0147] Figure 7 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 7 shown, in practice, the electronic device includes: a memory 21 and a processor 22. The memory 21 is used to store computer programs and can be configured to store various other data to support operations on the electronic device. Examples of these data include instructions for any application program or method for operating on the electronic device, data structures, contact data, phone book data, messages, pictures, videos, etc.

[0148] The processor 22 is coupled to the memory 21 and is used to execute the computer program in the memory 21 to implement the data processing method provided in the foregoing embodiment.

[0149] Further, as Figure 7 shown, the electronic device further includes: other components such as a communication component 23, a display 24, a power supply component 25, and an audio component 26. Figure 7 Only some components are schematically shown herein, and it does not mean that the electronic device only includes Figure 7 the components shown. The electronic device of this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, or an IOT device, or can also be a server device such as a conventional server, a cloud server, or a server array.

[0150] The foregoing memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof. The foregoing communication component is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner.

[0151] The above-mentioned display includes a screen, which may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user.

[0152] The above-mentioned power supply component provides power for various components of the device where the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located.

[0153] The above-mentioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC). When the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals can be further stored in a memory or sent via a communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals.

[0154] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor is enabled to implement the steps in the above method embodiment.

[0155] Among them, the computer-readable storage medium can be implemented by a volatile or non-volatile or a combination thereof, and can be removable or non-removable. Examples of the computer-readable storage medium include, but are not limited to, Phase-change Random Access Memory (PRAM), Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), other types of Random-Access Memory (RAM), or any other non-transmission medium.

[0156] Correspondingly, an embodiment of the present application further provides a computer program product, which includes a computer program or instruction. When the computer program or instruction is executed by a processor, the processor is enabled to implement the steps in the above method embodiment.

[0157] Figure 8 A data query device provided for an exemplary embodiment of the present application, as Figure 8 shown, the device includes: an identification module 31, a compression module 32, a query module 33, a first determination module 34, a second determination module 35, and a reading module 36.

[0158] An identification module 31, configured to receive a query request including query data sent by a client, and identify a target field corresponding to the query data in the distributed database.

[0159] A compression module 32, configured to perform compression processing on the query data to obtain first compressed sub-data in response to a determination result that the target field is not a sharding key field.

[0160] A query module 33, configured to query first compressed data corresponding to the first compressed sub-data from a reverse index dictionary, where the reverse index field stores compressed sub-data corresponding to query data and compressed data corresponding to a concatenation result including the query data, the first compressed data is obtained by performing compression processing on the concatenation result, the concatenation result is obtained by concatenating data corresponding to multiple high-frequency query fields, and the data corresponding to the multiple high-frequency query fields includes the query data.

[0161] A first determination module 34, configured to determine a target sharding key including the query data based on the first compressed data, where the target sharding key is obtained by merging the first compressed data and second compressed data, and the second compressed data is obtained by bitwise intercepting low-order field data with a preset bit width from the first compressed data.

[0162] A second determination module 35, configured to determine a first target physical table corresponding to the query data based on a logical sharding identifier in the target sharding key.

[0163] A reading module 36, configured to read target data corresponding to the query data from the first target physical table.

[0164] Optionally, the compression module 32 is specifically configured to: perform a non-encrypted hash operation on the query data to obtain a first hash value; perform a modulo operation on the first hash value to obtain a second hash value; and determine the second hash value as the first compressed sub-data.

[0165] Optionally, the first determination module 34 is specifically configured to: find a target sharding key corresponding to the first compressed data, where the target sharding key is a sharding key including the query data.

[0166] Optionally, the second determination module 35 is specifically configured to: parse the target sharding key to identify a numerical value of an identification bit corresponding to the logical sharding identifier in the target sharding key; find a first target physical table identifier corresponding to the numerical value of the identification bit from a preset mapping table, where the preset mapping table stores a mapping relationship between a logical sharding identifier and a physical table identifier; and determine the first target physical table to which the query data belongs based on the first target physical table identifier.

[0167] Optionally, the reading module 36 is specifically configured to: obtain the status information corresponding to the current first target physical table; in response to the determination result that the status information is in a migration state, determine the second target physical table corresponding to the first target physical table based on the migration direction; read the first query content corresponding to the query data from the first target physical table; read the second query content corresponding to the query data from the second target physical table; and merge the first query content and the second query content to obtain the target data.

[0168] Optionally, the device may further include a parsing module, and the parsing module is specifically configured to: in response to the determination result that the target field is a sharding key field, parse the query data to determine the logical sharding identifier corresponding to the query data; look up the second target physical table identifier corresponding to the logical sharding identifier from a preset mapping table; determine the third target physical table to which the query data belongs based on the second target physical table identifier; and read the target data corresponding to the query data from the third target physical table.

[0169] Figure 9 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 9 shown, in practice, the electronic device includes: a memory 41 and a processor 42.

[0170] The memory 41 is used to store computer programs and can be configured to store various other data to support operations on the electronic device. Examples of these data include instructions for any application program or method for operating on the electronic device, data structures, contact data, phone book data, messages, pictures, videos, etc.

[0171] The processor 42 is coupled to the memory 41 and is used to execute the computer program in the memory 41 to implement the data query method provided in the foregoing embodiment.

[0172] Further, as Figure 9 shown, the electronic device further includes: other components such as a communication component 43, a display 44, a power supply component 45, and an audio component 46. Figure 9 Only some components are schematically shown herein, and it does not mean that the electronic device only includes Figure 9 the components shown. The electronic device in this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, or an IOT device, or can also be a server device such as a conventional server, a cloud server, or a server array.

[0173] The above-mentioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc.

[0174] The above-mentioned communication component is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on a communication standard, such as 2G, 3G, 4G / LTE, 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel.

[0175] The above-mentioned display includes a screen, and the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operations.

[0176] The above-mentioned power supply component provides power for various components of the device where the power supply component is located. The power supply component can include a power management system, one or more power supplies, and other components associated with generating, managing and distributing power for the device where the power supply component is located.

[0177] The above-mentioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), and when the device where the audio component is located is in an operating mode, such as a call mode, a recording mode and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals can be further stored in the memory or sent via the communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals.

[0178] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to be able to implement the steps in the above method embodiments.

[0179] Among them, the computer-readable storage medium can be implemented by volatile or non-volatile or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices or any other non-transmission medium.

[0180] Accordingly, an embodiment of the present application further provides a computer program product, which includes a computer program or instruction, which, when executed by a processor, causes the processor to be able to implement the steps in the above method embodiments.

[0181] It should be understood that each process or a combination of multiple processes in the above method flow can be implemented by a computer program or instruction. In addition, these computer programs or instructions can be applied to the processors of general-purpose computers, special-purpose computers, embedded processors or other programmable data processing devices, so that the processors of general-purpose computers, special-purpose computers, embedded processors or other programmable data processing devices can be used as devices to implement the corresponding functions in the above method embodiments.

[0182] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A data processing method, characterized in that, Applied to a distributed database, the distributed database includes multiple physical tables, and the physical tables are obtained by logically partitioning the distributed database. The method includes: In response to a data write request triggered by a client for data to be written, determine multiple high-frequency query fields corresponding to the data to be written, and filter out the data corresponding to the multiple high-frequency query fields from the data to be written; Concatenate the data corresponding to the multiple high-frequency query fields, and perform compression processing on the concatenation result to obtain first compressed data; Intercept low-order field data with a preset bit width bit by bit from the first compressed data to generate second compressed data, and the second compressed data is used as a logical shard identifier; Perform a merging process on the first compressed data and the second compressed data to determine a target shard key corresponding to the data to be written; Query a preset mapping table to determine a target physical table corresponding to the second compressed data. The preset mapping table stores the mapping relationship between the logical shard identifier and the physical table identifier; Write the data to be written and the target shard key into the target physical table; Wherein, the first compressed data is used to determine the shard key containing the query data; the second compressed data is used to determine the target physical table corresponding to the query data.

2. The method according to claim 1, wherein The multiple high-frequency query fields include a first high-frequency query field and a second high-frequency query field; The step of concatenating the data corresponding to the multiple high-frequency query fields and performing compression processing on the concatenation result to obtain first compressed data includes: Perform normalization processing on the data corresponding to the first high-frequency query field and the data corresponding to the second high-frequency query field respectively to obtain the first normalized data and the second normalized data; Perform a concatenation process on the first normalized data and the second normalized data to obtain a concatenated string; Perform a non-encrypted hashing operation on the concatenated string to obtain a first hashing operation result; Perform a modulo operation on the first hashing operation result to obtain a second hashing operation result with a first preset length; Determine the second hashing operation result as the first compressed data, and the first compressed data is used as a combined query fragment.

3. The method according to claim 2, characterized in that, The step of performing a non-encrypted hashing operation on the concatenation result to obtain a first hashing operation result includes: Obtain an initial hash value with a second preset length; Convert the concatenated string into a byte stream; Perform an exclusive OR operation on the current byte in the byte stream and the least significant byte of the initial hash value to obtain an exclusive OR operation result; Shift the exclusive OR operation result to the left to obtain a left shift result; Multiply the left shift result by an FNV prime number to obtain a product value; Perform a modulo operation on the product value to obtain the hash value corresponding to the current byte; Based on the hash values corresponding to each byte in the byte stream, determine the first hashing operation result.

4. The method according to claim 2, characterized in that, The step of intercepting low-order field data with a preset bit width bit by bit from the first compressed data to generate second compressed data includes: Starting from the least significant bit of the first compressed data, intercept low-order field data with a preset bit width to the left to generate second compressed data.

5. The method according to claim 1, characterized in that, The merging process of the first compressed data and the second compressed data to determine the target sharding key corresponding to the data to be written includes: Obtain the reception timestamp corresponding to the query request; Generate a unique sequence corresponding to the data to be written by using an open-source tool class; Concatenate the reception timestamp, the first compressed data, the unique sequence, and the second compressed data in sequence from left to right to obtain a concatenated data of a third preset length; Determine the concatenated data as the target sharding key corresponding to the data to be written.

6. The method according to claim 1, characterized in that The querying of the preset mapping table to determine the target physical table corresponding to the second compressed data includes: Determine the second compressed data as the logical sharding identifier of the target logical shard corresponding to the data to be written; Search in the preset mapping table for the target physical table identifier corresponding to the logical sharding identifier; Based on the target physical table identifier, determine the target physical table corresponding to the data to be written.

7. A data query method, characterized in that, Applied to a distributed database, the distributed database includes multiple physical tables, and the physical tables are obtained by logically sharding the distributed database. The method includes: Receive a query request containing query data sent by a client, and identify the target field corresponding to the query data in the distributed database; In response to the determination result that the target field is not a sharding key field, perform compression processing on the query data to obtain a first compressed sub-data; Query the first compressed data corresponding to the first compressed sub-data from the reverse index dictionary. The reverse index dictionary stores the compressed sub-data corresponding to the query data and the compressed data corresponding to the concatenated result containing the query data. The first compressed data is obtained by performing compression processing on the concatenated result, and the concatenated result is obtained by concatenating the data corresponding to multiple high-frequency query fields. The data corresponding to the multiple high-frequency query fields includes the query data; Based on the first compressed data, determine the target sharding key containing the query data. The target sharding key is obtained by merging the first compressed data and the second compressed data. The second compressed data is obtained by bitwise intercepting the low-order field data of a preset bit width from the first compressed data; Based on the logical sharding identifier in the target sharding key, determine the first target physical table corresponding to the query data; Read the target data corresponding to the query data from the first target physical table.

8. The method according to claim 7, characterized in that The performing of compression processing on the query data to obtain a first compressed sub-data includes: Perform a non-encrypted hash operation on the query data to obtain a first hash value; Perform a modulo operation on the first hash value to obtain a second hash value; Determine the second hash value as the first compressed sub-data.

9. The method according to claim 7, characterized in that, The determining of the target sharding key containing the query data based on the first compressed data includes: Search for the target sharding key corresponding to the first compressed data. The target sharding key is the sharding key containing the query data; Among them, determining the first target physical table corresponding to the query data based on the logical shard identifier in the target shard key includes: Parsing the target shard key to identify the value of the identification bit corresponding to the logical shard identifier in the target shard key; Looking up the first target physical table identifier corresponding to the value of the identification bit from a preset mapping table, where the preset mapping table stores the mapping relationship between the logical shard identifier and the physical table identifier; Determining the first target physical table to which the query data belongs based on the first target physical table identifier.

10. The method according to claim 7, wherein Reading the target data corresponding to the query data from the first target physical table includes: Obtaining the status information corresponding to the current first target physical table; In response to the determination result that the status information is in the migration state, determining the second target physical table corresponding to the first target physical table based on the migration direction; Reading the first query content corresponding to the query data from the first target physical table; Reading the second query content corresponding to the query data from the second target physical table; Merging the first query content and the second query content to obtain the target data.

11. The method according to claim 7, wherein The method further includes: In response to the determination result that the target field is a shard key field, parsing the query data to determine the logical shard identifier corresponding to the query data; Looking up the second target physical table identifier corresponding to the logical shard identifier from a preset mapping table; Determining the third target physical table to which the query data belongs based on the second target physical table identifier; Reading the target data corresponding to the query data from the third target physical table.

12. An electronic device, characterized in that, It includes: A memory and a processor; The memory is used to store a computer program; the processor is coupled to the memory and is used to execute the computer program to implement the data processing method according to any one of claims 1-6 or the data query method according to any one of claims 7-11.

13. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the processor is caused to implement the data processing method according to any one of claims 1-6 or the data query method according to any one of claims 7-11.