User data query method and apparatus, and device and storage medium

Through the bucket storage method of bitmap identification, bucket identification and bucket content, the bucket identification and content associated with the query request are directly obtained for logical operations, solving the problem of time-consuming query of user data in the prior art, and achieving efficient user data query.

WO2025152694A1PCT designated stage expired Publication Date: 2025-07-24CHINA UNIONPAY

Patent Information

Application Number
PCT/CN2024/140233
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-19
Filing Date
2024-12-18
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

The prior art requires traversing all data when querying user data, resulting in inefficient query efficiency, especially in scenarios where the amount of user data is large and the user classification is large.

Method used

The bucket storage method of bitmap identifier, bucket identifier and bucket content is adopted. The associated bucket identifier and bucket content are directly obtained through logical operations indicated by query requests, avoid traversing all data, and use logical operations to process bucket content to obtain query results.

Benefits of technology

It shortens the user data query time and improves query efficiency, especially in large-scale user data and multi-classification scenarios, which significantly improves query speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024140233_24072025_PF_FP_ABST
    Figure CN2024140233_24072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application belongs to the field of data processing. Disclosed are a user data query method and apparatus, and a device and a storage medium. The method comprises: receiving a query request, wherein the query request comprises computation logic information; on the basis of a target bitmap identifier indicated by the computation logic information, determining a target bucket identifier in a target storage space, wherein the target storage space stores a bitmap identifier, a bucket identifier and bucket content, which have correspondences, the bitmap identifier is used for representing the classification of user data, and the bucket identifier and the bucket content are used for performing restoration to obtain the user data; and on the basis of a logic operation indicated by the computation logic information, processing bucket content, which corresponds to the target bucket identifier, in the target storage space, so as to obtain a query result and feed back same.
Need to check novelty before this filing date? Find Prior Art

Description

User data query method, device, equipment and storage medium

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to Chinese patent application No. 202410078863.4, filed on January 19, 2024, entitled “User Data Query Method, Apparatus, Device and Storage Medium,” and the entire contents of that application are incorporated herein by reference. Technical Field

[0003] The present application relates to the field of data processing, and in particular to a method, apparatus, device, and storage medium for querying user data. Background Art

[0004] With the continuous development of electronic information technology, more and more user data involved in business operations can be managed through electronic information technology. To facilitate user data processing, user data can be segmented and managed according to user characteristics. User data can be stored in a bitmap format based on user IDs. When querying user data, it is necessary to traverse all data stored in the bitmap format. As business continues to evolve, the dimensionality and volume of user data continue to expand, resulting in lengthy queries and low query efficiency. Summary of the Invention

[0005] The embodiments of the present application provide a method, apparatus, device, and storage medium for querying user data, which can improve the efficiency of querying user data.

[0006] In a first aspect, an embodiment of the present application provides a method for querying user data, comprising: receiving a query request, the query request including calculation logic information; determining a target bucket identifier in a target storage space according to a target bitmap identifier indicated by the calculation logic information, the target storage space storing corresponding bitmap identifiers, bucket identifiers, and bucket contents, the bitmap identifier being used to characterize the classification of user data, and the bucket identifier and bucket contents being used to restore the user data; processing the bucket contents corresponding to the target bucket identifier in the target storage space according to the logical operation indicated by the calculation logic information, obtaining a query result, and feeding it back.

[0007] In the second aspect, an embodiment of the present application provides a user data query device, including: a receiving module for receiving a query request, the query request including calculation logic information; a logic processing module for determining a target bucket identifier in a target storage space according to a target bitmap identifier indicated by the calculation logic information, the target storage space storing corresponding bitmap identifiers, bucket identifiers and bucket contents, the bitmap identifiers being used to characterize the classification of user data, and the bucket identifiers and bucket contents being used to restore the user data; and a module for processing the bucket contents corresponding to the target bucket identifier in the target storage space according to the logical operation indicated by the calculation logic information to obtain a query result; a sending module for feeding back the query result.

[0008] In a third aspect, an embodiment of the present application provides an electronic device comprising: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the method for querying user data of the first aspect is implemented.

[0009] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the method for querying user data of the first aspect is implemented.

[0010] The embodiments of the present application provide a method, apparatus, device and storage medium for querying user data, which can pre-store corresponding bitmap identifiers, bucket identifiers and bucket contents in the target storage space. The bitmap identifier is used to characterize the classification of user data, and the bucket identifier and bucket content can be used to restore the user data. According to the target bitmap identifier indicated by the calculation logic information in the query request, the target bucket identifier is determined in the target storage space, and the bucket content corresponding to the target bucket identifier is processed according to the logical operation indicated by the calculation logic information to obtain the query result corresponding to the query request. In the above-mentioned user data query process, the user data is stored in buckets through the bitmap identifier, bucket identifier and bucket content. Each query can directly obtain the bucket identifier and bucket content associated with the current query and perform the data operations required for the logical operation indicated by the query request, without traversing all the data. This can shorten the time required for user data query and improve query efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0012] FIG1 is a flowchart of a method for querying user data provided by an embodiment of the present application;

[0013] FIG2 is a schematic diagram of an example of bucket storage corresponding to part of the data in Table 2 provided in an embodiment of the present application;

[0014] FIG3 is a flowchart of a method for querying user data provided by another embodiment of the present application;

[0015] FIG4 is a flowchart of a method for querying user data provided by another embodiment of the present application;

[0016] FIG5 is a schematic diagram of an example of multiple ways of storing bucket identifiers and bucket contents provided in an embodiment of the present application;

[0017] FIG6 is a logical diagram of an example of an application architecture of a method for querying user data provided in an embodiment of the present application;

[0018] FIG7 is a schematic diagram of the structure of a user data query device provided in one embodiment of the present application;

[0019] FIG8 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0020] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by illustrating examples of the present application. It should be noted that the acquisition, storage, use, processing, etc. of information and data in the embodiments of the present application are authorized by the user or relevant agencies and comply with the relevant provisions of national laws and regulations.

[0021] With the continuous development of electronic information technology, more and more user data involved in business can be managed through electronic information technology. To facilitate the processing of user data, user data can be divided according to user characteristics and the divided user data can be managed. User data can be stored in a bitmap format based on the user's user ID. When querying user data, it is necessary to traverse all data stored in the bitmap format. As the business continues to upgrade, the dimension and magnitude of user data are also constantly expanding, resulting in longer query time and lower query efficiency. Even when using an efficient compressed bitmap (i.e., RoaringBitmap), when querying user data, it is still necessary to traverse all user data. Especially in scenarios with large amounts of user data and many types of user classifications, querying user data will take longer and the query efficiency will be lower.

[0022] The embodiments of the present application provide a method, apparatus, device and storage medium for querying user data, which can pre-store bitmap identifiers, bucket identifiers and bucket contents with corresponding relationships in the target storage space. The bitmap identifier can represent the classification of the user data, and the bucket identifier and bucket content can restore the user data. According to the query request, the bucket identifier corresponding to the bitmap identifier indicated by the query request can be obtained from the target storage space, and the bucket content corresponding to the obtained bucket identifier can be processed according to the logical operation indicated by the query request to obtain the query result corresponding to the query request. In the above query process, the user data is stored in buckets through the bitmap identifier, bucket identifier and bucket content. Each query can directly obtain the bucket identifier and bucket content associated with the current query and perform the data operations required for the logical operation indicated by the query request, without traversing all the data, shortening the time used for data input and output (IO), thereby shortening the time required for user data query and improving query efficiency.

[0023] The following describes the user data query method, device, equipment and storage medium provided by this application.

[0024] In a first aspect, the present application provides a method for querying user data, which can be applied to user selection scenarios under multiple classification labels, such as real-time user selection scenarios under conditions of a large number of users. The method for querying user data can be executed by a user data query device or equipment, and is not limited here. Figure 1 is a flowchart of a method for querying user data provided by an embodiment of the present application. As shown in Figure 1, the method for querying user data may include steps S101 to S103.

[0025] In step S101, a query request is received.

[0026] The query request is used to query user data. The query request includes calculation logic information. The calculation logic information may indicate a logical operation and a classification of user data associated with the logical operation. The calculation logic information may be regarded as the implementation of the query condition. For example, the query request indicates to query the user data of users aged between 20 and 40 in city A1. The logical operation indicated by the calculation logic information in the query request may include finding the intersection of user data of users under the city A1 classification and user data of users under the 20 to 40 age classification. Correspondingly, the classification of user data associated with the logical operation indicated by the calculation logic information in the query request includes the city A1 classification and the 20 to 40 age classification.

[0027] In step S102, a target bucket identifier is determined in the target storage space according to the target bitmap identifier indicated by the calculation logic information.

[0028] The target storage space stores corresponding bitmap identifiers, bucket identifiers, and bucket contents. The bitmap identifiers represent the classification of user data and can be set based on the user data's classification labels. For example, if user data is categorized by gender or age, the gender classification labels might include "male" and "female," with corresponding bitmap identifiers set as "man" and "woman." Age classification labels might include "0 to 20 years old," "21 to 30 years old," "31 to 40 years old," "41 to 50 years old," and "over 51 years old," with corresponding bitmap identifiers set as "age_0_20," "age_21_30," "age_31_40," "age_41_50," and "age_51." The bucket identifiers and bucket contents are used to restore the user data. Each user data item can be divided into two parts: one part forms the bucket identifier, and the other part forms the bucket contents. The bucket identifiers and bucket contents for the same user data are in a corresponding relationship. In some examples, user data may include high-order N bits of data and low-order M bits of data; the bucket identifier is obtained based on the high-order N bits of the user data, and the bucket content is obtained based on the low-order M bits of the user data. N and M are both positive integers, and their specific values ​​can be set based on scenarios, requirements, experience, etc. The values ​​of N and M can determine the data range that can be stored in the target storage space. For example, if N=M=16, the target storage space can store 32 bits of data, that is, the data range that can be stored in the target storage space is 0 to 4294967295. It should be noted that if the number of data bits of the user data is less than N+M bits, 0 can be added to the high bits of the user data to obtain user data with a data bit number that meets N+M bits, thereby obtaining the bucket identifier and bucket content based on the user data. For example, N=M=16, and the user data may include a user identifier. If the user identifier is c00000704004, the prefix c00 of the user identifier may be removed, and 000704004 may be converted into 32-bit binary data "00000000000010101011111000000100", and the upper 16-bit data "0000000000001010" may be converted into decimal data "10" as a bucket identifier, and the lower 16-bit data "1011111000000100" may be converted into decimal data "48644" as bucket content. If corresponding bucket identifiers and bucket contents are obtained, the user identifier may be restored.

[0029] In some examples, bitmap identifiers, bucket identifiers, and bucket contents having corresponding relationships can be stored according to the corresponding relationship between the bitmap identifier and the bucket identifier, and the corresponding relationship between the bitmap identifier and the bucket identifier and the bucket contents. In the corresponding relationship between the bitmap identifier and the bucket identifier, the bitmap identifier and the bucket identifier are regarded as two objects, and the corresponding relationship between the two objects is described. In the corresponding relationship between the bitmap identifier and the bucket identifier and the bucket contents, the bitmap identifier and the bucket identifier can be regarded as one object, and the bucket contents are regarded as another object, and the corresponding relationship between the two objects is described. Specifically, the target storage space stores pairs of keys (i.e., key) and values ​​(i.e., value). In the pairs of keys and values, when the key includes a bitmap identifier, the value includes the bucket identifier corresponding to the key; when the key includes a bitmap identifier and a bucket identifier, the value includes the bucket contents corresponding to the key.

[0030] For example, user data may be shown in Table 1 below:

[0031] Table 1

[0032] Table 1 shows the user ID, user gender, user age, and the bucket ID and bucket content obtained based on the user ID. If the user ID is classified by gender and age, the keys and values ​​stored in the target storage space can be shown in Table 2 below:

[0033] Table 2

[0034] Among them, "man" and "woman" are bitmap identifiers, indicating the user's gender as "male" and "female," respectively; "age_0_20," "age_21_30," "age_31_40," "age_41_50," and "age_51" are bitmap identifiers, indicating the user's age as "0 to 20 years old," "21 to 30 years old," "31 to 40 years old," "41 to 50 years old," and "over 51 years old," respectively. Keys that include bitmap identifiers and bucket identifiers are represented in Table 2 as "bitmap identifier:bucket identifier" format. For example, "man" in "man:0" is the bitmap identifier, and "0" in "man:0" is the bucket identifier. Other keys in Table 2 that include bitmap identifiers and bucket identifiers are not listed here. Figure 2 is a schematic diagram of an example of bucket storage corresponding to part of the data in Table 2 provided in an embodiment of the present application. As shown in Figure 2, the bucket identifiers corresponding to males include 488, 61, 46, 10 and 0, the bucket contents corresponding to the male bucket identifier 488 include 34434, the bucket contents corresponding to the male bucket identifier 61 include 12306, the bucket contents corresponding to the male bucket identifier 46 include 35347, the bucket contents corresponding to the male bucket identifier 10 include 48644, and the bucket contents corresponding to the male bucket identifier 0 include 1 and 1001; the bucket identifiers corresponding to females include 488, 46, 15 and 0, the bucket contents corresponding to the female bucket identifier 488 include 28434, the bucket contents corresponding to the female bucket identifier 46 include 35346, the bucket contents corresponding to the female bucket identifier 15 include 26961, and the bucket contents corresponding to the female bucket identifier 0 include 4001.

[0035] The bitmap identifier associated with the logical operation indicated by the calculation logic information is the target bitmap identifier indicated by the calculation logic information. According to the target bitmap identifier and the correspondence between the bitmap identifier and the bucket identifier in the target storage space, the bucket identifier corresponding to the target bitmap identifier, namely the target bucket identifier, can be determined. For example, the calculation logic information in the query request indicates the intersection of user data of users aged between 21 and 30 and user data of users whose gender is male. If the data stored in the target storage space is as shown in Table 2, the corresponding target bitmap identifiers include "man" and "age_21_30", and the target bucket identifiers corresponding to the target bitmap identifier may include "0", "10", "46", "61", and "488".

[0036] In step S103, the bucket content corresponding to the target bucket identifier in the target storage space is processed according to the logical operation indicated by the calculation logic information to obtain the query result and feed it back.

[0037] For different logical operations, the processing performed on the bucket content corresponding to the target bucket identifier is also different. That is, the processing performed on the bucket content corresponding to the target bucket identifier is related to the logical operation indicated by the calculation logic information. By processing the bucket content corresponding to the target bucket identifier, at least part of the bucket content corresponding to the target bucket identifier can be obtained. Based on the target bucket identifier and at least part of the bucket content obtained after processing, the user data can be restored. A query result can be generated based on the restored user data and fed back to the initiator of the query request. The query result may include the bucket content obtained after processing and the user data restored by the target bucket identifier. The query result may also include, but is not limited to, other information related to the restored user data, such as the quantity of restored user data.

[0038] In an embodiment of the present application, bitmap identifiers, bucket identifiers, and bucket contents with corresponding relationships can be stored in advance in the target storage space. The bitmap identifier is used to characterize the classification of user data, and the bucket identifier and bucket content can be used to restore the user data. According to the target bitmap identifier indicated by the calculation logic information in the query request, the target bucket identifier is determined in the target storage space, and the bucket content corresponding to the target bucket identifier is processed according to the logical operation indicated by the calculation logic information to obtain the query result corresponding to the query request. In the above-mentioned user data query process, user data is stored in buckets through bitmap identifiers, bucket identifiers, and bucket contents. Each query can directly obtain the bucket identifier and bucket content associated with this query and perform the data operations required for the logical operation indicated by the query request. There is no need to traverse all the data, which can shorten the time required for user data query and improve query efficiency.

[0039] In some embodiments, the complex logical operation indicated by the computational logic information can be split into binary logical operations, thereby determining the data operations that need to be performed on the bucket contents based on the binary logical operations and the target bucket identifiers associated with the binary logical operations, so as to obtain query results through the data operations. FIG3 is a flowchart of a method for querying user data provided by another embodiment of the present application. FIG3 differs from FIG1 in that step S103 in FIG1 can be specifically refined into steps S1031 to S1034 in FIG3.

[0040] In step S1031 , the logical operation indicated by the calculation logic information is split into binary logical operations.

[0041] In some cases, the logical operation indicated by the calculation logic information is relatively complex and has three or more associated target bitmap identifiers. In this case, to facilitate processing, the logical operation indicated by the calculation logic information can be split into binary logical operations, with each binary logical operation being associated with two target bitmap identifiers, which is more convenient for processing.

[0042] In step S1032 , a target bucket identifier associated with the binary logic operation is obtained.

[0043] There are two target bitmap identifiers associated with the binary logic operation, and the target bucket identifiers corresponding to the two obtained target bitmap identifiers are the target bucket identifiers associated with the binary logic operation.

[0044] In step S1033 , a data operation is determined according to the binary logical operation and the target bucket identifier associated with the binary logical operation.

[0045] Different binary logical operations may correspond to different data operations performed on the bucket contents of the target bucket identifier. In some examples, the data operation may be determined based on whether the binary logical operation and the target bucket identifier associated with the binary logical operation are the same. If the target bucket identifiers associated with the binary logical operation are the same, the data operation performed on the bucket contents corresponding to the same target bucket identifier may include the binary logical operation; if the target bucket identifiers associated with the binary logical operation are different, the data operation performed on the bucket contents corresponding to different target bucket identifiers may include additional operations, which may include but are not limited to copy operations, negation operations, and discard operations.

[0046] In step S1034, data operations are performed on the bucket contents corresponding to the target bucket identifier to obtain query results and provide feedback.

[0047] By performing the data operation determined in step S1033 on the bucket content corresponding to the target bucket identifier, at least part of the bucket content can be obtained. The obtained bucket content and the target bucket identifier corresponding to the obtained bucket content can be used to restore the user data, and the query result can be obtained based on the restored user data.

[0048] In some examples, to further improve user data query speed, when the logical operations indicated by the computational logic information include multiple binary logical operations, data operations can be performed in parallel on the bucket contents corresponding to the target bucket identifiers associated with the multiple binary logical operations. It should be noted that the binary logical operations corresponding to the parallelizable data operations are independent of each other. If one binary logical operation requires the result of another binary logical operation, the data operations corresponding to the two binary logical operations cannot be executed in parallel.

[0049] To facilitate understanding of the above content, a specific example is used here for illustration. The user data is shown in Table 1 above, and the storage method of the user data in the target storage space is shown in Table 2 above. If the logical operation indicated by the calculation logic information in the query request is to find the user data of users whose gender is male and whose age is 0 to 20 years old or 31 to 40 years old, then the logical operation can be split into two binary logical operations. The first binary logical operation is to find the union of the user data of users whose age is 0 to 20 years old and the user data of users whose age is 31 to 40 years old, and the second binary logical operation is to find the intersection of the result of the first binary logical operation and the user data of users whose gender is male.

[0050] The first binary logic operation process may include steps a1 to a3.

[0051] In step a1, the target bitmap identifier "age_0_20" corresponding to the age of 0 to 20 and the target bitmap identifier "age_31_40" corresponding to the age of 31 to 40 are obtained. The target bucket identifier corresponding to the target bitmap identifier "age_0_20" is "15", and the target bucket identifiers corresponding to the target bitmap identifier "age_31_40" are "0" and "46".

[0052] In step a2, the operator of the first binary logical operation is a union operation; for the same target bucket identifier, all bucket contents corresponding to the same target bucket identifier can be taken, that is, data processing of the union operation can be performed; for different target bucket identifiers, the bucket contents corresponding to different target bucket identifiers can be copied, that is, data processing of the copy operation can be performed.

[0053] In step a3, the target bucket IDs in this example are different. Based on the data processing determined in step a2, the bucket contents "26961" corresponding to "age_0_20:15," "4001" corresponding to "age_31_40:0," and "35347" corresponding to "age_31_40:46" are copied in parallel. The result of the first binary logic operation includes bucket ID "15" and bucket content "26961," bucket ID "0" and bucket content "4001," and bucket ID "46" and bucket content "35347."

[0054] The second binary logic operation process may include steps b1 to b3.

[0055] In step b1, a target bitmap identifier "man" of male gender is obtained, and the target bucket identifiers corresponding to the target bitmap identifier "man" are "0", "10", "46", "61" and "488".

[0056] In step b2, the operator of the second binary logical operation is an intersection operation; for a target bucket identifier that is the same as the bucket identifier in the result of the first binary logical operation, the bucket content corresponding to the same target bitmap identifier can be taken for intersection operation; for a target bucket identifier that is different from the bucket identifier in the result of the first binary logical operation, the bucket content corresponding to the target bucket identifier is discarded.

[0057] In step b3, the same target bucket identifiers in this example include "0" and "46". According to the data processing determined in step b2, the bucket content "4001" corresponding to "age_31_40:0" in the result of the first binary logical operation has no intersection with the bucket contents "1" and "1001" corresponding to "man:0". The intersection of the bucket content "35347" corresponding to "age_31_40:46" and the bucket content "35347" corresponding to "man:46" in the result of the first binary logical operation is "35347". In this example, the buckets corresponding to different target bitmap identifiers "10", "61" and "488" are discarded. The result of the second binary logical operation includes the bucket identifier "46" and the bucket content "35347".

[0058] Based on the result of the second binary logical operation, it can be determined that the query result includes the bucket identifier "46" and the user identifier "c00003050003" corresponding to the bucket content "35347". The query result can also include the number of user data found. In this example, the number of user data found in the query result is 1.

[0059] In some embodiments, the corresponding bitmap identifiers, bucket identifiers, and bucket contents in the target storage space may be updated based on the query results obtained from each query request. FIG4 is a flowchart of a method for querying user data provided in another embodiment of the present application. FIG4 differs from FIG1 in that the method for querying user data shown in FIG4 may further include steps S104 and S105.

[0060] In step S104, a new bitmap identifier is generated according to the calculation logic information, and a corresponding relationship is established between the bucket identifier and the bucket content corresponding to the query result and the new bitmap identifier.

[0061] In step S105 , the new bitmap identifier with the established corresponding relationship, the bucket identifier corresponding to the query result, and the bucket content are stored in the target storage space.

[0062] After obtaining the query results, a new bitmap identifier can be generated based on the computational logic information in the query request corresponding to the query results. The new bitmap identifier can represent the computational logic information. A corresponding relationship is established between the bucket identifier and bucket content corresponding to the query results and the new bitmap identifier, and the bucket content is stored in the target storage space. If a query request including the computational logic information is received again in the future, the bucket identifier and bucket content corresponding to the new bitmap identifier corresponding to the computational logic information can be directly obtained, thereby quickly obtaining the query results and further improving the speed and efficiency of user data queries. For example, in the above example, the computational logic information indicates that user data of users with male gender and ages between 0 and 20 years old or between 31 and 40 years old should be found. A new bitmap identifier "Crowd0001" can be generated. The user data that can be restored based on the bucket identifier and bucket content corresponding to the bitmap identifier "Crowd0001" includes user data of users with male gender and ages between 0 and 20 years old or between 31 and 40 years old.

[0063] In some embodiments, bucket identifiers and bucket contents may be stored in buckets, and according to the relationship between the amount of data of the bucket identifier, the amount of data of the bucket content, and the size of the pre-set amount of bucket data, the bucket identifiers and bucket contents may be stored in a corresponding storage method to reduce the space occupied by the bucket identifiers and bucket contents in the target storage space. For example, FIG5 is a schematic diagram of an example of multiple storage methods for bucket identifiers and bucket contents provided in an embodiment of the present application. As shown in FIG5 , the D2 data corresponding to the D1 data may be stored in a compressed form, a bitmap form, and an array form; in FIG5 , the D2 data corresponding to data No. 65534 in the D1 data is stored in a compressed form, wherein (0,10) means that data No. 0 to data No. 10 are stored in a compressed form, and similarly, (30,6000) means that data No. 30 to data No. 6000 are stored in a compressed form; in FIG5 , the D2 data corresponding to data No. 14672 in the D1 data is stored in a bitmap form; in FIG5 , the D2 data corresponding to data No. 2 in the D1 data is stored in an array form. In the case where the D1 data includes a bitmap identifier, the D2 data may include a bucket identifier; in the case where the D1 data includes a bucket identifier, the D2 data may include bucket contents.

[0064] In the target storage space, if the number of bucket identifiers corresponding to the bitmap identifier does not exceed the preset bucket data volume, the bucket identifiers corresponding to the bitmap identifier are stored in array form; if the number of bucket identifiers corresponding to the bitmap identifier exceeds the preset bucket data volume, the bucket identifiers corresponding to the bitmap identifier are stored in bitmap form. The preset bucket data volume can be determined based on the single bucket capacity of the bucket. Under the condition of single bucket capacity, if the number of bucket identifiers does not exceed the preset bucket data volume, the space occupied by storing the bucket identifiers in array form is less than the space occupied by storing the bucket identifiers in bitmap form; if the number of bucket identifiers exceeds the preset bucket data volume, the space occupied by storing the bucket identifiers in array form is greater than the space occupied by storing the bucket identifiers in bitmap form.

[0065] Similarly, in the target storage space, if the number of bucket contents corresponding to the bitmap identifier and the bucket identifier does not exceed the preset bucket data volume, the bucket contents corresponding to the bitmap identifier and the bucket identifier are stored in array form; if the number of bucket contents corresponding to the bitmap identifier and the bucket identifier exceeds the preset bucket data volume, the bucket contents corresponding to the bitmap identifier and the bucket identifier are stored in bitmap form. The preset bucket data volume corresponding to the bucket content and the preset bucket data volume corresponding to the bucket identifier may be the same. Under the condition of a single bucket capacity, if the number of bucket contents does not exceed the preset bucket data volume, the space occupied by storing the bucket contents in array form is less than the space occupied by storing the bucket contents in bitmap form; if the number of bucket contents exceeds the preset bucket data volume, the space occupied by storing the bucket contents in array form is greater than the space occupied by storing the bucket contents in bitmap form.

[0066] For example, if the capacity of a single bucket is 8KB (i.e., kilobytes), the preset bucket data size can be 4096. When the number of bucket identifiers or bucket contents is less than or equal to 4096, the bucket identifiers or bucket contents of short type variables can be stored in an array in an orderly manner, and each bucket can store 4096*2Byte (i.e., bytes) of data; when the number of bucket identifiers or bucket contents is greater than 4096, the bucket identifiers or bucket contents can be stored in a bitmap format, and each bucket can store 65536*1bit (i.e., bits) of data.

[0067] In some examples, among bucket identifiers stored in bitmap form, if there are consecutive bucket identifiers and the number of consecutive bucket identifiers exceeds the compression standard number, the consecutive bucket identifiers are stored in a compressed manner. Among bucket contents stored in bitmap form, if there are consecutive bucket contents and the number of consecutive bucket contents exceeds the compression standard number, the consecutive bucket contents are stored in a compressed manner. The compression standard number is the minimum value of the amount of data such that the space occupied by data stored in compressed form is less than the space occupied by data stored in uncompressed form. Bucket identifiers or bucket contents can only be stored in compressed form if they are consecutive, and in order to minimize the space occupied by data storage, when the number of consecutive bucket identifiers or the number of consecutive bucket contents exceeds the compression standard number, the bucket identifiers or bucket contents are stored in a compressed manner. In this case, the space occupied by storing the bucket identifiers or bucket contents in a compressed manner is less than the space occupied by storing the bucket identifiers or bucket contents in an uncompressed manner. Correspondingly, if the number of consecutive bucket identifiers or the number of consecutive bucket identifiers in the bucket identifiers or bucket contents stored in bitmap form does not exceed the compression standard number, in this case, the space occupied by storing the bucket identifiers or bucket contents in a non-compressed manner is smaller than the space occupied by storing the bucket identifiers or bucket contents in a compressed manner.

[0068] By taking into account the relationship between the number of bucket identifiers, the number of bucket contents and the preset bucket data volume, a storage method is selected that can reduce the space required to store bucket identifiers and bucket contents, and the space occupied by storing bucket identifiers and bucket contents that can represent user data is minimized. Especially in scenarios where the amount of user data is very large, such as scenarios with hundreds of millions of user data, the space occupied by user data storage under each bitmap identifier can be greatly reduced. For example, user data under a single bitmap identifier of 1 billion users requires hundreds of megabytes of space for storage if the storage method of bucket identifiers and bucket contents in the embodiment of the present application is not adopted. However, the space occupied by user data under a single bitmap identifier of 1 billion users can be reduced by more than 73% if the storage method of bucket identifiers and bucket contents in the embodiment of the present application is adopted, thereby greatly reducing the space occupied by user data.

[0069] In the above embodiments, the target storage space may include an in-memory database and / or a cache space. The in-memory database may store a full set of corresponding bitmap identifiers, bucket identifiers, and bucket contents. The cache space may store a portion of the corresponding bitmap identifiers, bucket identifiers, and bucket contents. In some examples, if the target storage space includes a cache space, the cache space stores corresponding bitmap identifiers, bucket identifiers, and bucket contents that meet predetermined conditions. The predetermined conditions may be set based on the scenario, needs, experience, etc. For example, the predetermined conditions may include the cache space storing a specified bitmap identifier and the bucket identifiers and bucket contents corresponding to the specified bitmap identifier. For another example, the predetermined conditions may include the cache space storing the L most frequently queried bitmap identifiers and the bucket identifiers and bucket contents corresponding to the L most frequently queried bitmap identifiers, where L is a positive integer. Storing the frequently queried bitmap identifiers and their corresponding bucket identifiers and bucket contents in the cache space can further improve the speed and efficiency of user data queries.

[0070] The following example illustrates the logical architecture of the user data query method provided in an embodiment of the present application. FIG6 is a logical diagram of an example of the application architecture of the user data query method provided in an embodiment of the present application. As shown in FIG6 , the user data query method may involve an offline database 21 , an in-memory database 22 , and a memory 23 .

[0071] Various types of user data can be obtained from the offline database 21, and corresponding bitmap identifiers, bucket identifiers, and bucket contents can be generated based on the user data. For example, as shown in Figure 6, user data such as the user basic information table 211 and the user transaction data table 212 can be obtained to generate a tag-bitmap data table 213 containing corresponding bitmap identifiers, bucket identifiers, and bucket contents. The offline database 21 can synchronize the corresponding bitmap identifiers, bucket identifiers, and bucket contents with the in-memory database 22.

[0072] The memory database 22 stores bitmap identifiers, bucket identifiers and bucket contents with corresponding relationships. The corresponding relationships among the bitmap identifiers, bucket identifiers and bucket contents in the memory database 22 can be reflected through the corresponding relationship 221 between the bitmap identifier and the bucket identifier, and the corresponding relationship 222 between the bitmap identifier-bucket identifier and the bucket content.

[0073] The memory 23 may include a request parsing function 231, a bitmap logical operation function 232, and a cache space 233. The cache space 233 stores a portion of the corresponding bitmap identifiers, bucket identifiers, and bucket contents obtained from the in-memory database 22. The memory 23 may receive a query request, parse the query request using the request parsing function 231, and pass the resulting logical operation to the bitmap logical operation function 232. Based on the logical operation, the bitmap logical operation function 232 obtains the bucket identifiers and bucket contents required for the logical operation from the in-memory database 22 or the cache space 233, performs the logical operation, and obtains the query result. The bitmap logical operation function 232 may feed the query result back to the request parsing function 231, which then feeds the query result back to the querying party. The request parsing function 231 may synchronize the new bitmap identifier corresponding to the query request with the in-memory database 22. The bitmap logical operation function 232 may also synchronize the query result with the in-memory database 22. The in-memory database 22 may store the new bitmap identifier and the bucket identifiers and bucket contents corresponding to the query result.

[0074] In an embodiment of the present application, user data can be stored separately according to bucket identifiers and bucket contents, and the efficient responsiveness of memory storage, the data sharding acquisition brought by bucket identifiers and bucket contents, and the convenient bit operations can be utilized to achieve efficient, accurate real-time, and multi-type user query and selection in massive user scenarios, thereby improving the effectiveness of refined management and analysis of user data.

[0075] FIG7 is a schematic diagram of the structure of the user data query device provided in one embodiment of the present application. As shown in FIG7 , the user data query device 300 may include a receiving module 301 , a logic processing module 302 , and a sending module 303 .

[0076] The receiving module 301 may be configured to receive a query request, wherein the query request includes computing logic information.

[0077] The logic processing module 302 can be used to determine the target bucket identifier in the target storage space according to the target bitmap identifier indicated by the calculation logic information. The target storage space stores corresponding bitmap identifiers, bucket identifiers and bucket contents. The bitmap identifier is used to represent the classification of user data, and the bucket identifier and bucket content are used to restore the user data; and it can be used to process the bucket content corresponding to the target bucket identifier in the target storage space according to the logical operation indicated by the calculation logic information to obtain the query result.

[0078] The sending module 303 may be used to feed back the query result.

[0079] In an embodiment of the present application, bitmap identifiers, bucket identifiers, and bucket contents with corresponding relationships can be stored in advance in the target storage space. The bitmap identifier is used to characterize the classification of user data, and the bucket identifier and bucket content can be used to restore the user data. According to the target bitmap identifier indicated by the calculation logic information in the query request, the target bucket identifier is determined in the target storage space, and the bucket content corresponding to the target bucket identifier is processed according to the logical operation indicated by the calculation logic information to obtain the query result corresponding to the query request. In the above-mentioned user data query process, user data is stored in buckets through bitmap identifiers, bucket identifiers, and bucket contents. Each query can directly obtain the bucket identifier and bucket content associated with this query and perform the data operations required for the logical operation indicated by the query request. There is no need to traverse all the data, which can shorten the time required for user data query and improve query efficiency.

[0080] In some embodiments, the logic processing module 302 can be specifically used to: split the logical operation indicated by the calculation logic information into binary logical operations; obtain the target bucket identifier associated with the binary logical operation; determine the data operation based on the binary logical operation and the target bucket identifier associated with the binary logical operation; perform data operations on the bucket content that has a corresponding relationship with the target bucket identifier, obtain query results and feedback.

[0081] In some examples, the logic processing module 302 may be specifically configured to determine a data operation based on a binary logic operation and whether target bucket identifiers associated with the binary logic operation are the same.

[0082] In some examples, the logic processing module 302 may be specifically configured to perform data operations on bucket contents corresponding to target bucket identifiers associated with the multiple binary logic operations in parallel when the logic operation indicated by the calculation logic information includes multiple binary logic operations.

[0083] In some embodiments, the logic processing module 302 can also be used to: generate a new bitmap identifier based on the calculation logic information, and establish a corresponding relationship between the bucket identifier and bucket content corresponding to the query result and the new bitmap identifier; store the new bitmap identifier with the corresponding relationship, the bucket identifier and bucket content corresponding to the query result in the target storage space.

[0084] In some examples, the target storage space stores pairs of keys and values. In the key and value pairs, when the key includes a bitmap identifier, the value includes a bucket identifier corresponding to the key; when the key includes a bitmap identifier and a bucket identifier, the value includes the bucket content corresponding to the key.

[0085] In some examples, the target storage space includes an in-memory database and / or a cache space. In the case where the target storage space includes a cache space, the cache space stores bitmap identifiers, bucket identifiers, and bucket contents that have corresponding relationships and meet predetermined conditions.

[0086] In some embodiments, the user data includes high N-bit data and low M-bit data, the bucket identifier is obtained based on the high N-bit data of the user data, and the bucket content is obtained based on the low M-bit data of the user data, where N and M are both positive integers.

[0087] In the target storage space, if the number of bucket identifiers corresponding to the bitmap identifier does not exceed the preset bucket data volume, the bucket identifiers corresponding to the bitmap identifier are stored in array form; if the number of bucket identifiers corresponding to the bitmap identifier exceeds the preset bucket data volume, the bucket identifiers corresponding to the bitmap identifier are stored in bitmap form.

[0088] In the target storage space, if the number of bucket contents corresponding to the bitmap identifier and the bucket identifier does not exceed the preset bucket data volume, the bucket contents corresponding to the bitmap identifier and the bucket identifier are stored in array form; if the number of bucket contents corresponding to the bitmap identifier and the bucket identifier exceeds the preset bucket data volume, the bucket contents corresponding to the bitmap identifier and the bucket identifier are stored in bitmap form.

[0089] In some examples, if there are consecutive bucket identifiers in the bucket identifiers stored in bitmap form and the number of consecutive bucket identifiers exceeds the compression standard, the consecutive bucket identifiers are stored in a compressed manner. If there are consecutive bucket contents in the bucket contents stored in bitmap form and the number of consecutive bucket contents exceeds the compression standard, the consecutive bucket contents are stored in a compressed manner.

[0090] In a third aspect, the present application further provides an electronic device. FIG8 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. As shown in FIG8 , the electronic device 400 includes a memory 401, a processor 402, and a computer program stored in the memory 401 and executable on the processor 402.

[0091] In some examples, the processor 402 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0092] The memory 401 may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Therefore, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method for querying user data in the embodiment of the present application.

[0093] The processor 402 reads the executable program code stored in the memory 401 to run a computer program corresponding to the executable program code, so as to implement the user data query method in the above embodiment.

[0094] In some examples, the electronic device 400 may further include a communication interface 403 and a bus 404. As shown in FIG8 , the memory 401, the processor 402, and the communication interface 403 are connected via the bus 404 and communicate with each other.

[0095] The communication interface 403 is mainly used to implement communication between the modules, devices, units and / or equipment in the embodiment of the present application. Input devices and / or output devices can also be connected through the communication interface 403.

[0096] The bus 404 includes hardware, software, or both that couples the components of the electronic device 400 to each other. By way of example, and not limitation, the bus 404 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-E) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of the above. Where appropriate, the bus 404 may include one or more buses. Although embodiments herein describe and illustrate a particular bus, this application contemplates any suitable bus or interconnect.

[0097] In a fourth aspect, the present application further provides a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the user data query method of the above-mentioned embodiment can be implemented, and the same technical effect can be achieved. To avoid repetition, the above-mentioned computer-readable storage medium may include a non-transitory computer-readable storage medium, such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., and is not limited here.

[0098] An embodiment of the present application provides a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the user data query method in the above embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0099] It should be understood that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. For device embodiments, equipment embodiments, computer-readable storage medium embodiments, and computer program product embodiments, the relevant parts can be referred to the description section of the method embodiment. This application is not limited to the specific steps and structures described above and shown in the figures. Those skilled in the art can make various changes, modifications and additions, or change the order of the steps after understanding the spirit of this application. In addition, for the sake of brevity, a detailed description of known method technologies is omitted here.

[0100] Aspects of the present application have been described above with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that these instructions executed via the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. This processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit. It is also understood that each box in the block diagram and / or the flowchart and the combination of the boxes in the block diagram and / or the flowchart can also be implemented by the dedicated hardware that performs the specified function or action, or can be implemented by the combination of dedicated hardware and computer instructions.

[0101] Those skilled in the art should understand that the above embodiments are illustrative rather than restrictive. Different technical features appearing in different embodiments can be combined to achieve beneficial effects. Based on a study of the drawings, the specification and the claims, those skilled in the art should be able to understand and implement other variations of the disclosed embodiments. In the claims, the term "comprising" does not exclude other devices or steps; the quantifier "one" does not exclude a plurality; the terms "first" and "second" are used to identify names rather than to indicate any specific order. Any figure marks in the claims should not be understood as limiting the scope of protection. The functions of multiple parts appearing in the claims can be implemented by a separate hardware or software module. The fact that certain technical features appear in different dependent claims does not mean that these technical features cannot be combined to achieve beneficial effects.

Claims

1. A method for querying user data, comprising: Receiving a query request, the query request including calculation logic information; Determining a target bucket identifier in a target storage space according to the target bitmap identifier indicated by the calculation logic information, where the target storage space stores bitmap identifiers, bucket identifiers, and bucket contents with corresponding relationships, the bitmap identifier is used to represent the classification of user data, and the bucket identifier and bucket content are used to restore user data; Processing the bucket content corresponding to the target bucket identifier in the target storage space according to the logical operation indicated by the calculation logic information, obtaining a query result and feeding it back.

2. The method according to claim 1, wherein, The processing the bucket content corresponding to the target bucket identifier in the target storage space according to the logical operation indicated by the calculation logic information, obtaining a query result and feeding it back includes: Splitting the logical operation indicated by the calculation logic information into binary logical operations; Obtaining the target bucket identifier associated with the binary logical operation; Determining a data operation according to the binary logical operation and the target bucket identifier associated with the binary logical operation; Performing the data operation on the bucket content corresponding to the target bucket identifier, obtaining a query result and feeding it back.

3. The method according to claim 2, wherein The determining a data operation according to the binary logical operation and the target bucket identifier associated with the binary logical operation includes: Determining a data operation according to whether the binary logical operation and the target bucket identifier associated with the binary logical operation are the same.

4. The method according to claim 2, wherein, The performing the data operation on the bucket content corresponding to the target bucket identifier includes: When the logical operation indicated by the calculation logic information includes multiple binary logical operations, performing the data operation on the bucket content corresponding to the target bucket identifiers associated with the multiple binary logical operations in parallel.

5. The method according to claim 1, further comprising: Generating a new bitmap identifier according to the calculation logic information, and establishing a corresponding relationship between the bucket identifier and bucket content corresponding to the query result and the new bitmap identifier; Storing the new bitmap identifier, the bucket identifier and bucket content corresponding to the query result with the established corresponding relationship into the target storage space.

6. The method according to claim 1, wherein The target storage space stores pairs of keys and values; Among the pairs of keys and values, when the key includes a bitmap identifier, the value includes the bucket identifier corresponding to the key, and when the key includes a bitmap identifier and a bucket identifier, the value includes the bucket content corresponding to the key.

7. The method according to claim 1, wherein The target storage space includes an in-memory database and / or a cache space; When the target storage space includes a cache space, the cache space stores bitmap identifiers, bucket identifiers, and bucket contents with corresponding relationships that meet a predetermined condition.

8. The method according to claim 1, wherein, The user data includes high N-bit data and low M-bit data, the bucket identifier is obtained based on the high N-bit data of the user data, the bucket content is obtained based on the low M-bit data of the user data, and both N and M are positive integers; In the target storage space, If the number of bucket identifiers corresponding to the bitmap identifier does not exceed the preset bucket data volume, the bucket identifiers corresponding to the bitmap identifier are stored in an array form. If the number of bucket identifiers corresponding to the bitmap identifier exceeds the preset bucket data volume, the bucket identifiers corresponding to the bitmap identifier are stored in a bitmap form. If the number of bucket contents corresponding to the bitmap identifier and the bucket identifier does not exceed the preset bucket data volume, the bucket contents corresponding to the bitmap identifier and the bucket identifier are stored in an array form. If the number of bucket contents corresponding to the bitmap identifier and the bucket identifier exceeds the preset bucket data volume, the bucket contents corresponding to the bitmap identifier and the bucket identifier are stored in a bitmap form.

9. The method according to claim 8, wherein Among the bucket identifiers stored in the bitmap form, if there are consecutive bucket identifiers and the number of consecutive bucket identifiers exceeds the compression standard number, the consecutive bucket identifiers are stored in a compressed manner. Among the bucket contents stored in the bitmap form, if there are consecutive bucket contents and the number of consecutive bucket contents exceeds the compression standard number, the consecutive bucket contents are stored in a compressed manner.

10. A query device for user data, comprising: a receiving module, configured to receive a query request, where the query request includes calculation logic information; a logic processing module, configured to determine a target bucket identifier in a target storage space according to the target bitmap identifier indicated by the calculation logic information. In the target storage space, there are stored a bitmap identifier, a bucket identifier, and a bucket content with a corresponding relationship. The bitmap identifier is used to represent the classification of user data, and the bucket identifier and the bucket content are used to restore the user data. And, configured to process the bucket content corresponding to the target bucket identifier in the target storage space according to the logical operation indicated by the calculation logic information to obtain a query result; a sending module, configured to feedback the query result.

11. An electronic device, comprising: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the query method for user data according to any one of claims 1 to 9 is implemented.

12. A computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the query method for user data according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Label data processing method, device and equipment and storage medium

    CN112015775A

  • Data storage method and device, query method, electronic equipment and readable medium

    CN112559522A

  • Data processing method and related equipment

    CN115408381A

  • User data query method and device, equipment and storage medium

    CN117951137A

  • Data processing method, apparatus, and device for federated feature engineering, and medium

    WO2023040429A1

Cited By

  • Goods source pushing method, electronic equipment, storage medium and program product

    CN120821920A