A Compression Method and Device for User Portrait Tag Data
By splitting the user ID into high 8 bits, medium 8 bits and low 16 bits, and adopting different data types of storage methods, the problem of sharp increase in user portrait tag data storage space is solved, and storage space optimization and cost savings are achieved.
Patent Information
- Application Number
- CN202510630927.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-16
AI Technical Summary
In the prior art, the increase in the number of user ids corresponding to the user portrait tag leads to a sharp increase in storage space, resulting in an increase in storage costs.
Split the user id into high 8 bits, medium and high 8 bits and low 16 bits, and use different data types (B, L, N data types) for compression storage. Use high 8 bits and medium and high 8 bits as indexes, and select the optimal storage type according to actual conditions.
It effectively reduces the space requirement for storing user IDs and reduces storage costs.
Smart Images

Figure CN120196635B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data compression, and particularly relates to a method and device for compressing user portrait tag data. Background Art
[0002] With the rapid development of information technology, users generate a vast amount of data during the use of various applications. By analyzing the data value, mining user characteristics and behaviors, and organizing them into user portrait tags, the user portrait tags can help enterprises accurately locate user needs, optimize product design, and improve market operation efficiency.
[0003] In the prior art, multiple corresponding user IDs are queried through user portrait tags, and the multiple user IDs corresponding to the user portrait tags are pre-stored as a group for later processing. Multiple user portrait tags correspond to multiple groups of user IDs. When the number of users is huge, the number of user IDs corresponding to the user portrait tags will also increase sharply, resulting in a sharp increase in the data space for storing user IDs, causing the problem of increased storage costs. Summary of the Invention
[0004] Therefore, the present invention provides a method and device for compressing user portrait tag data to solve the problem of increased storage costs caused by the sharp increase in the data space for storing user IDs in the prior art.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] In a first aspect, the present invention provides a method for compressing user portrait tag data, including:
[0007] Obtaining at least one field to be queried from a database according to a user portrait tag;
[0008] Obtaining a plurality of corresponding user IDs as user IDs to be stored according to the field to be queried; the user ID is of int type, and each user ID is represented by 32 bits;
[0009] Splitting each user ID in the user IDs to be stored into a high 8-bit part, a middle-high 8-bit part, and a low 16-bit part;
[0010] Compressively storing the user IDs to be stored according to the high 8-bit part, the middle-high 8-bit part, and the low 16-bit part.
[0011] Further, the compressing and storing the user IDs to be stored according to the high 8-bit part, the middle-high 8-bit part, and the low 16-bit part includes:
[0012] Obtaining a first mapping sequence according to the high 8-bit part of each user ID in the user IDs to be stored; the storage structure of the first mapping sequence is a bit array;
[0013] Obtain a plurality of second mapping sequences according to the high 8 bits and the medium-high 8 bits of each user ID in the user IDs to be stored; the storage structure of the second mapping sequence is a bit array;
[0014] Store the low 16 bits of each user ID in the user IDs to be stored as the B data type, the L data type, or the N data type; the storage structure of the B data type is a bit array; the storage structure of the L data type is a binary tuple; the storage structure of the N data type is an integer.
[0015] Further, the array length of the first mapping sequence is 256; the obtaining of the first mapping sequence according to the high 8 bits of each user ID in the user IDs to be stored includes:
[0016] Construct a first initial array; the first initial array is a bit array with an array length of 256;
[0017] Convert the high 8 bits of each user ID in the user IDs to be stored into decimal to obtain a plurality of first array subscripts;
[0018] After setting the bit value corresponding to each of the first array subscripts in the first initial array to 1, obtain the first mapping sequence.
[0019] Further, the array length of the second mapping sequence is 256; the obtaining of a plurality of second mapping sequences according to the high 8 bits and the medium-high 8 bits of each user ID in the user IDs to be stored includes:
[0020] Obtain the number of different high 8 bits as the first number;
[0021] Construct a second initial array with the first number; the second initial array is a bit array with an array length of 256;
[0022] Assign values to the second initial array according to the high 8 bits and the medium-high 8 bits to obtain the second mapping sequences with the first number.
[0023] Further, the assigning values to the second initial array according to the high 8 bits and the medium-high 8 bits to obtain the second mapping sequences with the first number includes:
[0024] Obtain one of the different high 8 bits as the first high 8 bit;
[0025] Obtain the medium-high 8 bits of a plurality of user IDs with the same first high 8 bit to obtain a plurality of first medium-high 8 bits;
[0026] After converting each of the first medium-high 8 bits into decimal, obtain the corresponding second array subscripts;
[0027] After setting the bit value corresponding to each second array subscript in the second initial array to 1, the second mapping sequence corresponding to the first high 8 bits is obtained.
[0028] Further, the maximum array length of the B data type is 65536; storing the lower 16 bits of each user ID in the user ID to be stored as the B data type, the L data type, or the N data type includes:
[0029] Obtain the lower 16 bits of user IDs with the same high 8 bits and middle high 8 bits to get the first lower 16 bits of the second quantity;
[0030] Calculate in sequence the storage spaces occupied after storing the first lower 16 bits of the second quantity as the B data type, the L data type, and the N data type;
[0031] If the storage space occupied by storing the first lower 16 bits of the second quantity as the L data type is the smallest, then store the first lower 16 bits of the second quantity as the L data type;
[0032] If the storage space occupied by storing the first lower 16 bits of the second quantity as the N data type is the smallest, then store the first lower 16 bits of the second quantity as the N data type;
[0033] If the storage space occupied by storing the first lower 16 bits of the second quantity as the B data type is the smallest, then store the first lower 16 bits of the second quantity as the B data type.
[0034] Further, storing the first lower 16 bits of the second quantity as the L data type includes:
[0035] Store the first lower 16 bits of the second quantity in a data structure of a binary tuple; the expression of the L data type is L = {(x1, y1), (x2, y2), …, (x n , y n )}; where, there are a total of n binary tuples, each binary tuple represents a continuous interval, and the binary tuple (x i , y i ) represents the i-th continuous interval, i ∈ (1, n), x i represents the first lower 16 bits at the start of the i-th continuous interval, y i represents the step size of the i-th continuous interval, and both x i and y i are represented in 16-bit binary.
[0036] Further, storing the first lower 16 bits of the second quantity as the N data type includes:
[0037] Store the second quantity of the first lower 16 bits directly in binary in sequence; each of the first lower 16 bits consists of 16-bit binary numbers.
[0038] Further, storing the second quantity of the first lower 16 bits as the B data type includes:
[0039] Convert the second quantity of the first lower 16 bits to decimal to obtain the second quantity of third array subscripts;
[0040] Construct a third initial array; the third initial array is a bit array; the array length of the third initial array is the subscript value of the largest third array subscript plus 1;
[0041] Set the bit value corresponding to each of the third array subscripts in the third initial array to 1.
[0042] In a second aspect, the present invention provides a compression device for user portrait tag data, and the device includes:
[0043] A query field module, configured to obtain at least one query field to be queried from a database according to a user portrait tag;
[0044] A user ID obtaining module, configured to obtain a plurality of corresponding user IDs as user IDs to be stored according to the query field to be queried; the user ID is of the int type, and each user ID is represented by 32 bit positions;
[0045] A splitting module, configured to split each user ID in the user ID to be stored into a high 8 bits, a middle high 8 bits, and a lower 16 bits;
[0046] A compression storage module, configured to compress and store the user ID to be stored according to the high 8 bits, the middle high 8 bits, and the lower 16 bits.
[0047] The present invention adopts the above technical solutions and at least has the following beneficial effects:
[0048] Provided is a compression method and device for user portrait tag data, obtaining at least one query field to be queried from a database according to a user portrait tag, obtaining a plurality of corresponding user IDs as user IDs to be stored according to the query field to be queried, the user ID is of the int type, each user ID is represented by 32 bit positions, splitting each user ID in the user ID to be stored into a high 8 bits, a middle high 8 bits, and a lower 16 bits, and compressing and storing the user ID to be stored according to the high 8 bits, the middle high 8 bits, and the lower 16 bits; based on compressing and storing a plurality of user IDs corresponding to user portrait tags by the high 8 bits, the middle high 8 bits, and the lower 16 bits, it can effectively reduce the storage space for storing user IDs and save storage costs.
[0049] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0051] Figure 1 is a flowchart of a method for compressing user portrait tag data shown in an exemplary embodiment of the present invention;
[0052] Figure 2 is a schematic diagram A of the storage structure of a method for compressing user portrait tag data shown in an exemplary embodiment of the present invention;
[0053] Figure 3 is a schematic diagram B of the storage structure of a method for compressing user portrait tag data shown in an exemplary embodiment of the present invention;
[0054] Figure 4 is a schematic block diagram of a device for compressing user portrait tag data shown in an exemplary embodiment of the present invention.
[0055] The present invention will be further described below in conjunction with the drawings and specific embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be described in detail below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other implementation manners obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0057] In the prior art, multiple corresponding users are queried through user portrait tags, and the multiple user IDs corresponding to the user portrait tags are stored and used for subsequent arithmetic processing. Storing each user ID generally occupies 32 bits. When the number of user IDs is very large, the data space for storing user IDs increases sharply, increasing the storage cost.
[0058] An embodiment of the present invention provides a method and apparatus for compressing user portrait tag data. Each user ID in the user IDs to be stored is split into a high 8-bit part, a medium-high 8-bit part, and a low 16-bit part. Based on the high 8-bit part, the medium-high 8-bit part, and the low 16-bit part, multiple user IDs corresponding to the user portrait tags are compressed and stored, which can effectively reduce the storage space for storing user IDs. Compared with the traditional storage where each user ID occupies 32 bits, the more the number of user IDs, the more obvious the effect of saving storage space.
[0059] The methods and apparatuses in the present invention will be described below through specific embodiments.
[0060] Please refer to Figure 1 , Figure 1 which is a flowchart of a method for compressing user portrait tag data shown in an exemplary embodiment of the present invention. Refer to Figure 1 The method includes:
[0061] Step S11: Obtain at least one field to be queried from the database according to the user portrait tag;
[0062] Step S12: Obtain multiple corresponding user IDs as the user IDs to be stored according to the field to be queried; the user ID is of int type, and each user ID is represented by 32 bits;
[0063] Step S13: Split each user ID in the user IDs to be stored into a high 8-bit part, a medium-high 8-bit part, and a low 16-bit part;
[0064] Step S14: Compress and store the user IDs to be stored according to the high 8-bit part, the medium-high 8-bit part, and the low 16-bit part.
[0065] It should be noted that the technical solution provided in this embodiment can be loaded and used in an existing user analysis system or application in the form of a small program or a plug-in in specific practice, or in the form of a separate application, and the function of compressing user portrait tag data can be realized through an external interface; applicable scenarios include but are not limited to: user ID compression.
[0066] It can be understood that for the method provided in this embodiment, at least one field to be queried is obtained from the database according to the user portrait tag, multiple corresponding user IDs are obtained as the user IDs to be stored according to the field to be queried, the user ID is of int type, and each user ID is represented by 32 bits. Each user ID in the user IDs to be stored is split into a high 8-bit part, a medium-high 8-bit part, and a low 16-bit part, and the user IDs to be stored are compressed and stored according to the high 8-bit part, the medium-high 8-bit part, and the low 16-bit part. Based on the high 8-bit part, the medium-high 8-bit part, and the low 16-bit part, multiple user IDs corresponding to the user portrait tags are compressed and stored, which can effectively reduce the storage space for storing user IDs and save storage costs.
[0067] In specific practice, for step S11, "obtaining at least one field to be queried from the database according to the user profile label", where the user profile label is the user feature stored in the database, specifically presented in the form of field values. For example, if the user profile label is male gender, the field of the user feature is gender; if the user profile label is male consumers, the fields of the user features are the gender field and the consumption field.
[0068] In specific practice, for step S12, "obtaining multiple corresponding user IDs as the user IDs to be stored according to the fields to be queried", it includes: The process of obtaining user IDs can use existing database query algorithms, that is, it can be queried using SQL statements. For example, if the user profile label is male consumers, in the gender field, the field value 0 represents female, and 1 represents male; in the consumption field, 1 represents high consumption, 2 represents medium consumption, and 3 represents low consumption. Then, according to the field value of 1 in the gender field and the field value of 2 in the consumption field, multiple user IDs corresponding to male consumers are queried as the user IDs to be stored.
[0069] In specific practice, for step S13, "splitting each user ID in the user IDs to be stored into a high 8-bit part, a middle high 8-bit part, and a low 16-bit part", it includes: The user ID is of int type, and each user ID is represented by 32 bits. The 32 bits are split into a high 8-bit part, a middle high 8-bit part, and a low 16-bit part. For example, for the splitting process of user ID 1: The int representation of 1 is 00000000 00000000 00000000 00000001, its high 8-bit part and middle high 8-bit part are both 00000000, and the low 16-bit part is 00000000 000000001.
[0070] In specific practice, for step S14, "compressing and storing the user IDs to be stored according to the high 8-bit part, the middle high 8-bit part, and the low 16-bit part", it includes: Obtaining a first mapping sequence according to the high 8-bit part of each user ID in the user IDs to be stored; the storage structure of the first mapping sequence is a bit array; obtaining multiple second mapping sequences according to the high 8-bit part and the middle high 8-bit part of each user ID in the user IDs to be stored; the storage structure of the second mapping sequence is a bit array; storing the low 16-bit part of each user ID in the user IDs to be stored as data type B, data type L, or data type N; the storage structure of data type B is a bit array; the storage structure of data type L is a binary tuple; the storage structure of data type N is an integer.
[0071] Specifically, the array length of the first mapping sequence is 256; the first mapping sequence is obtained based on the high 8 bits of each user ID in the user IDs to be stored, including: constructing a first initial array; the first initial array is a bit array with an array length of 256; converting the high 8 bits of each user ID in the user IDs to be stored into decimal to obtain a plurality of first array subscripts; after setting the bit value corresponding to each first array subscript in the first initial array to 1, the first mapping sequence is obtained.
[0072] It should be noted that as shown in Table 1 below, the array length of the first initial array is 256 (subscripts are 0 - 255), where the high 8 bits of multiple user IDs in the user IDs to be stored are 00000000, 00000101, and 10011011. Among them, 00000000, 00000101, and 10011011 are sequentially converted into decimal to obtain 0, 5, and 155. Then, after setting the bit values corresponding to the subscripts 0, 5, and 155 in the first initial array to 1, the first mapping sequence is obtained.
[0073] Table 1
[0074] 。
[0075] Specifically, the array length of the second mapping sequence is 256; a plurality of second mapping sequences are obtained based on the high 8 bits and the middle high 8 bits of each user ID in the user IDs to be stored, including: obtaining the number of different high 8 bits as the first number; constructing a second initial array with the first number; the second initial array is a bit array with an array length of 256; obtaining one of the different high 8 bits as the first high 8 bit; obtaining the middle high 8 bits of multiple user IDs with the same first high 8 bit to obtain a plurality of first middle high 8 bits; after converting each first middle high 8 bit into decimal, obtaining the corresponding second array subscript; after setting the bit value corresponding to each second array subscript in the second initial array to 1, the second mapping sequence corresponding to the first high 8 bit is obtained.
[0076] It should be noted that the data structure of the second mapping sequence is referred to Figure 2 , Figure 2It is a schematic storage structure diagram A of a compression method for user portrait tag data shown in an exemplary embodiment of the present invention. Among them, there are 3 different high 8 - bits in the first mapping sequence, that is, the first quantity is 3; 3 second initial arrays are constructed, and the array length of each second initial array is 256; a high 8 - bit 00000000 is obtained as the first high 8 - bit. The middle 8 - bits with the same first high 8 - bit are 00000000, 00000011, and 10011100. Among them, 00000000, 00000011, and 10011100 are converted to decimal numbers 0, 3, and 156 in sequence. Then, the bit values of the 0, 3, and 156 sub - scripts of the second initial array corresponding to the first high 8 - bit are set to 1 to obtain the second mapping sequence, and this second mapping sequence corresponds to the first high 8 - bit.
[0077] Specifically, the maximum array length of the B data type is 65536; the lower 16 - bits of each user ID in the user ID to be stored are stored as the B data type, the L data type, or the N data type, including: obtaining the lower 16 - bits of the user IDs with the same high 8 - bit and middle 8 - bits to get the first lower 16 - bits of the second quantity; calculating the storage space occupied by storing the first lower 16 - bits of the second quantity as the B data type, the L data type, and the N data type in sequence; if the storage space occupied by storing the first lower 16 - bits of the second quantity as the L data type is the smallest, then store the first lower 16 - bits of the second quantity as the L data type; if the storage space occupied by storing the first lower 16 - bits of the second quantity as the N data type is the smallest, then store the first lower 16 - bits of the second quantity as the N data type; if the storage space occupied by storing the first lower 16 - bits of the second quantity as the B data type is the smallest, then store the first lower 16 - bits of the second quantity as the B data type.
[0078] It should be noted that the data of the L data type is composed of multiple binary groups. Each binary group stores one first lower 16 - bit or multiple consecutive first lower 16 - bits. When storing one first lower 16 - bit as a continuous interval, the step size of the binary group is 1. When storing multiple consecutive first lower 16 - bits as a continuous interval, the step size of the binary group is the number of multiple consecutive first lower 16 - bits.
[0079] It should be noted that the storage spaces occupied by successively calculating the first lower 16 bits of the second quantity when stored as B data type, L data type, and N data type are calculated as follows: calculating the storage space occupied by storing as N data type: the second quantity × 16 = the number of bit positions occupied by storing as N data type, that is, the storage space occupied by storing as N data type; calculating the storage space occupied by storing as B data type: the maximum integer value after converting the lower 16 bits to an integer plus 1 = the number of bit positions occupied by storing as B data type, that is, the storage space occupied by storing as B data type; calculating the storage space occupied by storing as L data type: the number of pairs × 16 × 2 = the number of bit positions occupied by storing as L data type, that is, the storage space occupied by storing as L data type. The pair (x, y) represents a continuous interval, where both x and y are represented by 16-bit binary. x represents the initial first lower 16 bits of the continuous interval, and y represents the step size of the continuous interval.
[0080] Specifically, the first lower 16 bits of the second quantity are stored in the data structure of pairs; the expression of the L data type is L={(x1,y1),(x2,y2),…,(x n ,y n )}; where there are a total of n pairs, and each pair represents a continuous interval. The pair (x i ,y i ) represents the i-th continuous interval, i ∈ (1, n). x i represents the first lower 16 bits at the start of the i-th continuous interval, and y i represents the step size of the i-th continuous interval. Both x i and y i are represented by 16-bit binary.
[0081] Specifically, storing the first lower 16 bits of the second quantity as N data type includes: successively and directly storing the first lower 16 bits of the second quantity in binary; each first lower 16 bits consists of 16-bit binary.
[0082] Specifically, storing the first lower 16 bits of the second quantity as B data type includes: converting the first lower 16 bits of the second quantity to decimal to obtain the subscripts of the second quantity of the third array; constructing the third initial array; the third initial array is a bit array; the array length of the third initial array is the subscript value of the maximum third array subscript plus 1; setting the bit value corresponding to each third array subscript in the third initial array to 1.
[0083] It should be noted that, as shown in Table 2 below, taking the first lower 16 bits with 00000000 00000001, 00000000 00000011, 00000000 00000101, and 11111111 11111111 as examples, they are converted to decimal numbers 1, 3, 5, and 65535 in sequence. Then the subscript value of the largest third array subscript is 65535, and the array length of the third initial array is 65536 (subscripts range from 0 to 65535). Set the bit values corresponding to the subscripts 1, 3, 5, and 65535 in the third initial array to 1, that is, store them in the B data type.
[0084] Table 2
[0085] 。
[0086] In a specific embodiment, please refer to Figure 3 , Figure 3 which is a schematic storage structure diagram B of a compression method for user portrait tag data shown in an exemplary embodiment of the present invention. Refer to Figure 3 for the storage structure description:
[0087] Take multiple user IDs obtained based on a certain user portrait tag as the user IDs to be stored. The user IDs in the user IDs to be stored are: all odd numbers within the range of... 84082689 - 84148223,... 84344832,... 84346432,... 84404835,... 100597760 - 100597859,... 100663290 - 100663295;
[0088] The user IDs in the user IDs to be stored are of integer int type, that is, each user ID requires 32 bit positions to represent. When compressing and storing the user IDs to be stored, store the user IDs to be stored as multiple data segments; each data segment uses the high 8 bits and the middle high 8 bits of the user ID as the index. Specifically: multiple user IDs with the same high 8 bits and middle high 8 bits are stored in one data segment; multiple user IDs with the same high 8 bits but different middle high 8 bits are stored in multiple data segments, and each data segment corresponds to a different middle high 8 bit; the index of each data segment is the high 8 bits and the middle high 8 bits, and then store the lower 16 bits of the corresponding multiple user IDs in the data segment in the B data type, L data type, or N data type.
[0089] Store the lower 16 bits of the multiple user IDs corresponding to each data segment in the data segment in the B data type, L data type, or N data type. Specifically: in the first case, see Figure 3Stored in the middle bit data structure, the user IDs to be stored are all odd numbers within the range of 84082689 - 84148223 in the user ID to be stored, namely 84082689, 84082691, 84082693... 84148223. The high 8 bits of these user IDs are 00000101, and the middle high 8 bits are 00000011. Since the high 8 bits and the middle high 8 bits of these user IDs are the same, the low 16 bits of these user IDs are stored in a data segment; after calculating the storage space occupied by storing the low 16 bits of these user IDs as three data types, storing as the B data type saves the most space, so the odd numbers within the range of 84082689 - 84148223 in the user ID are stored in the bit data structure.
[0090] The first case of the business scenario: Assume that the odd user IDs in the range of 84082689 - 84148223 in the user ID represent male users, and the even user IDs represent female users; the user IDs with the user portrait label of male users are discontinuous, the user IDs are dense and scattered, and the low 16 bits of the user IDs with the user portrait label of male users are stored in multiple data segments. Storing multiple data segments using the B data structure saves the most space, that is, it is applicable to the situation where the user IDs corresponding to the user portrait label are dense and the continuous interval is small.
[0091] The second case, see Figure 3 Stored in the middle integer data structure, the user IDs to be stored are 84344832,..., 84346432,..., 84404835,... The high 8 bits of these user IDs are 00000101, and the middle high 8 bits are 00000111. Since the high 8 bits and the middle high 8 bits of these user IDs are the same, the low 16 bits of these user IDs are stored in a data segment; after calculating the storage space occupied by storing the low 16 bits of these user IDs as three data types, storing as the N data type saves the most space, so the low 16 bits of these user IDs are stored as integers 0,..., 1600,..., 60003,..., where each integer is represented in 16 - bit binary.
[0092] The second case of the business scenario: For example, the user portrait label is the high - consumption label. The number of user IDs with the high - consumption label is small, sparse and basically discontinuous. The low 16 bits of these user IDs are stored in multiple data segments, and storing multiple data using the N data structure saves the most space.
[0093] The third case, see Figure 3Medium-step data structure storage. The user IDs to be stored are 100597760 - 100597859, ……, 100663290 - 100663295 in the user ID to be stored. The high 8 bits of these user IDs are 00000101, and the middle high 8 bits are 11111111. Since the high 8 bits and the middle high 8 bits of these user IDs are the same, the low 16 bits of these user IDs are stored in a data segment. After calculating the storage space occupied by the low 16 bits of these user IDs stored as three data types, it is most space-saving to store them as the L data type. The consecutive low 16 bits are stored in a binary tuple structure in sequence, that is, the low 16 bits of these user IDs are stored as: L={(0,100), ……, (65530,6)}.
[0094] The third case business scenario: For example, the user portrait label is a scalper label, and the user IDs are continuously in an interval. Scalpers may use automated tools to centrally register a batch of users, and these user IDs are continuous; it is most space-saving to store the low 16 bits of these user IDs in the data segment using the L data structure.
[0095] It should be noted that during compression storage, there are multiple data segments indexed by the high 8 bits and the middle high 8 bits. For example, multiple user IDs with the same high 8 bits and middle high 8 bits are stored in one data segment, and multiple user IDs with the same high 8 bits but different middle high 8 bits are stored in multiple data segments. Each data segment corresponds to a different middle high 8 bit. The index of the data segment is the high 8 bits and the middle high 8 bits. Then, according to the above three cases, the low 16 bits of multiple user IDs are stored in the data segment, that is, the compression is completed.
[0096] In specific practice, decompression is the reverse process of compression. Locate the data segment by the high 8 bits and the middle high 8 bits, and then convert the data corresponding to the low 16 bits into binary and splice it with the high 8 bits and the middle high 8 bits bit by bit into 32 bits. After converting the 32-bit binary into an int type, the user ID is obtained.
[0097] Please refer to Figure 4 , Figure 4 which is a schematic block diagram of a compression device for user portrait label data shown in an exemplary embodiment of the present invention. Refer to Figure 4 , the compression device 100 for user portrait label data includes:
[0098] A query field module 101, configured to obtain at least one query field to be queried from a database according to a user portrait label;
[0099] A user ID acquisition module 102, configured to obtain a corresponding plurality of user IDs as user IDs to be stored according to the query field to be queried; the user ID is of int type, and each user ID is represented by 32 bits;
[0100] The splitting module 103 is used to split each user ID in the user IDs to be stored into the high 8 bits, the middle-high 8 bits, and the low 16 bits;
[0101] The compression storage module 104 is used to compress and store the user IDs to be stored according to the high 8 bits, the middle-high 8 bits, and the low 16 bits.
[0102] It should be noted that the applicable scenarios of the device provided in this embodiment include, but are not limited to: the compression storage of multiple user IDs corresponding to user portrait tags.
[0103] It can be understood that the device provided in this embodiment obtains at least one field to be queried from the database according to the user portrait tag, obtains multiple corresponding user IDs as the user IDs to be stored according to the field to be queried. The user ID is of int type, and each user ID is represented by 32 bits. Each user ID in the user IDs to be stored is split into the high 8 bits, the middle-high 8 bits, and the low 16 bits, and the user IDs to be stored are compressed and stored according to the high 8 bits, the middle-high 8 bits, and the low 16 bits; by compressing and storing multiple user IDs corresponding to the user portrait tag based on the high 8 bits, the middle-high 8 bits, and the low 16 bits, it can effectively reduce the storage space for storing user IDs and save the storage cost.
[0104] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. The non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. The volatile memory can include random access memory (RAM) or an external cache. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0105] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0106] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0107] Each embodiment in this specification is described in a related manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments.
[0108] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.
[0109] The above-described embodiments only represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application patent shall be subject to the appended claims.
Claims
1. A compression method for user portrait tag data, characterized in that, The method includes: Obtaining at least one field to be queried from a database according to user portrait tags; Obtaining a corresponding plurality of user IDs as user IDs to be stored according to the field to be queried; the user ID is of int type, and each user ID is represented by 32 bits; Splitting each user ID in the user IDs to be stored into a high 8 bits, a middle high 8 bits, and a low 16 bits; Compressively storing the user IDs to be stored according to the high 8 bits, the middle high 8 bits, and the low 16 bits; The compressively storing the user IDs to be stored according to the high 8 bits, the middle high 8 bits, and the low 16 bits includes: obtaining a first mapping sequence according to the high 8 bits of each user ID in the user IDs to be stored; the storage structure of the first mapping sequence is a bit array; obtaining a plurality of second mapping sequences according to the high 8 bits and the middle high 8 bits of each user ID in the user IDs to be stored; the storage structure of the second mapping sequence is a bit array; storing the low 16 bits of each user ID in the user IDs to be stored as a B data type, an L data type, or an N data type; the storage structure of the B data type is a bit array; the storage structure of the L data type is a binary tuple; the storage structure of the N data type is an integer; The array length of the first mapping sequence is 256; the obtaining the first mapping sequence according to the high 8 bits of each user ID in the user IDs to be stored includes: constructing a first initial array; the first initial array is a bit array with an array length of 256; converting the high 8 bits of each user ID in the user IDs to be stored into decimal to obtain a plurality of first array subscripts; setting the bit value corresponding to each first array subscript in the first initial array to 1 to obtain the first mapping sequence; The array length of the second mapping sequence is 256; the obtaining a plurality of second mapping sequences according to the high 8 bits and the middle high 8 bits of each user ID in the user IDs to be stored includes: obtaining the number of different high 8 bits as a first number; constructing a second initial array of the first number; the second initial array is a bit array with an array length of 256; assigning values to the second initial array according to the high 8 bits and the middle high 8 bits to obtain the second mapping sequences of the first number; The maximum array length of the B data type is 65536; storing the lower 16 bits of each user ID in the to-be-stored user IDs as the B data type, the L data type, or the N data type includes: obtaining the lower 16 bits of the user IDs with the same upper 8 bits and middle upper 8 bits to obtain a second quantity of first lower 16 bits; sequentially calculating the storage spaces occupied after storing the second quantity of first lower 16 bits as the B data type, the L data type, and the N data type; if the storage space occupied by storing the second quantity of first lower 16 bits as the L data type is the smallest, then storing the second quantity of first lower 16 bits as the L data type; if the storage space occupied by storing the second quantity of first lower 16 bits as the N data type is the smallest, then storing the second quantity of first lower 16 bits as the N data type; if the storage space occupied by storing the second quantity of first lower 16 bits as the B data type is the smallest, then storing the second quantity of first lower 16 bits as the B data type.
2. The compression method according to claim 1, wherein Assigning values to the second initial array according to the upper 8 bits and the middle upper 8 bits to obtain the first quantity of the second mapping sequences includes: Obtaining one of the different upper 8 bits as the first upper 8 bits; Obtaining a plurality of first middle upper 8 bits of the user IDs with the same first upper 8 bits; Converting each of the first middle upper 8 bits into a decimal number to obtain a corresponding second array subscript; Setting the bit value corresponding to each second array subscript in the second initial array to 1 to obtain the second mapping sequence corresponding to the first upper 8 bits.
3. The compression method according to claim 1, wherein Storing the second quantity of first lower 16 bits as the L data type includes: Store the first lower 16 bits of the second quantity in a data structure of a binary tuple; the expression of the L data type is L = {(x1, y1), (x2, y2), …, (x n , y n )}; where there are a total of n binary tuples, each binary tuple represents a continuous interval, and the binary tuple (x i , y i ) represents the i-th continuous interval, i ∈ (1, n), x i represents the first lower 16 bits at the start of the i-th continuous interval, y i represents the step size of the i-th continuous interval, and both x i and y i are represented in 16-bit binary.
4. The compression method according to claim 1, characterized in that Storing the second quantity of first lower 16 bits as the N data type includes: Sequentially storing the second quantity of first lower 16 bits directly in binary; each of the first lower 16 bits consists of 16 bits of binary.
5. The compression method according to claim 1, wherein Storing the second quantity of first lower 16 bits as the B data type includes: Converting the second quantity of first lower 16 bits into decimal numbers to obtain a second quantity of third array subscripts; Constructing a third initial array; the third initial array is a bit array; the array length of the third initial array is the subscript value of the largest third array subscript plus 1; Setting the bit value corresponding to each third array subscript in the third initial array to 1.
6. A compression device for user portrait tag data, characterized in that, The device includes: A query field module, configured to obtain at least one to-be-query field from a database according to a user portrait label; A user ID obtaining module, configured to obtain a plurality of corresponding user IDs as to-be-stored user IDs according to the to-be-query field; the user ID is of the int type, and each user ID is represented by 32 bits; A splitting module, configured to split each user ID in the to-be-stored user IDs into an upper 8 bits, a middle upper 8 bits, and a lower 16 bits; A compression storage module, which is used to compress and store the user ID to be stored according to the high 8 bits, the middle high 8 bits and the low 16 bits; obtain a first mapping sequence according to the high 8 bits of each user ID in the user ID to be stored; the storage structure of the first mapping sequence is a bit array; obtain a plurality of second mapping sequences according to the high 8 bits and the middle high 8 bits of each user ID in the user ID to be stored; the storage structure of the second mapping sequence is a bit array; store the low 16 bits of each user ID in the user ID to be stored as the B data type, the L data type or the N data type; the storage structure of the B data type is a bit array; the storage structure of the L data type is a binary tuple; the storage structure of the N data type is an integer; the array length of the first mapping sequence is 256; the obtaining the first mapping sequence according to the high 8 bits of each user ID in the user ID to be stored includes: constructing a first initial array; the first initial array is a bit array with an array length of 256; converting the high 8 bits of each user ID in the user ID to be stored into decimal to obtain a plurality of first array subscripts; setting the bit value corresponding to each first array subscript in the first initial array to 1 to obtain the first mapping sequence; the array length of the second mapping sequence is 256; the obtaining a plurality of second mapping sequences according to the high 8 bits and the middle high 8 bits of each user ID in the user ID to be stored includes: obtaining the number of different high 8 bits as the first number; constructing a second initial array of the first number; the second initial array is a bit array with an array length of 256; obtaining the second mapping sequences of the first number according to the high 8 bits and the middle high 8 bits by assigning values to the second initial array; the maximum array length of the B data type is 65536; the storing the low 16 bits of each user ID in the user ID to be stored as the B data type, the L data type or the N data type includes: obtaining the low 16 bits of the user IDs with the same high 8 bits and middle high 8 bits to obtain the first low 16 bits of the second number; calculating the storage space occupied by the first low 16 bits of the second number when stored as the B data type, the L data type and the N data type in sequence; if the storage space occupied by the first low 16 bits of the second number when stored as the L data type is the smallest, then store the first low 16 bits of the second number as the L data type; if the storage space occupied by the first low 16 bits of the second number when stored as the N data type is the smallest, then store the first low 16 bits of the second number as the N data type; if the storage space occupied by the first low 16 bits of the second number when stored as the B data type is the smallest, then store the first low 16 bits of the second number as the B data type.
Citation Information
Patent Citations
Data storage and query method and system and storage engine device
CN104462141A
Data compression method and device, computer equipment and readable storage medium
CN117478149A