Compression method and device for user portrait label data

By splitting the user ID into high 8 bits, medium 8 bits and low 16 bits, and using corresponding data structures for compression storage, the storage space and cost problems caused by the increase in the number of user IDs are solved, and effective storage space savings are achieved.

CN120196635AActive Publication Date: 2025-06-24SHENZHEN HUOLI TIAN HUI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510630927.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-24
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

In the prior art, the number of user ids corresponding to the user portrait tags has increased dramatically, resulting in a sharp increase in the data space for storing the user ids, causing the problem of increasing storage costs.

Method used

By splitting the user id to be stored into high 8 bits, medium 8 bits and low 16 bits, and compressing and storing them according to these bit fields, data structures such as bit arrays, binary groups or integers are used to reduce storage space.

Benefits of technology

It effectively reduces the storage space used to store user IDs and saves storage costs, especially when the number of user IDs is huge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196635A_ABST
    Figure CN120196635A_ABST
Patent Text Reader

Abstract

The invention relates to a compression method and device for user portrait label data, and belongs to the technical field of data compress.The method comprises the steps that at least one to-be-queried field is obtained from a database according to a user portrait label, a plurality of corresponding user ids are obtained according to the to-be-queried field to serve as to-be-stored user ids, the user ids are of an int type, and the to-be-stored user ids are stored according to the to-be-queried field; each user id is represented by 32 bits, each user id in the to-be-stored user id is divided into high 8 bits, middle-high 8 bits and low 16 bits, and the to-be-stored user id is compressed and stored according to the high 8 bits, the middle-high 8 bits and the low 16 bits; and the plurality of user ids corresponding to the user portrait label are compressed and stored based on high 8 bits, middle-high 8 bits and low 16 bits, so that the storage space for storing the user ids can be effectively reduced, and the storage cost is saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data compression, and particularly relates to a method and device for compressing user portrait label data. Background Art

[0002] With the rapid development of information technology, users generate a vast amount of data during the use of various applications. By analyzing the data value, mining user characteristics and behaviors, and organizing them into user portrait labels, the user portrait labels can help enterprises accurately locate user needs, optimize product design, and improve market operation efficiency.

[0003] In the prior art, multiple corresponding user IDs are queried through user portrait labels, and the multiple user IDs corresponding to the user portrait labels are pre-stored as a group for later processing. Multiple user portrait labels correspond to multiple groups of user IDs. When the number of users is huge, the number of user IDs corresponding to the user portrait labels will also increase sharply, resulting in a sharp increase in the data space for storing user IDs, causing the problem of increased storage costs. Summary of the Invention

[0004] Therefore, the present invention provides a method and device for compressing user portrait label data to solve the problem of increased storage costs caused by the sharp increase in the data space for storing user IDs in the prior art.

[0005] To achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a method for compressing user portrait label data, including: Obtaining at least one field to be queried from a database according to a user portrait label; Obtaining multiple corresponding user IDs as user IDs to be stored according to the field to be queried; the user ID is of int type, and each user ID is represented by 32 bits; Splitting each user ID in the user IDs to be stored into a high 8-bit part, a middle-high 8-bit part, and a low 16-bit part; Compressively storing the user IDs to be stored according to the high 8-bit part, the middle-high 8-bit part, and the low 16-bit part.

[0006] Further, the compressing and storing the user IDs to be stored according to the high 8-bit part, the middle-high 8-bit part, and the low 16-bit part includes: Obtaining a first mapping sequence according to the high 8-bit part of each user ID in the user IDs to be stored; the storage structure of the first mapping sequence is a bit array; Obtaining multiple second mapping sequences according to the high 8-bit part and the middle-high 8-bit part of each user ID in the user IDs to be stored; the storage structure of the second mapping sequence is a bit array; Store the lower 16 bits of each user ID in the to-be-stored user IDs as the B data type, the L data type, or the N data type; the storage structure of the B data type is a bit array; the storage structure of the L data type is a binary tuple; the storage structure of the N data type is an integer.

[0007] Further, the array length of the first mapping sequence is 256; obtaining the first mapping sequence according to the upper 8 bits of each user ID in the to-be-stored user IDs includes: Construct a first initial array; the first initial array is a bit array with an array length of 256; Convert the upper 8 bits of each user ID in the to-be-stored user IDs into decimal to obtain a plurality of first array subscripts; After setting the bit value corresponding to each of the first array subscripts in the first initial array to 1, obtain the first mapping sequence.

[0008] Further, the array length of the second mapping sequence is 256; obtaining a plurality of second mapping sequences according to the upper 8 bits and the middle upper 8 bits of each user ID in the to-be-stored user IDs includes: Obtain the number of different upper 8 bits as the first number; Construct a second initial array with the first number; the second initial array is a bit array with an array length of 256; Assign values to the second initial array according to the upper 8 bits and the middle upper 8 bits to obtain the second mapping sequences with the first number.

[0009] Further, the assigning values to the second initial array according to the upper 8 bits and the middle upper 8 bits to obtain the second mapping sequences with the first number includes: Obtain one of the different upper 8 bits as the first upper 8 bits; Obtain the middle upper 8 bits of a plurality of user IDs with the same first upper 8 bits to obtain a plurality of first middle upper 8 bits; After converting each of the first middle upper 8 bits into decimal, obtain the corresponding second array subscripts; After setting the bit value corresponding to each of the second array subscripts in the second initial array to 1, obtain the second mapping sequence corresponding to the first upper 8 bits.

[0010] Further, the maximum array length of the B data type is 65536; the storing the lower 16 bits of each user ID in the to-be-stored user IDs as the B data type, the L data type, or the N data type includes: Obtain the lower 16 bits of the user ID where the upper 8 bits and the upper-middle 8 bits are the same to get a second quantity of first lower 16 bits; Successively calculate the storage spaces occupied after storing the second quantity of first lower 16 bits as B data type, L data type, and N data type; If the storage space occupied by storing the second quantity of first lower 16 bits as L data type is the smallest, then store the second quantity of first lower 16 bits as L data type; If the storage space occupied by storing the second quantity of first lower 16 bits as N data type is the smallest, then store the second quantity of first lower 16 bits as N data type; If the storage space occupied by storing the second quantity of first lower 16 bits as B data type is the smallest, then store the second quantity of first lower 16 bits as B data type.

[0011] Further, the storing the second quantity of first lower 16 bits as L data type includes: Store the second quantity of first lower 16 bits in a data structure of a binary tuple; The expression of the L data type is L={(x1,y1),(x2,y2),…,(x n ,y n )}; where, there are a total of n binary tuples, each binary tuple represents a continuous interval, and the binary tuple (x i ,y i ) represents the i-th continuous interval, i∈(1, n), x i represents the first lower 16 bits at the start of the i-th continuous interval, y i represents the step size of the i-th continuous interval, and both x i and y i are represented in 16-bit binary.

[0012] Further, the storing the second quantity of first lower 16 bits as N data type includes: Successively store the second quantity of first lower 16 bits directly in binary; Each of the first lower 16 bits consists of 16-bit binary.

[0013] Further, the storing the second quantity of first lower 16 bits as B data type includes: Convert the second quantity of first lower 16 bits to decimal to obtain the subscripts of a second quantity of third arrays; Construct a third initial array; The third initial array is a bit array; The array length of the third initial array is the value of the subscript of the largest third array subscript plus 1; Set the bit value corresponding to each of the third array subscripts in the third initial array to 1.

[0014] Second aspect, the present invention provides a compression device for user portrait tag data, the device comprising: A query field module, configured to obtain at least one query field to be queried from a database according to user portrait tags; A user ID obtaining module, configured to obtain a plurality of corresponding user IDs as user IDs to be stored according to the query field to be queried; the user ID is of int type, and each user ID is represented by 32 bits; A splitting module, configured to split each user ID in the user IDs to be stored into a high 8 bits, a middle high 8 bits, and a low 16 bits; A compression storage module, configured to compress and store the user IDs to be stored according to the high 8 bits, the middle high 8 bits, and the low 16 bits.

[0015] Adopting the above technical solutions, the present invention has at least the following beneficial effects: Provided is a compression method and device for user portrait tag data. At least one query field to be queried is obtained from a database according to user portrait tags, a plurality of corresponding user IDs are obtained as user IDs to be stored according to the query field to be queried, the user ID is of int type, and each user ID is represented by 32 bits. Each user ID in the user IDs to be stored is split into a high 8 bits, a middle high 8 bits, and a low 16 bits, and the user IDs to be stored are compressed and stored according to the high 8 bits, the middle high 8 bits, and the low 16 bits. By compressing and storing a plurality of user IDs corresponding to user portrait tags based on the high 8 bits, the middle high 8 bits, and the low 16 bits, the storage space for storing user IDs can be effectively reduced, and the storage cost is saved.

[0016] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 is a flowchart of a compression method for user portrait tag data shown in an exemplary embodiment of the present invention; Figure 2 is a schematic storage structure diagram A of a compression method for user portrait tag data shown in an exemplary embodiment of the present invention; Figure 3It is a schematic diagram B of a storage structure of a compression method for user profile tag data shown in an exemplary embodiment of the present invention; Figure 4 It is a schematic block diagram of a compression device for user profile tag data shown in an exemplary embodiment of the present invention.

[0019] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. Specific Embodiments

[0020] In order to make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be described in detail below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope protected by the present invention.

[0021] In the prior art, multiple corresponding users are queried through user profile tags, and the multiple user IDs corresponding to the user profile tags are stored and used for later arithmetic processing. Storing each user ID generally occupies 32 bits. When the number of user IDs is very large, the data space for storing user IDs increases sharply, increasing the storage cost.

[0022] Embodiments of the present invention provide a compression method and device for user profile tag data. Each user ID in the user IDs to be stored is split into a high 8 bits, a medium-high 8 bits, and a low 16 bits. Based on the high 8 bits, the medium-high 8 bits, and the low 16 bits, multiple user IDs corresponding to the user profile tags are compressed and stored, which can effectively reduce the storage space for storing user IDs. Compared with the traditional storage where each user ID occupies 32 bits, the more the number of user IDs, the more obvious the effect of saving storage space.

[0023] The methods and devices in the present invention will be described below through specific embodiments.

[0024] Please refer to Figure 1 , Figure 1 It is a flowchart of a compression method for user profile tag data shown in an exemplary embodiment of the present invention. Refer to Figure 1 , the method includes: Step S11: Obtain at least one field to be queried from the database according to the user profile tag; Step S12: Obtain multiple corresponding user IDs as the user IDs to be stored according to the field to be queried; the user ID is of int type, and each user ID is represented by 32 bits; Step S13: Split each user ID in the user IDs to be stored into a high 8 bits, a medium-high 8 bits, and a low 16 bits; Step S14: Compress and store the user ID to be stored according to the high 8 bits, the medium-high 8 bits, and the low 16 bits.

[0025] It should be noted that the technical solution provided in this embodiment can be loaded and used in an existing user analysis system or application in the form of a small program or a plug-in in specific practice, or in the form of a separate application, and the user portrait label data compression function can be implemented through an external interface; the applicable scenarios include but are not limited to: user ID compression.

[0026] It can be understood that for the method provided in this embodiment, at least one field to be queried is obtained from the database according to the user portrait label, and a corresponding plurality of user IDs are obtained as the user IDs to be stored according to the field to be queried. The user ID is of int type, and each user ID is represented by 32 bits. Each user ID in the user ID to be stored is split into a high 8 bits, a medium-high 8 bits, and a low 16 bits, and the user ID to be stored is compressed and stored according to the high 8 bits, the medium-high 8 bits, and the low 16 bits; by compressing and storing a plurality of user IDs corresponding to the user portrait label based on the high 8 bits, the medium-high 8 bits, and the low 16 bits, the storage space for storing the user ID can be effectively reduced, and the storage cost can be saved.

[0027] In specific practice, in step S11, “obtain at least one field to be queried from the database according to the user portrait label”, where the user portrait label is the user feature stored in the database, and is specifically represented in the form of a field value. For example, the user portrait label is male gender, and the field of the user feature is gender; the user portrait label is male consumers, and the fields of the user features are the gender field and the consumption field.

[0028] In specific practice, step S12, “obtain a corresponding plurality of user IDs as the user IDs to be stored according to the field to be queried”, includes: the process of obtaining the user ID can use existing database query algorithms, that is, it can be queried using SQL statements. For example, the user portrait label is male consumers, the field value 0 in the gender field represents female, and 1 represents male; the field value 1 in the consumption field represents high consumption, 2 represents medium consumption, and 3 represents low consumption; then, according to the field value of 1 in the gender field and the field value of 2 in the consumption field, a query is performed to obtain a corresponding plurality of user IDs of male consumers as the user IDs to be stored.

[0029] In specific practice, step S13, "split each user ID in the user IDs to be stored into the high 8 bits, the middle-high 8 bits, and the low 16 bits", includes: the user ID is of int type, each user ID is represented by 32 bits, and the 32 bits are split into the high 8 bits, the middle-high 8 bits, and the low 16 bits. For example, in the splitting process of the user ID 1: the int representation of 1 is 00000000 00000000 00000000 00000001, its high 8 bits and middle-high 8 bits are both 00000000, and the low 16 bits are 00000000 00000001.

[0030] In specific practice, step S14, "compress and store the user IDs to be stored according to the high 8 bits, the middle-high 8 bits, and the low 16 bits", includes: obtaining a first mapping sequence according to the high 8 bits of each user ID in the user IDs to be stored; the storage structure of the first mapping sequence is a bit array; obtaining multiple second mapping sequences according to the high 8 bits and the middle-high 8 bits of each user ID in the user IDs to be stored; the storage structure of the second mapping sequence is a bit array; storing the low 16 bits of each user ID in the user IDs to be stored as the B data type, the L data type, or the N data type; the storage structure of the B data type is a bit array; the storage structure of the L data type is a binary tuple; the storage structure of the N data type is an integer.

[0031] Specifically, the array length of the first mapping sequence is 256; obtaining the first mapping sequence according to the high 8 bits of each user ID in the user IDs to be stored includes: constructing a first initial array; the first initial array is a bit array with an array length of 256; converting the high 8 bits of each user ID in the user IDs to be stored into decimal to obtain multiple first array subscripts; setting the bit value corresponding to each first array subscript in the first initial array to 1 to obtain the first mapping sequence.

[0032] It should be noted that as shown in Table 1 below, the array length of the first initial array is 256 (the subscripts are 0-255), and among them, the high 8 bits of multiple user IDs in the user IDs to be stored are 00000000, 00000101, and 10011011. Among them, 00000000, 00000101, and 10011011 are converted into decimal to obtain 0, 5, and 155 in sequence. Then, setting the bit values corresponding to the 0, 5, and 155 subscripts in the first initial array to 1 to obtain the first mapping sequence.

[0033] Table 1 。

[0034] Specifically, the array length of the second mapping sequence is 256; multiple second mapping sequences are obtained based on the high 8 bits and the middle high 8 bits of each user ID in the user ID to be stored, including: obtaining the number of different high 8 bits as the first number; constructing a second initial array with the first number, where the second initial array is a bit array with an array length of 256; obtaining one of the different high 8 bits as the first high 8 bit; obtaining multiple first middle high 8 bits of the user IDs with the same first high 8 bit; converting each first middle high 8 bit to decimal to obtain the corresponding second array subscript; and setting the bit value corresponding to each second array subscript in the second initial array to 1 to obtain the second mapping sequence corresponding to the first high 8 bit.

[0035] It should be noted that the data structure of the second mapping sequence is referred to Figure 2 , Figure 2 Figure A is a schematic storage structure diagram of a compression method for user portrait label data shown in an exemplary embodiment of the present invention. Among them, there are 3 different high 8 bits in the first mapping sequence, that is, the first number is 3; 3 second initial arrays are constructed, and the array length of each second initial array is 256; obtaining a high 8 bit 00000000 as the first high 8 bit, and the middle high 8 bits of the user IDs with the same first high 8 bit are 00000000, 00000011, and 10011100. Among them, 00000000, 00000011, and 10011100 are converted to decimal to obtain 0, 3, and 156 in sequence. Then, the bit values of the 0, 3, and 156 subscripts of the second initial array corresponding to the first high 8 bit are set to 1 to obtain the second mapping sequence, and this second mapping sequence corresponds to the first high 8 bit.

[0036] Specifically, the maximum array length of the B data type is 65536; storing the low 16 bits of each user ID in the user ID to be stored as the B data type, the L data type, or the N data type, including: obtaining the low 16 bits of the user IDs with the same high 8 bits and middle high 8 bits to obtain the second number of first low 16 bits; calculating the storage space occupied by storing the second number of first low 16 bits as the B data type, the L data type, and the N data type in sequence; if the storage space occupied by storing the second number of first low 16 bits as the L data type is the smallest, then storing the second number of first low 16 bits as the L data type; if the storage space occupied by storing the second number of first low 16 bits as the N data type is the smallest, then storing the second number of first low 16 bits as the N data type; if the storage space occupied by storing the second number of first low 16 bits as the B data type is the smallest, then storing the second number of first low 16 bits as the B data type.

[0037] It should be noted that the data of the L data type is composed of multiple binary tuples. Each binary tuple stores one first lower 16 bits or multiple consecutive first lower 16 bits. When storing one first lower 16 bits as a continuous interval, the step size of the binary tuple is 1. When storing multiple consecutive first lower 16 bits as a continuous interval, the step size of the binary tuple is the number of multiple consecutive first lower 16 bits.

[0038] It should be noted that calculate the storage space occupied after storing the first lower 16 bits of the second quantity as the B data type, the L data type, and the N data type in sequence, including: calculating the storage space occupied when storing as the N data type: the second quantity × 16 = the number of bit positions occupied when storing as the N data type, that is, the storage space occupied when storing as the N data type; calculating the storage space occupied when storing as the B data type: the maximum integer after converting the lower 16 bits to an integer plus 1 = the number of bit positions occupied when storing as the B data type, that is, the storage space occupied when storing as the B data type; calculating the storage space occupied when storing as the L data type: the number of binary tuples × 16 × 2 = the number of bit positions occupied when storing as the L data type, that is, the storage space occupied when storing as the L data type. The binary tuple (x, y) represents a continuous interval, where both x and y are represented by 16 - bit binary. x represents the initial first lower 16 bits of the continuous interval, and y represents the step size of the continuous interval.

[0039] Specifically, store the first lower 16 bits of the second quantity in the data structure of binary tuples; the expression of the L data type is L = {(x1, y1), (x2, y2), …, (x n , y n )}; where there are a total of n binary tuples, and each binary tuple represents a continuous interval. The i - th continuous interval represented by the binary tuple (x i , y i ), i ∈ (1, n), x i represents the first lower 16 bits at the start of the i - th continuous interval, and y i represents the step size of the i - th continuous interval. Both x i and y i are represented by 16 - bit binary.

[0040] Specifically, store the first lower 16 bits of the second quantity as the N data type, including: storing the first lower 16 bits of the second quantity directly in binary in sequence; each first lower 16 bits is composed of 16 - bit binary.

[0041] Specifically, store the first lower 16 bits of the second quantity as the B data type, including: converting the first lower 16 bits of the second quantity into decimal to obtain the subscripts of the third array for the second quantity; constructing the third initial array; the third initial array is a bit array; the array length of the third initial array is the subscript value of the largest third array subscript plus 1; setting the bit value corresponding to each third array subscript in the third initial array to 1.

[0042] It should be noted that as shown in Table 2 below, taking the first lower 16 bits of 00000000 00000001, 00000000 00000011, 00000000 00000101, and 11111111 11111111 as examples, converting them into decimal in sequence gives 1, 3, 5, and 65535. Then the subscript value of the largest third array subscript is 65535, and the array length of the third initial array is 65536 (subscripts range from 0 to 65535). Set the bit values corresponding to the subscripts of 1, 3, 5, and 65535 in the third initial array to 1, that is, store them as the B data type.

[0043] Table 2 。

[0044] In a specific embodiment, please refer to Figure 3 , Figure 3 which is the schematic diagram B of the storage structure of a compression method for user portrait tag data shown in an exemplary embodiment of the present invention. See Figure 3 for the storage structure description: Take multiple user IDs obtained based on a certain user portrait tag as the user IDs to be stored. The user IDs to be stored include: all odd numbers within the range of... 84082689 - 84148223,... 84344832,... 84346432,... 84404835,... 100597760 - 100597859,... 100663290 - 100663295; The user IDs to be stored are of the integer int type, that is, each user ID requires 32 bit positions to represent. When compressing and storing the user IDs to be stored, store the user IDs to be stored as multiple data segments; each data segment uses the high 8 bits and the mid - high 8 bits of the user ID as the index. Specifically: multiple user IDs with the same high 8 bits and mid - high 8 bits are stored in one data segment; multiple user IDs with the same high 8 bits but different mid - high 8 bits are stored in multiple data segments, and each data segment corresponds to a different mid - high 8 bit; the index of each data segment is the high 8 bits and the mid - high 8 bits, and then store the lower 16 bits of the corresponding multiple user IDs in the data segment in the B data type, L data type, or N data type.

[0045] Store the lower 16 bits of multiple user IDs corresponding to each data segment in the data segment in B data type, L data type or N data type. Specifically, in the first case, see Figure 3 The storage in the bit data structure. The user IDs to be stored are all odd numbers in the range of 84082689 - 84148223, that is, 84082689, 84082691, 84082693... 84148223. The high 8 bits of these user IDs are 00000101, and the middle high 8 bits are 00000011. Since the high 8 bits and the middle high 8 bits of these user IDs are the same, the lower 16 bits of these user IDs are stored in one data segment. After calculating the storage space occupied by storing the lower 16 bits of these user IDs in three data types, storing in B data type saves the most space. Therefore, all odd numbers in the range of 84082689 - 84148223 are stored as the bit data structure.

[0046] The business case of the first case: Assume that odd user IDs in the range of 84082689 - 84148223 represent male users, and even user IDs represent female users; the user IDs with the user portrait label of male users are discontinuous, the user IDs are dense and scattered, and the lower 16 bits of the user IDs with the user portrait label of male users are stored in multiple data segments. Storing multiple data segments using B data structure saves the most space, that is, it is applicable to the situation where the user IDs corresponding to the user portrait label are dense and the continuous interval is small.

[0047] In the second case, see Figure 3 The storage in the integer data structure. The user IDs to be stored are 84344832,..., 84346432,..., 84404835,... The high 8 bits of these user IDs are 00000101, and the middle high 8 bits are 00000111. Since the high 8 bits and the middle high 8 bits of these user IDs are the same, the lower 16 bits of these user IDs are stored in one data segment. After calculating the storage space occupied by storing the lower 16 bits of these user IDs in three data types, storing in N data type saves the most space. Therefore, the lower 16 bits of these user IDs are stored as integers 0,..., 1600,..., 60003,..., where each integer is represented in 16 - bit binary.

[0048] The business case of the second case: For example, the user portrait label is the high - consumption label. The number of user IDs with the high - consumption label is small, sparse and basically discontinuous. The lower 16 bits of these user IDs are stored in multiple data segments. Storing multiple data segments using N data structure saves the most space.

[0049] In the third case, see Figure 3Medium-step data structure storage. The user IDs to be stored are in the range of 100597760 - 100597859, ……, 100663290 - 100663295 in the user ID to be stored. The high 8 bits of these user IDs are 00000101, and the middle high 8 bits are 11111111. Since the high 8 bits and the middle high 8 bits of these user IDs are the same, the low 16 bits of these user IDs are stored in a data segment. After calculating the storage space occupied by the low 16 bits of these user IDs stored as three data types, storing as the L data type saves the most space. The consecutive low 16 bits are stored in a binary tuple structure in sequence, that is, the low 16 bits of these user IDs are stored as: L={(0,100), ……, (65530,6)}.

[0050] The third case of the business scenario: For example, the user portrait label is a scalper label, and the user IDs are continuously in an interval. Scalpers may use automated tools to centrally register a batch of users, and these user IDs are continuous; the low 16 bits of these user IDs are stored in the data segment using the L data structure, which saves the most space.

[0051] It should be noted that during compression storage, there are multiple data segments indexed by the high 8 bits and the middle high 8 bits. For example, multiple user IDs with the same high 8 bits and middle high 8 bits are stored in one data segment, and multiple user IDs with the same high 8 bits but different middle high 8 bits are stored in multiple data segments. Each data segment corresponds to a different middle high 8 bit. The index of the data segment is the high 8 bits and the middle high 8 bits. Then, according to the above three cases, the low 16 bits of multiple user IDs are stored in the data segment, that is, the compression is completed.

[0052] In specific practice, decompression is the reverse process of compression. Locate the data segment by the high 8 bits and the middle high 8 bits. Then, after converting the data corresponding to the low 16 bits into binary, it is bit-concatenated with the high 8 bits and the middle high 8 bits into 32 bits. After converting the 32-bit binary into an int type, the user ID is obtained.

[0053] Please refer to Figure 4 , Figure 4 which is a schematic block diagram of a compression device for user portrait label data shown in an exemplary embodiment of the present invention. Refer to Figure 4 The compression device 100 for user portrait label data includes: The query field module 101 is used to obtain at least one query field to be queried from the database according to the user portrait label; The user ID acquisition module 102 is used to obtain a corresponding plurality of user IDs as user IDs to be stored according to the query field to be queried; the user ID is of int type, and each user ID is represented by 32 bits; The splitting module 103 is used to split each user ID in the user ID to be stored into high 8 bits, middle high 8 bits, and low 16 bits; The compression storage module 104 is used to compress and store the user ID to be stored according to the high 8 bits, the middle-high 8 bits, and the low 16 bits.

[0054] It should be noted that the applicable scenarios of the device provided in this embodiment include, but are not limited to: the compression storage of multiple user IDs corresponding to user portrait tags.

[0055] It can be understood that the device provided in this embodiment obtains at least one field to be queried from the database according to the user portrait tag, obtains multiple corresponding user IDs as the user IDs to be stored according to the field to be queried. The user ID is of int type, and each user ID is represented by 32 bits. Each user ID in the user IDs to be stored is split into the high 8 bits, the middle-high 8 bits, and the low 16 bits, and the user IDs to be stored are compressed and stored according to the high 8 bits, the middle-high 8 bits, and the low 16 bits; by compressing and storing multiple user IDs corresponding to the user portrait tag based on the high 8 bits, the middle-high 8 bits, and the low 16 bits, the storage space for storing user IDs can be effectively reduced, and the storage cost can be saved.

[0056] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to the memory, storage, database, or other media used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0057] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0058] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0059] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment.

[0060] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.

[0061] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be pointed out that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application patent shall be subject to the appended claims.

Claims

1. A method for compressing user portrait label data, characterized in that: The method comprises: Obtain at least one field to be queried from the database according to the user portrait tag; According to the to-be-queried field, a plurality of corresponding user IDs are obtained as the to-be-stored user IDs; the user ID is of int type, and each user ID is represented by 32 bits; Split each user ID in the user ID to be stored into high 8 bits, middle high 8 bits and low 16 bits; The user ID to be stored is compressed and stored according to the high 8 bits, the middle high 8 bits and the low 16 bits.

2. The compression method according to claim 1, characterized in that: The method of compressing and storing the user ID to be stored according to the high 8 bits, the middle high 8 bits and the low 16 bits includes: Obtaining a first mapping sequence according to the upper 8 bits of each user ID in the user ID to be stored; the storage structure of the first mapping sequence is a bit array; A plurality of second mapping sequences are obtained according to the upper 8 bits and the middle upper 8 bits of each user ID in the user ID to be stored; the storage structure of the second mapping sequence is a bit array; The lower 16 bits of each user ID in the user ID to be stored are stored as a B data type, an L data type or an N data type; the storage structure of the B data type is a bit array; the storage structure of the L data type is a tuple; the storage structure of the N data type is an integer.

3. The compression method according to claim 2, characterized in that: The array length of the first mapping sequence is 256; the first mapping sequence is obtained according to the upper 8 bits of each user ID in the user ID to be stored, including: Constructing a first initial array; the first initial array is a bit array with an array length of 256; Convert the upper 8 bits of each user ID in the user ID to be stored into decimal to obtain a plurality of first array subscripts; The first mapping sequence is obtained by setting the bit value corresponding to each first array subscript in the first initial array to 1.

4. The compression method according to claim 3, characterized in that: The array length of the second mapping sequence is 256; the multiple second mapping sequences are obtained according to the high 8 bits and the middle high 8 bits of each user ID in the user ID to be stored, including: Obtain the number of different upper 8 bits as the first number; Constructing a second initial array of the first quantity; the second initial array is a bit array with an array length of 256; The second initial array is assigned values ​​according to the upper 8 bits and the middle upper 8 bits to obtain the first number of the second mapping sequences.

5. The compression method according to claim 4, characterized in that: The step of assigning the second initial array according to the high 8 bits and the middle high 8 bits to obtain the first number of the second mapping sequences includes: Get one of the different upper 8 bits as the first upper 8 bits; Acquire the middle and high 8 bits of multiple user IDs with the same first high 8 bits to obtain multiple first middle and high 8 bits; Convert each of the first high 8 bits into decimal to obtain the corresponding second array subscript; The second mapping sequence corresponding to the first high 8 bits is obtained by setting the bit value corresponding to each second array subscript in the second initial array to 1.

6. The compression method according to claim 5, characterized in that: The maximum array length of the B data type is 65536; the storing the lower 16 bits of each user ID in the user IDs to be stored as the B data type, the L data type or the N data type includes: Obtain the lower 16 bits of the user ID whose upper 8 bits and middle upper 8 bits are the same to obtain the first lower 16 bits of the second quantity; sequentially calculating the storage space occupied by the first lower 16 bits of the second quantity after being stored as the B data type, the L data type, and the N data type; If the storage space occupied by storing the first lower 16 bits of the second quantity as the L data type is the smallest, then storing the first lower 16 bits of the second quantity as the L data type; If the storage space occupied by storing the first lower 16 bits of the second quantity as the N data type is the smallest, then storing the first lower 16 bits of the second quantity as the N data type; If the storage space occupied by storing the first lower 16 bits of the second quantity as the B data type is the smallest, the first lower 16 bits of the second quantity are stored as the B data type.

7. The compression method according to claim 6, characterized in that: The storing the first lower 16 bits of the second quantity as the L data type comprises: The first lower 16 bits of the second quantity are stored in a two-tuple data structure; the expression of the L data type is L={(x1,y1),(x2,y2),…,(x n ,y n )}; There are a total of n tuples, each of which represents a continuous interval. i ,y i ) represents the i-th continuous interval, i∈(1,n), x i Indicates the first low 16 bits of the start of the i-th continuous interval, y i represents the step size of the i-th continuous interval, x i and i All are represented by 16-bit binary.

8. The compression method according to claim 6, characterized in that: The storing the first lower 16 bits of the second quantity as the N data type comprises: The first lower 16 bits of the second quantity are directly stored in binary in sequence; each of the first lower 16 bits consists of 16 bits of binary.

9. The compression method according to claim 6, characterized in that: The storing the first lower 16 bits of the second quantity as a B data type comprises: Convert the first low 16 bits of the second number into decimal to obtain the second number of third array subscripts; Constructing a third initial array; the third initial array is a bit array; the array length of the third initial array is the subscript value of the largest third array subscript plus 1; The bit value corresponding to each third array subscript in the third initial array is set to 1.

10. A device for compressing user portrait label data, characterized in that: The device comprises: A query field module is used to obtain at least one field to be queried from the database according to the user portrait tag; A user ID acquisition module is used to acquire multiple corresponding user IDs as user IDs to be stored according to the to-be-queried field; the user ID is of int type, and each user ID is represented by 32 bits; A splitting module, used for splitting each user ID in the user ID to be stored into high 8 bits, middle high 8 bits and low 16 bits; The compression storage module is used to compress and store the user ID to be stored according to the high 8 bits, the middle high 8 bits and the low 16 bits.

Citation Information

Patent Citations

  • Data storage and query method and system and storage engine device

    CN104462141A

  • Data compression method and device, computer equipment and readable storage medium

    CN117478149A

  • Method and system for identifying air ticket abnormity search user based on user portrait and clustering technology

    CN117520994A

  • MIP Map Compression

    US20180020223A1