Data writing, querying methods and devices

By using a Bloom filter to distribute data across shards in a distributed cache system, the method addresses slow data writing and querying issues in Hbase, improving performance and user experience through optimized data distribution and reduced cache bandwidth usage.

CN113760837BActive Publication Date: 2025-07-15BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011165418.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-27
Publication Date
2025-07-15
Estimated Expiration
2040-10-27

AI Technical Summary

Technical Problem

The prior art data writing time in Hbase is too long and the QPS is low, which cannot meet the query needs during big promotions, affecting the user experience.

Method used

The Bloom filter is combined with the distributed cache system, and the data items are written into the distributed cache system through hashing operations. When generating the Bloom filter, the error judgment rate is considered to avoid generating too large bitmaps and reduce the use of cache bandwidth.

Benefits of technology

It improves the speed of data writing and query, reduces the impact on system performance, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113760837B_ABST
    Figure CN113760837B_ABST
Patent Text Reader

Abstract

The present invention discloses a data writing and querying method and device, relating to the field of computer technology. Among them, the data writing method includes: obtaining a data file to be written and the line number information of the data file to be written; generating a Bloom filter according to the line number information of the data file to be written and the shard quantity information of the distributed cache system; performing a hash operation on the data items in the data file to be written based on the Bloom filter; and writing the data file to be written into the distributed cache system according to the result of the hash operation. Through the above steps, the data writing speed can be improved, the impact of data writing on the system performance can be reduced, and the user experience can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method and device for data writing and querying. Background Art

[0002] In many application scenarios, data writing and querying are required. For example, when performing targeted marketing information delivery for a specific user group, it is necessary to write the data of the user group into Hbase. After a user logs in to an e-commerce platform, it is determined whether the user hits the user group by querying Hbase. If the user hits the user group, coupons, advertisements, or other marketing information are targeted to the user.

[0003] In the process of implementing the present invention, the inventors of the present invention found that the existing data writing and querying solutions have the following defects: First, the time for writing data into Hbase is too long. For example, if a user group contains 20 million user identifiers, it takes about 2 hours to write them into Hbase. Second, the QPS (queries per second) of the Hbase interface is relatively low, and it cannot meet the query requirements of the foreground service during major promotions. Summary of the Invention

[0004] In view of this, the present invention provides a method and device for data writing and querying, which can improve the data writing and querying speed, reduce the impact of data writing and querying on system performance, and improve the user experience.

[0005] To achieve the above object, according to the first aspect of the present invention, a data writing method is provided.

[0006] The data writing method of the present invention includes: obtaining a data file to be written and the line number information of the data file to be written; generating a Bloom filter according to the line number information of the data file to be written and the shard number information of the distributed cache system; performing a hash operation on the data items in the data file to be written based on the Bloom filter; and writing the data file to be written into the distributed cache system according to the result of the hash operation.

[0007] Optionally, the obtaining a data file to be written and the line number information of the data file to be written includes: obtaining the information of the data file to be written according to the file identifier of the data file to be written; where the information of the file to be written includes: the storage address information of the data file to be written and the line number information of the data file to be written; and downloading the data file to be written to the local according to the storage address information of the data file to be written.

[0008] Optionally, generating the Bloom filter according to the line number information of the data file to be written and the shard number information of the distributed cache system includes: calculating the ratio of a preset value to the shard number of the distributed cache system, determining multiple line number value ranges according to the ratio; determining the line number value range where the line number information of the data file to be written is located, and using the misjudgment rate corresponding to the line number value range where it is located as the misjudgment rate of the Bloom filter.

[0009] Optionally, performing a hash operation on the data items in the data file to be written based on the Bloom filter includes: determining the shard identifier corresponding to the data item in the data file to be written; generating a key corresponding to the data item according to the shard identifier corresponding to the data item and the file identifier of the data file to be written; performing a hash operation on the key corresponding to the data item.

[0010] Optionally, determining the shard identifier corresponding to the data item in the data file to be written includes: traversing the data file to be written to obtain the identifiers of each data item in the data file to be written; determining the shard identifier corresponding to the data item according to the identifier of the data item, the shard number of the distributed cache system, and the file identifier of the data file to be written.

[0011] To achieve the above object, according to the second aspect of the present invention, a data query method is provided.

[0012] The data query method of the present invention includes: determining the shard identifier corresponding to the data item to be detected, obtaining a Bloom filter according to the shard identifier and the identifier of the target data file; performing a hash operation on the data item to be detected according to the Bloom filter; querying the distributed cache system according to the result of the hash operation to determine whether the data item to be detected is in the target data file according to the query result.

[0013] Optionally, determining the shard identifier corresponding to the data item to be detected includes: determining the shard identifier corresponding to the data item to be detected according to the identifier of the data item to be detected, the shard number of the distributed cache system, and the identifier of the target data file.

[0014] Optionally, obtaining a Bloom filter according to the shard identifier and the identifier of the target data file includes: obtaining a Bloom filter from the local cache according to the shard identifier and the identifier of the target data file; if the Bloom filter cannot be obtained from the local cache, obtaining a Bloom filter from the distributed cache system according to the shard identifier and the identifier of the target data file, and writing the Bloom filter into the local cache.

[0015] To achieve the above object, according to the third aspect of the present invention, a data writing device is provided.

[0016] The data writing device of the present invention includes: an acquisition module for acquiring the data file to be written and the line number information of the data file to be written; a generation module for generating a Bloom filter according to the line number information of the data file to be written and the shard number information of the distributed cache system; a hash operation module for performing a hash operation on the data items in the data file to be written based on the Bloom filter; and a writing module for writing the data file to be written into the distributed cache system according to the result of the hash operation.

[0017] To achieve the above object, according to the fourth aspect of the present invention, a data query device is provided.

[0018] The data query device of the present invention includes: an acquisition module for determining the shard identifier corresponding to the data item to be detected, and acquiring a Bloom filter according to the shard identifier and the identifier of the target data file; a hash operation module for performing a hash operation on the data item to be detected according to the Bloom filter; and a query module for querying the distributed cache system according to the result of the hash operation to determine whether the data item to be detected is in the target data file according to the query result.

[0019] To achieve the above object, according to the fifth aspect of the present invention, an electronic device is provided.

[0020] The electronic device of the present invention includes: one or more processors; and a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the data writing method or the data query method of the present invention.

[0021] To achieve the above object, according to the sixth aspect of the present invention, a computer-readable medium is provided.

[0022] The computer-readable medium of the present invention stores a computer program thereon, and when the program is executed by a processor, the data writing method or the data query method of the present invention is implemented.

[0023] One embodiment of the above invention has the following advantages or beneficial effects: By acquiring the data file to be written and the line number information of the data file to be written, generating a Bloom filter according to the line number information of the data file to be written and the shard number information of the distributed cache system, performing a hash operation on the data items in the data file to be written based on the Bloom filter, and writing the data file to be written into the distributed cache system according to the result of the hash operation, the data writing speed can be improved, the impact of data writing on system performance can be reduced, and the user experience can be improved.

[0024] The further effects of the above-mentioned non-conventional alternative ways will be described below in conjunction with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:

[0026] Figure 1 is an exemplary system architecture diagram to which the embodiments of the present invention can be applied;

[0027] Figure 2 is a schematic diagram of the main process of the data writing method according to the first embodiment of the present invention;

[0028] Figure 3 is a schematic diagram of the main process of the data query method according to the second embodiment of the present invention;

[0029] Figure 4 is a schematic diagram of the main modules of the data writing device according to the third embodiment of the present invention;

[0030] Figure 5 is a schematic diagram of the main modules of the data query device according to the fourth embodiment of the present invention;

[0031] Figure 6 is a schematic diagram of the structure of the computer system of the electronic device suitable for implementing the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] The following describes exemplary embodiments of the present invention with reference to the drawings. Various details of the embodiments of the present invention are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted below.

[0033] It should be noted that, without affecting the implementation of the present invention, the various embodiments in the present invention and the technical features in the embodiments can be combined with each other.

[0034] Before introducing the embodiments of the present invention in detail, some technical terms related to the embodiments of the present invention are first described.

[0035] Bloom Filter: Bloom Filter, proposed by Bloom in 1970, is actually a very long binary vector and a series of random mapping functions.

[0036] Figure 1An exemplary system architecture 100 to which the data writing / querying method or data writing / querying device according to embodiments of the present invention can be applied is shown.

[0037] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0038] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0039] The terminal devices 101, 102, 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0040] The server 105 may be a server providing various services, such as a background management server that supports shopping websites browsed by users using the terminal devices 101, 102, 103. For example, the background management server may process data writing requests or data query requests, etc., sent by the terminal device through the network, and feedback the processing results to the terminal device.

[0041] It should be noted that the data writing / querying method provided by the embodiments of the present invention is generally executed by the server 105. Correspondingly, the data writing / querying device is generally set in the server 105.

[0042] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the server in

[0043] Figure 2 is a schematic main flow diagram of the data writing method according to the first embodiment of the present invention. As Figure 2 shown, the data writing method according to the embodiments of the present invention includes:

[0044] Step S201: Obtain the data file to be written and the line number information of the data file to be written.

[0045] In an alternative example, step S201 specifically includes: obtaining information of the data file to be written according to the file identifier of the data file to be written; wherein, the information of the file to be written includes: storage address information of the data file to be written and the number of lines information of the data file to be written; downloading the data file to be written to the local according to the storage address information of the data file to be written.

[0046] For example, when the data file to be written is a user group data file, its file identifier can be the id (identifier) of the user group. In step S201, configuration information such as the storage address of the user group data file and the number of lines information of the user group data file can be obtained according to the id of the user group, and then the user group data file is downloaded from the storage address to the local file.

[0047] In another alternative example, step S201 specifically includes: obtaining the storage address information of the data file to be written according to the file identifier of the data file to be written; downloading the data file to be written to the local according to the storage address information of the data file to be written; traversing the data file to be written to determine the number of lines information of the data file to be written.

[0048] For example, when the data file to be written is a user group data file, the storage address of the user group data file can be obtained according to the id of the user group, and then the user group data file is downloaded from the storage address to the local file, and the number of lines information of the file is obtained by traversing the user group data file.

[0049] Step S202: Generate a Bloom filter according to the number of lines information of the data file to be written and the shard number information of the distributed cache system.

[0050] Exemplarily, step S202 specifically includes: calculating the ratio of a preset value to the number of shards of the distributed cache system, determining a plurality of line number value ranges according to the ratio; determining the line number value range where the data file to be written is located according to the number of lines information of the data file to be written, and using the misjudgment rate corresponding to the line number value range where it is located as the misjudgment rate of the Bloom filter. Wherein, the distributed cache system can be Redis (an open-source log-type, Key-Value database written in ANSI C language, supporting network, memory-based and persistent), JIMDB (a distributed memory storage system) or other distributed cache systems.

[0051] For example, assume that the number of lines to be written to the data file is "lineCount", the number of shards of the distributed cache system is "shardNum", and the preset values are 5,000,000, 10,000,000, and 20,000,000. Then, the values of 5,000,000 / shardNum, 10,000,000 / shardNum, and 20,000,000 / shardNum can be calculated. Then, according to the above calculation results, the following value ranges are determined: (0, 5,000,000 / shardNum), and the misjudgment rate corresponding to this value range is 0.01; [5,000,000 / shardNum, 10,000,000 / shardNum), and the misjudgment rate corresponding to this value range is 0.05; [10,000,000 / shardNum, 20,000,000 / shardNum), and the misjudgment rate corresponding to this value range is 0.1; (20,000,000 / shardNum, +∞), and the misjudgment rate corresponding to this value range is 0.2. Then, it is determined which of the above value ranges the number of lines lineCount of the data file to be written falls into, and the misjudgment rate corresponding to the value range where it is located is used as the misjudgment rate of the Bloom filter. For example, assume that the number of lines lineCount of the data file to be written falls into the value range [10,000,000 / shardNum, 20,000,000 / shardNum), then the misjudgment rate of the Bloom filter is 0.1.

[0052] In the embodiment of the present invention, by generating a Bloom filter according to the number of lines information of the data file to be written and the number of shards information of the distributed cache system, the bitmap can be scattered onto the shards of the distributed cache system, avoiding the bitmap of the generated Bloom filter from being too large. Furthermore, when writing data subsequently, it is possible to avoid writing data into one bitmap, reducing the occupancy of the cache bandwidth of the distributed cache system by data writing, which helps to further improve the data writing efficiency and reduce the impact of data writing on the system performance.

[0053] Step S203: Perform a hash operation on the data items in the data file to be written based on the Bloom filter.

[0054] Exemplarily, step S203 specifically includes: determining the shard identifier corresponding to the data item in the data file to be written; generating a key corresponding to the data item according to the shard identifier corresponding to the data item and the file identifier of the data file to be written; performing a hash operation on the key corresponding to the data item.

[0055] Further, in the above example, the shard identifier corresponding to the data item to be written into the data file can be determined in the following manner: traverse the data file to be written to obtain the identifiers of each data item in the data file to be written; determine the shard identifier corresponding to the data item according to the identifier of the data item, the number of shards of the distributed cache system, and the file identifier of the data file to be written.

[0056] For example, when the data file to be written is a user group data file, its file identifier can be the id (identifier) of the user group, and the data item identifier can be a user identifier such as a username or a user mobile phone number. In step S203, the user group data file can be traversed first to obtain all user identifiers in the file; for each user identifier, the shard identifier corresponding to it can be determined in the following manner: shardCode = groupId + pin.hashcode & (shardNum - 1)). Where shardCode represents the shard identifier, groupId represents the id of the user group, pin represents the user identifier, pin.Hashcode represents the hash operation on the user identifier, and shardNum represents the number of shards. After obtaining the shard identifiers corresponding to each user identifier, keys corresponding to each user identifier can be generated according to the shard identifiers corresponding to the user identifiers and the id of the user group, and multiple hash operations can be performed on the keys corresponding to each user identifier based on the Bloom filter generated in step S202 to obtain an array composed of hash operation values.

[0057] Step S204: Write the data file to be written into the distributed cache system according to the result of the hash operation.

[0058] Exemplarily, the result of the hash operation can specifically be an array composed of hash operation values corresponding to each data item. In this example, the array can be traversed, and each data item in the data file to be written can be written into the distributed cache system through the pipelelineClietnt.setBit method. For example, assuming that the data item is specifically "zhangsan", 3 hash operation values, that is, 3 offsets "1, 2, 3" can be obtained after 3 hash operations. Next, the pipelelineClietnt.setBit method can be used to set the positions with offsets 1, 2, and 3 of the key corresponding to "zhangsan" in the bitmap to true, thereby realizing writing "zhangsan" into the distributed cache system.

[0059] In the embodiments of the present invention, by combining a Bloom filter with a distributed cache system, and through such as Figure 2The process shown can achieve data writing, which can improve the data writing speed, reduce the impact of data writing on system performance, and improve the user experience. Additionally, by generating a Bloom filter based on the number of lines of the data file to be written and the number of shards of the distributed cache system, the bitmap of the Bloom filter can be scattered across the shards of the distributed cache system, avoiding an overly large bitmap (bit map) of the generated Bloom filter. Furthermore, when writing data subsequently, it is possible to avoid writing data into a single bitmap, reducing the occupancy of the cache bandwidth of the distributed cache system for data writing and data reading, which helps to further improve the data writing and reading efficiency and reduce the impact of data writing and reading on system performance.

[0060] Furthermore, in addition to achieving data writing through the Figure 2 process shown, the present invention also provides a data query method based on the Figure 2 basis shown.

[0061] Figure 3 is a schematic diagram of the main process of the data query method according to the second embodiment of the present invention. As Figure 3 shown, the data query method of the embodiment of the present invention includes:

[0062] Step S301: Determine the shard identifier corresponding to the data item to be detected, and obtain a Bloom filter based on the shard identifier and the identifier of the target data file.

[0063] In specific implementation, step S301 can be executed after receiving an interface call request. Among them, the interface call request may include: the identifier of the data item to be detected, the identifier of the target data file. For example, in the application scenario of detecting whether the currently logged-in user hits a certain user group, the foreground can initiate a corresponding interface call request, and this interface call request may include: the user identifier of the currently logged-in user, the id of the user group.

[0064] In an optional example, the shard identifier corresponding to the data item to be detected can be determined in the following manner: Determine the shard identifier corresponding to the data item to be detected based on the identifier of the data item to be detected, the number of shards of the distributed cache system, and the identifier of the target data file. Among them, the distributed cache system can be Redis (an open-source Key-Value database written in ANSI C language, supporting networking, memory-based and persistent log-based), JIMDB (a distributed memory storage system), or other distributed cache systems.

[0065] For example, when the target data file is specifically a user group data file (a data file composed of multiple user identifiers) and the data item to be detected is specifically a certain user identifier, the shard identifier corresponding to the data item to be detected can be determined based on the data item to be detected, the number of shards of the distributed cache system, and the id of the user group. For example, the shard identifier corresponding to the data item to be detected can be determined in the following manner: shardCode = groupId + pin.hashcode & (shardNum - 1)). Here, shardCode represents the shard identifier corresponding to the data item to be detected, groupId represents the id of the user group, pin represents the data item to be detected, pin.Hashcode represents the hash operation on the data item to be detected, and shardNum represents the number of shards.

[0066] After determining the shard identifier corresponding to the data item to be detected, a key (Key) of the data item to be detected can be generated based on the shard identifier and the identifier of the target data file, and then a Bloom filter can be obtained according to the key of the data item to be detected.

[0067] In an alternative example, a Bloom filter can be obtained in the following manner: a key (Key) of the data item to be detected is generated based on the shard identifier and the identifier of the target data file, and then the Bloom filter is obtained from the distributed cache system according to the key of the data item to be detected. It should be noted that the obtained Bloom filter specifically refers to the configuration information of the Bloom filter, that is, the configuration of the Bloom filter generated based on the number of data file rows and the false positive rate, rather than the bitmap of the Bloom filter.

[0068] In another alternative example, a Bloom filter can be obtained in the following manner: the Bloom filter is obtained from the local cache based on the shard identifier and the identifier of the target data file; if the Bloom filter cannot be obtained from the local cache, the Bloom filter is obtained from the distributed cache system based on the shard identifier and the identifier of the target data file, and the Bloom filter is written into the local cache. By first obtaining the Bloom filter from the local cache and then obtaining the Bloom filter from the distributed cache system when it cannot be obtained, the network overhead required to obtain the Bloom filter can be reduced and the acquisition efficiency can be improved.

[0069] Step S302: Perform a hash operation on the data item to be detected according to the Bloom filter.

[0070] Exemplarily, step S302 may specifically include: performing multiple hash operations on the key corresponding to the data item to be detected based on the Bloom filter to obtain the hash operation values corresponding to the data item to be detected. For example, assume that the data item to be detected is specifically "zhangsan". After performing 3 hash operations on the key corresponding to "zhangsan", 3 hash operation values, that is, 3 offsets "1, 2, 3", can be obtained.

[0071] Step S303: Query the distributed cache system according to the result of the hash operation, and determine whether the data item to be detected is located in the target data file according to the query result.

[0072] Exemplarily, after obtaining the hash operation values (i.e., offsets) corresponding to the data item to be detected according to step S202, the distributed cache system can be queried according to these hash operation values. If the values at all positions corresponding to the offsets are true, it indicates that the data item to be detected is located in the target data file; otherwise, it indicates that the data item to be detected is not located in the target data file.

[0073] Further, after determining whether the data item to be detected is located in the target data file through step S303, the detection result can also be returned to the interface caller.

[0074] In the embodiments of the present invention, by combining the Bloom filter with the distributed cache system and implementing data query through the Figure 3 shown process, the data query speed can be improved, the impact of data query on system performance can be reduced, and the user experience can be improved. In addition, since the obtained Bloom filter is generated according to the line number information of the data file to be written and the shard quantity information of the distributed cache system, the bitmap can be scattered onto the shards of the distributed cache system, avoiding the bitmap of the generated Bloom filter from being too large, thereby helping to reduce the occupancy of the cache bandwidth of the distributed cache system during data reading, further improving the data reading efficiency, and reducing the impact of data reading on system performance.

[0075] Figure 4 is a schematic diagram of the main modules of the data writing device according to the third embodiment of the present invention. As Figure 4 shown, the data writing device 400 in the embodiments of the present invention includes: an acquisition module 401, a generation module 402, a hash operation module 403, and a writing module 404.

[0076] The acquisition module 401 is used to acquire the data file to be written and the line number information of the data file to be written.

[0077] Exemplarily, the obtaining module 401 obtaining the data file to be written and the line number information of the data file to be written specifically includes: the obtaining module 401 obtaining the information of the data file to be written according to the file identifier of the data file to be written; wherein, the information of the file to be written includes: the storage address information of the data file to be written and the line number information of the data file to be written; the obtaining module 401 downloading the data file to be written to the local according to the storage address information of the data file to be written.

[0078] For example, when the data file to be written is a user group data file, its file identifier can be the id (identifier) of the user group. The obtaining module 401 can obtain configuration information such as the storage address of the user group data file and the line number information of the user group data file according to the id of the user group, and then download the user group data file from the storage address to the local file.

[0079] The generating module 402 is configured to generate a Bloom filter according to the line number information of the data file to be written and the shard number information of the distributed cache system.

[0080] Exemplarily, the generating module 402 generating a Bloom filter may specifically include: the generating module 402 calculating the ratio of a preset value to the shard number of the distributed cache system, determining a plurality of line number value ranges according to the ratio; the generating module 402 determining the line number value range where it is located according to the line number information of the data file to be written, and using the misjudgment rate corresponding to the line number value range where it is located as the misjudgment rate of the Bloom filter. Wherein, the distributed cache system can be Redis (an open-source key-value database written in ANSI C, supporting networking, memory-based and persistent log-type), JIMDB (a distributed memory storage system) or other distributed cache systems.

[0081] For example, assume that the number of lines in the data file to be written is "lineCount", the number of shards in the distributed cache system is "shardNum", and the preset values are 5,000,000, 10,000,000, and 20,000,000. Then, the values of 5,000,000 / shardNum, 10,000,000 / shardNum, and 20,000,000 / shardNum can be calculated. Next, the following value ranges can be determined based on the above calculation results: (0, 5,000,000 / shardNum), and the misjudgment rate corresponding to this value range is 0.01; [5,000,000 / shardNum, 10,000,000 / shardNum), and the misjudgment rate corresponding to this value range is 0.05; [10,000,000 / shardNum, 20,000,000 / shardNum), and the misjudgment rate corresponding to this value range is 0.1; (20,000,000 / shardNum, +∞), and the misjudgment rate corresponding to this value range is 0.2. Then, it is determined which of the above value ranges the number of lines lineCount in the data file to be written falls into, and the misjudgment rate corresponding to the value range where it is located is used as the misjudgment rate of the Bloom filter. For example, assume that the number of lines lineCount in the data file to be written falls into the value range [10,000,000 / shardNum, 20,000,000 / shardNum), then the misjudgment rate of the Bloom filter is 0.1.

[0082] In the embodiment of the present invention, the generation module 402 generates a Bloom filter according to the line number information of the data file to be written and the shard number information of the distributed cache system, which can scatter the bitmap to the shards of the distributed cache system, avoiding the bitmap (bit map) of the generated Bloom filter from being too large. Furthermore, when writing data subsequently, it is possible to avoid writing data into one bitmap, reducing the occupancy of the cache bandwidth of the distributed cache system by data writing, which helps to further improve the data writing efficiency and reduce the impact of data writing on system performance.

[0083] The hash operation module 403 is used to perform a hash operation on the data items in the data file to be written based on the Bloom filter.

[0084] Exemplarily, the hash operation module 403 performing a hash operation on the data items in the data file to be written based on the Bloom filter may specifically include: the hash operation module 403 determines the shard identifier corresponding to the data item in the data file to be written; the hash operation module 403 generates a key corresponding to the data item according to the shard identifier corresponding to the data item and the file identifier of the data file to be written; the hash operation module 403 performs a hash operation on the key corresponding to the data item.

[0085] Further, in the above example, the hash operation module 403 may determine the shard identifier corresponding to the data item to be written into the data file in the following manner: The hash operation module 403 traverses the data file to be written to obtain the identifiers of each data item in the data file to be written; the hash operation module 403 determines the shard identifier corresponding to the data item according to the identifier of the data item, the number of shards of the distributed cache system, and the file identifier of the data file to be written.

[0086] For example, when the data file to be written is a user group data file, its file identifier may be the id (identifier) of the user group, and the data item identifier may be a user identifier such as a username or a user mobile phone number. The hash operation module 403 may first traverse the user group data file to obtain all user identifiers in the file; further, for each user identifier, the hash operation module 403 may determine the corresponding shard identifier in the following manner: shardCode = groupId + pin.hashcode & (shardNum - 1)). Where shardCode represents the shard identifier, groupId represents the id of the user group, pin represents the user identifier, pin.Hashcode represents the hash operation on the user identifier, and shardNum represents the number of shards. After obtaining the shard identifiers corresponding to each user identifier, the hash operation module 403 may generate keys corresponding to each user identifier according to the shard identifier corresponding to the user identifier and the id of the user group, and perform multiple hash operations on the keys corresponding to each user identifier based on the generated Bloom filter to obtain an array composed of hash operation values.

[0087] The writing module 404 is configured to write the data file to be written into the distributed cache system according to the result of the hash operation.

[0088] Exemplarily, the result of the hash operation may specifically be an array composed of hash operation values corresponding to each data item. In this example, the writing module 404 may traverse the array and write each data item in the data file to be written into the distributed cache system through the pipelelineClietnt.setBit method. For example, assuming that the data item is specifically "zhangsan", after 3 hash operations, 3 hash operation values, that is, 3 offsets "1, 2, 3" can be obtained. Next, the writing module 404 may set the positions with offsets 1, 2, and 3 of the key corresponding to "zhangsan" in the bitmap to true through the pipelelineClietnt.setBit method, thereby realizing writing "zhangsan" into the distributed cache system.

[0089] In an embodiment of the present invention, by combining a Bloom filter with a distributed cache system and implementing data writing through the modules as shown in Figure 4 , the data writing speed can be improved, the impact of data writing on system performance can be reduced, and the user experience can be enhanced. Additionally, by generating a Bloom filter based on the line number information of the data file to be written and the shard number information of the distributed cache system, the bitmap of the Bloom filter can be scattered across the shards of the distributed cache system, preventing the generated bitmap of the Bloom filter from being too large. Furthermore, when writing data subsequently, it is possible to avoid writing data into a single bitmap, reducing the occupancy of the cache bandwidth of the distributed cache system for data writing and data reading, which helps to further improve the data writing and reading efficiency and reduce the impact of data writing and reading on system performance.

[0090] Furthermore, in addition to implementing data writing through the device as shown in Figure 4 , the present invention also provides a data query device.

[0091] Figure 5 is a schematic diagram of the main modules of the data query device according to the fourth embodiment of the present invention. As shown in Figure 5 , the data query device 500 in the embodiment of the present invention includes: an acquisition module 501, a hash operation module 502, and a query module 503.

[0092] The acquisition module 501 is configured to determine the shard identifier corresponding to the data item to be detected, and obtain a Bloom filter according to the shard identifier and the identifier of the target data file.

[0093] In an alternative example, the acquisition module 501 can determine the shard identifier corresponding to the data item to be detected in the following manner: The acquisition module 501 determines the shard identifier corresponding to the data item to be detected according to the identifier of the data item to be detected, the shard number of the distributed cache system, and the identifier of the target data file. Among them, the distributed cache system can be Redis (an open-source key-value database written in ANSI C, supporting networking, memory-based and persistent logging), JIMDB (a distributed memory storage system), or other distributed cache systems.

[0094] For example, when the target data file is specifically a user group data file (a data file composed of multiple user identifiers) and the data item to be detected is specifically a certain user identifier, the obtaining module 501 can determine the shard identifier corresponding to the user identifier to be detected according to the user identifier to be detected, the number of shards of the distributed cache system, and the id of the user group. For example, the obtaining module 501 can determine the shard identifier corresponding to the user identifier to be detected in the following manner: shardCode = groupId + pin.hashcode & (shardNum - 1)). Here, shardCode represents the shard identifier corresponding to the user identifier to be detected, groupId represents the id of the user group, pin represents the user identifier to be detected, pin.Hashcode represents the hash operation on the user identifier to be detected, and shardNum represents the number of shards.

[0095] After determining the shard identifier corresponding to the data item to be detected, the obtaining module 501 can generate the key of the data item to be detected according to the shard identifier and the identifier of the target data file, and then obtain the Bloom filter according to the key of the data item to be detected.

[0096] In an optional example, the obtaining module 501 can obtain the Bloom filter in the following manner: the obtaining module 501 generates the key of the data item to be detected according to the shard identifier and the identifier of the target data file, and then obtains the Bloom filter from the distributed cache system according to the key of the data item to be detected. It should be noted that the Bloom filter obtained by the obtaining module 501 specifically refers to the configuration information of the Bloom filter, that is, the configuration of the Bloom filter generated based on the number of data file rows and the false positive rate, rather than the bitmap of the Bloom filter.

[0097] In another optional example, the obtaining module 501 can obtain the Bloom filter in the following manner: the obtaining module 501 obtains the Bloom filter from the local cache according to the shard identifier and the identifier of the target data file; if the Bloom filter cannot be obtained from the local cache, the obtaining module 501 obtains the Bloom filter from the distributed cache system according to the shard identifier and the identifier of the target data file, and writes the Bloom filter into the local cache. By first obtaining the Bloom filter from the local cache and then obtaining it from the distributed cache system when it cannot be obtained, the network overhead required to obtain the Bloom filter can be reduced and the obtaining efficiency can be improved.

[0098] The hash operation module 502 is used to perform a hash operation on the data item to be detected according to the Bloom filter.

[0099] Exemplarily, the hash operation module 502 performs multiple hash operations on the key corresponding to the data item to be detected based on the Bloom filter to obtain the hash operation value corresponding to the data item to be detected. For example, assuming that the data item to be detected is specifically "zhangsan", the hash operation module 502 can perform 3 hash operations on the key corresponding to "zhangsan", thereby obtaining 3 hash operation values, that is, 3 offsets "1, 2, 3".

[0100] The query module 503 is used to query the distributed cache system according to the result of the hash operation, so as to judge whether the data item to be detected is located in the target data file according to the query result.

[0101] Exemplarily, after obtaining the hash operation value corresponding to the data item to be detected, that is, the offset, the query module 503 can query the distributed cache system according to these hash operation values. If the values at all positions corresponding to the offsets are true, it means that the data item to be detected is located in the target data file; otherwise, it means that the data item to be detected is not located in the target data file.

[0102] Furthermore, after judging whether the data item to be detected is located in the target data file, the query module 503 can also return the detection result to the interface caller.

[0103] In the embodiments of the present invention, by combining the Bloom filter with the distributed cache system and implementing data query through the modules as Figure 5 shown, the data query speed can be improved, the impact of data query on system performance can be reduced, and the user experience can be improved. In addition, since the Bloom filter obtained by the acquisition module is generated according to the line number information of the data file to be written and the shard number information of the distributed cache system, the bitmap can be scattered to the shards of the distributed cache system, avoiding the bitmap (bitmap) of the generated Bloom filter from being too large, thereby helping to reduce the occupancy of the cache bandwidth of the distributed cache system for data reading, further improving the data reading efficiency, and reducing the impact of data reading on system performance.

[0104] Next, refer to Figure 6 , which shows a schematic structural diagram of a computer system 600 of an electronic device suitable for implementing the embodiments of the present invention. Figure 6 The shown computer system is only an example and should not bring any limitation to the functions and usage scopes of the embodiments of the present invention.

[0105] As Figure 6As shown, computer system 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded into a random access memory (RAM) 603 from a storage section 608. In the RAM 603, various programs and data required for the operation of the system 600 are also stored. The CPU 601, ROM 602, and RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0106] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read from it can be installed into the storage section 608 as needed.

[0107] Specifically, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 609, and / or installed from the removable medium 611. When the computer program is executed by the central processing unit (CPU) 601, the above functions defined in the system of the present invention are executed.

[0108] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0109] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0110] The modules involved in the embodiments of the present invention can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, it can be described as: a processor includes an acquisition module, a generation module, a hashing operation module, and a writing module. Among them, the names of these modules do not constitute a limitation to the module itself in some cases. For example, the acquisition module can also be described as "a module for acquiring the data file to be written and the line number information of the data file to be written".

[0111] As another aspect, the present invention also provides a computer-readable medium, which can be included in the device described in the above embodiments; or can exist alone without being assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the device, the device is caused to execute the following data writing process: acquire the data file to be written and the line number information of the data file to be written, generate a Bloom filter according to the line number information of the data file to be written and the shard quantity information of the distributed cache system, perform a hashing operation on the data items in the data file to be written based on the Bloom filter, and write the data file to be written into the distributed cache system according to the result of the hashing operation; or cause the device to execute the following data query process: determine the shard identifier corresponding to the data item to be detected, acquire the Bloom filter according to the shard identifier and the identifier of the target data file; perform a hashing operation on the data item to be detected according to the Bloom filter; query the distributed cache system according to the result of the hashing operation to determine whether the data item to be detected is in the target data file according to the query result.

[0112] According to the technical solution of the embodiments of the present invention, it is possible to improve the data writing and query speeds, reduce the impact of data writing and query on system performance, and improve the user experience.

[0113] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A data writing method, characterized in that, The method includes: Obtaining a data file to be written and line number information of the data file to be written; Calculating a ratio of a preset value to the number of shards of the distributed cache system, and determining a plurality of line number value ranges according to the ratio; determining the line number value range where the data file to be written is located according to the line number information of the data file to be written, and using the misjudgment rate corresponding to the line number value range where it is located as the misjudgment rate of the Bloom filter; Determining a shard identifier corresponding to a data item in the data file to be written; generating a key corresponding to the data item according to the shard identifier corresponding to the data item and the file identifier of the data file to be written; performing a hash operation on the key corresponding to the data item; Writing the data file to be written into the distributed cache system according to the result of the hash operation.

2. The method according to claim 1, characterized in that, The obtaining a data file to be written and line number information of the data file to be written includes: Obtaining information of the data file to be written according to the file identifier of the data file to be written; wherein, the information of the data file to be written includes: storage address information of the data file to be written and line number information of the data file to be written; Downloading the data file to be written to the local according to the storage address information of the data file to be written.

3. The method according to claim 1, characterized in that, The determining a shard identifier corresponding to a data item in the data file to be written includes: Traversing the data file to be written to obtain identifiers of each data item in the data file to be written; determining the shard identifier corresponding to the data item according to the identifier of the data item, the number of shards of the distributed cache system, and the file identifier of the data file to be written.

4. A data query method, characterized in that, The method includes: Determining a shard identifier corresponding to a data item to be detected, and obtaining a Bloom filter according to the shard identifier and the identifier of the target data file; wherein, the target data file is written into the distributed cache system by using the data writing method described in any one of claims 1-3; Performing a hash operation on the data item to be detected according to the Bloom filter; Querying the distributed cache system according to the result of the hash operation to determine whether the data item to be detected is located in the target data file according to the query result.

5. The method according to claim 4, wherein The determining a shard identifier corresponding to a data item to be detected includes: Determining the shard identifier corresponding to the data item to be detected according to the identifier of the data item to be detected, the number of shards of the distributed cache system, and the identifier of the target data file.

6. The method according to claim 4, wherein The obtaining a Bloom filter according to the shard identifier and the identifier of the target data file includes: Obtaining a Bloom filter from the local cache according to the shard identifier and the identifier of the target data file; if the Bloom filter cannot be obtained from the local cache, obtaining the Bloom filter from the distributed cache system according to the shard identifier and the identifier of the target data file, and writing the Bloom filter into the local cache.

7. A data writing device, characterized in that, The apparatus includes: An obtaining module, configured to obtain a data file to be written and line number information of the data file to be written; A generation module, configured to calculate a ratio of a preset value to the number of shards of a distributed cache system, determine a plurality of row number value ranges according to the ratio; determine the row number value range where the row number information of the data file to be written is located, and use the misjudgment rate corresponding to the row number value range where it is located as the misjudgment rate of the Bloom filter; A hash operation module, configured to determine a shard identifier corresponding to a data item in the data file to be written; generate a key corresponding to the data item according to the shard identifier corresponding to the data item and the file identifier of the data file to be written; perform a hash operation on the key corresponding to the data item; A writing module, configured to write the data file to be written into the distributed cache system according to the result of the hash operation.

8. A data query device, characterized in that, The apparatus includes: An acquisition module, configured to determine a shard identifier corresponding to a data item to be detected, and acquire a Bloom filter according to the shard identifier and the identifier of a target data file; wherein, the target data file is written into the distributed cache system by using the data writing method described in any one of claims 1-3; A hash operation module, configured to perform a hash operation on the data item to be detected according to the Bloom filter; A query module, configured to query the distributed cache system according to the result of the hash operation, so as to determine whether the data item to be detected is located in the target data file according to the query result.

9. An electronic device, characterized in that, It includes: One or more processors; A storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any one of claims 1-6.

10. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method described in any one of claims 1-6.

Citation Information

Patent Citations

  • Data query request processing method and apparatus, and electronic device

    CN106445944A