Data Processing Method, Apparatus, Computer-Readable Medium, and Electronic Device

By hashing and superimposing storage of the to-process data, combining sorting and non-target data generation, the problem of low efficiency in big data storage and processing is solved, and efficient data processing and storage is achieved.

CN115114283BActive Publication Date: 2025-06-24TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210576032.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-25
Publication Date
2025-06-24
Estimated Expiration
2042-05-25

AI Technical Summary

Technical Problem

With the advent of the big data era, the number of original data has increased, and the storage method has led to a large amount of storage space occupied and reduced data processing efficiency.

Method used

The hashing operation process is performed on the to-process data in the target scene by at least one hashing algorithm, and the hash value is obtained, and the to-process data is superimposed and stored according to the hash value. Then the superimposed data is sorted, the target superimposed data matching the set indicator is obtained, and non-target data is generated, and the data matching the set indicator among the multiple pending data is finally calculated.

Benefits of technology

This greatly reduces the storage space requirements during data processing, improves data processing efficiency, and saves resources required for index calculation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115114283B_ABST
    Figure CN115114283B_ABST
Patent Text Reader

Abstract

The present application discloses a data processing method, apparatus, computer-readable medium, and electronic device. The method includes: performing hash operation processing on each piece of data to be processed in a target scenario through at least one hash algorithm to obtain at least one hash value corresponding to each piece of data to be processed; adding the data to be processed to the data at the corresponding positions in a preset data list corresponding to the hash algorithm and the hash value to obtain superimposed data corresponding to each hash algorithm; sorting the superimposed data corresponding to each hash algorithm in the preset data list, and obtaining target superimposed data and non-target data according to the sorting result; calculating the data that matches a set index among multiple pieces of data to be processed according to the target superimposed data and the non-target data. The technical solution of the present application enables the data to be processed to be transformed from the original value space data into data that occupies less storage space, thereby greatly reducing the requirement for storage space during data processing and saving the resources required for index calculation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of data processing, and particularly relates to a data processing method, apparatus, computer-readable medium, and electronic device. Background Art

[0002] With the advent of the big data era, the amount of raw data to be processed during data analysis is increasing. Generally, raw data is recorded in a sequential storage manner according to the data generation time. When performing data processing, the raw data is retrieved from the corresponding data storage location for calculation. For example, the raw data is accumulated according to certain conditions, and then the calculation results are stored together for subsequent use. However, the more raw data there is, the more storage space this sequential storage method occupies. At the same time, the number of data storage locations to be accessed during data processing will also increase, which may reduce the data processing efficiency.

[0003] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of this application, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0004] The purpose of this application is to provide a data processing method, apparatus, computer-readable medium, and electronic device to optimize the problem of large storage space occupied by data processing in related technologies.

[0005] Other features and advantages of this application will become apparent through the following detailed description, or will be learned in part through the practice of this application.

[0006] According to one aspect of the embodiments of this application, a data processing method is provided, including:

[0007] Performing hash operation processing on each piece of data to be processed in a target scenario through at least one hash algorithm to obtain at least one hash value corresponding to each piece of data to be processed;

[0008] According to the at least one hash value corresponding to the data to be processed, superimposing the data to be processed on the data at the position corresponding to the hash algorithm and the hash value in a preset data list to obtain superimposed data corresponding to each hash algorithm;

[0009] Sorting the superimposed data corresponding to each hash algorithm in the preset data list, obtaining target superimposed data corresponding to each hash algorithm and matching a set index according to the sorting result, and generating non-target data corresponding to each hash algorithm according to the other superimposed data except the target superimposed data in the superimposed data corresponding to each hash algorithm;

[0010] Calculate the data in the multiple data to be processed that matches the set metric according to the target superimposed data and the non-target data, so as to obtain the data processing result corresponding to the target scenario.

[0011] According to one aspect of the embodiments of the present application, a data processing device is provided, including:

[0012] A hash operation module, configured to perform hash operation processing on each data to be processed in the target scenario through at least one hash algorithm, and obtain at least one hash value corresponding to each data to be processed;

[0013] A data superimposing module, configured to superimpose the data to be processed on the data at the position corresponding to the hash algorithm and the hash value in a preset data list according to the at least one hash value corresponding to the data to be processed, and obtain the superimposed data corresponding to each hash algorithm;

[0014] A data calculation module, configured to sort the superimposed data corresponding to each hash algorithm in the preset data list, obtain the target superimposed data corresponding to each hash algorithm and matching the set metric according to the sorting result, and generate non-target data corresponding to each hash algorithm according to the other superimposed data except the target superimposed data in the superimposed data corresponding to each hash algorithm;

[0015] A result generation module, configured to calculate the data in the multiple data to be processed that matches the set metric according to the target superimposed data and the non-target data, so as to obtain the data processing result corresponding to the target scenario.

[0016] In an embodiment of the present application, the data to be processed includes data in the form of key-value pairs; the hash operation module is specifically configured to: perform hash operation processing on the keys of each data to be processed in the target scenario through at least one hash algorithm;

[0017] The data superimposing module is specifically configured to: superimpose the value of the data to be processed on the data at the position corresponding to the hash algorithm and the hash value in the preset data list according to the at least one hash value corresponding to the data to be processed.

[0018] In an embodiment of the present application, the hash algorithm includes hash function operation and modulo operation; the hash operation module includes:

[0019] A hash calculation unit, configured to perform hash calculation on the keys of each data to be processed in the target scenario through at least one hash function operation, and obtain at least one hash result of each data to be processed;

[0020] A modulo operation unit that performs a modulo operation on at least one hash result of each of the data to be processed against a preset number of hash buckets, and uses the result of the modulo operation as at least one hash value corresponding to each of the data to be processed; the preset number of hash buckets is used to indicate the size of the storage space occupied by the preset data list.

[0021] In an embodiment of the present application, the superimposed data corresponding to the hash algorithm includes superimposed data stored in multiple hash buckets, and one hash bucket represents a storage location in the preset data list; the data calculation module includes:

[0022] A target superimposed data generation unit for obtaining, according to the sorting result, the superimposed data stored in a set number of hash buckets corresponding to each of the hash algorithms and matching a set index; summing the superimposed data stored in the set number of hash buckets as the target superimposed data corresponding to each of the hash algorithms.

[0023] In an embodiment of the present application, the data calculation module includes:

[0024] A numerical expectation calculation unit for generating a numerical expectation of non-target data according to the superimposed data other than the target superimposed data corresponding to the hash algorithm and the amount of data to be processed corresponding to the other superimposed data;

[0025] A quantity expectation calculation unit for calculating the quantity expectation of non-target data according to the amount of data to be processed corresponding to the target superimposed data;

[0026] A non-target data calculation unit for generating non-target data corresponding to the hash algorithm according to the product of the numerical expectation of the non-target data and the quantity expectation of the non-target data.

[0027] In an embodiment of the present application, the numerical expectation calculation unit includes:

[0028] A non-target superimposed data generation subunit for summing the superimposed data other than the target superimposed data corresponding to the hash algorithm to obtain non-target superimposed data corresponding to the hash algorithm;

[0029] A data deduplication subunit for determining the number of deduplicated data to be processed corresponding to the other superimposed data according to the amount of data to be processed corresponding to the other superimposed data;

[0030] A numerical expectation calculation subunit for obtaining the numerical expectation of non-target data according to the ratio of the non-target superimposed data to the number of deduplicated data to be processed corresponding to the other superimposed data.

[0031] In one embodiment of the present application, the data deduplication subunit is specifically configured to:

[0032] Deduplicate multiple pieces of data to be processed in the target scenario according to the key of the data to be processed, and obtain the total number of deduplicated data to be processed;

[0033] Deduplicate multiple pieces of data to be processed corresponding to the target superimposed data according to the key of the data to be processed, and obtain the number of deduplicated data to be processed corresponding to the target superimposed data;

[0034] Obtain the number of deduplicated data to be processed corresponding to the other superimposed data according to the difference between the total number of deduplicated data to be processed and the number of deduplicated data to be processed corresponding to the target superimposed data.

[0035] In one embodiment of the present application, the quantity expectation calculation unit includes:

[0036] A first calculation subunit, configured to generate the quantity expectation of the target superimposed data according to a preset number of hash buckets, the total number of deduplicated data to be processed corresponding to the target scenario, and a preset fitting function;

[0037] A second calculation subunit, configured to obtain the quantity expectation of the non-target data according to the difference between the quantity expectation of the target superimposed data and the value of the set index.

[0038] In one embodiment of the present application, the quantity expectation calculation unit further includes:

[0039] A fitting function construction unit, configured to construct a fitting function related to the number of hash buckets and the total number of deduplicated data to be processed according to undetermined fitting coefficients;

[0040] A training unit, configured to train the fitting function through sample data to obtain the target value of the undetermined fitting coefficients; the sample data includes the sample number of hash buckets and the total number of deduplicated sample data corresponding to the target scenario;

[0041] A preset fitting function generation unit, configured to generate the preset fitting function according to the target value of the undetermined fitting coefficients.

[0042] In one embodiment of the present application, the training unit is specifically configured to:

[0043] Randomly generate an initial value of the undetermined fitting coefficients;

[0044] Calculate the predicted quantity expectation of the sample data through the fitting function with the initial value of the undetermined fitting coefficients;

[0045] Adjust the initial value of the to-be-determined fitting coefficient according to the difference between the predicted quantity expectation and the actual quantity expectation of the sample data until the difference is less than a preset threshold, and obtain the target value of the to-be-determined fitting coefficient.

[0046] In an embodiment of the present application, the result generation module includes:

[0047] A data processing unit, configured to obtain data that matches the set index among the multiple pieces of data to be processed corresponding to each hash algorithm according to the difference between the target superimposed data and the non-target data corresponding to each hash algorithm;

[0048] A statistics unit, configured to perform statistical processing on the data that matches the set index among the multiple pieces of data to be processed corresponding to each hash algorithm, and obtain a data processing result corresponding to the target scenario.

[0049] In an embodiment of the present application, the statistics unit is specifically configured to:

[0050] Calculate the expected value of the data that matches the set index among the multiple pieces of data to be processed corresponding to each hash algorithm, and use it as the data processing result corresponding to the target scenario.

[0051] According to one aspect of the embodiments of the present application, there is provided a computer-readable medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the data processing method in the above technical solution.

[0052] According to one aspect of the embodiments of the present application, there is provided an electronic device, which includes: a processor; and a memory, configured to store executable instructions of the processor; wherein, when the processor executes the executable instructions, the electronic device executes the data processing method in the above technical solution.

[0053] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the data processing method in the above technical solution.

[0054] In the technical solution provided by the embodiments of the present application, each piece of data to be processed in the target scenario is subjected to a hashing operation through at least one hashing algorithm to obtain at least one hash value corresponding to each piece of data to be processed, and the data to be processed is stored in an overlay manner according to the hashing algorithm and the hash value. The hashing algorithm enables the data to be processed to be transformed from the original value space data into data that occupies less storage space, thereby greatly reducing the storage space requirements during data processing; then, the overlay data corresponding to each hashing algorithm is sorted, and based on the sorting result, target overlay data that matches the set metrics is obtained, and non-target data is generated based on other overlay data except the target overlay data; finally, the data to be processed that matches the set metrics is obtained according to the target overlay data and the non-target data, and a data processing result is generated. This is equivalent to using a small amount of computing resources to implement the processing and calculation of large-scale data, obtaining the corresponding metric data, and saving the resources required for metric calculation.

[0055] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0057] Figure 1A Schematically shows an exemplary system architecture block diagram applying the technical solution of the present application.

[0058] Figure 1B Schematically shows a schematic diagram of an application scenario of the technical solution of the present application.

[0059] Figure 1C Schematically shows a schematic diagram of another application scenario of the technical solution of the present application.

[0060] Figure 2 Schematically shows a flowchart of a data processing method provided by an embodiment of the present application.

[0061] Figure 3 Schematically shows a flowchart of a data processing method provided by an embodiment of the present application.

[0062] Figure 4 Schematically shows a schematic diagram of a preset data list provided by an embodiment of the present application.

[0063] Figure 5A schematic diagram showing the sorting result of the superimposed data provided by an embodiment of the present application is shown.

[0064] Figure 6 A graph showing a preset fitting function constructed by a linear fitting method provided by an embodiment of the present application is schematically shown.

[0065] Figure 7 A structural block diagram of a data processing device provided by an embodiment of the present application is schematically shown.

[0066] Figure 8 A computer system structural block diagram of an electronic device suitable for implementing the embodiments of the present application is schematically shown. Detailed implementation manners

[0067] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.

[0068] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.

[0069] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0070] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the contents and operations / steps, nor do they necessarily have to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0071] Figure 1A A schematic block diagram of an exemplary system architecture applying the technical solution of the present application is shown.

[0072] As Figure 1AAs shown, the system architecture 100 may include a terminal device 110, a network 120, and a server 130. The terminal device 110 may include a smart phone, a tablet computer, a laptop computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, and so on. The server 130 may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The network 120 may be a communication medium of various connection types capable of providing a communication link between the terminal device 110 and the server 130. For example, it may be a wired communication link or a wireless communication link.

[0073] According to the implementation requirements, the system architecture in the embodiments of the present application may have any number of terminal devices, networks, and servers. For example, the server 130 may be a server group composed of multiple server devices. In addition, the technical solutions provided in the embodiments of the present application may be applied to the terminal device 110, or may be applied to the server 130, or may be jointly implemented by the terminal device 110 and the server 130. The present application does not make special limitations on this.

[0074] In an embodiment of the present application, the data processing method provided in the embodiments of the present application is implemented by the server 130. Specifically: the server 130 performs a hashing operation on each piece of data to be processed in the target scenario through at least one hashing algorithm, and obtains at least one hash value corresponding to each piece of data to be processed; wherein, one piece of data to be processed obtains one hash value through one hashing algorithm, and the data to be processed may be sent from the terminal device 110 to the server 130. Then, the server 130 stacks the data to be processed onto the data at the corresponding positions of the preset data list corresponding to the hashing algorithm and the hash value according to the at least one hash value corresponding to the data to be processed, and obtains the stacked data corresponding to each hashing algorithm; that is, the data to be processed is stored in the preset data list in a stacked storage manner. Next, the server 130 sorts the stacked data corresponding to each hashing algorithm in the preset data list, obtains the target stacked data corresponding to each hashing algorithm and matching the set index according to the sorting result, and generates non-target data corresponding to each hashing algorithm according to the other stacked data except the target stacked data in the stacked data corresponding to each hashing algorithm; wherein, the set index is usually used to limit the selection range of data. For example, the set index indicates to select the top five sorted data. Finally, the server 130 calculates the data in the multiple pieces of data to be processed that match the set index according to the target stacked data and the non-target data, so as to obtain the data processing result corresponding to the target scenario. The technical solution of the present application is equivalent to calculating according to the stacked storage data of the data to be processed to obtain the data in the original data to be processed that matches the set index. The server 130 may send the data processing result to the terminal device 110 so that the terminal device 110 can perform a visual display on the data processing result.

[0075] In an embodiment of the present application, taking the transfer scenario of virtual resources as the target scenario as an example to illustrate the implementation process of the technical solution of the present application, the set indicator is the resource transfer volume of the top three accounts in terms of the total virtual resource transfer. In the transfer scenario of virtual resources, a data to be processed represents a transfer record, which represents the amount of virtual resources transferred during a settlement process of an account, and can be expressed as (account identifier, resource transfer volume). For example, if an account identifier is 123 and the resource transfer volume is 100, the data to be processed is expressed as (123, 100).

[0076] In the transfer scenario of virtual resources, an account can conduct multiple settlements, that is, it has multiple transfer records. Then, an account can correspond to multiple data to be processed. For example, the data to be processed corresponding to an account includes (account identifier, resource transfer volume 1), (account identifier, resource transfer volume 2), (account identifier, resource transfer volume 3), etc. Denote the data to be processed in the transfer scenario of virtual resources as: (account identifier 1, resource transfer volume 1), (account identifier 2, resource transfer volume 2)…(account identifier n, resource transfer volume n), where the numbers of the account identifiers are mainly used to distinguish transfer records, and the same account identifier may exist among these account identifiers with different numbers. Exemplarily, Figure 1B A schematic diagram of an application scenario of the technical solution of the present application is schematically shown. As Figure 1B shown, on the one hand, account 1 of virtual resources transfers a certain amount of virtual resources to account 2 of virtual resources, and this resource transfer forms a transfer record (account identifier 1, resource transfer volume 1). On the other hand, account 1 of virtual resources transfers a certain amount of virtual resources to account 3 of virtual resources, and this resource transfer forms a transfer record (account identifier 2, resource transfer volume 2). Both of these two transfer records are data to be processed. It can be seen that the two transfer records are distinguished by account identifier 1 and account identifier 2, but account identifier 1 and account identifier 2 are the same account identifier, both of which are the account identifiers of account 1 of virtual resources.

[0077] First, the server 130 performs hash operation processing on each data to be processed in the target scenario through at least one hash algorithm to obtain at least one hash value corresponding to each data to be processed. Assuming that f(x) represents the hash algorithm, then the hash value obtained by performing hash operation processing on the data to be processed can be expressed as: f(account identifier) = hash value. If there are f1(x), f2(x)…f n (x) for the hash algorithms, then for the data to be processed (account identifier 1, resource transfer volume 1), f1(account identifier 1) = hash value 1, f2(account identifier 1) = hash value 2, …, f n(Account identifier 1) = hash value n. That is, a hash algorithm can obtain a hash value through one hash operation on the data to be processed. A data to be processed can obtain at least one hash value through the hash operations of at least one hash algorithm. Exemplarily, as Figure 1B shown, the data to be processed (Account identifier 1, Resource transfer amount 1) is processed by hash algorithms f1(x), f2(x)…f n (x) to obtain n hash values; the data to be processed (Account identifier 2, Resource transfer amount 2) is processed by hash algorithms f1(x), f2(x)…f n (x) to obtain n hash values.

[0078] Then, the server 130 superimposes the data to be processed on the data at the position corresponding to the hash algorithm and the hash value in the preset data list according to at least one hash value corresponding to the data to be processed, to obtain the superimposed data corresponding to each hash algorithm. Exemplarily, as Figure 1B shown, the hash value obtained by the data to be processed (Account identifier 1, Resource transfer amount 1) through the hash algorithm f1(x) is 1, then the Resource transfer amount 1 is superimposed on the position indicated by (1, 1) in the preset data list. Assume that the data stored at (1, 1) is 200 and the Resource transfer amount 1 is 100. Superimposing means adding the data 200 stored at (1, 1) and the Resource transfer amount 100, and the obtained superimposed data is 300. For the data to be processed (Account identifier 2, Resource transfer amount 2), as Figure 1B shown, since the Account identifier 2 is actually equal to the Account identifier 1, therefore, f1(Account identifier 2) = f1(Account identifier 1) = 1, then the storage position of the Resource transfer amount 2 is the same as that of the Resource transfer amount 1, that is, the Resource transfer amount 2 is superimposed on the position indicated by (1, 1) in the preset data list. Assume that after storing the Resource transfer amount 1, the data at (1, 1) is 300 and the Resource transfer amount 2 is 50. After superimposing and storing the Resource transfer amount 2, the superimposed data at (1, 1) is 350.

[0079] Next, the server 130 sorts the superimposed data corresponding to each hash algorithm in the preset data list, obtains the target superimposed data corresponding to each hash algorithm and matching the set metrics according to the sorting result, and generates non-target data corresponding to each hash algorithm based on the other superimposed data except the target superimposed data in the superimposed data corresponding to each hash algorithm. As in the previous example, the set metric is the resource transfer volume of the top three accounts in terms of the total virtual resource transfer. Then the target superimposed data is generated based on the top three superimposed data in the superimposed data corresponding to each hash algorithm. The target superimposed data includes the resource transfer volume of the top three accounts in terms of the total virtual resource transfer and the resource transfer volume of the accounts other than the top three in terms of the total virtual resource transfer. The non-target data is generated based on the other superimposed data except the top three superimposed data, and the non-target data includes the resource transfer volume of the accounts other than the top three in terms of the total virtual resource transfer.

[0080] Finally, the server 130 calculates the data that matches the set metrics among the multiple data to be processed based on the target superimposed data and the non-target data, so as to obtain the data processing result corresponding to the target scenario. Subtracting the non-target data from the target superimposed data can obtain the resource transfer volume of the top three accounts in terms of the total virtual resource transfer, and thus the data processing result is obtained. Exemplarily, as Figure 1B shown, the target superimposed data and the non-target data are obtained according to the superimposed data stored in the preset data list, and then the data processing result is generated based on the target superimposed data and the non-target data.

[0081] In an embodiment of the present application, the target scenario may also be a network access scenario, and the data to be processed may be a record of the website and the corresponding number of visits to the website. For example, the data to be processed is represented as (website identifier, number of visits). Exemplarily, if the IP address (Internet Protocol Address) of a certain website is 1.2.3.4 and the number of visits is 50 times, then the data to be processed is represented as (1.2.3.4, 50). Exemplarily, Figure 1C schematically shows a schematic diagram of an application scenario of the technical solution of the present application. As Figure 1CAs shown in the figure, the number of clicks on the website is counted by day. For website 1, the number of clicks in two days can respectively form the data to be processed: (website identifier 1, number of visits 1) and (website identifier 2, number of visits 2). Similarly, for website 2, the number of clicks in two days can respectively form the data to be processed: (website identifier 3, number of visits 3) and (website identifier 4, number of visits 4). It can be seen that although website identifier 1 and website identifier 2 are different numbers, the specific identifier information corresponding to the two is the same, which is the identifier information of website 1 (for example, both are the IP addresses of website 1). Similarly, although website identifier 3 and website identifier 4 are different numbers, the specific identifier information corresponding to the two is the same, which is the identifier information of website 2 (for example, both are the IP addresses of website 2).

[0082] Suppose the set indicator is the top five websites in terms of total number of visits. Then, server 130 processes the website access data according to the data processing method provided in the embodiment of the present application, and obtains the relevant data of the top five websites in terms of total number of visits. The specific implementation process can refer to the description of the relevant process in the virtual resource transfer scenario, or refer to the description of the subsequent embodiments, which will not be elaborated here.

[0083] In an embodiment of the present application, the target scenario can also be an information recommendation scenario, and the data to be processed can be the record of information and the number of clicks corresponding to the information. For example, the data to be processed is represented as (information identifier, number of visits). Exemplarily, if an information identifier is 111 and the number of clicks is 10 times, then the data to be processed is represented as (111, 10). The set indicator is the top ten information in terms of total number of clicks. Then, server 130 processes the information click data according to the data processing method provided in the embodiment of the present application, and obtains the relevant data of the top ten information in terms of total number of clicks. The specific implementation process can refer to the description of the relevant process in the virtual resource transfer scenario, or refer to the description of the subsequent embodiments, which will not be elaborated here.

[0084] The following makes a detailed description of the data processing provided by the present application in combination with the specific implementation manners.

[0085] Figure 2 Schematically shows the flowchart of the data processing method provided by an embodiment of the present application, as Figure 2 shown, this method includes steps 210 to 240, specifically as follows:

[0086] Step 210: Perform hash operation processing on each data to be processed in the target scenario through at least one hash algorithm, and obtain at least one hash value corresponding to each data to be processed.

[0087] Specifically, the target scenario can be any scenario that requires batch data processing or analysis, such as virtual resource transfer scenarios, network access scenarios, information recommendation scenarios, market data analysis scenarios, etc. The data to be processed in the target scenario is the original record data of the relevant information in this scenario. For example, the data to be processed in the virtual resource transfer scenario is the record data of each resource transfer information, and the data to be processed in the network access scenario is the record data of the access information of each website.

[0088] The hash algorithm can map data that occupies a large space to data that occupies a smaller space. The data processed by the hash algorithm is more compact, so the storage space occupied is reduced. Performing hash operation processing on each data to be processed through at least one hash algorithm means that for any data to be processed, it needs to go through the operations of various hash algorithms. A data to be processed obtains a hash value through one hash algorithm. Then, after being processed by at least one hash algorithm, at least one hash value is obtained.

[0089] In an embodiment of the present application, the data to be processed is in the form of key-value pairs, presented as the form of (key, value). The key is the key of the data to be processed, and the value is the value of the data to be processed. Generally, the corresponding value can be found according to the key. In the embodiment of the present application, when performing hash operation processing on the data to be processed through the hash algorithm, it is to use the hash algorithm to operate on the key in the data to be processed to obtain the corresponding hash value.

[0090] Step 220: According to at least one hash value corresponding to the data to be processed, stack the data to be processed on the data at the position corresponding to the hash algorithm and the hash value in the preset data list to obtain the stacked data corresponding to each hash algorithm.

[0091] Specifically, after obtaining the hash value corresponding to the data to be processed, the data to be processed is stored in the preset data list in a stacked storage manner. Stacked storage means that the data to be processed is stacked with the data already stored at the corresponding position in the preset list to obtain the new storage data at this position.

[0092] The storage location of the data to be processed in the preset data list is jointly determined by the hash algorithm and the hash value. The preset data list can be regarded as a data storage table composed of multiple rows and columns. The hash algorithm and the hash value are respectively used to determine one of the row and column of the storage location of the data to be processed. For example, a row of data in the preset data list is calculated corresponding to the same hash algorithm, then the hash algorithm number can represent the row number of the storage location of the data to be processed; a column of data in the preset data list corresponds to the same hash value, then the hash value can represent the column number of the storage location of the data to be processed. Of course, it is also possible to represent the column number of the storage location of the data to be processed by the hash algorithm number and represent the row number of the storage location of the data to be processed by the hash value.

[0093] In an embodiment of the present application, the data to be processed is in the form of key-value pairs. When performing superimposed storage, the value in the data to be processed is superimposed and stored at the corresponding storage location in the preset data list.

[0094] In an embodiment of the present application, Figure 4 A schematic diagram of a preset data list is schematically shown. Assume that there are m hash algorithms, denoted as f1(x), f2(x)…f m (x). Using the hash algorithm number to represent the row number of the storage location of the data to be processed and the hash value to represent the column number of the storage location of the data to be processed, the storage location of the data to be processed is represented as (row number, column number). Among them, the total number of columns in the preset data list is preset.

[0095] Such as Figure 4 shown, for the data to be processed (key, value), f1(key) = 4, then a storage location of this data to be processed is (1, 4), and the value of the data to be processed is superimposed and stored with the data at the position of the 1st row and the 4th column, that is, the data at the position (1, 4) + value. f2(key) = 2, then a storage location of this data to be processed is (2, 2), and the value of the data to be processed is superimposed and stored with the data at the position of the 2nd row and the 2nd column, that is, the data at the position (2, 2) + value. f m (key) = 5, then a storage location of this data to be processed is (m, 5), and the value of the data to be processed is superimposed and stored with the data at the position of the mth row and the 5th column, that is, the data at the position (m, 5) + value.

[0096] It can be seen that after the data to be processed (key, value) is operated by m hash algorithms, it will be respectively stored in each row corresponding to each hash algorithm in the preset data list. The data corresponding to each hash algorithm in the preset data list includes superimposed data at multiple storage locations (that is, superimposed data in multiple columns).

[0097] Step 230: Sort the superimposed data corresponding to each hash algorithm in the preset data list, obtain the target superimposed data corresponding to each hash algorithm and matching the set metrics according to the sorting result, and generate non-target data corresponding to each hash algorithm based on the other superimposed data except the target superimposed data in the superimposed data corresponding to each hash algorithm.

[0098] Specifically, the sorting is performed on the multiple superimposed data corresponding to each hash algorithm. The multiple superimposed data corresponding to each hash algorithm can be arranged from largest to smallest or from smallest to largest. The set metrics are usually used to limit the selection range of data. For example, the set metrics are the header metric data of a certain type of data. Exemplarily, the set metrics are the sum of the access volumes of the websites ranked top 3 in terms of total access times, which is also expressed as the total access volume of the top 3 websites in terms of access times.

[0099] After sorting, the multiple superimposed data corresponding to each hash algorithm are arranged in size. Then, it is relatively convenient to obtain the target superimposed data matching the set metrics from the multiple superimposed data corresponding to each hash algorithm. For example, if the multiple superimposed data corresponding to each hash algorithm are arranged from largest to smallest, then according to the sorting result, the top 3 superimposed data are obtained to get the target superimposed data.

[0100] In the embodiment of the present application, since the superimposed data at each storage location in the preset data list is the result of superimposed storage of multiple data to be processed, the target superimposed data not only includes the data to be processed matching the set metrics, but also includes some data to be processed not matching the set metrics. And this part of the data to be processed not matching the set metrics is the data to be discarded during the data processing process. For example, the data to be processed is website access data. The target superimposed data corresponding to each hash algorithm includes the sum of the access volumes of the websites ranked top 3 in terms of total access times, and also includes the sum of the access volumes of some websites ranked after the top 3 in terms of total access times. Exemplarily, for the hash algorithm f1(x), for different key1 and key2, it is possible that f1(key1) = f1(key2), then the corresponding value1 and value2 are superimposed and stored in the same location. If key1 is a website ranked top 3 in terms of total access times, the data to be processed matching the set metrics is value1, and the target superimposed data contains the superimposed data at this storage location, that is, the sum of value1 + value2. Therefore, value2 needs to be discarded from the target superimposed data to obtain the data to be processed value1 matching the set metrics.

[0101] For other superimposed data in the superimposed data corresponding to each hashing algorithm except the target superimposed data, these data do not match the set metrics. Then, based on these data that do not match the set metrics, the data to be processed in the target superimposed data that does not match the set metrics can be estimated, and this data is the non-target data.

[0102] Step 240: Calculate the data that matches the set metrics among multiple data to be processed based on the target superimposed data and the non-target data, so as to obtain the data processing result corresponding to the target scenario.

[0103] According to the foregoing analysis, the non-target data is the data that needs to be discarded in the target superimposed data. Then, subtracting the non-target data from the target superimposed data gives the target data, and this target data is the data to be processed that matches the set metrics.

[0104] Since the superimposed data corresponding to each hashing algorithm is obtained by performing hashing operation processing on all the data to be processed corresponding to the target scenario, that is to say, the target data obtained according to a hashing algorithm is actually the data to be processed that matches the set metrics in the target scenario, namely the data processing result corresponding to the target scenario. Therefore, when multiple hashing algorithms are used, one can be selected from the target data obtained from each hashing algorithm as the data processing result corresponding to the target scenario. Optionally, in order to improve the data processing accuracy, statistical processing can also be performed on the target data obtained from each hashing algorithm, and then the result of the statistical processing is used as the data processing result corresponding to the target scenario. For example, the mean value of the target data obtained from each hashing algorithm is used as the data processing result corresponding to the target scenario.

[0105] In the technical solution provided in the embodiments of the present application, each data to be processed in the target scenario is subjected to hashing operation processing through at least one hashing algorithm to obtain at least one hash value corresponding to each data to be processed, and the data to be processed is stored in a superimposed manner according to the hashing algorithm and the hash value. The hashing algorithm enables the data to be processed to be transformed from the original value space data into data that occupies less storage space, thereby greatly reducing the requirement for storage space during data processing; then, the superimposed data corresponding to each hashing algorithm is sorted, and the target superimposed data that matches the set metrics is obtained based on the sorting result, and non-target data is generated based on other superimposed data except the target superimposed data; finally, the data to be processed that matches the set metrics is obtained based on the target superimposed data and the non-target data, and the data processing result is generated, which is equivalent to using a small amount of computing resources to implement the processing and calculation of large-scale data, obtaining the corresponding index data, and saving the resources required for index calculation.

[0106] Figure 3The flowchart of the data processing method provided by an embodiment of the present application is schematically shown. This method is a further refinement of the above embodiment. As Figure 3 shown, the data processing method provided by the embodiment of the present application includes steps 310 to 390, specifically as follows:

[0107] Step 310: Perform a hashing operation on the keys of each piece of data to be processed in the target scenario through at least one hashing algorithm, and obtain at least one hash value corresponding to each piece of data to be processed.

[0108] Specifically, the data to be processed includes data in the form of key-value pairs, that is, the data to be processed consists of a key and a value, expressed as (key, value).

[0109] Exemplarily, the data to be processed in the target scenario includes:

[0110] (key1, value1), (key2, value2)…(key n , value n )

[0111] Among them, the numbers of the keys are different to reflect that two pieces of data to be processed are different data records, and the specific values of the keys can be the same. For example, in the network access scenario, the key represents the website identifier, and the value represents the number of website accesses. Then, it is possible that key1 = key2, which means that (key1, value1) and (key2, value2) are records for the same website. For example, (key1, value1) records the number of accesses to website 1 in the previous hour, and (key2, value2) records the number of accesses to website 1 in the next hour.

[0112] Generally, when processing the data to be processed in the target scenario, the values with the same key in the data to be processed are accumulated to obtain an accumulated value, as shown in the following formula (1):

[0113]

[0114] Where, v i represents value i , s k represents the sum of the values of the data to be processed with key k. The meaning of formula (1) is that when key = k, sum all the corresponding values.

[0115] Then, obtain the sum of the top N accumulated values (briefly recorded as TOPN accumulation) from multiple accumulated values, denoted as:

[0116]

[0117] Among them, TOPN means the top N. TOPN is accumulated to obtain the final data processing result.

[0118] In one embodiment of the present application, a hash algorithm includes a hash function operation and a modulus operation. The hash function operations corresponding to each hash algorithm are different, but the modulus operations are the same. The specific calculation process of the hash value includes: performing hash calculation on the keys of each data to be processed in the target scene through at least one hash function operation to obtain a hash result corresponding to each data to be processed; performing a modulus operation on the hash result corresponding to each data to be processed with a preset hash bucket number to obtain at least one hash value corresponding to each data to be processed.

[0119] Let each hash function be h1(x), h2(x)…h m (x), the modulus operation of the preset hash bucket number is mod b, where b represents the preset hash bucket number. The definition of the preset hash bucket number refers to the relevant description in step 320. The calculation formula of the hash value is shown in the following formula (2):

[0120] δ i,j =h j (key i )mod b (2)

[0121] Among them, h j (key i ) represents the key of the i-th data to be processed through the j-th hash function i The hash result after hash calculation. i,j Indicates the key of the i-th data to be processed i After the jth hash function is calculated, the hash value obtained by taking the modulus of the preset hash bucket number b is δ i,j It is also called a sketch bit. It can be seen that the hash value corresponding to the data to be processed is an integer less than or equal to the preset number of hash buckets.

[0122] Since the hash function operations of various hash algorithms are different, but the modulus operations are the same, different hash functions can also be used to represent different hash operations. For example, the hash function h1(x) can be used to represent the hash operation f1(x).

[0123] Step 320: According to at least one hash value corresponding to the data to be processed, the value of the data to be processed is superimposed on the data at the position corresponding to the hash algorithm and the hash value in the preset data list to obtain superimposed data corresponding to each hash algorithm.

[0124] Specifically, consider the preset data list as a data storage table composed of multiple rows and columns. The number of rows of the preset data list is determined by the number of hash algorithms, and the number of columns of the preset data list is determined by the preset number of hash buckets (of course, it can also be that the hash algorithm determines the number of columns and the preset number of hash buckets determines the number of rows, which will not be elaborated here). The preset number of hash buckets refers to the number of pre-set hash buckets. A hash bucket is equivalent to a storage linked list corresponding to a fixed hash value, and the storage linked list stores the superimposed data obtained by processing through each hash algorithm. Since the storage space occupied by a hash bucket is fixed, the preset number of hash buckets actually reflects the size of the storage space occupied by the preset data list.

[0125] Exemplarily, Figure 4 Schematically shows a schematic diagram of a preset data list. In this preset data list structure, the preset number of hash buckets is b, which is equivalent to that the preset data list has b columns; there are m hash algorithms in total, denoted as f1(x), f2(x)... f m (x), which is equivalent to that the preset data list has m rows.

[0126] In an embodiment of the present application, the calculation formula of the accumulated data is as shown in the following formula (3):

[0127]

[0128] where hs j,l The j-th hash algorithm, and the accumulated data corresponding to the hash value of l, which is equivalent to the superimposed data at the j-th row and l-th column in the preset data list. What formula (3) represents is that when the j-th hash algorithm calculates the hash value of l, sum the values of all the data to be processed corresponding thereto.

[0129] For example, for Figure 4 the data to be processed (key, value) shown, f1(key) = 4, then superimpose the value to the position (1, 4); f2(key) = 2, then superimpose the value to the position (2, 2); f m (key) = 5, then superimpose the value to the position (m, 5).

[0130] This data storage method does not need to store all the data to be processed item by item, but only stores the intermediate results (i.e., the superimposed data), so as to achieve the effect of calculating a large amount of header index data through a small amount of intermediate results (equivalent to limited storage resources).

[0131] Step 330: Sort the superimposed data corresponding to each hash algorithm in the preset data list, and obtain the superimposed data stored in a set number of hash buckets corresponding to each hash algorithm and matching the set index according to the sorting result.

[0132] Specifically, this step mainly extracts the superimposed data in the TOPN hash buckets from the superimposed data corresponding to each hashing algorithm. Exemplarily, as Figure 4 shown in the list of data to be processed, for each cell in each row, it is equivalent to a hash bucket of the corresponding hashing algorithm. The superimposed data in the TOPN hash buckets is to obtain the superimposed data of the first N hash buckets.

[0133] Exemplarily, Figure 5 schematically shows a schematic diagram of the sorting result of multiple superimposed data corresponding to a certain hashing algorithm. As Figure 5 shown, each bar represents the superimposed data of a hash bucket. Assuming TOPN is TOP3, the superimposed data of the 3 hash buckets matching the set metrics is Figure 5 the superimposed data represented by the 3 bars within the dashed box in

[0134] Step 340: Sum the superimposed data stored in a set number of hash buckets as the target superimposed data corresponding to each hashing algorithm.

[0135] Specifically, sum the superimposed data in the extracted TOPN hash buckets to obtain the target superimposed data, which can be expressed as:

[0136]

[0137] where hs l represents the superimposed data in the l-th hash bucket.

[0138] Exemplarily, in the Figure 5 shown sorting result, the target superimposed data corresponding to TOP3 is the sum of the data of the 3 bars within the dashed box.

[0139] In the Figure 5 shown bar chart, a bar is composed of at least one square. In fact, a square represents the sum of all values corresponding to a key. Taking the website access scenario as an example, a square represents the sum of the access times corresponding to a website. The keys corresponding to each square are different. In Figure 5 , the sum of the data corresponding to the 3 squares with the largest area within the dashed box ( Figure 5 shown as the shaded part in

[0140] Step 350: Generate the numerical expectation of non-target data based on other superimposed data except the target superimposed data in the superimposed data corresponding to the hash algorithm and the quantity of the data to be processed corresponding to the other superimposed data.

[0141] Specifically, in the superimposed data corresponding to the hash algorithm, the data except the target superimposed data is other superimposed data, such as Figure 5 the data represented by the bar outside the dashed box in. In this application, the numerical expectation of non-target data is calculated through other superimposed data and the quantity of the data to be processed corresponding to them. The numerical expectation of non-target data refers to the expectation of value in non-target data.

[0142] Taking a hash algorithm as an example, the calculation process of non-target data is described below. The calculation process of the numerical expectation of non-target data includes: summing other superimposed data except the target superimposed data in the superimposed data corresponding to the hash algorithm to obtain the non-target superimposed data corresponding to the hash algorithm; determining the distinct number of the data to be processed corresponding to the other superimposed data according to the quantity of the data to be processed corresponding to the other superimposed data; and obtaining the numerical expectation of non-target data according to the ratio of the non-target superimposed data to the distinct number of the data to be processed corresponding to the other superimposed data.

[0143] Specifically, the sum of other superimposed data is used as the non-target superimposed data, and the non-target superimposed data can be expressed as:

[0144]

[0145] where hs l represents the superimposed data stored in the l-th hash bucket. When the l-th hash bucket does not belong to the TOPN hash bucket, the superimposed data of this hash bucket is other superimposed data. Summing the superimposed data of all hash buckets that do not belong to the TOPN hash bucket, the non-target superimposed data is obtained.

[0146] The distinct number of the data to be processed refers to the quantity of the data to be processed after deduplication according to the key, that is, the number of keys after deduplication of the key. The distinct number of the data to be processed is equivalent to the type of the key. The distinct number of the data to be processed corresponding to the other superimposed data is the number of keys obtained after deduplication of the keys in the total quantity of the data to be processed corresponding to the other superimposed data. For example, if 100 data to be processed correspond to 100 keys and 50 keys are obtained after deduplication of the keys, then the distinct number of the data to be processed is 50.

[0147] In one embodiment of the present application, the calculation method for the deduplication count of the data to be processed corresponding to other superimposed data is as follows: Based on the key of the data to be processed, deduplicate the multiple data to be processed in the target scenario to obtain the total deduplication count of the data to be processed; based on the key of the data to be processed, deduplicate the multiple data to be processed corresponding to the target superimposed data to obtain the deduplication count of the data to be processed corresponding to the target superimposed data; based on the difference between the total deduplication count of the data to be processed and the deduplication count of the data to be processed corresponding to the target superimposed data, obtain the deduplication count of the data to be processed corresponding to other superimposed data.

[0148] That is to say, first obtain the deduplication count of the data to be processed corresponding to the target superimposed data that is opposite to other superimposed data, and then subtract the deduplication count of the data to be processed corresponding to the target superimposed data from the total deduplication count of the data to be processed in the target scenario, and the deduplication count of the data to be processed corresponding to other superimposed data can be obtained. This is because the deduplication count of the data to be processed corresponding to other superimposed data is usually greater than the deduplication count of the data to be processed corresponding to the target superimposed data (because the target superimposed data is TOPN data, while other superimposed data is data outside TOPN). Calculating the deduplication count of the data to be processed corresponding to other superimposed data through the deduplication count of the data to be processed corresponding to the target superimposed data will handle a smaller amount of data than directly deduplicating the data to be processed corresponding to other superimposed data, thereby improving the data processing speed.

[0149] Exemplarily, the calculation formula for the numerical expectation of non-target data is shown in the following formula (4):

[0150]

[0151] Wherein, K represents the total deduplication count of the data to be processed in the target scenario; represents the deduplication count of the data to be processed corresponding to the target superimposed data, that is, the deduplication count of the keys in the TOPN hash bucket; represents the deduplication count of the data to be processed corresponding to other superimposed data, that is, the deduplication count of the keys in the hash bucket outside TOPN.

[0152] Step 360: Calculate the numerical expectation of non-target data according to the amount of data to be processed corresponding to the target superimposed data.

[0153] Specifically, the numerical expectation of non-target data refers to the expectation of the amount of data to be processed corresponding to non-target data. The numerical expectation of non-target data can be obtained by subtracting the set index value from the numerical expectation of the target superimposed data, that is, subtracting N from the numerical expectation of the TOPN hash bucket to obtain the numerical expectation of non-target data, as shown in the following formula (5):

[0154]

[0155] Wherein, It represents the expected number of target superimposed data, and N represents the specific value of the set index.

[0156] In an embodiment of the present application, the expected number of target superimposed data is generated according to the preset number of hash buckets, the total number of duplicate data to be processed corresponding to the target scenario, and the preset fitting function. Exemplarily, K represents the total number of duplicate data to be processed corresponding to the target scenario; b represents the preset number of hash buckets, and the preset fitting function can be expressed as Substitute the total number of duplicate data to be processed and the preset number of hash buckets in the embodiment of the present application into the preset fitting function to obtain the expected number of target superimposed data

[0157] In an embodiment of the present application, the generation process of the preset fitting function includes: constructing a fitting function related to the number of hash buckets and the total number of duplicate data to be processed according to the preset number of hash buckets, the total number of duplicate data to be processed corresponding to the target scenario, and the undetermined fitting coefficients; training the fitting function with sample data to obtain the values of the undetermined fitting coefficients; generating the preset fitting function according to the preset number of hash buckets, the total number of duplicate data to be processed corresponding to the target scenario, and the values of the undetermined fitting coefficients.

[0158] Exemplarily, the preset fitting function is represented as shown in the following formula (6):

[0159]

[0160] Among them, K represents the total number of duplicate data to be processed corresponding to the target scenario; b represents the number of hash buckets; a1, a2, a3 are undetermined fitting coefficients. The undetermined fitting coefficients can be obtained by fitting the fitting function with sample data.

[0161] In an embodiment of the present application, the fitting calculation of the sample data for the fitting function includes: randomly generating the initial values of the undetermined fitting coefficients; calculating the sample data through the fitting function with the initial values of the undetermined fitting coefficients to obtain the predicted expected number of the sample data; adjusting the initial values of the undetermined fitting coefficients according to the difference between the predicted expected number of the sample data and the actual expected number of the sample data until the difference is less than the preset threshold to obtain the values of the undetermined fitting coefficients.

[0162] First, generate the initial values of the to-be-determined fitting coefficients a1, a2, and a3. Then, substitute the sample data into the to-be-determined fitting coefficients with the initial values to calculate the expected predicted quantity corresponding to the sample data. The sample data includes the number of sample hash buckets and the total number of deduplicated sample data corresponding to the target scenario, that is, substitute the initial values of the to-be-determined fitting coefficients, the number of sample hash buckets, and the total number of deduplicated sample data corresponding to the target scenario into the above formula (6) to obtain the expected predicted quantity corresponding to the sample data. In the stage of constructing the preset fitting function, the actual expected quantity corresponding to the sample data is known. Next, adjust the initial values of the to-be-determined fitting coefficients according to the difference between the expected predicted quantity and the actual expected quantity corresponding to the sample data. Then, calculate the sample data according to the fitting function after adjusting the initial values to obtain a new preset expected quantity, calculate the expected difference, and adjust the values of the to-be-determined fitting coefficients according to the difference. Calculate in this way in a loop until the difference is less than the preset threshold to obtain the target values of the to-be-determined fitting coefficients. Based on the target values of the to-be-determined fitting coefficients, the preset fitting function is obtained.

[0163] Substitute the target values of the to-be-determined fitting coefficients, the preset number of hash buckets, and the total number of deduplicated data to be processed corresponding to the target scenario into formula (6) to obtain the expected quantity of the target superimposed data in the embodiments of the present application.

[0164] Exemplarily, Figure 6 Schematically shows a graph of constructing a preset fitting function by linear fitting. It can be seen that the difference between the fitting value and the actual value of the preset fitting function is very small. Therefore, the fitting value of the preset fitting function can be used as the expected quantity of the target superimposed data.

[0165] Step 370: Generate non-target data corresponding to the hash algorithm according to the product of the numerical expectation of the non-target data and the quantity expectation of the non-target data.

[0166] Specifically, multiply the numerical expectation by the quantity expectation to obtain non-target data, as shown in the following formula (7):

[0167]

[0168] The non-target data represents the sum of the data to be processed that does not match the set index in the target superimposed data, that is, in the target superimposed data corresponding to the TOPN hash bucket, the sum of the data to be processed that does not belong to TOPN.

[0169] Step 380: Obtain the data that matches the set index in the multiple data to be processed corresponding to each hash algorithm according to the difference between the target superimposed data and the non-target data corresponding to each hash algorithm.

[0170] Specifically, subtract the non-target data from the target superimposed data to obtain the target data that matches the set index, that is, the sum of the TOPN data to be processed, namely TOPN accumulation.

[0171] Step 390: Statistically process the data that matches the set index among the multiple data to be processed corresponding to each hash algorithm to obtain the data processing result corresponding to the target scenario.

[0172] Specifically, calculate the expectation of the TOPN accumulation for each hash algorithm to obtain the TOPN accumulation in the data to be processed in the target scenario, as shown in the following formula (8):

[0173]

[0174] Among them, H represents the total number of hash functions, which is equivalent to the total number of hash algorithms. E h∈H represents calculating the expectation of the results of all hash algorithms, and E h∈H The content in the parentheses after that is the calculation content of a certain hash algorithm.

[0175] The technical solution provided by the embodiments of the present application stores the data to be processed in a superimposed manner through a preset number of hash buckets. Since the storage space occupied by the hash buckets is certain, this data storage method does not need to store the data to be processed item by item, but only stores the superimposed data obtained from the intermediate calculation. Therefore, the effect of calculating a large amount of head index data with limited storage resources is achieved. At the same time, during the storage process of the superimposed data, the data to be processed is not discarded, which is equivalent to retaining the global information of the data to be processed. Data with relatively scattered data to be processed but large superimposed data can be taken into consideration, making the data used for index calculation more comprehensive. Thus, even if the occupation of storage resources is reduced, the calculation accuracy will not decrease.

[0176] It should be noted that although the steps of the method in the present application are described in a specific order in the drawings, this does not require or imply that these steps must be executed in this specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution, etc.

[0177] The following introduces the device embodiments of the present application, which can be used to execute the data processing method in the above embodiments of the present application. Figure 7 Schematically shows the structural block diagram of the data processing device provided by the embodiments of the present application. As Figure 7 shown, the data processing device provided by the embodiments of the present application includes:

[0178] A hash operation module 710 is configured to perform hash operation processing on each piece of data to be processed in a target scenario through at least one hash algorithm, and obtain at least one hash value corresponding to each piece of data to be processed;

[0179] A data superposition module 720 is configured to superpose the data to be processed on the data at the position corresponding to the hash algorithm and the hash value in a preset data list according to the at least one hash value corresponding to the data to be processed, so as to obtain superposed data corresponding to each hash algorithm;

[0180] A data calculation module 730 is configured to sort the superposed data corresponding to each hash algorithm in the preset data list, obtain target superposed data corresponding to each hash algorithm and matching a set index according to the sorting result, and generate non-target data corresponding to each hash algorithm according to other superposed data except the target superposed data in the superposed data corresponding to each hash algorithm;

[0181] A result generation module 740 is configured to calculate the data in multiple pieces of data to be processed that match the set index according to the target superposed data and the non-target data, so as to obtain a data processing result corresponding to the target scenario.

[0182] In an embodiment of the present application, the data to be processed includes data in the form of key-value pairs; the hash operation module 710 is specifically configured to perform hash operation processing on the keys of each piece of data to be processed in the target scenario through at least one hash algorithm;

[0183] The data superposition module 720 is specifically configured to superpose the values of the data to be processed on the data at the position corresponding to the hash algorithm and the hash value in the preset data list according to the at least one hash value corresponding to the data to be processed.

[0184] In an embodiment of the present application, the hash algorithm includes hash function operation and modulo operation; the hash operation module 710 includes:

[0185] A hash calculation unit is configured to perform hash calculation on the keys of each piece of data to be processed in the target scenario through at least one hash function operation, and obtain at least one hash result of each piece of data to be processed;

[0186] A modulo operation unit is configured to perform modulo operation on the at least one hash result of each piece of data to be processed with respect to a preset number of hash buckets, and use the result of the modulo operation as the at least one hash value corresponding to each piece of data to be processed; the preset number of hash buckets is used to indicate the size of the storage space occupied by the preset data list.

[0187] In one embodiment of the present application, the superimposed data corresponding to the hash algorithm includes the superimposed data stored in multiple hash buckets, and one hash bucket represents a storage location in the preset data list; the data calculation module 730 includes:

[0188] A target superimposed data generation unit, configured to obtain, according to the sorting result, the superimposed data stored in a set number of hash buckets corresponding to each hash algorithm and matching a set index; sum the superimposed data stored in the set number of hash buckets as the target superimposed data corresponding to each hash algorithm.

[0189] In one embodiment of the present application, the data calculation module 730 includes:

[0190] A numerical expectation calculation unit, configured to generate a numerical expectation of non-target data according to the superimposed data other than the target superimposed data in the superimposed data corresponding to the hash algorithm and the amount of data to be processed corresponding to the other superimposed data;

[0191] A quantity expectation calculation unit, configured to calculate a quantity expectation of non-target data according to the amount of data to be processed corresponding to the target superimposed data;

[0192] A non-target data calculation unit, configured to generate non-target data corresponding to the hash algorithm according to the product of the numerical expectation of the non-target data and the quantity expectation of the non-target data.

[0193] In one embodiment of the present application, the numerical expectation calculation unit includes:

[0194] A non-target superimposed data generation subunit, configured to sum the superimposed data other than the target superimposed data in the superimposed data corresponding to the hash algorithm to obtain non-target superimposed data corresponding to the hash algorithm;

[0195] A data deduplication subunit, configured to determine the number of duplicates of the data to be processed corresponding to the other superimposed data according to the amount of data to be processed corresponding to the other superimposed data;

[0196] A numerical expectation calculation subunit, configured to obtain a numerical expectation of non-target data according to the ratio of the non-target superimposed data to the number of duplicates of the data to be processed corresponding to the other superimposed data.

[0197] In one embodiment of the present application, the data deduplication subunit is specifically configured to:

[0198] Deduplicate multiple data to be processed in the target scenario according to the key of the data to be processed to obtain the total number of deduplicated data to be processed;

[0199] Deduplicate a plurality of data to be processed corresponding to the target superimposed data according to the key of the data to be processed, and obtain the number of deduplicated data to be processed corresponding to the target superimposed data;

[0200] Obtain the number of deduplicated data to be processed corresponding to the other superimposed data according to the difference between the total number of deduplicated data to be processed and the number of deduplicated data to be processed corresponding to the target superimposed data.

[0201] In an embodiment of the present application, the quantity expectation calculation unit includes:

[0202] A first calculation subunit, configured to generate the quantity expectation of the target superimposed data according to a preset number of hash buckets, the total number of deduplicated data to be processed corresponding to the target scenario, and a preset fitting function;

[0203] A second calculation subunit, configured to obtain the quantity expectation of the non-target data according to the difference between the quantity expectation of the target superimposed data and the value of the set index.

[0204] In an embodiment of the present application, the quantity expectation calculation unit further includes:

[0205] A fitting function construction unit, configured to construct a fitting function related to the number of hash buckets and the total number of deduplicated data to be processed according to undetermined fitting coefficients;

[0206] A training unit, configured to train the fitting function through sample data to obtain the target value of the undetermined fitting coefficient; the sample data includes the sample number of hash buckets and the total number of deduplicated sample data corresponding to the target scenario;

[0207] A preset fitting function generation unit, configured to generate the preset fitting function according to the target value of the undetermined fitting coefficient.

[0208] In an embodiment of the present application, the training unit is specifically configured to:

[0209] Randomly generate an initial value of the undetermined fitting coefficient;

[0210] Calculate the sample data through the fitting function with the initial value of the undetermined fitting coefficient to obtain the predicted quantity expectation of the sample data;

[0211] Adjust the initial value of the undetermined fitting coefficient according to the difference between the predicted quantity expectation of the sample data and the actual quantity expectation of the sample data until the difference is less than a preset threshold, and obtain the target value of the undetermined fitting coefficient.

[0212] In an embodiment of the present application, the result generation module 740 includes:

[0213] A data processing unit, configured to obtain data that matches the set metric among the multiple pieces of data to be processed corresponding to the respective hash algorithms according to the difference between the target superimposed data and the non-target data corresponding to the respective hash algorithms;

[0214] A statistical unit, configured to perform statistical processing on the data that matches the set metric among the multiple pieces of data to be processed corresponding to the respective hash algorithms, to obtain a data processing result corresponding to the target scenario.

[0215] In an embodiment of the present application, the statistical unit is specifically configured to:

[0216] Calculate the expected value of the data that matches the set metric among the multiple pieces of data to be processed corresponding to the respective hash algorithms, as the data processing result corresponding to the target scenario.

[0217] The specific details of the data processing device provided in the embodiments of the present application have been described in detail in the corresponding method embodiments, and will not be elaborated here.

[0218] Figure 8 Schematically shows a block diagram of a computer system of an electronic device for implementing the embodiments of the present application.

[0219] It should be noted that Figure 8 The computer system 800 of the electronic device shown is only an example, and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0220] As Figure 8 shown, the computer system 800 includes a central processing unit 801 (Central Processing Unit, CPU), which can perform various appropriate actions and processes according to the program stored in the read-only memory 802 (Read-Only Memory, ROM) or the program loaded from the storage section 808 into the random access memory 803 (Random Access Memory, RAM). In the random access memory 803, various programs and data required for system operation are also stored. The central processing unit 801, the read-only memory 802, and the random access memory 803 are connected to each other through a bus 804. The input / output interface 805 (Input / Output interface, i.e., I / O interface) is also connected to the bus 804.

[0221] The following components are connected to the input / output interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a local area network card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 810 as needed so that a computer program read therefrom can be installed into the storage section 808 as needed.

[0222] Specifically, according to an embodiment of the present application, the processes described in each of the method flowcharts can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by the central processing unit 801, various functions defined in the system of the present application are executed.

[0223] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0224] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and the combination of blocks in a block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0225] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0226] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0227] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application.

[0228] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A data processing method, characterized in that, Including: Performing hash operation processing on each piece of data to be processed in a target scenario through at least one hash algorithm to obtain at least one hash value corresponding to each piece of data to be processed; According to the at least one hash value corresponding to the data to be processed, superimposing the data to be processed on the data at the position corresponding to the hash algorithm and the hash value in a preset data list to obtain superimposed data corresponding to each hash algorithm; Sorting the superimposed data corresponding to each hash algorithm in the preset data list, obtaining target superimposed data corresponding to each hash algorithm and matching a set index according to the sorting result, and generating non-target data corresponding to each hash algorithm according to the other superimposed data except the target superimposed data in the superimposed data corresponding to each hash algorithm; Calculating the data in the multiple pieces of data to be processed that match the set index according to the target superimposed data and the non-target data to obtain the data processing result corresponding to the target scenario.

2. The data processing method according to claim 1, wherein The data to be processed includes data in the form of key-value pairs; The performing hash operation processing on each piece of data to be processed in a target scenario through at least one hash algorithm includes: performing hash operation processing on the keys of each piece of data to be processed in a target scenario through at least one hash algorithm; The superimposing the data to be processed on the data at the position corresponding to the hash algorithm and the hash value in a preset data list according to the at least one hash value corresponding to the data to be processed includes: superimposing the value of the data to be processed on the data at the position corresponding to the hash algorithm and the hash value in a preset data list according to the at least one hash value corresponding to the data to be processed.

3. The data processing method according to claim 2, characterized in that The hash algorithm includes hash function operation and modulo operation; the performing hash operation processing on the keys of each piece of data to be processed in a target scenario through at least one hash algorithm to obtain at least one hash value corresponding to each piece of data to be processed includes: Performing hash calculation on the keys of each piece of data to be processed in the target scenario through at least one hash function operation to obtain at least one hash result of each piece of data to be processed; Performing modulo operation on the at least one hash result of each piece of data to be processed with respect to a preset number of hash buckets, and using the result of the modulo operation as the at least one hash value corresponding to each piece of data to be processed; the preset number of hash buckets is used to indicate the size of the storage space occupied by the preset data list.

4. The data processing method according to claim 1, wherein The superimposed data corresponding to the hash algorithm includes superimposed data stored in multiple hash buckets, and one hash bucket represents a storage position in the preset data list; Obtaining target superimposed data corresponding to each hash algorithm and matching a set index according to the sorting result includes: Obtaining the superimposed data stored in a set number of hash buckets corresponding to each hash algorithm and matching the set index according to the sorting result; Summing the superimposed data stored in the set number of hash buckets as the target superimposed data corresponding to each hash algorithm.

5. The data processing method according to claim 1, wherein Generating non-target data corresponding to each of the hash algorithms based on other superimposed data in the superimposed data corresponding to each of the hash algorithms except the target superimposed data, includes: Generating a numerical expectation of the non-target data based on other superimposed data in the superimposed data corresponding to the hash algorithm except the target superimposed data, and the amount of data to be processed corresponding to the other superimposed data; Calculating an expected number of the non-target data based on the amount of data to be processed corresponding to the target superimposed data; Generating the non-target data corresponding to the hash algorithm based on the product of the numerical expectation of the non-target data and the expected number of the non-target data; 6. The data processing method according to claim 5, wherein Generating a numerical expectation of the non-target data based on other superimposed data in the superimposed data corresponding to the hash algorithm except the target superimposed data, and the amount of data to be processed corresponding to the other superimposed data, includes: Summing other superimposed data in the superimposed data corresponding to the hash algorithm except the target superimposed data to obtain non-target superimposed data corresponding to the hash algorithm; Determining a deduplication number of the data to be processed corresponding to the other superimposed data based on the amount of data to be processed corresponding to the other superimposed data; Obtaining a numerical expectation of the non-target data based on a ratio of the non-target superimposed data and the deduplication number of the data to be processed corresponding to the other superimposed data; 7. The data processing method according to claim 6, wherein The data to be processed includes data in the form of key-value pairs; determining a deduplication number of the data to be processed corresponding to the other superimposed data based on the amount of data to be processed corresponding to the other superimposed data, includes: Deduplicating multiple data to be processed in the target scenario according to the key of the data to be processed to obtain a total deduplication number of the data to be processed; Deduplicating multiple data to be processed corresponding to the target superimposed data according to the key of the data to be processed to obtain a deduplication number of the data to be processed corresponding to the target superimposed data; Obtaining a deduplication number of the data to be processed corresponding to the other superimposed data based on a difference between the total deduplication number of the data to be processed and the deduplication number of the data to be processed corresponding to the target superimposed data; 8. The data processing method according to claim 5, characterized in that Calculating an expected number of the non-target data based on the amount of data to be processed corresponding to the target superimposed data, includes: Generating an expected number of the target superimposed data based on a preset number of hash buckets, the total deduplication number of the data to be processed corresponding to the target scenario, and a preset fitting function; Obtaining an expected number of the non-target data based on a difference between the expected number of the target superimposed data and a numerical value of the set index; 9. The data processing method according to claim 8, wherein Before generating an expected number of the target superimposed data based on a preset number of hash buckets, the total deduplication number of the data to be processed corresponding to the target scenario, and a preset fitting function, the method further includes: Constructing a fitting function related to the number of hash buckets and the total deduplication number of the data to be processed according to undetermined fitting coefficients; Training the fitting function with sample data to obtain a target value of the undetermined fitting coefficients; the sample data includes a sample number of hash buckets and a total deduplication number of sample data corresponding to the target scenario; Generating the preset fitting function based on the target value of the undetermined fitting coefficients.

10. The data processing method according to claim 9, wherein Training the fitting function with sample data to obtain the target values of the fitting coefficients to be determined, including: Randomly generating initial values of the fitting coefficients to be determined; Calculating the predicted quantity expectations of the sample data by using the fitting function with the initial values of the fitting coefficients to be determined; Adjusting the initial values of the fitting coefficients to be determined according to the difference between the predicted quantity expectations of the sample data and the actual quantity expectations of the sample data until the difference is less than a preset threshold to obtain the target values of the fitting coefficients to be determined.

11. The data processing method according to any one of claims 1-10, characterized in that, Calculating the data in the multiple data to be processed that matches the set index according to the target superimposed data and the non-target data to obtain the data processing result corresponding to the target scenario, including: Obtaining the data in the multiple data to be processed that matches the set index corresponding to each hash algorithm according to the difference between the target superimposed data and the non-target data corresponding to each hash algorithm; Performing statistical processing on the data in the multiple data to be processed that matches the set index corresponding to each hash algorithm to obtain the data processing result corresponding to the target scenario.

12. The data processing method according to claim 11, wherein Performing statistical processing on the data in the multiple data to be processed that matches the set index corresponding to each hash algorithm to obtain the data processing result corresponding to the target scenario, including: Calculating the expected values of the data in the multiple data to be processed that matches the set index corresponding to each hash algorithm as the data processing result corresponding to the target scenario.

13. A data processing device, characterized in that, Including: A hash operation module, configured to perform hash operation processing on each data to be processed in a target scenario through at least one hash algorithm to obtain at least one hash value corresponding to each data to be processed; A data superimposing module, configured to superimpose the data to be processed on the data at the position corresponding to the hash algorithm and the hash value in a preset data list according to the at least one hash value corresponding to the data to be processed to obtain the superimposed data corresponding to each hash algorithm; A data calculation module, configured to sort the superimposed data corresponding to each hash algorithm in the preset data list, obtain the target superimposed data corresponding to each hash algorithm and matching the set index according to the sorting result, and generate the non-target data corresponding to each hash algorithm according to the other superimposed data except the target superimposed data in the superimposed data corresponding to each hash algorithm; A result generation module, configured to calculate the data in the multiple data to be processed that matches the set index according to the target superimposed data and the non-target data to obtain the data processing result corresponding to the target scenario.

14. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the data processing method according to any one of claims 1 to 12.

15. An electronic device, characterized in that, Including: A processor; And A memory, configured to store the executable instructions of the processor; Wherein, when the processor executes the executable instructions, the electronic device executes the data processing method according to any one of claims 1 to 12.

16. A computer program product, characterized in that, The computer program product includes a computer program which is stored in a computer-readable storage medium. A processor of an electronic device reads and executes the computer program, so that the electronic device executes the data processing method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Data processing method and device, medium and computing equipment

    CN114064840A

  • Hash encoding method, apparatus and device, and readable storage medium

    WO2021232752A1