A method and apparatus for data processing, a computer device, and a storage medium
By associating the to-processed data with the initial bit array, replacing the position parameters of the initial bit array with modulus values, we can determine whether the to-processed data is duplicate data, and solve the problem of repeated message sending in real-time data statistics by Flink and other streaming frameworks, and realize efficient data processing and accurate statistical results.
Patent Information
- Application Number
- CN202111240901.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-25
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-10-25
AI Technical Summary
When performing real-time data statistics, Flink and other streaming processing frameworks will have the problem of repeated messages being sent, resulting in deviations in statistical results. The existing methods can determine whether it is duplicated by recording each piece of data, resulting in the need to access external storage for each piece of data, increasing storage resource consumption and reducing data processing efficiency.
By associating the pending data with the initial bit array, replacing the position parameters of the initial bit array with the modulus value, we can judge whether the pending data is duplicate data, and avoid accessing each piece of data from external storage.
It effectively reduces the consumption of storage resources in large-scale data processing scenarios, improves the efficiency of real-time data processing, and ensures the accuracy of statistical results.
Smart Images

Figure CN113934767B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data statistical analysis, and particularly to a method and device for data processing, a computer device, and a storage medium. Background Art
[0002] In the digital age, the application scope and boundaries of the Internet are constantly expanding. Many Internet companies and traditional enterprises are gradually accelerating the update and iteration of their own application systems, including computer terminals and mobile terminals, leveraging the technological power of the Internet to empower their services and products. For the upgrade and optimization of application systems, operational behavior statistical analysis is the most important technical support. Through operational behavior statistical analysis, the wishes, characteristics, and demands of users can be fully reflected, and big data technology can be used to assist enterprises in product design, interface optimization, and precision marketing, improving the user experience. Generally, streaming processing frameworks such as Flink are used for large-scale real-time data statistics. However, due to the problem of duplicate message sending when Flink processes data, there are significant deviations in the final statistical results. Therefore, it is necessary to deduplicate data during the statistical analysis process to eliminate the duplicate data generated by unreliable data sources and make the final statistical results more accurate. Currently, it is possible to determine whether data is duplicate by recording each piece of data. However, using the above method will cause each piece of data to access external storage, so in scenarios with a large amount of data, it will increase the consumption of storage resources and thus reduce the real-time data processing efficiency. Therefore, how to process data more efficiently has become an urgent problem to be solved. Summary of the Invention
[0003] The embodiments of this application provide a method and device for data processing, a computer device, and a storage medium. Since the initial digit is associated with the data to be processed one by one, that is, a data to be processed is mapped into an initial digit, the modulo value corresponding to each data to be processed in each initial digit array is used to replace the position parameter of the initial digit associated with the data to be processed, so as to obtain the position parameter of the target digit associated with the data to be processed. When the position parameter of the target digit associated with the data to be processed is different from the initial value, it is determined that the data to be processed is duplicate data, and it is not necessary to access external storage for each piece of data. In scenarios with a large amount of data, it can reduce the consumption of storage resources and thus improve the real-time data processing efficiency.
[0004] In view of this, the first aspect of this application provides a method for data processing, including:
[0005] Obtain a set of data to be processed, where the set of data to be processed includes L pieces of data to be processed, and L is an integer greater than 1;
[0006] Create an initial digit array set based on the data set to be processed, where the initial digit array set includes N initial digit arrays, each initial digit array includes L initial digits, the initial digits are associated with the data to be processed one by one, and the position parameter of each initial digit is an initial value, and N is an integer greater than 1;
[0007] Process each data to be processed in the data to be processed based on the initial digit array set, and obtain the modulo values corresponding to each data to be processed in each initial digit array, where the modulo values are in one-to-one correspondence with one of the L initial digits included in the initial digit array;
[0008] Replace the position parameter of each initial digit in the N initial digit arrays with the modulo value corresponding to each data to be processed in each initial digit array, and obtain a target digit array set, where the target digit array set includes N target digit arrays, each target digit array includes L target digits, the target digits are associated with the data to be processed one by one, and the position parameter of each target digit is the modulo value corresponding to each data to be processed in each initial digit array;
[0009] Determine the target data from the L data to be processed based on the target digit array set, where the position parameters of the N target digits associated with the target data are different from the initial values, and the target data is duplicate data.
[0010] The second aspect of this application provides a data processing device, including:
[0011] An acquisition module for acquiring a data set to be processed, where the data set to be processed includes L data to be processed, and L is an integer greater than 1;
[0012] A creation module for creating an initial digit array set based on the data set to be processed, where the initial digit array set includes N initial digit arrays, each initial digit array includes L initial digits, the initial digits are associated with the data to be processed one by one, and the position parameter of each initial digit is an initial value, and N is an integer greater than 1;
[0013] A processing module for processing each data to be processed in the data to be processed based on the initial digit array set, and obtaining the modulo values corresponding to each data to be processed in each initial digit array, where the modulo values are in one-to-one correspondence with one of the L initial digits included in the initial digit array;
[0014] A replacement module for replacing the position parameter of each initial digit in N initial digit arrays with the modulo value corresponding to each data to be processed in each initial digit array, to obtain a set of target digit arrays, where the set of target digit arrays includes N target digit arrays, each target digit array includes L target digits, the target digits are associated with the data to be processed one by one, and the position parameter of each target digit is the modulo value corresponding to each data to be processed in each initial digit array;
[0015] A determination module for determining target data from L data to be processed based on the set of target digit arrays, where the position parameters of the N target digits associated with the target data are different from the initial values, and the target data is duplicate data.
[0016] In a possible implementation manner, the processing module is specifically configured to create a hash function corresponding to the set of initial digit arrays;
[0017] Create a modulo function corresponding to each initial digit array;
[0018] Based on the hash function corresponding to the set of initial digit arrays and the modulo function corresponding to each initial digit array, process each data to be processed in the data to be processed, to obtain the modulo value corresponding to each data to be processed in each initial digit array.
[0019] In a possible implementation manner, the processing module is specifically configured to perform parsing processing on L data to be processed in the set of data to be processed, to obtain the key value corresponding to each data to be processed;
[0020] Use the hash function corresponding to the set of initial digit arrays to perform hash calculation on the key value corresponding to each data to be processed, to obtain the hash value of the key value corresponding to each data to be processed;
[0021] Use the modulo function corresponding to each initial digit array to perform modulo processing on the hash value of the key value corresponding to each data to be processed, to obtain the modulo value corresponding to each data to be processed in each initial digit array.
[0022] In a possible implementation manner, the acquisition module is further configured to perform parsing processing on L data to be processed in the set of data to be processed, to obtain the non-key value corresponding to each data to be processed;
[0023] The processing module is further configured to discard the target data after the determination module determines the target data from L data to be processed based on the set of target digit arrays.
[0024] In a possible implementation manner, the data processing device further includes a reading and writing module;
[0025] The determination module is further configured to determine cached data from the L to-be-processed data based on the set of target bit arrays, where the position parameter of each target digit in the N target digits associated with the cached data is an initial value, and the cached data is non-duplicate data;
[0026] The read-write module is configured to write the key value corresponding to the cached data and the non-key value corresponding to the cached data into the cache.
[0027] In a possible implementation manner, the data processing device further includes a statistics module;
[0028] The determination module is further configured to determine the key value corresponding to the cached data as a non-duplicate key value;
[0029] The processing module is further configured to perform a counting and statistical process on the cached data existing in the L to-be-processed data to obtain the number of non-duplicate key values, where the number of non-duplicate key values is the number of all cached data in the L to-be-processed data;
[0030] The statistics module is further configured to count the modulo values corresponding to all cached data in the L to-be-processed data in each initial bit array;
[0031] The determination module is further configured to determine the summary data of the non-duplicate key values based on the number of non-duplicate key values and the modulo values corresponding to all cached data in the L to-be-processed data in each initial bit array, where the summary data is used to perform data analysis processing on the to-be-processed data set.
[0032] In a possible implementation manner, the acquisition module is specifically configured to acquire an initial data set, where the initial data set includes M initial data, each initial data corresponds to a theme type one by one, and the M initial data are from at least two data sources, and M is an integer greater than L;
[0033] Perform parsing processing on each initial data in the initial data set to obtain the theme type corresponding to each initial data;
[0034] Extract the initial data with the theme type being the target theme type to generate a to-be-processed data set.
[0035] The third aspect of the present application provides a computer-readable storage medium, in which instructions are stored, and when they run on a computer, the computer is made to execute the methods described in the above aspects.
[0036] A fourth aspect of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of the methods provided in the above aspects.
[0037] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages:
[0038] In an embodiment of the present application, a data processing method is provided. First, a set of data to be processed including L data to be processed is obtained, where L is an integer greater than 1. Then, an initial bit array set is created based on the set of data to be processed. The initial bit array set includes N initial bit arrays, and each initial bit array includes L initial bits. The initial bits are associated with the data to be processed one by one, and the position parameter of each initial bit is an initial value. N is an integer greater than 1. Based on this, each data to be processed in the set of data to be processed is processed based on the initial bit array set to obtain the modulo values corresponding to each data to be processed in each initial bit array. The modulo values correspond to one of the L initial bits included in the initial bit array one by one. Then, the position parameter of each initial bit in the N initial bit arrays is replaced with the modulo value corresponding to each data to be processed in each initial bit array to obtain a target bit array set. The target bit array set includes N target bit arrays, and each target bit array includes L target bits. The target bits are associated with the data to be processed one by one, and the position parameter of each target bit is the modulo value corresponding to each data to be processed in each initial bit array. Finally, the target data is determined from the L data to be processed based on the target bit array set. Among them, the position parameters of the N target bits associated with the target data are different from the initial values, and the target data is duplicate data. In the above manner, since the initial bits are associated with the data to be processed one by one, that is, one data to be processed is mapped to one initial bit, the modulo value corresponding to each data to be processed in each initial bit array is used to replace the position parameter of the initial bit associated with the data to be processed to obtain the position parameter of the target bit associated with the data to be processed. When the position parameter of the target bit associated with the data to be processed is different from the initial value, it is determined that the data to be processed is duplicate data. It is not necessary to access external storage for each piece of data, and the consumption of storage resources can be reduced in a scenario with a large amount of data, thereby improving the real-time data processing efficiency. Description of the Drawings
[0039] Figure 1 It is a schematic diagram of the system architecture of a data processing method provided by an embodiment of the present application;
[0040] Figure 2Schematic flowchart of a data processing method provided by an embodiment of the present application;
[0041] Figure 3 Schematic diagram of an embodiment of a data processing method provided by an embodiment of the present application;
[0042] Figure 4 Schematic diagram of an embodiment of an initial digit array set provided by an embodiment of the present application;
[0043] Figure 5 Schematic diagram of an embodiment of a target digit array set provided by an embodiment of the present application;
[0044] Figure 6 Schematic diagram of an embodiment of determining target data provided by an embodiment of the present application;
[0045] Figure 7 Schematic diagram of an embodiment of determining cached data provided by an embodiment of the present application;
[0046] Figure 8 Schematic structural diagram of a data processing device provided by an embodiment of the present application;
[0047] Figure 9 Schematic diagram of an embodiment of a server in an embodiment of the present application. Detailed implementation manners
[0048] An embodiment of the present application provides a data processing method, device, computer device, and storage medium. Since the initial digits are associated with the data to be processed one by one, that is, a data to be processed is mapped into an initial digit, the modulo value corresponding to each data to be processed in each initial digit array is used to replace the position parameter of the initial digit associated with the data to be processed, so as to obtain the position parameter of the target digit associated with the data to be processed. When the position parameter of the target digit associated with the data to be processed is different from the initial value, it is determined that the data to be processed is duplicate data, and it is not necessary to access the external storage for each piece of data, which can reduce the consumption of storage resources in the scenario of a large amount of data, thereby improving the real-time data processing efficiency.
[0049] In the specification, claims and the above-mentioned drawings of this application, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0050] In the digital age, the application scope and boundaries of the Internet are constantly expanding. Many Internet companies and traditional enterprises are gradually accelerating the update and iteration of their own application systems, including computer terminals and mobile terminals, leveraging the technological power of the Internet to empower their services and products. For the upgrade and optimization of application systems, operational behavior statistical analysis is the most important technical support. Through operational behavior statistical analysis, the wishes, characteristics, and demands of users can be fully reflected, and big data technology can be used to assist enterprises in product design, interface optimization, and precision marketing, improving the user experience. Generally, a streaming processing framework such as Flink is used for large-scale real-time data statistics. However, since there will be a problem of duplicate message sending when Flink processes data, the final statistical results will have a large deviation. Therefore, it is necessary to deduplicate the data during the statistical analysis process to eliminate the duplicate data generated by unreliable data sources and make the final statistical results more accurate. Currently, it is possible to determine whether data is duplicate by recording each piece of data. However, using the above method will result in each piece of data needing to access external storage, so in the scenario of a large amount of data, it will increase the consumption of storage resources and thus reduce the real-time data processing efficiency. Therefore, how to process data more efficiently has become an urgent problem to be solved. To solve the foregoing problems, the embodiments of this application provide a data processing method that can determine the data to be processed as duplicate data when the position parameter of the target digit associated with the data to be processed is different from the initial value, without each piece of data accessing external storage, and can reduce the consumption of storage resources in the scenario of a large amount of data, thereby improving the real-time data processing efficiency.
[0051] First, for the convenience of understanding, some terms or concepts involved in this application will be explained first.
[0052] I. Streaming Computing
[0053] Streaming computing refers to the input and processing of data in the form of data streams, and then output. In contrast to batch processing, batch processing systems often require a batch of data to arrive before starting to process, while streaming computing usually processes each piece of data or each small batch of data immediately when it arrives.
[0054] II. Flink
[0055] Flink is a unified framework for stream processing and batch processing. Since pipeline data is transmitted between parallel tasks, Flink supports stream processing and batch processing at runtime.
[0056] III. Bloom Filter
[0057] A Bloom filter is a probabilistic data structure composed of hash functions and arrays, characterized by efficient insertion and query. The Bloom filter can obtain the result of "something must not exist or may exist".
[0058] IV. Hbase
[0059] Hbase is a distributed, column-oriented open-source database that can be used to quickly write and query data based on keys.
[0060] V. Redis
[0061] Redis is an open-source database written in C language, supporting networking, and can be memory-based or persistent log-based.
[0062] VI. Duplicate Consumption
[0063] During the process of log generation and real-time computing, due to dirty data and service restart, the situation where the same piece of data is processed multiple times is called duplicate consumption. In this application, the data that will be processed multiple times (i.e., duplicate consumption) is defined as duplicate data.
[0064] Based on this, the application scenarios of the embodiments of this application will be introduced below. It can be understood that the data processing method is executed by the server. Here, taking the server as the execution entity, the data processing method provided by the embodiments of this application will be introduced. Please refer to Figure 1 , Figure 1 which is the system architecture diagram of a data processing method provided by the embodiments of this application. As Figure 1As shown in the figure, the video processing system includes a server and terminal devices. The server obtains multiple pieces of data to be processed from multiple terminal devices, and processes the multiple pieces of data to be processed through the data processing method provided by the embodiments of the present application. Thus, when the position parameter of the target digit associated with the data to be processed is different from the initial value, it is determined that the data to be processed is duplicate data, and it is not necessary to access the external storage for each piece of data. In the scenario of a large amount of data, the consumption of storage resources can be reduced, thereby improving the real-time data processing efficiency.
[0065] It should be noted that the server involved in the present application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (Content Delivery Network, CDN), as well as big data and artificial intelligence platforms. Terminal devices include, but are not limited to, mobile phones, computers, intelligent voice interaction devices, intelligent home appliances, vehicle-mounted terminals, etc. And the terminal devices and the server can communicate through wireless networks, wired networks or removable storage media. Among them, the above-mentioned wireless networks use standard communication technologies and / or protocols. The wireless network is usually the Internet, but it can also be any network, including but not limited to any combination of Bluetooth, local area network (Local Area Network, LAN), metropolitan area network (Metropolitan Area Network, MAN), wide area network (Wide Area Network, WAN), mobile, private network or virtual private network). In some embodiments, customized or dedicated data communication technologies can be used to replace or supplement the above data communication technologies. The removable storage medium can be a universal serial bus (Universal Serial Bus, USB) flash drive, a mobile hard disk or other removable storage media, etc.
[0066] Although Figure 1 only five terminal devices and one server are shown in the figure, it should be understood that Figure 1 the examples in the figure are only used to understand the present solution, and the specific numbers of terminal devices and servers should be flexibly determined according to the actual situation.
[0067] Secondly, the embodiments of the present invention can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc. Among them, cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and network within a wide area network or a local area network to achieve data computing, storage, processing, and sharing. Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model, which can form a resource pool, be used on demand, and be flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the background system for logical processing. Data at different levels will be processed separately, and various industry data requires the support of a powerful system background, which can only be achieved through cloud computing.
[0068] Cloud computing refers to the delivery and usage model of IT infrastructure, which means obtaining the required resources in a on-demand and easily scalable manner through the network; in a broad sense, cloud computing refers to the delivery and usage model of services, which means obtaining the required services in a on-demand and easily scalable manner through the network. Such services can be related to IT and software, the Internet, or other services. Cloud computing is the product of the development and integration of traditional computer and network technologies such as grid computing, distributed computing, parallel computing, utility computing, network storage technologies, virtualization, and load balance.
[0069] With the development of the Internet, real-time data streams, and the diversification of connected devices, as well as the promotion of demands such as search services, social networks, mobile commerce, and open collaboration, cloud computing has developed rapidly. Different from the previous parallel distributed computing, the emergence of cloud computing will, in concept, drive a revolutionary change in the entire Internet model and enterprise management model.
[0070] The so-called artificial intelligence cloud service is generally also called AIaaS (AI as a Service, which is "AI as a service" in Chinese). This is a currently mainstream service mode of artificial intelligence platforms. Specifically, the AIaaS platform will split several common AI services and provide independent or packaged services in the cloud. This service mode is similar to opening an AI-themed mall: all developers can access and use one or more artificial intelligence services provided by the platform through the API interface. Some senior developers can also use the AI frameworks and AI infrastructure provided by the platform to deploy and operate their own exclusive cloud artificial intelligence services.
[0071] Figure 2 The flowchart of a data processing method provided by an embodiment of this application is shown as Figure 2 As shown, the specific data processing process in the embodiment of this application includes parsing the initial data set to obtain the data set to be processed, creating N initial bit arrays based on the data set to be processed, creating hash functions corresponding to the N initial bit arrays and modulo functions corresponding to each initial bit array, obtaining the target bit array set, and determining the target data based on the target bit array set. The functions and processes of each part will be introduced below. Specifically:
[0072] In step A1, the initial data set is parsed to obtain the data set to be processed. Specifically, the operation behaviors and item data from different terminal devices are collected through software development kit (SDK) logging and reported to the server. The aforementioned operation behaviors and item data are the initial data. The server side parses and processes the initial data of different terminal devices and sets at least once mode to ensure that the initial data will not be lost, but there may be duplicate data. Based on this, the server parses the initial data from different terminal devices to obtain the theme type corresponding to each initial data. The initial data of different theme types will be split for calculation. In the embodiment of this application, the data of the target theme type is processed. Then the server can extract the initial data with the theme type of the target theme type and generate the data set to be processed, and the data set to be processed includes L data to be processed, where L is an integer greater than 1.
[0073] In step A2, the server creates N initial bit arrays based on the set of data to be processed. Specifically, since the set of data to be processed includes L pieces of data to be processed, in order to determine whether each piece of data to be processed is duplicate data, the server creates N initial bit arrays with a length of L, that is, each initial bit array includes L initial bits. An initial bit is associated with a piece of data to be processed, and the position parameter of each initial bit is initialized to an initial value, which is 0 in the embodiments of the present application.
[0074] In step A3, after the creation of N initial bit arrays in step A2 is completed, a hash function common to the N initial bit arrays is created, and a modulo function corresponding to each initial bit array is created. The specific hash function and modulo function will be introduced in detail in the subsequent embodiments.
[0075] In step A4, each piece of data to be processed in the set of data to be processed is processed through the hash function common to the N initial bit arrays created in step A3 and the modulo function corresponding to each initial bit array, to obtain the modulo value corresponding to each piece of data to be processed in each initial bit array. The modulo value corresponds one-to-one with one of the L initial bits included in the initial bit array. Then, the position parameter of each initial bit in the N initial bit arrays is replaced with the modulo value corresponding to each piece of data to be processed in each initial bit array, to obtain a set of target bit arrays. The set of target bit arrays includes N target bit arrays, each target bit array includes L target bits, and a target bit is associated with a piece of data to be processed. The position parameter of each target bit is the modulo value corresponding to each piece of data to be processed in each initial bit array, and the modulo value is actually the remainder obtained by taking the remainder of each piece of data to be processed in different initial bit arrays.
[0076] In step A5, target data is determined from the foregoing L pieces of data to be processed based on the set of target bit arrays obtained in step A4. The position parameters of the N target bits associated with the target data are different from the initial value, and the target data is the duplicate data described in the embodiments of the present application.
[0077] Combined with the above introduction, taking the server as the execution entity as an example, the method for data processing in the present application will be introduced below. Please refer to Figure 3 , Figure 3 is a schematic diagram of an embodiment of a method for data processing provided by an embodiment of the present application. As shown in Figure 3 The method includes:
[0078] 101. Obtain a set of data to be processed.
[0079] In this embodiment, the server obtains a set of data to be processed, where the set of data to be processed includes L pieces of data to be processed, and L is an integer greater than 1. Specifically, multiple pieces of data to be processed in the aforementioned set of data to be processed are at least from two different terminal devices. For example, the set of data to be processed includes 3 pieces of data to be processed, namely data to be processed A, data to be processed B, and data to be processed C, and data to be processed A and data to be processed B are from terminal device A, while data to be processed C is from terminal device B. The foregoing example is only for understanding the source of the data to be processed in this solution and should not be construed as a limitation of this solution.
[0080] 102. Create an initial bit array set based on the set of data to be processed.
[0081] In this embodiment, the server creates an initial bit array set based on the set of data to be processed. Based on step 101, it is known that the set of data to be processed includes L pieces of data to be processed. The purpose of this application is to determine whether the data to be processed is duplicate data. Therefore, the server will create N initial bit arrays with a length of L, that is, each initial bit array includes L initial bits, and one initial bit is associated with one piece of data to be processed. The aforementioned N is an integer greater than 1. For example, in the initial bit array A, there are initial bits 1, initial bit 2, and initial bit 3, and data to be processed 1 is associated with initial bit 1, data to be processed 2 is associated with initial bit 2, and data to be processed 3 is associated with initial bit 3. Secondly, the server also needs to initialize the position parameter of each initial bit to an initial value, which is 0 in the embodiment of this application.
[0082] For ease of understanding, please refer to Figure 4 , Figure 4 which is a schematic diagram of an embodiment of the initial bit array set provided by the embodiment of this application. As Figure 4 shown, B11, B12, B21, and B22 all indicate initial bits. Based on this, the initial bit array B1 includes multiple initial bits such as initial bit B11 and initial bit B12, and data to be processed 1 is associated with initial bit B11, data to be processed 2 is associated with initial bit B12. Secondly, the position parameter of initial bit B11 is 0, and the position parameter of initial bit B12 is 0. Similarly, the initial bit array B2 includes multiple initial bits such as initial bit B21 and initial bit B22, and data to be processed 1 is associated with initial bit B21, data to be processed 2 is associated with initial bit B22. Secondly, the position parameter of initial bit B21 is 0, the position parameter of initial bit B12 is 0, the position parameter of initial bit B21 is 0, and the position parameter of initial bit B22 is 0. It should be understood that the foregoing examples are only for understanding this solution and should not be construed as a limitation of this solution.
[0083] 103. Process each piece of data to be processed in the initial digit array set, and obtain the modulo values corresponding to each piece of data to be processed in each initial digit array.
[0084] In this embodiment, the server processes each piece of data to be processed in the initial digit array set, and obtains the modulo values corresponding to each piece of data to be processed in each initial digit array. The modulo value corresponds one-to-one with one of the L initial digits included in the initial digit array. For example, based on the example in step 102, the initial digit array A includes the initial digit 1, the initial digit 2, and the initial digit 3, and the data to be processed 1 is associated with the initial digit 1, the data to be processed 2 is associated with the initial digit 2, and the data to be processed 3 is associated with the initial digit 3. Then, the modulo value of the data to be processed 1 is associated with the initial digit 1, the modulo value of the data to be processed 2 is associated with the initial digit 2, and the modulo value of the data to be processed 3 is associated with the initial digit 3. Specifically, the modulo value is actually the remainder obtained by performing modulo processing on each piece of data to be processed based on the initial digit array.
[0085] 104. Replace the position parameters of each initial digit in the N initial digit arrays with the modulo values corresponding to each piece of data to be processed in each initial digit array, and obtain the target digit array set.
[0086] In this embodiment, the server replaces the position parameters of each initial digit in the N initial digit arrays with the modulo values corresponding to each piece of data to be processed in each initial digit array, and obtains the target digit array set. Specifically, the server maps the modulo values corresponding to each piece of data to be processed in each initial digit array to the position parameters of each initial digit associated with the data to be processed.
[0087] Specifically, the obtained target digit array set includes N target digit arrays, and each target digit array includes L target digits. The target digits are associated with the data to be processed one-to-one, and the position parameter of each target digit is the modulo value corresponding to each piece of data to be processed in each initial digit array. For ease of understanding, based on Figure 4 the shown initial digit array set as an example for further illustration, please refer to Figure 5 , Figure 5 is a schematic diagram of an embodiment of the target digit array set provided by the embodiment of the present application. As shown in Figure 5As shown, C11, C12, C21, and C22 all indicate the initial digits, C31, C32, C41, and C42 all indicate the target digits, C51 indicates the modulo value corresponding to the data to be processed 1 in the initial digit array C1, C52 indicates the modulo value corresponding to the data to be processed 1 in the initial digit array C2, C53 indicates the modulo value corresponding to the data to be processed 2 in the initial digit array C1, and C54 indicates the modulo value corresponding to the data to be processed 5 in the initial digit array C2.
[0088] Based on this, if the modulo value C52 corresponding to the data to be processed 1 in the initial digit array C1 is 1, and the modulo value C52 corresponding to it in the initial digit array C2 is 0, then the modulo value C51 can be mapped to the position parameter of the initial data C11, and the modulo value C52 can be mapped to the position parameter of the initial data C21. Similarly, if the modulo value C53 corresponding to the data to be processed 2 in the initial digit array C1 is 0, and the modulo value C54 corresponding to it in the initial digit array C2 is 1, then the modulo value C53 can be mapped to the position parameter of the initial data C21, and the modulo value C54 can be mapped to the position parameter of the initial data C22. Thus, the target digits C3 and C4 are obtained, and the data to be processed 1 is associated with the target digits C31 and C41, and the data to be processed 2 is associated with the target digits C41 and C42. Secondly, the position parameter of the target digit C31 is 1, the position parameter of the target digit C41 is 0, the position parameter of the target digit C32 is 2, and the position parameter of the target digit C42 is 1. It should be understood that the foregoing examples are all for understanding this solution and should not be construed as limitations of this solution.
[0089] 105. Determine the target data from the L data to be processed based on the target digit array set.
[0090] In this embodiment, the server determines the target data from the L data to be processed based on the target digit array set. The position parameters of the N target digits associated with the target data are different from the initial values, that is, the target data is duplicate data. Specifically, the target digit array set can be obtained through step 104, and the position parameter of each target data in the target digit array has been replaced with the modulo value of the associated data to be processed. Based on the Bloom filter mentioned in the foregoing embodiment, based on the operating logic of the Bloom filter, it can be determined that if the position parameters of all the target digits associated with the data to be processed are all 1, then the data to be processed is in the set, that is, the data to be processed has been processed at least once, and thus it is determined that the data to be processed is the target data (duplicate data).
[0091] Exemplarily, please refer to Figure 6 , Figure 6A schematic diagram of an embodiment for determining target data provided by an embodiment of the present application, as Figure 6 shown, if the target bit array set includes target bit array D1 and target bit array D2, target bit array D1 includes target digit D11 associated with data to be processed 1, and the position parameter of target digit D11 is "1", and target bit array D2 includes target digit D21 associated with data to be processed 1, and the position parameter of target digit D21 is "1". Therefore, the position parameters of target digit D11 and target digit D21 associated with data to be processed 1 are both different from the initial value. Therefore, data to be processed 1 can be determined as the target data. It should be understood that in this embodiment, the position data is judged to be 1. Since the modulo value is the remainder in actual applications, and the remainder can be positive integers such as 0, 1, 2, etc., as long as the position parameters of the target digits are all different from the initial value (0 in this application), the target data can be determined, and there may be cases where the position parameters of the target digits associated with the target data in different target bit arrays are different. Therefore, this solution does not limit this.
[0092] In an embodiment of the present application, a data processing method is provided. Through the above method, since the initial digit is associated with the data to be processed one by one, that is, a data to be processed is mapped to an initial digit, the modulo value corresponding to each data to be processed in each initial bit array is used to replace the position parameter of the initial digit associated with the data to be processed to obtain the position parameter of the target digit associated with the data to be processed. When the position parameter of the target digit associated with the data to be processed is different from the initial value, the data to be processed is determined as duplicate data, and it is not necessary to access external storage for each piece of data, which can reduce the consumption of storage resources in the scenario of a large amount of data, thereby improving the real-time data processing efficiency.
[0093] Optionally, on the basis of the above Figure 3 corresponding embodiment, in an optional embodiment of the data processing method provided by the embodiment of the present application, each data to be processed in the data to be processed is processed based on the initial bit array set to obtain the modulo value corresponding to each data to be processed in each initial bit array, specifically including:
[0094] Create a hash function corresponding to the initial bit array set;
[0095] Create a modulo function corresponding to each initial bit array;
[0096] Based on the hash function corresponding to the initial bit array set and the modulo function corresponding to each initial bit array, each data to be processed in the data to be processed is processed to obtain the modulo value corresponding to each data to be processed in each initial bit array.
[0097] In this embodiment, the server needs to specifically create a hash function corresponding to the initial bit array set, that is, create a corresponding hash function for the initial bit array set. And it also needs to create a modulo function corresponding to each initial bit array, that is, create a corresponding modulo function for all the initial bit arrays in the initial bit array set. For example, the hash function 1 corresponding to the initial bit array set, the modulo function 1 corresponding to the initial bit array 1 in the initial bit array set, the modulo function 2 corresponding to the initial bit array 2, and the modulo function 3 corresponding to the initial bit array 3. Thus, based on the hash function corresponding to the initial bit array set and the modulo function corresponding to each initial bit array, the server processes each piece of data to be processed in the data to be processed, and obtains the modulo value corresponding to each piece of data to be processed in each initial bit array. Next, how to obtain the modulo value based on the hash function and the modulo function will be specifically introduced.
[0098] Optionally, based on the above Figure 3 corresponding embodiment, in an alternative embodiment of the data processing method provided by the embodiments of the present application, based on the hash function corresponding to the initial bit array set and the modulo function corresponding to each initial bit array, each piece of data to be processed in the data to be processed is processed to obtain the modulo value corresponding to each piece of data to be processed in each initial bit array, which specifically includes:
[0099] Parse and process the L pieces of data to be processed in the data set to be processed, and obtain the key value corresponding to each piece of data to be processed;
[0100] Use the hash function corresponding to the initial bit array set to perform hash calculation on the key value corresponding to each piece of data to be processed, and obtain the hash value of the key value corresponding to each piece of data to be processed;
[0101] Use the modulo function corresponding to each initial bit array to perform modulo processing on the hash value of the key value corresponding to each piece of data to be processed, and obtain the modulo value corresponding to each piece of data to be processed in each initial bit array.
[0102] In this embodiment, the server first parses and processes the L pieces of data to be processed in the data set to be processed, and obtains the key value (key) corresponding to each piece of data to be processed. Then, it uses the hash function corresponding to the initial bit array set to perform hash calculation on the key value corresponding to each piece of data to be processed, and obtains the hash value hash(key) of the key value corresponding to each piece of data to be processed. Based on this, it then uses the modulo function corresponding to each initial bit array to perform modulo processing on the hash value hash(key) of the key value corresponding to each piece of data to be processed, and obtains the modulo value corresponding to each piece of data to be processed in each initial bit array. The above-mentioned modulo processing to obtain the modulo value corresponding to each piece of data to be processed in each initial bit array is specifically performed through the following formula (1):
[0103] yi = mod i (hash(key))
[0104] y i ∈ {1,..., L}; (1)
[0105] where y i refers to the modulo value corresponding to the data to be processed in the initial bit array i, hash(key) refers to the hash value of the key value corresponding to the data to be processed, L refers to the number of data to be processed in the data set to be processed, i refers to the initial bit array i, and i belongs to N.
[0106] In an embodiment of the present application, another data processing method is provided. Through the above method, a hash function can be used to obtain the hash value of the key value corresponding to the data to be processed. However, the generated hash value of the key value is modulo-calculated using the modulo functions corresponding to multiple initial bit arrays, so that the obtained modulo value can be associated with the initial digit position parameter subsequently, thereby ensuring the feasibility of data processing. Secondly, it is not necessary to perform multiple hash calculations on the key value corresponding to each data to be processed, thereby reducing the running time of hash calculations, and thus improving the efficiency of data processing.
[0107] Optionally, on the basis of the above Figure 3 corresponding embodiment, in an optional embodiment of the data processing method provided in the embodiment of the present application, the data processing method further includes:
[0108] Parse and process the L data to be processed in the data set to be processed to obtain the non-key value corresponding to each data to be processed;
[0109] And after determining the target data from the L data to be processed based on the target bit array set, the data processing method further includes:
[0110] Discard the target data.
[0111] In this embodiment, the server can also parse and process the L data to be processed in the data set to be processed to obtain the non-key value corresponding to each data to be processed. It should be understood that the server can parse to obtain the non-key value corresponding to the data to be processed either simultaneously with obtaining the key value corresponding to the data to be processed or after determining that the data to be processed is not the target data and then parsing to obtain the non-key value corresponding to the data to be processed. The specific method is not limited here. Secondly, after the server determines the target data from the L data to be processed based on the target bit array set, the target data will be discarded. Thus, duplicate data can be filtered out with a sufficiently small amount of storage resources, and the duplicate data is discarded to avoid waste of storage space. Thereby, it can also ensure the reasonable allocation of storage resources in the storage space, thereby reducing the operation and maintenance cost of the storage system.
[0112] Optionally, based on the corresponding embodiments above, in an optional embodiment of the data processing method provided by the embodiments of the present application, the data processing method further includes: Figure 3 Determining cached data from L pieces of data to be processed based on the target bit array set, wherein the position parameter of each target bit among the N target bits associated with the cached data is an initial value, and the cached data is non-duplicate data;
[0113] Writing the key value corresponding to the cached data and the non-key value corresponding to the cached data into the cache.
[0114] In this embodiment, the server can always determine the target data in step 105. Then, the server can also determine the cached data from L pieces of data to be processed based on the target bit array set. The position parameter of each target bit among the N target bits associated with the cached data is an initial value. Based on the Bloom filter mentioned in the foregoing embodiments, based on the operation logic of the Bloom filter, it can be determined that if there is a 0 in the position parameters of all target bits associated with the data to be processed, then it has not been processed yet, and thus it is determined that the data to be processed is cached data (non-duplicate data). Further, since the server obtains the non-key value corresponding to each piece of data to be processed based on the method described in the foregoing embodiments, the server can also write the key value corresponding to the cached data and the non-key value corresponding to the cached data into the cache.
[0115] Exemplarily, please refer to
[0116] which is a schematic diagram of an embodiment for determining cached data provided by the embodiments of the present application. As Figure 7 , Figure 7 Figure 7 As shown, if the set of target bit arrays includes target bit array E1 and target bit array E2, target bit array E1 includes target digit E11 associated with data to be processed 1 and target digit E12 associated with data to be processed 2, and the position parameter of target digit E11 is "0", and the position parameter of target digit E12 is "1". Secondly, target bit array E2 includes target digit E21 associated with data to be processed 1 and target digit E22 associated with data to be processed 2, and the position parameter of target digit E21 is "1", and the position parameter of target digit E22 is "0". Therefore, although the position parameter of target digit E21 associated with data to be processed 1 is different from the initial value, there is a target digit E11 associated with data to be processed 1 whose position parameter is consistent with the initial value. Therefore, data to be processed 1 can be determined as cached data. Similarly, it can be known that the position parameter of target digit E12 associated with data to be processed 2 is different from the initial value. Although there is a target digit E22 associated with data to be processed 2 whose position parameter is consistent with the initial value, data to be processed 2 can also be determined as cached data. It should be understood that the foregoing example is only for understanding this solution. In practical applications, regardless of the value of the position parameter of the target digit associated with the data to be processed, as long as there is a situation where the position parameter of the target digit is the initial value (0 in this application), the data to be processed can be used as cached data. Therefore, the foregoing example should not be construed as a limitation of this solution.
[0117] In an embodiment of the present application, another data processing method is provided. Through the above method, waste of storage space is avoided by discarding duplicate data, and thus reasonable allocation of storage resources in the storage space can also be ensured, thereby reducing the operation and maintenance cost of the storage system. Secondly, by writing the key values and non-key values corresponding to the non-duplicate cached data into the cache, when the storage space needs to process the cached data, the cached data can be directly obtained from the cache address, which improves the data processing efficiency on the basis of ensuring the reliability of data processing.
[0118] Optionally, on the basis of the above Figure 3 corresponding embodiment, in an optional embodiment of the data processing method provided by the embodiment of the present application, the data processing method further includes:
[0119] Determine the key value corresponding to the cached data as a non-duplicate key value;
[0120] Perform counting and statistical processing on the cached data existing in L pieces of data to be processed to obtain the number of non-duplicate key values, where the number of non-duplicate key values is the number of all cached data in L pieces of data to be processed;
[0121] Statistically calculate the modulo values corresponding to all cached data in L pieces of data to be processed in each initial bit array;
[0122] Determine the summary data of the non-repeating key values based on the number of non-repeating key values and the modulo values corresponding to all cached data in each of the L data to be processed in each initial bit array, where the summary data is used for data analysis and processing of the data set to be processed.
[0123] In this embodiment, the server can also determine the key values corresponding to the cached data as non-repeating key values, and then perform a counting and statistical process on the cached data existing in the L data to be processed to obtain the number of non-repeating key values. Since all cached data can be determined from the L data to be processed through the foregoing embodiment, the number of key values (non-repeating key values) corresponding to all cached data is equal to the number of all cached data in the L data to be processed. Based on this, the server can also obtain the modulo value corresponding to each data to be processed in each initial bit array through the foregoing embodiment. Therefore, based on the situation of determining all cached data from the L data to be processed, the server can also count the modulo values corresponding to all cached data in each initial bit array among the L data to be processed. Finally, based on the number of non-repeating key values and the modulo values corresponding to all cached data in each initial bit array among the L data to be processed, determine the summary data of the non-repeating key values, and this summary data is used for data analysis and processing of the data set to be processed, such as performing data monitoring on the data to be processed in the data set to be processed, or abnormal behavior monitoring, or obtaining user model features based on the data set to be processed, etc., and specific details are not limited here.
[0124] Specifically, the server determines the summary data of the non-repeating key values through the following formula (2):
[0125]
[0126] where s is the summary data of the non-repeating key values, K refers to the number of non-repeating key values, and refers to the modulo values corresponding to all cached data in each initial bit array among the L data to be processed.
[0127] In the embodiment of the present application, another data processing method is provided. Through the above method, by discarding duplicate data and performing cache processing on non-duplicate data, finally all areas in the storage space are non-duplicate and can be processed. Therefore, by statistically obtaining the summary data of the relevant information of these non-duplicate data, it can enable the operation and maintenance personnel to perform data analysis and processing on the data set to be processed based on this summary data, thereby ensuring the stability of the system.
[0128] Optionally, based on the foregoing Figure 3 corresponding embodiment, in an optional embodiment of the data processing method provided in the embodiment of the present application, obtaining a data set to be processed specifically includes:
[0129] Obtain an initial data set, where the initial data set includes M initial data, each initial data corresponds to a topic type one by one, and the M initial data are from at least two data sources, and M is an integer greater than L;
[0130] Perform parsing processing on each initial data in the initial data set to obtain the topic type corresponding to each initial data;
[0131] Extract the initial data with the topic type being the target topic type to generate a data set to be processed.
[0132] In this embodiment, the operation behaviors and item data from different terminal devices are collected through SDK embedding and reported to the server. The foregoing operation behaviors and item data are the initial data. The server side parses and processes the initial data of different terminal devices, and sets the at least once mode to ensure that the initial data will not be lost, but there may be duplicate data. Based on this, the server parses the foregoing initial data from different terminal devices to obtain the topic type corresponding to each initial data. The initial data of different topic types will be shunted for calculation. In this application embodiment, the data of the target topic type is processed. Then the server can extract the initial data with the topic type being the target topic type to generate a data set to be processed.
[0133] In the embodiment of the present application, another data processing method is provided. Through the above method, by processing the initial data set from different terminal devices, the initial data of different topic types are shunted to different servers for processing, thereby ensuring the reliability and efficiency of each server in processing data.
[0134] Figure 8 It is a schematic structural diagram of a data processing device provided in an embodiment of the present application, as Figure 8 shown. The data processing device includes:
[0135] An obtaining module 801, configured to obtain a data set to be processed, where the data set to be processed includes L data to be processed, and L is an integer greater than 1;
[0136] A creating module 802, configured to create an initial bit array set based on the data set to be processed, where the initial bit array set includes N initial bit arrays, each initial bit array includes L initial bits, the initial bits are associated with the data to be processed one by one, and the position parameter of each initial bit is an initial value, and N is an integer greater than 1;
[0137] A processing module 803, configured to process each piece of data to be processed in the data to be processed based on the initial digit array set, and obtain the modulo values corresponding to each piece of data to be processed in each initial digit array, where the modulo values are in one-to-one correspondence with one of the L initial digits included in the initial digit array;
[0138] A replacement module 804, configured to replace the position parameter of each initial digit in the N initial digit arrays with the modulo value corresponding to each piece of data to be processed in each initial digit array, to obtain a target digit array set, where the target digit array set includes N target digit arrays, each target digit array includes L target digits, the target digits are associated with the data to be processed one by one, and the position parameter of each target digit is the modulo value corresponding to each piece of data to be processed in each initial digit array;
[0139] A determination module 805, configured to determine target data from the L pieces of data to be processed based on the target digit array set, where the position parameters of the N target digits associated with the target data are different from the initial values, and the target data is duplicate data.
[0140] Optionally, based on the embodiment corresponding to the above Figure 8 In another embodiment of the data processing apparatus 800 provided by the embodiment of the present application, the processing module 803 is specifically configured to create a hash function corresponding to the initial digit array set;
[0141] Create a modulo function corresponding to each initial digit array;
[0142] Based on the hash function corresponding to the initial digit array set and the modulo function corresponding to each initial digit array, process each piece of data to be processed in the data to be processed, and obtain the modulo values corresponding to each piece of data to be processed in each initial digit array.
[0143] Optionally, based on the embodiment corresponding to the above Figure 8 In another embodiment of the data processing apparatus 800 provided by the embodiment of the present application, the processing module 803 is specifically configured to perform parsing processing on the L pieces of data to be processed in the data to be processed set, and obtain the key value corresponding to each piece of data to be processed;
[0144] Use the hash function corresponding to the initial digit array set to perform hash calculation on the key value corresponding to each piece of data to be processed, and obtain the hash value of the key value corresponding to each piece of data to be processed;
[0145] Use the modulo function corresponding to each initial digit array to perform modulo processing on the hash value of the key value corresponding to each piece of data to be processed, and obtain the modulo values corresponding to each piece of data to be processed in each initial digit array.
[0146] Optionally, based on the embodiment corresponding to the above Figure 8Based on the corresponding embodiment, in another embodiment of the data processing device 800 provided by the embodiments of the present application, the obtaining module 801 is further configured to perform parsing processing on L pieces of data to be processed in the data set to be processed, and obtain non-critical values corresponding to each piece of data to be processed;
[0147] The processing module 803 is further configured to discard the target data after the determining module 805 determines the target data from the L pieces of data to be processed based on the target bit array set.
[0148] Optionally, based on the corresponding embodiment above Figure 8 In another embodiment of the data processing device 800 provided by the embodiments of the present application, the data processing device further includes a reading and writing module 806;
[0149] The determining module 805 is further configured to determine cached data from the L pieces of data to be processed based on the target bit array set, where the position parameter of each target digit in the N target digits associated with the cached data is an initial value, and the cached data is non-duplicate data;
[0150] The reading and writing module 806 is configured to write the key value corresponding to the cached data and the non-critical value corresponding to the cached data into the cache.
[0151] Optionally, based on the corresponding embodiment above Figure 8 In another embodiment of the data processing device 800 provided by the embodiments of the present application, the data processing device further includes a statistics module 807;
[0152] The determining module 805 is further configured to determine the key value corresponding to the cached data as a non-duplicate key value; and determine summary data of the non-duplicate key values based on the number of non-duplicate key values and the modulo values corresponding to all the cached data in each initial bit array among the L pieces of data to be processed, where the summary data is used to perform data analysis processing on the data set to be processed;
[0153] The processing module 803 is further configured to perform counting and statistical processing on the cached data existing in the L pieces of data to be processed to obtain the number of non-duplicate key values, where the number of non-duplicate key values is the number of all the cached data among the L pieces of data to be processed;
[0154] The statistics module 807 is further configured to count the modulo values corresponding to all the cached data in each initial bit array among the L pieces of data to be processed.
[0155] Optionally, based on the corresponding embodiment above Figure 8Based on the corresponding embodiments, in another embodiment of the data processing device 800 provided by the embodiments of the present application, the acquisition module 801 is specifically configured to acquire an initial data set, where the initial data set includes M initial data, each initial data corresponds to a theme type one by one, and the M initial data are derived from at least two data sources, and M is an integer greater than L;
[0156] Parse each initial data in the initial data set to obtain the theme type corresponding to each initial data;
[0157] Extract the initial data with the theme type being the target theme type to generate a data set to be processed.
[0158] The embodiments of the present application also provide another data processing device. The data processing devices are all deployed on the server. In the present application, the example of the data processing device being deployed on the server is used for illustration. Please refer to Figure 9 , Figure 9 This is a schematic diagram of an embodiment of the server in the embodiments of the present application. As Figure 9 shown, the server 1000 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 1022 (for example, one or more processors) and a memory 1032, and one or more storage media 1030 (for example, one or more mass storage devices) for storing application programs 1042 or data 1044. Among them, the memory 1032 and the storage media 1030 may be transient storage or persistent storage. The program stored in the storage media 1030 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 1022 may be configured to communicate with the storage media 1030 and execute a series of instruction operations in the storage media 1030 on the server 1000.
[0159] The server 1000 may further include one or more power supplies 1026, one or more wired or wireless network interfaces 1050, one or more input / output interfaces 1058, and / or one or more operating systems 1041, such as Windows Server TM ,Mac OS X TM ,Unix TM ,Linux TM ,FreeBSD TM and so on.
[0160] The steps executed by the server in the above embodiments may be based on the Figure 9 server structure shown.
[0161] The CPU 1022 included in the server is used to execute the embodiments as Figure 3 shown and Figure 3 the corresponding respective embodiments.
[0162] In an embodiment of the present application, a computer-readable storage medium is further provided. A computer program is stored in the computer-readable storage medium. When it runs on a computer, it causes the computer to execute the steps performed by the server in the method described in the foregoing Figure 3 embodiments as shown.
[0163] In an embodiment of the present application, a computer program product including a program is further provided. When it runs on a computer, it causes the computer to execute the steps performed by the server in the method described in the foregoing Figure 3 embodiments as shown.
[0164] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0165] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, at least two units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, indirect couplings or communication connections of devices or units, and can be in electrical, mechanical, or other forms.
[0166] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or can be distributed to at least two network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0167] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0168] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0169] As described above, the above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of this application.
Claims
1. A method for data processing, characterized in that, Including: Obtain a set of data to be processed, where the set of data to be processed includes L pieces of data to be processed, and L is an integer greater than 1; Create an initial bit array set based on the set of data to be processed, where the initial bit array set includes N initial bit arrays, each initial bit array includes L initial bits, the initial bits are associated with the data to be processed one by one, and the position parameter of each initial bit is an initial value, and N is an integer greater than 1; Process each piece of data to be processed in the set of data to be processed based on the initial bit array set, and obtain the modulo values corresponding to each piece of data to be processed in each initial bit array, where the modulo values are in one-to-one correspondence with one of the L initial bits included in the initial bit array; Replace the position parameter of each initial bit in the N initial bit arrays with the modulo value corresponding to each piece of data to be processed in each initial bit array, and obtain a target bit array set, where the target bit array set includes N target bit arrays, each target bit array includes L target bits, the target bits are associated with the data to be processed one by one, and the position parameter of each target bit is the modulo value corresponding to each piece of data to be processed in each initial bit array; Determine target data from the L pieces of data to be processed based on the target bit array set, where the position parameters of the N target bits associated with the target data are different from the initial value, and the target data is duplicate data; Determine cached data from the L pieces of data to be processed based on the target bit array set, where the position parameter of each of the N target bits associated with the cached data is the initial value, and the cached data is non-duplicate data; Write the key value corresponding to the cached data and the non-key value corresponding to the cached data into the cache.
2. The method according to claim 1, wherein The step of processing each piece of data to be processed in the set of data to be processed based on the initial bit array set to obtain the modulo value corresponding to each piece of data to be processed in each initial bit array includes: Create a hash function corresponding to the initial bit array set; Create a modulo function corresponding to each initial bit array; Process each piece of data to be processed in the set of data to be processed based on the hash function corresponding to the initial bit array set and the modulo function corresponding to each initial bit array, and obtain the modulo value corresponding to each piece of data to be processed in each initial bit array.
3. The method according to claim 2, characterized in that, The step of processing each piece of data to be processed in the set of data to be processed based on the hash function corresponding to the initial bit array set and the modulo function corresponding to each initial bit array to obtain the modulo value corresponding to each piece of data to be processed in each initial bit array includes: Perform parsing processing on the L pieces of data to be processed in the set of data to be processed, and obtain the key value corresponding to each piece of data to be processed; Use the hash function corresponding to the initial bit array set to perform hash calculation on the key value corresponding to each piece of data to be processed, and obtain the hash value of the key value corresponding to each piece of data to be processed; Using the modulo function corresponding to each of the initial bit arrays, perform a modulo operation on the hash value of the key value corresponding to each data to be processed, to obtain the modulo value corresponding to each data to be processed in each initial bit array.
4. The method according to claim 3, wherein The method further includes: Performing a parsing process on the L data to be processed in the data set to be processed, to obtain a non-key value corresponding to each data to be processed; And after determining the target data from the L data to be processed based on the target bit array set, the method further includes: Discarding the target data.
5. The method according to claim 1, characterized in that, The method further includes: Determining the key value corresponding to the cached data as a non-repeating key value; Performing a counting and statistical process on the cached data existing in the L data to be processed, to obtain the quantity of non-repeating key values, where the quantity of non-repeating key values is the quantity of all cached data in the L data to be processed; Statistically calculating the modulo values corresponding to all cached data in the L data to be processed in each initial bit array; Based on the quantity of non-repeating key values, and the modulo values corresponding to all cached data in the L data to be processed in each initial bit array, determining summary data of non-repeating key values, where the summary data is used for performing a data analysis process on the data set to be processed.
6. The method according to claim 1, characterized in that, The obtaining of the data set to be processed includes: Obtaining an initial data set, where the initial data set includes M initial data, each initial data corresponds to a theme type one by one, and the M initial data are from at least two data sources, and M is an integer greater than L; Performing a parsing process on each initial data in the initial data set, to obtain the theme type corresponding to each initial data; Extracting the initial data with the theme type being the target theme type, to generate the data set to be processed.
7. A data processing device, characterized in that, The data processing device includes: An obtaining module, configured to obtain a data set to be processed, where the data set to be processed includes L data to be processed, and L is an integer greater than 1; A creating module, configured to create an initial bit array set based on the data set to be processed, where the initial bit array set includes N initial bit arrays, each initial bit array includes L initial bits, the initial bits are associated with the data to be processed one by one, and the position parameter of each initial bit is an initial value, and N is an integer greater than 1; A processing module, configured to process each data to be processed in the data to be processed based on the initial bit array set, to obtain the modulo value corresponding to each data to be processed in each initial bit array, where the modulo value corresponds to one of the L initial bits included in the initial bit array one by one; A replacement module, configured to replace the position parameter of each initial digit in the N initial digit arrays with the modulo value corresponding to each data to be processed in each initial digit array, so as to obtain a set of target digit arrays, where the set of target digit arrays includes N target digit arrays, each target digit array includes L target digits, the target digits are associated with the data to be processed one by one, and the position parameter of each target digit is the modulo value corresponding to each data to be processed in each initial digit array; A determination module, configured to determine target data from the L data to be processed based on the set of target digit arrays, where the position parameters of the N target digits associated with the target data are different from the initial values, and the target data is duplicate data; The determination module is further configured to determine cached data from the L data to be processed based on the set of target digit arrays, where the position parameter of each target digit among the N target digits associated with the cached data is the initial value, and the cached data is non-duplicate data; A read-write module, configured to write the key value corresponding to the cached data and the non-key value corresponding to the cached data into the cache.
8. The device according to claim 7, characterized in that, The processing module is specifically configured to: Create a hash function corresponding to the set of initial digit arrays; Create a modulo function corresponding to each initial digit array; Based on the hash function corresponding to the set of initial digit arrays and the modulo function corresponding to each initial digit array, process each data to be processed in the data to be processed, so as to obtain the modulo value corresponding to each data to be processed in each initial digit array.
9. The device according to claim 8, characterized in that, The processing module is specifically configured to: Perform parsing processing on the L data to be processed in the set of data to be processed, and obtain the key value corresponding to each data to be processed; Use the hash function corresponding to the set of initial digit arrays to perform hash calculation on the key value corresponding to each data to be processed, so as to obtain the hash value of the key value corresponding to each data to be processed; Use the modulo function corresponding to each initial digit array to perform modulo processing on the hash value of the key value corresponding to each data to be processed, so as to obtain the modulo value corresponding to each data to be processed in each initial digit array.
10. The apparatus according to claim 9, wherein The acquisition module is further configured to perform parsing processing on the L data to be processed in the set of data to be processed, and obtain the non-key value corresponding to each data to be processed; The processing module is further configured to discard the target data after the determination module determines the target data from the L data to be processed based on the set of target digit arrays.
11. The device according to claim 7, characterized in that, The apparatus further includes: a statistics module; The determination module is further configured to determine the key value corresponding to the cached data as a non-duplicate key value; The processing module is further configured to perform counting and statistical processing on the cached data existing in the L data to be processed, so as to obtain the quantity of non-duplicate key values, where the quantity of non-duplicate key values is the quantity of all cached data in the L data to be processed; The statistical module is configured to count the modulo values corresponding to all cached data in each initial bit array among the L data to be processed; The determination module is further configured to determine summary data of non-repeating key values based on the number of non-repeating key values and the modulo values corresponding to all cached data in each initial bit array among the L data to be processed, wherein the summary data is used for data analysis and processing of the data set to be processed.
12. The device according to claim 7, wherein The obtaining module is specifically configured to: Obtain an initial data set, where the initial data set includes M initial data, each initial data corresponds to a subject type one by one, and the M initial data are from at least two data sources, and M is an integer greater than L; Perform parsing processing on each initial data in the initial data set to obtain the subject type corresponding to each initial data; Extract the initial data with the subject type being the target subject type to generate the data set to be processed.
13. A computer device, characterized in that, It includes: A memory, a transceiver, a processor, and a bus system; Wherein, the memory is used to store programs; The processor is used to execute the programs in the memory to implement the method according to any one of claims 1 to 6; The bus system is used to connect the memory and the processor to enable communication between the memory and the processor.
14. A computer-readable storage medium, characterized in that, It includes instructions that, when running on a computer, cause the computer to execute the method according to any one of claims 1 to 6.
15. A computer program product, characterized in that, The computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Event processing method and device, terminal equipment and storage medium
CN113419792A
Efficient lookup in multiple bloom filters
US20170154099A1