Mass data processing method and system
The method addresses uneven data distribution and model optimization issues in large-scale data processing by using distributed message queues, entropy-based slicing, and a hybrid deduplication model with convergence criteria, enhancing efficiency and accuracy.
Patent Information
- Application Number
- CN202510551682.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional large-batch data processing systems are difficult to realize real-time data collection and processing when facing multi-source heterogeneous data. Uneven data sharding leads to node overload, and complex data analysis tasks lack effective convergence criteria, resulting in excessive iteration or suboptimal solutions.
Data is collected in real time by a distributed message queue, entropy-weight dynamic sharding algorithm is used for uniform sharding, combined with HyperLogLog++ and Bloom filter for deduplication, Δ-convergence criteria are used for iterative calculation, and the analysis results are stored in columnar storage format.
It realizes efficient data collection and transmission, avoids node overload, improves data processing efficiency and resource utilization, and ensures data purity and accuracy of analysis results.
Smart Images

Figure CN120316104A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large - scale data processing, and particularly to a large - scale data processing method and system. Background Art
[0002] Large - scale data processing technology refers to a series of technologies and methods used to process massive data sets, aiming to improve the speed, accuracy, and scalability of data processing. Therefore, how to use advanced technical means to improve the intelligent level and security of large - scale data processing has become one of the urgent problems to be solved currently.
[0003] In the field of large - scale data processing, traditional data processing systems often have difficulty in realizing real - time data collection and processing when facing multi - source heterogeneous data. During the big data processing process, if the data sharding is not uniform enough, it will cause some nodes to be overloaded, thereby reducing the processing efficiency of the entire system. Moreover, for complex data analysis tasks, the optimization of model parameters is a key but time - consuming process, lacking effective convergence criteria, which is prone to over - iteration or sub - optimal solutions. Summary of the Invention
[0004] In view of the above - mentioned existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a large - scale data processing method to solve the problems that during the big data processing process, if the data sharding is not uniform enough, it will cause some nodes to be overloaded, thereby reducing the processing efficiency of the entire system, and for complex data analysis tasks, the optimization of model parameters is a key but time - consuming process, lacking effective convergence criteria, which is prone to over - iteration or sub - optimal solutions.
[0006] To solve the above - mentioned technical problems, the present invention provides the following technical solutions: In the first aspect, the present invention provides a large - scale data processing method, which includes: Real - time collecting multi - source heterogeneous data by using a distributed message queue to obtain an original data stream; Performing sharding processing on the original data stream by using an entropy - weight dynamic sharding algorithm to obtain uniformly - distributed data shards; Constructing a hybrid model combining HyperLogLog++ and Bloom filters based on the original data stream and removing duplicates from the data shards to obtain a unique data set; Performing iterative calculation on the unique data set by using a Δ - convergence criterion to obtain a final analysis result; Writing the analysis result into a distributed file platform by using column - based storage.
[0007] As a preferred solution of the large - scale data processing method of the present invention, where: the multi - source heterogeneous data is collected in real - time by using a distributed message queue to obtain an original data stream. The specific steps are as follows: Use Apache Kafka to subscribe to and receive data generated by data producers from different sources, and ensure an efficient data transmission mechanism by configuring the Producer API of Kafka to obtain initial data messages collected from each data source; Use the Consumer API of Kafka to process the received data messages; According to the preset topic and partition rules, classify and store the data messages in the Kafka cluster to achieve the preliminary collation of multi - source heterogeneous data, and obtain an original data stream organized by topic and partition; Use a time - window - based strategy to filter the original data stream in the Kafka cluster. By setting a fixed time window, trigger the data extraction operation at the end of each time window to obtain an original data stream that meets the time - window requirements; Use the Kafka Connect tool to export the original data stream that meets the time - window requirements to an external database, and realize data transfer by configuring connectors to obtain the final original data stream.
[0008] As a preferred solution of the large - scale data processing method of the present invention, where: the original data stream is segmented by using an entropy - weight dynamic segmentation algorithm to obtain evenly - distributed data segments. The specific steps are as follows: Use a pre - processing module to perform preliminary cleaning and formatting on the original data stream obtained from the Kafka cluster, remove noise data and convert it into a unified data structure to obtain a standardized data set; Use the information entropy formula to calculate the information entropy of each data block to obtain the information entropy value of each data block; Use the entropy - weight method to calculate the weights of each data block to obtain a set of data blocks based on information - entropy weights; Use the dynamic programming algorithm to perform an optimal segmentation of the original data stream according to the calculated data - block weights, minimize the difference between the segmented data segments, and at the same time ensure that the data within each segment is as evenly distributed as possible, and obtain evenly - distributed data segments.
[0009] As a preferred solution of the large - scale data processing method of the present invention, where: a hybrid model combining HyperLogLog++ and Bloom filters is constructed based on the original data stream, and duplicate data in the data segments is removed to obtain a unique data set. The specific steps are as follows: The HyperLogLog++ algorithm is used to estimate the cardinality of the elements in each data shard, and the cardinality estimation values of each data shard are obtained; The elements in the data shard are preprocessed using a Bloom filter, and a bit array of size and independent hash functions are initialized; For each input element, its positions in the bit array are calculated through all hash functions and set to 1, establishing a structure for detecting the existence of elements, and a Bloom filter is obtained; The Bloom filter is used to deduplicate the elements in each data shard, and each element in the data shard is traversed.
[0010] As a preferred solution of the large - batch data processing method described in the present invention, wherein: when using the Bloom filter to deduplicate the elements in each data shard and traversing each element in the data shard, the specific steps are as follows: Use all hash functions to check whether all the bits corresponding to this element are 1. When there is at least one bit that is 0, this element is considered a new element and the Bloom filter is updated; Otherwise, it is regarded as a duplicate element and excluded, and the data shard after preliminary deduplication is obtained; A combined model of HyperLogLog++ and Bloom filter is used to perform secondary verification on the data shard after preliminary deduplication; For the false - positive situation generated by the Bloom filter, HyperLogLog++ is used to confirm the true status of each suspected duplicate element, and an accurate deduplication result is obtained; The false - positive situation is misjudging a non - member as a member; A merge operation is used to integrate all the data shards that have passed the secondary verification into a complete and unique data set.
[0011] As a preferred solution of the large - batch data processing method described in the present invention, wherein: when using the Δ - convergence criterion to perform iterative calculations on the unique data set to obtain the final analysis result, the specific steps are as follows: An initialization module is used to make an initial setting for the unique data set; The initial setting includes determining an initial parameter vector, setting a maximum number of iterations, and a convergence threshold; A target function is used to quantify the error between each data point in the unique data set and its corresponding predicted value; The gradient descent method is used to update the parameter vector , until the maximum number of iterations is satisfied; The Δ - convergence criterion is used to determine whether the iteration should stop, and the change in the objective function value between two adjacent iterations is calculated; The validation set is input into the model, and the difference between the model output and the actual value is compared to evaluate the model performance, and the final analysis result is obtained.
[0012] As a preferred solution of the large - volume data processing method described in the present invention, wherein: the analysis result is written into the distributed file platform by column - based storage, and the specific steps are as follows: The final analysis result is format - converted to meet the requirements of column - based storage; The format conversion includes numerical type conversion, field sorting, and compression processing to obtain an analysis result that conforms to the column - based storage standard; The converted analysis result is encoded using the Apache Parquet column - based storage format, and the data type, compression algorithm, and encoding method of each column are defined to achieve data compression and support for fast query, obtaining an analysis result encoded in the column - based storage format; The encoded analysis result is batch - written into the distributed file platform using the batch loading technology, and an appropriate block size of 256MB and the number of concurrent write threads are set to optimize the write performance, obtaining an analysis result successfully stored in the distributed file platform.
[0013] In a second aspect, the present invention provides a large - volume data processing system, including: A data acquisition module, a data sharding module, a deduplication processing module, a data analysis module, and a data storage module; The data acquisition module is used to perform real - time subscription and reception of data from different sources using Apache Kafka, and ensure an efficient data transmission mechanism by configuring the Producer and Consumer API of Kafka, realizing the preliminary collation of multi - source heterogeneous data and the time - window strategy screening, and finally exporting the qualified original data stream to external storage; The data sharding module is used to pre - process the original data stream obtained from the Kafka cluster, calculate the information entropy of each data block and determine the weight, and apply the dynamic programming algorithm to optimally split the data according to the weight to obtain evenly - distributed data shards; The deduplication processing module is used to estimate the cardinality of the elements in the data shards using the HyperLogLog++ algorithm, and perform pre - processing and deduplication operations on these elements through a Bloom filter, combining the advantages of the two technologies to confirm the true state of the suspected duplicate elements, thereby obtaining a unique data set; The data analysis module is used to initialize analysis parameters, define the objective function to quantify errors, and use the gradient descent method to optimize the model parameters until the Δ-convergence criterion is met. Finally, the performance of the model is evaluated through a validation set to ensure accurate and reliable analysis results are output; The data storage module is used to convert the final analysis results into a format suitable for columnar storage, encode them using the Apache Parquet columnar storage format, and then efficiently write them to a distributed file platform using bulk loading technology for subsequent querying and analysis.
[0014] In a third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, any step of the large-scale data processing method described in the first aspect of the present invention is implemented.
[0015] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and: when the computer program is executed by the processor, any step of the large-scale data processing method described in the first aspect of the present invention is implemented.
[0016] The beneficial effects of the present invention are as follows: By using a distributed message queue to collect multi-source heterogeneous data in real time, an efficient data collection and transmission mechanism is realized. The design enables the method to easily handle the inflow of large-scale data, enhancing the scalability and stability of the method. By applying the entropy weight dynamic sharding algorithm to process the original data stream, effective data segmentation and uniform distribution are achieved. The method effectively avoids the processing bottleneck caused by excessive data volume in some shards, improving the efficiency and resource utilization rate of the entire data processing process. By constructing a hybrid model combining HyperLogLog++ and Bloom filters to deduplicate data shards, an efficient and accurate data deduplication operation is realized. The method significantly reduces the storage space requirements and improves the query efficiency, which is particularly important for data management in a big data environment, ensuring that the subsequent data analysis is based on the purest data set and improving the accuracy and reliability of the analysis results. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for description in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 It is a flowchart of the large-scale data processing method in Embodiment 1.
[0019] Figure 2Schematic diagram of the large - scale data processing system in Embodiment 1. Detailed implementation manners
[0020] To make the above - mentioned objects, features and advantages of the present invention more obvious and understandable, the following will describe in detail the specific implementation manners of the present invention with reference to the accompanying drawings of the specification.
[0021] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0022] Secondly, the so - called "one embodiment" or "embodiment" herein refers to a specific feature, structure or characteristic that can be included in at least one implementation manner of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or selectively exclusive embodiments from other embodiments.
[0023] Embodiment 1, referring to Figure 1 and Figure 2 , is the first embodiment of the present invention. This embodiment provides a method for processing a large amount of data, including the following steps: S1. Use a distributed message queue to collect multi - source heterogeneous data in real - time to obtain an original data stream; Furthermore, use Apache Kafka to subscribe to and receive data generated by data producers from different sources, and ensure an efficient data transmission mechanism by configuring the Producer API of Kafka to obtain initial data messages collected from each data source; Use the Consumer API of Kafka to process the received data messages; According to the preset topic and partition rules, classify and store the data messages in the Kafka cluster to achieve preliminary sorting of multi - source heterogeneous data, and obtain an original data stream organized by topic and partition; Use a time - window - based strategy to filter the original data stream in the Kafka cluster. By setting a fixed time window, trigger a data extraction operation at the end of each time window to obtain an original data stream that meets the time - window requirements; Use the Kafka Connect tool to export the original data stream that meets the time - window requirements to an external database, and implement data transfer by configuring connectors to obtain the final original data stream; It should be noted that by using Apache Kafka as a distributed message queue, not only can a large amount of data streams be processed, but also the high availability and fault tolerance of data can be ensured. The persistence feature of Kafka enables data to be guaranteed not to be lost even in case of system failures. At the same time, its efficient publish - subscribe mechanism supports the requirements of real - time data processing, which is particularly important for application scenarios that require quick responses.
[0024] S2. Use the entropy - weight dynamic sharding algorithm to perform sharding processing on the original data stream to obtain evenly distributed data shards; Furthermore, use a pre - processing module to perform preliminary cleaning and formatting on the original data stream obtained from the Kafka cluster, remove noise data and convert it into a unified data structure to obtain a standardized data set; Use the information entropy formula to calculate the information entropy of each data block to obtain the information entropy value of each data block. The expression is: ; Among them, represents the information entropy of the data block, is the probability that the th data item appears, and is the base of the logarithm; Use the entropy - weight method to calculate the weights of each data block to obtain a set of data blocks based on information - entropy weights. The expression is: ; Among them, represents the weight of the th data block, is its information entropy value, is the total number of data blocks; Use the dynamic programming algorithm to perform optimal segmentation on the original data stream according to the calculated weights of the data blocks, minimize the differences between the segmented data shards, and at the same time ensure that the data within each shard is as evenly distributed as possible, and obtain evenly distributed data shards. The expression is: ; Among them, represents the th shard, represents the entire data set, is the predetermined number of shards; It should be noted that the entropy - weight dynamic sharding algorithm evaluates the complexity of each data block by calculating its information entropy and assigns weights based on this, so as to achieve more reasonable data segmentation. This method can not only effectively avoid the situation of some shards being overloaded, but also ensure the consistency and balance of the data within each shard, providing a guarantee for subsequent efficient processing.
[0025] S3: Build a hybrid model combining HyperLogLog++ and Bloom filter based on the original data stream, and deduplicate the data shards to obtain a unique data set; Furthermore, the HyperLogLog++ algorithm is used to estimate the cardinality of the elements in each data shard, and the estimated cardinality of each data shard is obtained, which is expressed as: ; in, is the estimated base number, According to the number of barrels The adjusted constant, It is The value of the bucket; Use Bloom filter to preprocess the elements in the data slice and initialize a size of The bit array and Independent hash functions; For each input element, calculate its position in the bit array through all hash functions and set it to 1, establish a structure for detecting whether the element exists, and obtain a Bloom filter; Use Bloom filter to deduplicate elements in each data shard, traversing each element in the data shard; Use all hash functions to check whether all bits corresponding to the element are 1. If at least one bit is 0, the element is considered to be a new element and the Bloom filter is updated. Otherwise, it is considered as a duplicate element and excluded, and the data shards after preliminary deduplication are obtained; A combination model of HyperLogLog++ and Bloom filter is used to perform secondary verification on the data shards after initial deduplication. For the false positives generated by the Bloom filter, HyperLogLog++ is used to confirm the true status of each suspected duplicate element to obtain accurate deduplication results; A false positive situation is when a non-member is misidentified as a member; Use the merge operation to integrate all the data shards that have been verified twice into a complete and unique data set; It should be noted that the combined application of HyperLogLog++ and Bloom filter can provide efficient cardinality estimation and deduplication capabilities while maintaining low memory usage. The method is particularly suitable for large-scale data sets and can significantly reduce storage requirements and computing time without affecting accuracy, greatly improving data processing efficiency and resource utilization.
[0026] S4, using the Δ-convergence criterion to iteratively calculate the unique data set to obtain the final analysis result; Further, an initialization module is used to make an initial setting for the unique data set; The initial setting includes determining an initial parameter vector, setting a maximum number of iterations, and a convergence threshold; The objective function is used to quantify the error between each data point in the unique data set and its corresponding predicted value, and the expression is: ; where, is the objective function, is the actual value of the th data point, is the prediction model of the data point based on the current parameter vector ; The gradient descent method is used to update the parameter vector until the maximum number of iterations is satisfied; The Δ-convergence criterion is used to judge whether the iteration should stop, and the change in the objective function value between two adjacent iterations is calculated. The expression is: ; If < , it is considered that the optimal solution has been found. Otherwise, continue the iteration until is reached to obtain a parameter vector that satisfies the convergence condition; The validation set is input into the model, and the difference between the model output and the actual value is compared to evaluate the model performance, and the final analysis result is obtained; It should be noted that the Δ-convergence criterion ensures the stability and accuracy of model training by precisely controlling the parameter update during the iteration process. The method not only reduces unnecessary iteration times and computational costs, but also guarantees the quality of the model through strict convergence conditions, making the final analysis result more reliable and credible.
[0027] S5. The analysis result is written into the distributed file platform using columnar storage; Further, the final analysis result is format-converted to adapt to the requirements of columnar storage; The format conversion includes numerical type conversion, field sorting, and compression processing to obtain an analysis result that conforms to the columnar storage standard; The converted analysis result is encoded using the Apache Parquet columnar storage format, and the data type, compression algorithm, and encoding method of each column are defined to achieve data compression and fast query support, and an analysis result encoded in the columnar storage format is obtained; The encoded analysis results are batch-written to the distributed file platform using the batch loading technique. The appropriate block size of 256 MB and the number of concurrent write threads are set to optimize the write performance, and the analysis results successfully stored in the distributed file platform are obtained; It should be noted that the columnar storage format not only optimizes the utilization of storage space but also significantly improves the query performance. This storage method is particularly suitable for big data analysis scenarios because it allows only the required columns to be read instead of the entire row, thus greatly accelerating the data analysis speed and reducing the cost of I / O operations. In addition, the write efficiency is further improved through the batch loading technique, ensuring that the data can be quickly and stably stored in the distributed file platform.
[0028] This embodiment also provides a large-scale data processing system, including: A data acquisition module, a data sharding module, a deduplication processing module, a data analysis module, and a data storage module; The data acquisition module is used to perform real-time subscription and reception of data from different sources using Apache Kafka, and ensure an efficient data transmission mechanism by configuring the Producer and Consumer API of Kafka, realize the preliminary collation of multi-source heterogeneous data and the time window strategy screening, and finally export the qualified original data stream to external storage; The data sharding module is used to preprocess the original data stream obtained from the Kafka cluster, calculate the information entropy of each data block and determine the weight, and apply the dynamic programming algorithm to optimally split the data according to the weight to obtain evenly distributed data shards; The deduplication processing module is used to estimate the cardinality of the elements in the data shards using the HyperLogLog++ algorithm, and perform preprocessing and deduplication operations on these elements through a Bloom filter, and combine the advantages of the two technologies to confirm the true state of the suspected duplicate elements, thereby obtaining a unique data set; The data analysis module is used to initialize the analysis parameters, define the objective function to quantify the error, and use the gradient descent method to optimize the model parameters until the Δ-convergence criterion is met, and finally evaluate the model performance through the validation set to ensure the output of accurate and reliable analysis results; The data storage module is used to convert the final analysis results into a format suitable for columnar storage, encode them using the Apache Parquet columnar storage format, and then efficiently write them to the distributed file platform using the batch loading technique for subsequent query and analysis.
[0029] This embodiment also provides a computer device, which is applicable to the case of a large - volume data processing method, and includes: a memory and a processor; the memory is used to store computer - executable instructions, and the processor is used to execute the computer - executable instructions to implement the large - volume data processing method proposed in the above - mentioned embodiment.
[0030] The computer device can be a terminal. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non - volatile storage medium and an internal memory. The non - volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non - volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad set on the outer shell of the computer device, or an external keyboard, a touchpad, or a mouse, etc.
[0031] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the large - volume data processing method proposed in the above - mentioned embodiment; the storage medium can be implemented by any type of volatile or non - volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM for short), Electrically Erasable Programmable Read - Only Memory (EEPROM for short), Erasable Programmable Read Only Memory (EPROM for short), Programmable Read - Only Memory (PROM for short), Read - Only Memory (ROM for short), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0032] In summary, the present invention realizes an efficient data collection and transmission mechanism by using a distributed message queue to collect multi-source heterogeneous data in real time. The design enables the method to easily handle the inflow of large-scale data, enhancing the scalability and stability of the method. By applying the entropy weight dynamic sharding algorithm to process the original data stream, the effective segmentation and uniform distribution of data are achieved. The method effectively avoids the processing bottleneck caused by excessive data volume in some shards, improving the efficiency and resource utilization rate of the entire data processing process. By constructing a hybrid model combining HyperLogLog++ and Bloom filters to deduplicate data shards, an efficient and accurate data deduplication operation is realized. The method significantly reduces the storage space requirement and improves the query efficiency, which is particularly important for data management in a big data environment, ensuring that the subsequent data analysis is based on the purest data set and improving the accuracy and reliability of the analysis results.
[0033] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A method for processing a large amount of data, characterized in that: Including: Using a distributed message queue to collect multi-source heterogeneous data in real time to obtain an original data stream; Using an entropy weight dynamic sharding algorithm to perform sharding processing on the original data stream to obtain evenly distributed data shards; Constructing a hybrid model combining HyperLogLog++ and Bloom filter based on the original data stream, and performing deduplication on the data shards to obtain a unique data set; Using the Δ-convergence criterion to perform iterative calculation on the unique data set to obtain the final analysis result; Using columnar storage to write the analysis result into a distributed file platform.
2. The method for processing a large amount of data according to claim 1, wherein: The specific steps of using a distributed message queue to collect multi-source heterogeneous data in real time to obtain an original data stream are as follows: Using Apache Kafka to subscribe to and receive data generated by data producers from different sources, and ensuring an efficient data transmission mechanism by configuring the Producer API of Kafka to obtain initial data messages collected from each data source; Using the Consumer API of Kafka to process the received data messages; Classifying and storing the data messages in the Kafka cluster according to the pre-set topic and partition rules to achieve preliminary sorting of multi-source heterogeneous data, and obtaining an original data stream organized by topic and partition; Using a time window-based strategy to filter the original data stream in the Kafka cluster, triggering data extraction operations at the end of each fixed time window by setting a fixed time window to obtain an original data stream that meets the time window requirements; Using the Kafka Connect tool to export the original data stream that meets the time window requirements to an external database, and implementing data transfer by configuring connectors to obtain the final original data stream.
3. The mass data processing method according to claim 2, wherein: The specific steps of using an entropy weight dynamic sharding algorithm to perform sharding processing on the original data stream to obtain evenly distributed data shards are as follows: Using a preprocessing module to perform preliminary cleaning and formatting on the original data stream obtained from the Kafka cluster, removing noise data and converting it into a unified data structure to obtain a standardized data set; Calculating the information entropy of each data block using the information entropy formula to obtain the information entropy value of each data block; Calculating the weights of each data block using the entropy weight method to obtain a set of data blocks based on information entropy weights; Using a dynamic programming algorithm to perform optimal segmentation on the original data stream according to the calculated data block weights, minimizing the difference between the divided data shards, and at the same time ensuring that the data within each shard is as evenly distributed as possible, and obtaining evenly distributed data shards.
4. The method for processing a large amount of data according to claim 3, wherein: The specific steps of constructing a hybrid model combining HyperLogLog++ and Bloom filter based on the original data stream, and performing deduplication on the data shards to obtain a unique data set are as follows: Using the HyperLogLog++ algorithm to estimate the cardinality of the elements in each data shard to obtain the cardinality estimation value of each data shard; Preprocess the elements in the data shard using a Bloom filter, and initialize a bit array of size and independent hash functions; For each input element, calculating its position in the bit array through all hash functions and setting it to 1 to establish a structure for detecting the existence of elements to obtain a Bloom filter; Use a Bloom filter to deduplicate the elements in each data shard and traverse each element in the data shard.
5. The method for processing a large amount of data according to claim 4, wherein: The steps of using a Bloom filter to deduplicate the elements in each data shard and traverse each element in the data shard are as follows: Use all hash functions to check whether all bits corresponding to the element are 1. When there is at least one bit that is 0, the element is considered a new element and the Bloom filter is updated; Otherwise, it is regarded as a duplicate element and excluded, and the data shard after preliminary deduplication is obtained; Use a combined model of HyperLogLog++ and Bloom filter to perform secondary verification on the data shard after preliminary deduplication; For the false positive situation generated by the Bloom filter, use HyperLogLog++ to confirm the true status of each suspected duplicate element to obtain an accurate deduplication result; The false positive situation refers to misjudging a non-member as a member; Use a merge operation to integrate all data shards that have passed secondary verification into a complete and unique data set.
6. The method for processing a large amount of data according to claim 5, wherein: The steps of using the Δ-convergence criterion to perform iterative calculations on the unique data set to obtain the final analysis result are as follows: Use an initialization module to make an initial setting for the unique data set; The initial setting includes determining the initial parameter vector, setting the maximum number of iterations, and the convergence threshold; Quantify the error between each data point in the unique data set using the objective function and its corresponding predicted value ; Update the parameter vector using the gradient descent method until the maximum number of iterations is reached; Use the Δ-convergence criterion to judge whether the iteration should stop, and calculate the change in the objective function value between two adjacent iterations; Input the validation set into the model, compare the difference between the model output and the actual value to evaluate the model performance, and obtain the final analysis result.
7. The mass data processing method according to claim 6, wherein: The steps of using columnar storage to write the analysis result into a distributed file platform are as follows: Perform format conversion on the final analysis result to make it meet the requirements of columnar storage; The format conversion includes numerical type conversion, field sorting, and compression processing to obtain an analysis result that meets the columnar storage standard; Use the Apache Parquet columnar storage format to encode the converted analysis result, define the data type, compression algorithm, and encoding method of each column, and implement data compression and fast query support to obtain the analysis result encoded in the columnar storage format; Use the batch loading technology to batch write the encoded analysis result into the distributed file platform, set an appropriate block size of 256MB and the number of concurrent write threads to optimize the write performance, and obtain the analysis result successfully stored in the distributed file platform.
8. A large-scale data processing system, based on the large-scale data processing method according to any one of claims 1 to 7, characterized in that: Including: Data acquisition module, data sharding module, deduplication processing module, data analysis module, and data storage module; The data acquisition module is used to perform real-time subscription and reception of data from different sources using Apache Kafka, and ensure an efficient data transmission mechanism by configuring the Producer and Consumer API of Kafka, realize the preliminary collation of multi-source heterogeneous data and the time window strategy screening, and finally export the qualified original data stream to external storage; The data sharding module is used to preprocess the original data stream obtained from the Kafka cluster, calculate the information entropy of each data block and determine the weight, and apply the dynamic programming algorithm to optimally split the data according to the weight to obtain evenly distributed data shards; The duplicate removal processing module is used to estimate the cardinality of the elements in the data shards by using the HyperLogLog++ algorithm, and perform preprocessing and duplicate removal operations on these elements through a Bloom filter, and combine the advantages of the two technologies to confirm the true state of the suspected duplicate elements, so as to obtain a unique data set; The data analysis module is used to initialize the analysis parameters, define the objective function to quantify the error, and use the gradient descent method to optimize the model parameters until the Δ-convergence criterion is met. Finally, the model performance is evaluated through the validation set to ensure accurate and reliable analysis results are output; The data storage module is used to convert the final analysis result into a format suitable for columnar storage, encode it using the Apache Parquet columnar storage format, and then use the bulk loading technology to efficiently write it into the distributed file platform for subsequent query and analysis.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the large-scale data processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the large-scale data processing method according to any one of claims 1 to 7.
Citation Information
Cited By
Heterogeneous energy data processing method and device and storage medium
CN120821734A