Method for constructing large-scale data set based on multi-source data aggregation
Through standardized interfaces, hashing algorithms, automated cleaning and distributed computing frameworks, the heterogeneity and efficiency problems in multi-source data integration are solved, and efficient and flexible large-scale data processing and storage are achieved to meet the needs of real-time analysis.
Patent Information
- Application Number
- CN202510398561.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-08-05
AI Technical Summary
The prior art has problems such as heterogeneity, uneven data quality, low processing efficiency, insufficient system flexibility and scalability in multi-source data integration, which is difficult to meet the real-time computing and analysis needs of large-scale data.
Standardized interface technology is used to access multi-source data, deduplication is performed through hashing algorithms and Bloom filters, missing values are processed in combination with mean filling and interpolation, noise is removed using automated data cleaning tools and rules engines, data merging is used using hash connection and partitioning operations, parallel calculation is performed in combination with distributed computing framework, and data storage is optimized using efficient distributed databases and columnar storage formats.
It realizes efficient access and standardization of multi-source data, improves data accuracy and completeness, reduces computing delays, improves data processing efficiency and system scalability, and supports real-time decision-making analysis and different business needs.
Smart Images

Figure CN120429346A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of constructing large-scale data sets, and in particular, relates to a method for constructing large-scale data sets based on multi-source data aggregation. Background Art
[0002] In the era of big data, data sources are becoming increasingly diverse. Building high-quality, large-scale datasets has become crucial for data analysis, business decision-making, and machine learning model training. However, existing technologies for integrating multi-source data still face numerous challenges, primarily due to the heterogeneity of data sources, uneven data quality, low processing efficiency, and insufficient system flexibility and scalability.
[0003] First, there are significant differences in data formats, storage structures, data models, and semantics among different data sources. For example, there are structured data in relational databases, semi-structured data in log files, and unstructured data such as social media, images, and videos. This heterogeneity makes traditional ETL (Extract-Transform-Load) tools inefficient in the data collection, transformation, and integration process, making it difficult to ensure data consistency and availability.
[0004] Secondly, multi-source data often contains a large amount of noise, missing values, redundant data, and erroneous data, such as unusual fluctuations in sensor data, incorrectly formatted user input, and incomplete crawled data. While existing data cleaning technologies can remove noise, fill missing values, and standardize data formats to a certain extent, they still struggle to comprehensively and automatically improve data quality. This is especially true in large-scale data scenarios. Traditional data cleaning methods often require manual rule adjustments or rely on the expertise of specific domain experts, lacking versatility and intelligent processing capabilities.
[0005] Furthermore, faced with the ever-increasing volume of data, existing data processing tools are facing increasingly prominent bottlenecks in computing and storage. Traditional centralized data processing models struggle to meet the real-time computing and analysis needs of large-scale data. While distributed computing frameworks (such as Hadoop and Spark) have improved processing capabilities to a certain extent, they still have limitations when it comes to complex data cleaning, aggregation, and storage optimization. These limitations include issues like inefficient computing resource scheduling, data transmission overhead, and data query latency. These issues result in low overall processing efficiency, making it difficult to support the high-concurrency demands of businesses and real-time decision-making and analysis.
[0006] Finally, existing data integration solutions lack flexibility and scalability, making it difficult for the systems to adapt to dynamically changing business needs and growing data volumes. Many traditional solutions rely on fixed data models and predefined processing flows. When new data sources are added, data formats change, or data volumes rapidly expand, these solutions often require extensive manual adjustments or even the reconstruction of the entire data processing flow, significantly increasing maintenance costs and technical complexity.
[0007] In view of this, the present invention is proposed. Summary of the Invention
[0008] The technical problem to be solved by the present invention is to overcome the deficiencies of the prior art and provide a method for constructing a large-scale data set based on multi-source data aggregation, thereby solving the problems raised in the above-mentioned background technology.
[0009] In order to solve the above technical problems, the basic concept of the technical solution adopted by the present invention is:
[0010] A method for constructing a large-scale dataset based on multi-source data aggregation, comprising the following steps:
[0011] S1: Access multiple data sources through standardized interface technology, and use message queues for real-time data collection and transmission for real-time data sources;
[0012] S2: Perform format conversion, fill in missing values, remove duplicate data, and perform preliminary data normalization on the incoming data. Use data conversion tools to convert data in different formats into a unified format.
[0013] S3: Use hashing algorithms and Bloom filters to deduplicate data, and use mean filling and interpolation to handle missing values.
[0014] S4: Use automated data cleaning tools and rule engines to remove noise, verify data formats, standardize formats, and handle outliers.
[0015] S5: Based on business needs and data types, different aggregation methods are used for data integration, and distributed computing frameworks are used to perform parallel computing on large-scale data.
[0016] S6: Uses an efficient distributed database to store cleaned and aggregated data, and combines columnar storage format and compression technology to optimize data storage structure, improving storage efficiency and query performance.
[0017] S7: Provides RESTful API through API gateway, supporting data query, analysis and export.
[0018] Optionally, perform format conversion, fill in missing values, remove duplicate data, and perform preliminary data normalization on the received data. The steps for converting data in different formats into a unified format using data conversion tools are as follows:
[0019] Use data conversion tools to convert data in different formats into a unified format standard , where the conversion process is expressed as: F standard =T(F source ), where F source Represents the original data format, T is the format conversion function;
[0020] Deduplication is performed based on the hash algorithm to remove redundant data. Let the data set be D = {d1, d2, ..., d n}, the hash deduplication process is expressed as: Here, H(x) represents a hash function. If the hash values of two data are the same, they are considered duplicate data and need to be further compared, confirmed, and removed.
[0021] Optionally, use a hash algorithm and Bloom filter to deduplicate data, and combine mean filling and interpolation to handle missing values. The steps are as follows:
[0022] For missing values X miss , using mean filling and interpolation method for processing. Mean filling is expressed as: Among them, X i is the complete data sample, n is the total number of data samples;
[0023] Interpolation filling adopts linear interpolation, and its expression is: Among them, t miss is the time corresponding to the missing value, X i-1 and X i+1 are the known data points adjacent to the missing values;
[0024] Normalize the numerical data and convert it to the [0,1] interval using the Min-Max normalization method: Among them, X min and X max Represents the minimum and maximum values of the data respectively.
[0025] Optionally, the steps of using automated data cleaning tools and rule engines to remove noise, verify data formats, standardize formats, and process outliers on the data are:
[0026] In order to deal with random errors and noise in the data, the sliding average method is used for smoothing;
[0027] Based on the rule engine, data format verification rules are defined to verify the format of input data. Suppose the data format rule set is R = {r1, r2, ..., r m}, for data D = {d1, d2, ..., d n}, the verification process is expressed as: like It is marked as abnormal data and format converted;
[0028] Use regularization methods to unify the format of text data, remove special characters, convert uppercase and lowercase characters; unify the time format T standard YYYY-MM-DD:T standard =f(T source ) Where f is a normalization function that converts time data in different formats into a unified format;
[0029] Based on the 3σ rule, outliers are identified and eliminated. For the mean μ and standard deviation σ of the data set X, the 3σ rule defines the outlier threshold as follows: X 异常 ={x|x<μ-3σ or x>μ+3σ}.
[0030] Optionally, use a hash join method to merge data from different data sources and aggregate them based on the common key values in the data. The specific steps are:
[0031] A hash value is generated for each data row through a hash algorithm, and the hash value is used to assign the data to the corresponding partition for association. The mathematical formula for hash join aggregation is: Join Hash(D1, D2) = {(d1, d2) | h(d1) = h(d2)}, where D1 and D2 are two data sets, d1 and d2 are elements in the data sets, and h is a hash function;
[0032] Use partitioning to group data and perform group statistical analysis. Partitioning is done based on certain attributes of the data, dividing the data into multiple subsets, and performing statistical calculations on each subset. The formula for group statistics is as follows: Where D is the data set, A is the basis for grouping (such as a field), and f(d) is the operation on each data item d. It is an aggregation operation on the data in group a;
[0033] Based on the attribute fusion strategy of the latest value, maximum value, and average value, different aggregation methods are selected to fuse the attributes according to the business needs of the data. Subsequently, a distributed computing framework is used to process large-scale data in parallel to improve data processing efficiency. Then, the computing process is accelerated by splitting the data into multiple data blocks and performing aggregation operations on multiple nodes simultaneously.
[0034] Optionally, for numeric data, aggregation strategies include selecting the latest value, maximum value, and average value. Specific strategies can use the following formulas:
[0035] Latest value aggregation: X latest =Latest(D), where D is the data set, X latest Indicates taking the latest value in data set D;
[0036] Maximum Aggregation: X max =max(D) where X max Represents the maximum value in the data set D;
[0037] Average aggregation: Among them, X i is each data item in the data set, n is the total number of data sets, X avg Represents the mean of the data set.
[0038] Optionally, use an efficient distributed database to store cleaned and aggregated data, and combine columnar storage format and compression technology to optimize the data storage structure to improve storage efficiency and query performance. The steps are as follows:
[0039] The data is divided into multiple blocks, each of which is stored in a different node. Then, a columnar storage format is used to optimize the storage structure. Subsequently, compression technology is combined to compress the data.
[0040] After adopting the above technical solution, the present invention has the following beneficial effects compared with the prior art. Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described below at the same time:
[0041] 1. Utilize standardized interface technologies and automated data cleaning tools to achieve efficient access, format conversion, missing value filling, deduplication, and preliminary normalization of multi-source data. Improve the accuracy and efficiency of data deduplication through hashing algorithms and Bloom filters. Optimize missing value processing and effectively enhance data integrity through strategies such as mean imputation and interpolation.
[0042] 2. Use the sliding average method to remove data noise, and combine it with the rule engine to define data format verification rules to ensure that data meets expected standards. Regularization methods are used to standardize the format of text data, and outliers are detected and processed based on the 3σ rule, thereby comprehensively improving data accuracy and reliability, providing a high-quality data foundation for subsequent data analysis and machine learning.
[0043] 3. Hash joins are used to efficiently merge multi-source data, supporting fast key-value-based associative queries and data aggregation. Integrating with distributed computing frameworks (such as Apache Spark and Flink), partitioning and grouping statistical strategies improve the parallel processing capabilities of data computation. Attribute fusion strategies based on the latest value, maximum value, and average value are supported, flexibly adapting to different business needs while reducing computational latency and improving data processing efficiency.
[0044] 4. Use efficient distributed databases (such as Hadoop HDFS, Cassandra, and HBase) to store cleaned and aggregated data, and optimize data structures by combining them with columnar storage formats (such as Parquet and ORC). Use compression technologies like Snappy and Zlib to reduce data storage space and improve query speed. Furthermore, combining index optimization and data partitioning techniques enables efficient retrieval and query of large-scale data, further improving system performance and scalability.
[0045] 5. Through a modular architecture, the system can quickly adapt to new data sources and data input in different formats. Combined with an API gateway, it provides a RESTful API interface, enabling external data query, analysis, and export capabilities, facilitating integration into diverse business systems. Support for distributed deployment and dynamic expansion ensures high performance even as data scale grows, meeting the big data processing needs of diverse scenarios.
[0046] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The drawings described below are only some embodiments. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0048] In the picture:
[0049] Figure 1 It is a structural diagram.
[0050] It should be noted that these drawings and textual descriptions are not intended to limit the conceptual scope of the present invention in any way, but rather to illustrate the concept of the present invention for those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0051] The present invention will now be described in further detail with reference to the accompanying drawings.
[0052] See also Figure 1 As shown, in this embodiment, a method for constructing a large-scale data set based on multi-source data aggregation is provided, including the following steps:
[0053] Access multiple data sources through standardized interface technologies, including API interfaces, JDBC connections, FTP, Web crawling, etc., supporting the access of structured data, semi-structured data, and unstructured data. Message queues (such as Kafka) are used for real-time data collection and transmission for real-time data sources.
[0054] Perform format conversion, fill in missing values, remove duplicate data, and perform preliminary data normalization on the incoming data. Use data conversion tools (such as Apache Nifi, Talend) to convert data in different formats into a unified format;
[0055] Use hashing algorithms and Bloom filters to deduplicate data, and combine mean filling and interpolation methods to handle missing values;
[0056] Use automated data cleaning tools (such as OpenRefine and DataCleaner) and rule engines (such as Drools) to remove noise, verify data formats, standardize formats, and handle outliers.
[0057] Based on business needs and data types, different aggregation methods are used for data integration, including association aggregation based on hash joins, group statistics based on partition operations, and attribute fusion strategies based on latest values, maximum values, and average values. Distributed computing frameworks (such as Apache Spark and Flink) are also used to perform parallel computing on large-scale data.
[0058] Use efficient distributed databases (such as Hadoop HDFS, Cassandra, and HBase) to store cleaned and aggregated data, and combine columnar storage formats (such as Parquet and ORC) and compression technologies (such as Snappy and Zlib) to optimize data storage structure and improve storage efficiency and query performance.
[0059] Provide RESTful API through API gateway (such as Kong, NGINX), supporting data query, analysis and export.
[0060] Standardized interface technologies and automated data cleaning tools are used to achieve efficient access, format conversion, missing value filling, duplication removal, and preliminary normalization of multi-source data. Technologies such as hashing algorithms and Bloom filters are used to improve the accuracy and efficiency of data deduplication. Strategies such as mean imputation and interpolation are combined to optimize missing value processing and effectively improve data integrity.
[0061] We use a sliding average method to remove data noise, and integrate a rule engine to define data format verification rules to ensure that data meets expected standards. We use regularization methods to standardize the format of text data, and detect and process outliers based on the 3σ rule, thereby comprehensively improving data accuracy and reliability, providing a high-quality data foundation for subsequent data analysis and machine learning.
[0062] Hash joins are used to efficiently merge multi-source data, supporting fast key-value-based relational queries and data aggregation. Integrating with distributed computing frameworks (such as Apache Spark and Flink), partitioning and grouping statistical strategies enhance the parallel processing capabilities of data computations. Attribute fusion strategies based on the latest value, maximum value, and average value are supported, flexibly adapting to diverse business needs while reducing computational latency and improving data processing efficiency.
[0063] We use efficient distributed databases (such as Hadoop HDFS, Cassandra, and HBase) to store cleaned and aggregated data, and optimize data structures using columnar storage formats (such as Parquet and ORC). We use compression technologies like Snappy and Zlib to reduce data storage space and improve query speed. Furthermore, we combine index optimization and data partitioning to achieve efficient retrieval and query of large-scale data, further improving system performance and scalability.
[0064] Through modular architecture design, the system can quickly adapt to new data sources and data input in different formats. Combined with the API gateway, it provides a RESTful API interface to achieve external data query, analysis and export capabilities, making it easy to integrate into different business systems. Supports distributed deployment and dynamic expansion to ensure that the system maintains high performance even when the data scale grows, meeting the big data processing needs in different scenarios.
[0065] In this embodiment, the steps of converting the format of the received data, filling missing values, removing duplicate data, and performing preliminary data normalization, and converting the data in different formats into a unified format using a data conversion tool (such as Apache Nifi or Talend) are as follows:
[0066] Use data conversion tools (such as Apache Nifi, Talend) to convert data in different formats (such as CSV, JSON, XML, Parquet) into a unified format. standard , where the conversion process is expressed as: F standard =T(F source ), where F source Represents the original data format, T is the format conversion function;
[0067] Deduplication is performed based on the hash algorithm to remove redundant data. Let the data set be D = {d1, d2, ..., d n}, the hash deduplication process is expressed as: Here, H(x) represents a hash function. If the hash values of two data are the same, they are considered to be duplicate data and need to be further compared, confirmed, and removed.
[0068] In this embodiment, the steps of using the hash algorithm and Bloom filter to deduplicate data and combining mean filling and interpolation to handle missing values are as follows:
[0069] For missing values X miss , using mean filling and interpolation method for processing. Mean filling is expressed as: Among them, X i is the complete data sample, n is the total number of data samples;
[0070] Interpolation filling adopts linear interpolation, and its expression is: Among them, t miss is the time corresponding to the missing value, X i-1 and X i+1 are the known data points adjacent to the missing values;
[0071] Normalize the numerical data and convert it to the [0,1] interval using the Min-Max normalization method: Among them, X min and X max They represent the minimum and maximum values of the data respectively to ensure consistent data distribution across different data sources and improve the availability of data analysis.
[0072] In this embodiment, the steps of using the automated data cleaning tool and the rule engine to remove noise, verify the data format, standardize the format, and process outliers on the data are as follows:
[0073] For random errors and noise in the data, the sliding average method is used for smoothing. For example, for time series data X = {x1, x2, ..., x n}, the expression for sliding average calculation is: Among them, k is the sliding window size, x i ′ is the data point after denoising;
[0074] Based on the rule engine (such as Drools), define data format verification rules to verify the format of input data. Suppose the data format rule set is R = {r1, r2, ..., r m}, for data D = {d1, d2, ..., dn}, the verification process is expressed as: like It is marked as abnormal data and format converted;
[0075] Regularization methods are used to unify the format of text data, remove special characters, and convert uppercase and lowercase letters; for example, the time format T is unified. standard For YYYY-MM-DD:
[0076] T standard =f(T source ) Where f is a normalization function that converts time data in different formats into a unified format;
[0077] Based on the 3σ rule, outliers are identified and eliminated. For the mean μ and standard deviation σ of the data set X, the 3σ rule defines the outlier threshold as follows: X 异常 ={x|x<μ-3σ or x>μ+3σ}, where outliers can be filled or deleted by the median to ensure the rationality of data distribution and improve data quality.
[0078] In this embodiment, a hash join method is used to merge data from different data sources and aggregate them based on the common key values in the data. The specific steps are as follows:
[0079] A hash algorithm generates a hash value for each data row, which is then used to assign the data to the corresponding partition for association. The mathematical formula for a hash join aggregation is: Join Hash(D1, D2) = {(d1, d2) | h(d1) = h(d2)}, where D1 and D2 are two data sets, d1 and d2 are elements in the data sets, and h is a hash function. A hash join determines which data rows should be associated by calculating the hash value h(d).
[0080] Use partitioning to group data and perform group statistical analysis. Partition data based on certain attributes (such as date, region, etc.), divide the data into multiple subsets, and perform statistical calculations (such as sum, average, maximum value, etc.) on each subset. The formula for group statistics is as follows: Where D is the data set, A is the basis for grouping (such as a field), and f(d) is the operation on each data item d (such as sum). It is an aggregation operation on the data in group a;
[0081] Based on the attribute fusion strategy of the latest value, maximum value, and average value, different aggregation methods are selected to fuse attributes according to the business needs of the data;
[0082] A distributed computing framework is used to process large amounts of data in parallel to improve data processing efficiency. The data is then divided into multiple blocks and aggregated simultaneously on multiple nodes to accelerate the computation process. For example, in Apache Spark, the parallel operation of distributed computing can be expressed as follows: Among them, D is the data set, D i For each subset in the dataset, Reduce(D i ) means for each subset D i The final result of the aggregation operation performed is the summary of the aggregation results of all subsets.
[0083] In this embodiment, for numerical data, the aggregation strategy includes selecting the latest value, the maximum value, and the average value. The specific strategy can use the following formula:
[0084] Latest value aggregation: X latest =Latest(D), where D is the data set, X latest Indicates taking the latest value in data set D;
[0085] Maximum Aggregation: X max =max(D) where X max Represents the maximum value in the data set D;
[0086] Average aggregation: Among them, X i is each data item in the data set, n is the total number of data sets, X avg Represents the mean of the data set.
[0087] In this embodiment, the steps of using an efficient distributed database to store cleaned and aggregated data and combining columnar storage format and compression technology to optimize the data storage structure and improve storage efficiency and query performance are as follows:
[0088] Data is divided into multiple blocks, each stored on a different node. A columnar storage format is then used to optimize the storage structure. Compression technology is then used to compress the data. In columnar storage, data is stored by column rather than row. This improves query efficiency for specific columns, especially when only a subset of columns is required. This further optimizes storage space and improves data transmission efficiency. Compression technology reduces data storage space requirements, speeds up data transmission, and reduces storage and bandwidth costs.
[0089] Compression technology optimizes storage space, improving storage and query efficiency. Combining a distributed database with a columnar storage format ensures high scalability and stability. Distributed databases enable horizontal data expansion, supporting large-scale data storage and management. Columnar storage reduces disk I / O, making queries more efficient and ensuring the system maintains high efficiency and stability when processing large amounts of data.
[0090] Glossary
[0091] Multi-source data aggregation: refers to obtaining data from multiple different sources (such as databases, sensors, social media, etc.) and integrating this data into a unified large-scale dataset.
[0092] Data cleaning: refers to the processing of raw data to remove or correct errors, missing values and noise in the data to ensure data quality.
[0093] Data preprocessing: preliminary processing steps performed before data analysis, including data format conversion, normalization, missing value filling, etc.
[0094] Data storage and management: refers to storing processed data in a database and providing efficient query and management functions.
[0095] Distributed database: A database system in which data is distributed across multiple physical nodes and stored and managed over a network, with high scalability and high availability.
[0096] ETL (Extract, Transform, Load): refers to the process of extracting data from the data source, converting the data format, and loading the data into the target database.
[0097] API interface: Application program interface that allows data and functions to be exchanged between different software systems
[0098] The present invention is not limited to the above-described embodiments. Any structural changes made under the guidance of the present invention, which have the same or similar technical solutions as the present invention, should be understood to fall within the scope of protection of the present invention. The technologies, shapes, and structural parts not described in detail in the present invention are all well-known technologies.
Claims
1. A method for constructing a large-scale dataset based on multi-source data aggregation, characterized in that: The following steps are involved: S1: Access multiple data sources through standardized interface technology, and use message queues for real-time data collection and transmission for real-time data sources; S2: Perform format conversion, fill in missing values, remove duplicate data, and perform preliminary data normalization on the incoming data. Use data conversion tools to convert data in different formats into a unified format. S3: Use hashing algorithms and Bloom filters to deduplicate data, and use mean filling and interpolation to handle missing values. S4: Use automated data cleaning tools and rule engines to remove noise, verify data formats, standardize formats, and handle outliers. S5: Based on business needs and data types, different aggregation methods are used for data integration, and distributed computing frameworks are used to perform parallel computing on large-scale data. S6: Uses an efficient distributed database to store cleaned and aggregated data, and combines columnar storage format and compression technology to optimize data storage structure, improving storage efficiency and query performance. S7: Provides RESTful API through API gateway, supporting data query, analysis and export.
2. The method for constructing a large-scale data set based on multi-source data aggregation according to claim 1, characterized in that: The steps for converting the format of the incoming data, filling in missing values, removing duplicate data, and performing preliminary data normalization are as follows: Use data conversion tools to convert data in different formats into a unified format standard , where the conversion process is expressed as: F standard =T(F source ), where F source Represents the original data format, T is the format conversion function; Deduplication is performed based on the hash algorithm to remove redundant data. Let the data set be D = {d1, d2, ..., d n }, the hash deduplication process is expressed as: Here, H(x) represents a hash function. If the hash values of two data are the same, they are considered duplicate data and need to be further compared, confirmed, and removed.
3. The method for constructing a large-scale data set based on multi-source data aggregation according to claim 1, characterized in that: The steps for using hash algorithms and Bloom filters to deduplicate data and combining mean filling and interpolation to handle missing values are as follows: For missing values X miss , the mean filling and interpolation methods are used for processing, and the mean filling is expressed as: Among them, X i is the complete data sample, n is the total number of data samples; Interpolation filling adopts linear interpolation, and its expression is: Among them, t miss is the time corresponding to the missing value, X i-1 and X i+1 are the known data points adjacent to the missing values; Normalize the numerical data and convert it to the [0,1] interval using the Min-Max normalization method: Among them, X min and X max Represents the minimum and maximum values of the data respectively.
4. The method for constructing a large-scale data set based on multi-source data aggregation according to claim 1, characterized in that: The steps to use automated data cleaning tools and rule engines to remove noise, verify data format, standardize format and handle outliers are as follows; In order to deal with random errors and noise in the data, the sliding average method is used for smoothing; Based on the rule engine, data format verification rules are defined to verify the format of input data. Suppose the data format rule set is R={r1, r2, ..., r m }, for data D = {d1, d2, ..., d n }, the verification process is expressed as: The format is correct, if It is marked as abnormal data and format converted; Use regularization methods to unify the format of text data, remove special characters, convert uppercase and lowercase characters; unify the time format T standard YYYY-MM-DD:T standard =f(T source ) Where f is a normalization function that converts time data in different formats into a unified format; Based on the 3σ rule, outliers are identified and eliminated. For the mean μ and standard deviation σ of the data set X, the 3σ rule defines the outlier threshold as follows: X 异常 ={x|x<μ-3σ or x>μ+3σ}.
5. The method for constructing a large-scale data set based on multi-source data aggregation according to claim 1, characterized in that: Use the hash join method to merge data from different data sources and aggregate them based on the common key values in the data. The specific steps are: A hash value is generated for each data row through a hash algorithm, and the hash value is used to assign the data to the corresponding partition for association. The mathematical formula for hash join aggregation is: Join Hash(D1, D2) = {(d1, d2) | h(d1) = h(d2)}, where D1 and D2 are two data sets, d1 and d2 are elements in the data sets, and h is a hash function; Use partitioning to group data and perform group statistical analysis. Partitioning is done based on certain attributes of the data, dividing the data into multiple subsets, and performing statistical calculations on each subset. The formula for group statistics is as follows: Where D is the data set, A is the basis for grouping (such as a field), and f(d) is the operation on each data item d. It is an aggregation operation on the data in group a; Based on the attribute fusion strategy of the latest value, maximum value, and average value, different aggregation methods are selected to fuse the attributes according to the business needs of the data. Subsequently, a distributed computing framework is used to process large-scale data in parallel to improve data processing efficiency. Then, the computing process is accelerated by splitting the data into multiple data blocks and performing aggregation operations on multiple nodes simultaneously.
6. The method for constructing a large-scale data set based on multi-source data aggregation according to claim 1, characterized in that: For numerical data, aggregation strategies include selecting the latest value, maximum value, and average value. Specific strategies can use the following formulas: Latest value aggregation: X latest =Latest(D), where D is the data set, X latest Indicates taking the latest value in data set D; Maximum Aggregation: X max =max(D) where X max Represents the maximum value in the data set D; Average aggregation: Among them, X i is each data item in the data set, n is the total number of data sets, X avg Represents the mean of the data set.
7. The method for constructing a large-scale data set based on multi-source data aggregation according to claim 1, characterized in that: Use an efficient distributed database to store cleaned and aggregated data, and combine columnar storage format and compression technology to optimize the data storage structure. The steps to improve storage efficiency and query performance are as follows: The data is divided into multiple blocks, each of which is stored in a different node. Then, a columnar storage format is used to optimize the storage structure. Subsequently, compression technology is combined to compress the data.