Big data processing method and system and storage medium
Through a big data processing method, including data type identification, preliminary cleaning, data quality evaluation, outlier value and noise repair, multi-source data fusion and deduplication, the problems of low data quality and low processing efficiency in science and technology big data processing are solved, and efficient and unified data processing and storage are achieved.
Patent Information
- Application Number
- CN202510543549.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The existing technology is difficult to effectively clean science and technology big data from multiple data sources, resulting in low data quality, low processing efficiency, and a lack of a unified processing framework.
A big data processing method is adopted, including data type identification, preliminary cleaning, data quality evaluation, outlier and noise repair, multi-source data fusion and deduplication, and the processing results are stored in a designated storage medium.
Through automated cleaning and multi-source data fusion technology, the quality and accuracy of science and technology innovation big data have been significantly improved, efficient and unified processing and storage of data have been achieved, manual intervention has been reduced, and data cleaning costs have been reduced.
Smart Images

Figure CN120068003A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data processing, and particularly to a big data processing method, system and storage medium. Background Art
[0002] With the rapid development of big data technology, especially in the fields of scientific research, technological innovation, enterprise management, etc., the scale and complexity of scientific and technological innovation big data have increased exponentially. However, these data often come from different data sources, including experimental data, sensor data, literature data, historical data, etc. These data not only differ in terms of structure, format, timeliness, etc., but are also often affected by problems such as noise, missing values, outliers, etc., resulting in low data quality and seriously affecting the accuracy and efficiency of subsequent analysis and applications.
[0003] Traditional data processing methods usually rely on manual intervention or rule-based cleaning processes. This method has problems such as low data processing efficiency, poor cleaning effect, and high labor costs. In addition, since the data comes from different sources, how to effectively integrate the information in multiple data sources and remove redundant information to ensure the unity and consistency of the data is also a major challenge faced by current technologies. Therefore, how to improve the quality and accuracy of scientific and technological innovation big data through automated cleaning and multi-source data fusion technologies has become a technical problem to be solved urgently.
[0004] Most existing solutions are limited to the cleaning of a single data source, lack a unified processing framework for multi-source data, and do not fully consider the relevance and heterogeneity between data during the data cleaning process. This makes them limited in the face of large-scale and high-dimensional scientific and technological innovation data. Therefore, how to efficiently clean the data from multiple data sources based on automated algorithms and achieve effective fusion and deduplication of the data while ensuring data quality has become a key technical problem in improving the processing ability of scientific and technological innovation big data. Summary of the Invention
[0005] The present invention provides a big data processing method, system and storage medium to solve the problem of how to improve the quality and accuracy of scientific and technological innovation big data through automated cleaning and multi-source data fusion technologies, ensure that the data from different data sources can be processed and stored efficiently and uniformly, and provide reliable basic data for subsequent data analysis, decision support and intelligent applications.
[0006] To solve the above technical problems, the present invention provides a big data processing method, including: Obtaining original data and classifying it based on a data type recognition module to obtain a data type label sequence; Based on the data type label sequence, performing preliminary cleaning on the original data, including denoising and filling missing values, to generate a preliminarily cleaned data set; Based on the preliminarily cleaned dataset, a multi-dimensional data quality assessment model is used to evaluate the integrity, accuracy, and consistency of the data, generating a data quality score sequence; according to the data quality score sequence, outliers and noise in the data are identified, and the outliers and noise are repaired to generate a repaired dataset after quality assessment; Based on the repaired dataset, multi-source data fusion and deduplication are performed to remove duplicate data and integrate the cleaning results of different data sources, generating a fused dataset after deduplication; The fused dataset after deduplication is stored in a specified storage medium and provided through a data interface for subsequent data analysis and applications.
[0007] Further, the step of performing data quality assessment based on the preliminarily cleaned dataset includes: Evaluating the integrity of data items through a missing rate calculation formula and calculating the missing rate of each piece of data ; Based on the standard score Z and Kullback-Leibler divergence Evaluating the accuracy of the data; Using weighted cosine similarity and a time alignment error model to evaluate the consistency between data items and generating a consistency score.
[0008] Further, the step of generating the data quality score sequence includes: Combining the evaluation results of integrity, accuracy, and consistency, generating a quality score for each piece of data based on a weighted calculation formula ; Through weight coefficients , , Determining the contribution ratio of integrity, accuracy, and consistency to the data quality score and generating a data quality score sequence .
[0009] Further, the step of obtaining the original data includes: Obtaining original data from multiple data sources, where the multiple data sources include sensor data, historical data, experimental data, and literature data.
[0010] Further, the step of automatically classifying based on the data type recognition module includes: Automatically identifying and labeling the types of original data through preset classification rules or machine learning algorithms, generating a label sequence corresponding to the data types.
[0011] Further, the preliminary cleaning step includes: Process the original data using a denoising algorithm to remove noise data, where the noise data are data with abnormal fluctuations and deviations from the normal pattern. Use an imputation algorithm to fill in the missing values in the data, where the imputation algorithm includes interpolation or a machine learning-based missing value prediction method.
[0012] Further, the data quality assessment step includes: Evaluate the preliminarily cleaned data set based on a statistical model or a data distribution model, identify outliers, noise, and inconsistent data in the data, and generate a data quality report.
[0013] Further, the deduplicated and fused data set stored in the specified storage medium includes: The data that has been cleaned and deduplicated, where the data set contains multi-source data such as sensor data, historical data, and experimental data, and stores metadata related to the data, and the metadata includes data source, timestamp, and cleaning status.
[0014] Further, the efficient data output interface includes: Provide a standardized API interface that supports output in multiple data formats, where the data formats include CSV, JSON, and XML formats.
[0015] Further, a scientific and technological innovation big data processing system based on automated cleaning, which is applied to the big data processing method described in any one of the above, includes: A data acquisition module (10) that obtains data in the field of scientific and technological innovation from multiple data sources.
[0016] A data cleaning module (20) that preprocesses the collected original data.
[0017] A multi-source data fusion and deduplication module (30) that merges the cleaned data from different data sources and removes duplicate data items.
[0018] A data storage module (40) that stores the deduplicated and fused data set in a specified storage medium.
[0019] A data output and interface module (50) that provides an efficient data output interface.
[0020] A data analysis and decision support module (60) that performs in-depth analysis based on the deduplicated and fused data set and generates corresponding reports or decision support information.
[0021] A system monitoring and adaptive module (70) that continuously monitors the running status of the entire system and optimizes system parameters through an adaptive algorithm.
[0022] The key innovation points of the present invention include: (1) Automated data cleaning algorithm: Developed an automated data cleaning algorithm based on machine learning and rules, which can intelligently identify and repair noise, missing values, and outliers in big data of science and technology innovation.
[0023] (2) Multi-source heterogeneous data processing: The algorithm is applicable to structured, semi-structured, and unstructured data, and supports the cleaning of various data types such as text, images, audio, and video.
[0024] (3) Efficient data quality improvement: Through parallel computing and optimized algorithms, improve the efficiency and accuracy of data cleaning, and significantly enhance data quality.
[0025] Through the automated cleaning technology and multi-source data fusion algorithm of the present invention, combined with the optimization module in the cloud computing platform, the big data of science and technology innovation is intelligently processed and repaired, effectively solving the problems of poor data quality, low processing efficiency, and scattered data sources existing in traditional data cleaning methods. Compared with traditional data cleaning methods, the present invention can monitor and automatically repair missing values, outliers, and noise in data in real time, ensuring the integrity, accuracy, and consistency of data, thereby providing reliable basic data for subsequent data analysis and intelligent decision-making. At the same time, the optimized data set can be continuously updated through the cloud platform to form a closed-loop feedback mechanism, further improving the data processing efficiency, reducing manual intervention, and lowering the data cleaning cost. This method not only significantly improves data quality and processing accuracy, but also realizes the intelligence, automation, and high efficiency of the data processing process, promoting the intelligent application and precise analysis of big data of science and technology innovation. Brief Description of the Drawings
[0026] Figure 1 It is a schematic flow chart of a big data processing method provided by an embodiment of the present application; Figure 2 It is a structural block diagram of a big data processing system provided by an embodiment of the present application. Detailed Embodiments
[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above description of the drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and are not used to describe a specific order.
[0028] References to "embodiments" in this specification mean that the particular features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0029] Embodiment 1: Refer to Figure 1 , which is a schematic flowchart of a big data processing method provided by an embodiment of the present invention. This process can at least include steps S100 - S500: S100. Obtain the original data and automatically classify it based on the data type recognition module to obtain a data type tag sequence; S200. Based on the data type tag sequence, perform preliminary cleaning on the original data and generate a preliminary cleaned data set; S300. Based on the preliminary cleaned data set, perform data quality assessment, identify and repair outliers and noise, and obtain a repaired data set after quality assessment; S400. Based on the repaired data set, perform multi - source data fusion and deduplication, remove duplicate data and integrate the cleaning results of different data sources to obtain a deduplicated fusion data set; S500. Store the deduplicated fusion data set in a specified storage medium and provide an efficient data output interface for subsequent data analysis.
[0030] Step S100 at least includes steps S110 - S130: S110: Obtain the original data and perform preliminary processing to obtain an initial data set.
[0031] First, the system obtains the original data from multiple data sources. Assume the data source set is , where can be a database, sensor, document, web crawler, image acquisition device, or audio / video input device, etc.
[0032] Data formatting: All the original data collected from the data sources undergoes unified formatting processing and is converted into a standard data structure. For example: Text data is converted into a standard text format (such as UTF - 8 encoding), Image data is converted into a digital image matrix, Audio data is converted into a time - domain signal or frequency - domain feature, Video data is converted into an image sequence or a feature vector of video frames, etc.
[0033] Generate the initial data set: Assume the initial data set is , where represents each original data item, . The initial dataset contains all the original data obtained from different data sources. Each data item can be one of structured data, semi-structured data, or unstructured data. This dataset will serve as the basis for subsequent analysis and processing.
[0034] S120: Analyze the initial dataset based on the data type recognition algorithm to obtain a data type label sequence.
[0035] First, the system uses a trained machine learning model (such as a deep neural network, support vector machine, etc.) to classify and identify each data item . This algorithm determines the type of the data item by analyzing the characteristics of the data item (such as word frequency of text, pixel values of images, spectral characteristics of audio, etc.).
[0036] Furthermore, assume that each original data is identified as a certain data type , and the data type can be one of the following:
[0037]
[0038]
[0039]
[0040] Furthermore, after classifying each data item in the preliminary dataset , a data type label sequence is obtained, where is the type label corresponding to the data item .
[0041] For each data item , extract the feature vector through the classification model , and then use the model to predict its data type label :
[0042] Among them, represents classifying the feature vector through the trained classification model to obtain the type label of the data item .
[0043] The result of this step is a sequence of data type tags , which will be used as the input for subsequent data cleaning and processing.
[0044] S130: Post-process the sequence of data type tags to generate the final classification result.
[0045] First, the system first performs a consistency check to ensure that the tags for each data item are accurate. If it is found that the tags for some data items are inconsistent or ambiguous, the system will correct them through the following methods: Furthermore, if it is found that the tags for some data items are ambiguous tags such as "image / video", the system will further analyze the characteristics of the data items to confirm their specific types. For example, by analyzing the pixel characteristics of the data item or the frame rate of the video, to clarify whether the data item is an image or a video.
[0046] For those ambiguous tags, such as "audio / video", the system refines the tags through further feature analysis (such as spectral analysis, video frame feature extraction, etc.) to ensure that the tags for each data item are clear.
[0047] If the system cannot determine the type of a data item during the automatic classification process, or there is a classification error, the system will introduce a manual review mechanism, and a human will review and correct the uncertain data items. This process will ensure the high accuracy of the sequence of data type tags .
[0048] After post-processing, the system finally generates an accurate and cleaned sequence of data type tags . Among them, represents the final data type tag for each data item .
[0049] This sequence of tags will be used as the input for subsequent data cleaning and processing steps to ensure that the data can be correctly classified and cleaned.
[0050] During the post-processing, if there is uncertainty or error in the tags, the system will recalculate the tags of the data items through a refinement or correction process. Assuming that the tag after correction is , then:
[0051] Among them, represents the correction process, which ensures the accuracy of the tags by further analyzing the characteristics of the data items.
[0052] S200 includes at least steps S210 - S230: S210: Denoise the data based on the data type tag sequence.
[0053] According to the final data type tag sequence generated in S130 , the system starts to denoise the data. According to the characteristics of each data type, the system applies different denoising algorithms.
[0054] Text data Denoising: The system uses natural language processing (NLP) techniques to remove noise words such as invalid characters, punctuation marks, stop words, and spelling mistakes in the text. The most representative words in the text can be identified through the TF-IDF algorithm, and then the irrelevant noise words are removed.
[0055] For each piece of text data , its denoised text can be expressed as:
[0056] where represents applying NLP techniques to remove the noise information in the text and return the denoised text.
[0057] Image data Denoising: For image data, the system can use common image denoising algorithms such as median filtering and mean filtering to remove the noise points in the image.
[0058] For each image data , its denoised image can be expressed as:
[0059] where represents applying an image denoising algorithm to process the image data and remove the noise points in the image.
[0060] Audio data (Tᵢ' = audio) denoising: The system removes the background noise in the audio data through techniques such as spectrum analysis and filtering, and retains the effective voice or audio information.
[0061] For each piece of audio data , its denoised audio can be expressed as:
[0062] where represents applying an audio denoising algorithm to process the audio data and remove the background noise in the audio.
[0063] Video data Denoising: For video data, the system uses image denoising techniques combined with time-domain filtering algorithms to remove noise from the video.
[0064] For each segment of video data , its denoised video can be expressed as:
[0065] where represents the application of video denoising algorithms to remove noise from the video.
[0066] Through the above denoising steps, the system generates a dataset with noise removed , where represents the data item after denoising.
[0067] S220: Fill in missing values for the data based on the data type label sequence.
[0068] After completing the denoising process, the system further fills in the missing values in the denoised dataset .
[0069] Text data Missing value filling: For text data, the system uses context-based filling methods, such as word embedding models (e.g., Word2Vec or GloVe), or text generation models (such as GPT-3 or BERT) to infer and fill in missing words or sentence parts.
[0070] Missing value filling for image data (Tᵢ' = image): For image data, the system uses image inpainting algorithms (such as image interpolation, generative adversarial networks (GAN), etc.) to restore the missing parts of the image.
[0071] For each image data , if it contains missing regions, the filled image can be expressed as:
[0072] where represents the use of image inpainting algorithms to restore the missing parts of the image.
[0073] Audio data Missing value filling: For audio data, the system uses algorithms such as time-domain interpolation or spectral interpolation to restore the missing audio parts.
[0074] For each segment of audio data , if it contains missing parts, the filled audio can be expressed as:
[0075] where represents filling the missing parts in the audio through time-domain interpolation or spectral interpolation algorithms.
[0076] Missing value filling for video data (Tᵢ' = video): For video data, the system uses time-frame interpolation or video repair technology to fill the missing frames based on the information of the previous and next frames.
[0077] For each segment of video data , if it contains missing frames, the filled video can be expressed as:
[0078] where represents using time-frame interpolation or video repair technology to fill the missing frames.
[0079] After these missing value fillings, the system generates a preliminary cleaned dataset , which already contains the data after denoising and filling missing values.
[0080] S230: Generate a preliminary cleaned dataset and prepare for subsequent cleaning steps.
[0081] After completing the data filling in S220, the system generates a preliminary cleaned dataset . This dataset has already gone through preliminary cleaning steps such as denoising and filling missing values and is ready for subsequent more complex data cleaning and data fusion operations.
[0082] After that, the system performs a simple quality check on to confirm the integrity of the cleaned data. The check contents include: Validity check: Confirm whether the data conforms to the expected format and range.
[0083] Accuracy check: Confirm whether the data conforms to the actual situation and has no obvious errors.
[0084] Consistency check: Confirm whether the data is consistent among different data items.
[0085] For any abnormal data (such as still having outliers or inconsistent formats), the system will trigger further cleaning operations. These abnormal data will be marked and passed to the subsequent processing stage.
[0086] Finally, the generated As a dataset after preliminary cleaning, it will be transmitted to the subsequent advanced data processing stage. Specifically, the dataset will be used as the input for subsequent steps such as data deduplication and data standardization, preparing for further refined cleaning.
[0087] Step S300 includes at least steps S310 - S330: S310: Conduct data quality assessment based on the dataset after preliminary cleaning.
[0088] First, the system conducts a comprehensive data quality assessment on the dataset after preliminary cleaning starting from three dimensions: data integrity, accuracy, and consistency, and calculating the quality score sequence for each piece of data .
[0089] ① Data integrity assessment. For the dataset , the missing rate of each piece of data is evaluated , and it is compared with the integrity threshold τ of the entire dataset. The formula is as follows:
[0090] where represents the missing rate of data item ; represents the value of data item in the j - th field; m is the total number of fields of the data; 1(⋅) is an indicator function that takes the value 1 when the condition is true and 0 otherwise.
[0091] If Pi > τ, then mark as low - integrity data.
[0092] ② Data accuracy assessment. For numerical data, Z - Score and distribution deviation are jointly used to evaluate data accuracy. The specific formula is as follows:
[0093] where is the standard score, representing the deviation of the data point from the mean; and are the mean and standard deviation of the dataset respectively; is the Kullback - Leibler (KL) divergence, used to measure the difference between the data distribution P(x) and the expected distribution Q(x).
[0094] If | | > 3 or > δ, then mark as abnormal data.
[0095] ③ Data consistency evaluation. For multi-source data, calculate the consistency score between the data , and adopt the joint evaluation formula of weighted cosine similarity and time alignment error:
[0096] where represents the consistency score of data items and ; and are the values of the data item on the k-th field; and are the timestamps of the data item; and are the weight coefficients, is the decay factor of time alignment.
[0097] If < ζ, it is determined that there is a consistency problem with the data and .
[0098] ④ Calculate the quality score sequence. Based on the comprehensive evaluation results of integrity, accuracy, and consistency, calculate the quality score qi of each data item. The formula is as follows:
[0099] where , , are the weight coefficients of each dimension, satisfying + + = 1.
[0100] After the above evaluation, generate the quality score sequence of each data item, which is used as the input basis for subsequent repair and fusion operations.
[0101] S320: Identify and repair outliers.
[0102] In S310, we have obtained the quality scores of the data items. Next, in S320, based on identify and repair outliers. The key to outlier detection is to use and (interquartile range) method.
[0103] Use to detect outliers in numerical data. If , then this data item is regarded as an outlier. In addition, the system also adopts method to detect outliers, and the formula is as follows:
[0104] Among them, and are the first quartile and the third quartile of the data set respectively. If is less than or greater than , it is regarded as an outlier.
[0105] For the detected outliers, the system uses interpolation methods for repair. For example, linear interpolation is used to repair missing numerical data :
[0106] Among them, is the interpolation coefficient, and are valid data adjacent to .
[0107] For time series data, the system uses trend analysis methods to repair outliers. Assuming the time series data is , outliers can be repaired through regression analysis.
[0108] The repaired data is updated to:
[0109] And a quality assessment result after repair is generated:
[0110] S330: Identify and repair noisy data.
[0111] After the outlier repair is completed in S320, the system then processes the noisy data, usually using signal processing techniques to identify and repair the noisy data.
[0112] The system uses frequency domain analysis techniques to detect noise. Taking audio data as an example, the system uses the Fourier transform to perform spectral analysis on the audio signal:
[0113] Among them, is the time domain signal is the frequency, is the frequency domain signal.
[0114] For image data, the system uses wavelet transform to detect noise. The wavelet transform formula is:
[0115] Among them, is the original signal, is a wavelet function.
[0116] For the detected noise, the system repairs it through a denoising algorithm. For example, a denoising convolutional neural network (CNN) is used to repair the image data:
[0117] Among them, is the original noise data, is the data after denoising.
[0118] For audio data, the system uses a filter to smooth and repair the signal, and the output of the filter is:
[0119] Among them, are the filter coefficients, is the input signal.
[0120] Step S400 includes at least steps S410 - S430: S410: Multi-source data fusion.
[0121] In S410, we perform multi-source data fusion based on the repaired dataset D3D_3D3. Assume that in the previous steps, we have obtained the cleaned datasets from multiple data sources respectively, as follows:
[0122]
[0123] The datasets are from different data sources respectively and have completed preliminary data cleaning (denoising, filling in missing values, and repairing outliers). Our goal is to merge them into a unified dataset to provide a basis for the deduplication and integration steps.
[0124] In the first step of multi-source data fusion, we need to align the cleaned data items from different data sources. Data item alignment is usually based on some common features (such as the identifier of the data item, timestamp, or other unique identifiers). For each data item , we need to ensure that they can correspond in the feature space. For example, for time series data, alignment needs to be performed according to time.
[0125] This process can be achieved by merging all data sources to obtain the fused dataset
[0126]
[0127] Among them, Represents the combined result of the datasets after cleaning all data sources.
[0128] The aligned data source set May contain duplicate data items. Therefore, we need to use a feature matching algorithm to determine whether these data items represent the same entity. For example, use cosine similarity to measure the similarity between data items:
[0129] Where and Are the values of the data items and On the Dimension respectively, Is the dimension of the data item. If the similarity exceeds the set threshold, it is considered that these two data items come from the same entity.
[0130] For data items considered to be the same (i.e., items with high similarity), we need to fuse their features into a unified data item. The fusion method can use simple average, weighted average or other data synthesis methods. For example, use weighted average to fuse the values of two data items: Where and Are the weight coefficients, which can usually be determined based on the quality score or credibility of the data item.
[0131] S420: Data deduplication.
[0132] In S420, we perform data deduplication based on the fused dataset The purpose of data deduplication is to ensure that each data item in the final dataset is unique and avoid the interference of redundant data on subsequent processing.
[0133] To identify duplicate data items, we need to compare the similarity of each data item with other data items in the dataset. For example, using the aforementioned cosine similarity metric, if the similarity of two data items exceeds a certain set threshold , then they are considered duplicate data items:
[0134] In this case, the data items and Are considered duplicate. At this time, the system will mark them and prepare for deduplication processing.
[0135] For the identified duplicate data items, the system will select one data item to retain and discard the other duplicates according to certain rules. Usually, we will determine which data item is more trustworthy based on the quality score Q of the data item or other relevant features (such as timestamp). For example, select the data item with a higher quality score as the retained item and discard the other data items.
[0136] The specific operation can be described by the following formula:
[0137] where is the deduplicated data set.
[0138] S430: Integrate the cleaning results.
[0139] In S430, we perform data integration operations based on the deduplicated data set to generate the final integrated data set . The goal of data integration is to ensure that the data items cleaned from multiple data sources can be integrated through certain rules, so as to obtain a comprehensive and accurate data set.
[0140] For each data item after deduplication , the system integrates its features with other relevant data items. The integration method is usually based on the quality score of the data item, the credibility of the source, or other weighting mechanisms. For example, use the weighted average to combine data items from different sources:
[0141] where the function represents the feature integration method, represents the integrated data item.
[0142] The integrated data items will be summarized to finally generate the integrated data set . This data set will be used as the basis for processing in the next stage (such as advanced data analysis or modeling). The integrated data set is a high-quality data set that has been deduplicated and integrates data items from different sources.
[0143]
[0144] Step S500 includes at least steps S510 - S530: S510: Store the data set to the specified storage medium.
[0145] In step S510, first we will store the deduplicated integrated data set Store it in a specified storage medium. The storage medium can be a traditional database system (such as SQL, NoSQL databases) or a distributed storage system (such as Hadoop HDFS, cloud storage). Depending on the scale of the data and future query requirements, the choice of storage medium will affect the storage efficiency and subsequent processing performance.
[0146] It is necessary to format the data set into a format suitable for storage. For example, for structured data, we can use format, which is suitable for efficient storage and query of large-scale data. For unstructured data, format can be used as a flexible storage option.
[0147] For data items , we will serialize their features to ensure that their format meets the requirements of the storage medium:
[0148] Store the formatted data set in the specified storage medium. Assuming the data is stored in a distributed storage system, the storage process can be completed through batch processing or stream processing:
[0149] where represents the selected storage medium.
[0150] S520: Data storage optimization.
[0151] In S520, we optimize the data storage process to ensure the efficient use of the storage medium and provide a fast access path. Data storage optimization is not only for saving storage space but also for efficiently supporting subsequent data analysis tasks.
[0152] First, build an index for the stored data set so that data items can be quickly located during subsequent analysis. The index can be built based on the primary key of the data or other key feature fields, such as timestamps, categories, etc. For each data item , we need to build an index:
[0153] To improve storage efficiency and retrieval performance, the data set can be compressed. Through data compression, the storage space occupied is reduced, and the I / O efficiency during query is improved. In addition, dividing the data set into multiple partitions (based on factors such as time, region, etc.) can also improve the access performance:
[0154]
[0155] S530: Provide an efficient data output interface.
[0156] In S530, we provide an efficient data output interface for the dataset stored in the specified storage medium These interfaces will support subsequent data analysis tasks such as data visualization, machine learning model training, etc. The data interfaces not only need to support common data query languages (such as SQL), but also need to support API interfaces and bulk data transfer functions.
[0157] To facilitate subsequent data access, we provide a set of standardized data query interfaces. For example, use SQL interfaces or NoSQL interfaces to access the data stored in the database. Assume using the SQL interface to query the dataset Its query operation is as follows:
[0158] In addition to the traditional database query method, we also provide RESTful API interfaces for the dataset to support other systems to call data through the API. Assume the data interface is REST API, and the API call can be as follows:
[0159] To support large-scale data output, a bulk data transfer interface is provided. For example, the dataset can be or formatted and bulk output to other systems or analysis tools:
[0160] The key innovation points of the present invention include: (1) Automated data cleaning algorithm: Developed an automated data cleaning algorithm based on machine learning and rules, which can intelligently identify and repair noise, missing values, and outliers in the scientific and technological innovation big data.
[0161] (2) Multi-source heterogeneous data processing: The algorithm is applicable to structured, semi-structured, and unstructured data, and supports the cleaning of various data types such as text, images, audio, and video.
[0162] (3) Efficient data quality improvement: Through parallel computing and optimized algorithms, improve the efficiency and accuracy of data cleaning, and significantly improve data quality.
[0163] Through the automated cleaning technology and multi-source data fusion algorithm, combined with the optimization module in the cloud computing platform, the present invention intelligently processes and repairs the big data in scientific and technological innovation, effectively solving the problems of poor data quality, low processing efficiency, and scattered data sources existing in traditional data cleaning methods. Compared with traditional data cleaning methods, the present invention can monitor and automatically repair missing values, outliers, and noises in the data in real time, ensuring the integrity, accuracy, and consistency of the data, thereby providing reliable basic data for subsequent data analysis and intelligent decision-making. At the same time, the optimized data set can be continuously updated through the cloud platform to form a closed-loop feedback mechanism, further improving the data processing efficiency, reducing manual intervention, and lowering the data cleaning cost. This method not only significantly improves the data quality and processing accuracy but also realizes the intelligentization, automation, and high efficiency of the data processing process, promoting the intelligent application and precise analysis of the big data in scientific and technological innovation.
[0164] Embodiment Two: Figure 2 The structural block diagram of a big data processing system according to an embodiment of the present invention is shown. As Figure 2 shown, the structure may include: The data acquisition module (10) is responsible for obtaining data in the field of scientific and technological innovation from multiple data sources, including but not limited to patent data, scientific research papers, technical reports, experimental data, etc. These data are collected in real time through web crawlers, open APIs, or enterprise internal database interfaces, etc., and the extensiveness and diversity of the data sources are ensured. Specifically, the data acquisition module collects data in different formats, such as text, images, tables, etc., through an automated process, providing rich input data for subsequent data processing.
[0165] The data cleaning module (20) performs preprocessing on the collected raw data, including denoising, filling missing values, and repairing outliers. This module uses advanced automated cleaning algorithms, such as classification models based on machine learning, to identify and repair noises and missing values in the data. The cleaned data can be stored in structured, semi-structured, or unstructured formats to ensure that the data quality meets the requirements of subsequent analysis.
[0166] Denoising processing: Adopt a data denoising algorithm to identify and eliminate invalid or duplicate data items.
[0167] Missing value filling: Fill in the missing data according to different data characteristics using interpolation, regression models, or multiple imputation methods.
[0168] Outlier repair: Identify and repair outliers through rule-based or statistical analysis methods to ensure data consistency.
[0169] The multi-source data fusion and deduplication module (30) is responsible for merging the cleaned data from different data sources and removing duplicate data items. This module uses feature matching algorithms, such as cosine similarity metric, to judge the similarity between data items, identify duplicate entities, and generate unique fused data items through fusion strategies (such as weighted average).
[0170] Data alignment and fusion: Align the data in multiple data sources to ensure that the same entities from different sources can be matched in the feature space.
[0171] Deduplication operation: Compare the similarity of data items, remove duplicate data items whose similarity exceeds the set threshold, and retain the unique data items with higher quality.
[0172] The data storage module (40) is responsible for storing the deduplicated fused data set in the specified storage medium. This module uses an efficient database system to support the efficient access and query of large-scale data. To adapt to future data analysis requirements, the data storage module can provide a flexible storage structure to support incremental updates and backups of data.
[0173] Database support: Provide support for relational databases, NoSQL databases, and big data storage platforms (such as Hadoop, Spark, etc.).
[0174] Efficient query: Adopt indexing technology and optimized storage structure to ensure efficient data retrieval and analysis operations.
[0175] The data output and interface module (50) provides an efficient data output interface for subsequent data analysis and modeling. This module includes functions such as data export, API interface, and data visualization, and can provide clear and accurate data output for users or other systems. The output data can be used for further decision support, algorithm modeling, report generation, etc.
[0176] Data export: Support the export of common data formats such as CSV, JSON, and XML for the convenience of users and other systems.
[0177] API interface: Provide RESTful API interfaces to support the integration with other systems or applications to ensure real-time sharing and invocation of data.
[0178] The data analysis and decision support module (60) is responsible for conducting in-depth analysis based on the deduplicated fused data set and generating corresponding reports or decision support information. This module uses machine learning and data mining technologies to extract valuable patterns and trends from the data and provide scientific and reasonable decision-making basis for users.
[0179] Data Mining and Modeling: Using techniques such as regression analysis, clustering analysis, and classification algorithms to deeply explore the potential patterns in the data.
[0180] Report Generation: Automatically generate decision-making reports based on the analysis results to assist users in making accurate decisions.
[0181] The System Monitoring and Adaptive Module (70) continuously monitors the operating status of the entire system and optimizes system parameters through adaptive algorithms to ensure the system operates efficiently and stably in various complex environments. This module can automatically adjust the data cleaning and fusion strategies to ensure the robustness and efficiency of the data processing process.
[0182] Real-time Monitoring: Monitor the system operating status, detect, and report any potential faults or bottlenecks.
[0183] Adaptive Optimization: Automatically adjust data processing and storage parameters based on real-time feedback data to optimize system performance.
[0184] The following are its main beneficial effects: (1) Automatic Classification and Cleaning: Through automatic classification by the data type recognition module, customized processing can be carried out for different types of data. Compared with the traditional manual intervention method, it greatly reduces human errors and improves the efficiency and accuracy of data cleaning.
[0185] (2) Denoising and Filling Missing Values: In the initial data cleaning stage, automatically remove noise and fill in missing values to ensure the integrity of the subsequent analysis data. Compared with traditional methods, it avoids the deviation caused by manual judgment and improves the integrity and consistency of the data.
[0186] (3) Data Quality Assessment and Repair: Through a multi-level data quality assessment and repair mechanism, it can accurately identify outliers and noise to ensure that the final dataset meets high-quality standards. The automated processing of this process not only improves the accuracy of data repair but also significantly shortens the time for data cleaning.
[0187] (4) Multi-source Data Fusion and Duplicate Removal: In the process of multi-source data fusion and duplicate removal, the present invention can efficiently integrate data from different sources, remove duplicate data, avoid the influence of redundant data on the analysis results, and provide a more accurate and efficient basis for data analysis. The automation and intelligence of this process greatly improve the data processing efficiency and avoid the duplication and conflict problems existing in traditional manual fusion.
[0188] (5) Efficient Storage and Output Interface: By storing the deduplicated dataset in the specified storage medium and providing an efficient data output interface, the present invention not only ensures the stability of data storage but also greatly improves the response speed of subsequent data analysis and intelligent applications.
[0189] An embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program runs on a computer, the method provided in the embodiment of the present application is implemented.
[0190] An embodiment of the present application further provides a chip, which includes a processor for calling and running instructions stored in a memory, so that a communication device equipped with the chip executes the method provided in the embodiment of the present application.
[0191] An embodiment of the present application further provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, the output interface, the processor, and the memory are connected through an internal connection path. The processor is used to execute the code in the memory. When the code is executed, the processor is used to execute the method provided in the embodiment of the application.
[0192] It should be understood that the above-mentioned processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the advanced reduced instruction set machine (ARM) architecture.
[0193] Further, optionally, the above-mentioned memory may include a read-only memory and a random access memory, and may further include a non-volatile random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).
[0194] In the above embodiments, it may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium.
[0195] Obviously, the embodiments described above are only a part of the embodiments of this application, rather than all of them. The preferred embodiments of this application are shown in the accompanying drawings, but they do not limit the patent scope of this application. This application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of this application more thorough and comprehensive. Although this application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions recorded in the foregoing specific embodiments, or perform equivalent replacements on some of the technical features. Any equivalent structure made by using the content of this application's specification and drawings, directly or indirectly applied in other related technical fields, is similarly within the scope of patent protection of this application.
Claims
1. A method for processing big data, characterized in that: The method comprises: Obtain the original data and classify it based on the data type recognition module to obtain a data type label sequence; Based on the data type label sequence, the raw data is preliminarily cleaned, including denoising and filling missing values, to generate a preliminarily cleaned data set; Based on the data set after preliminary cleaning, a multi-dimensional data quality assessment model is used to assess the integrity, accuracy and consistency of the data to generate a data quality score sequence; based on the data quality score sequence, outliers and noise in the data are identified, and the outliers and noise are repaired to generate a repaired data set after quality assessment; Based on the repaired data set, perform multi-source data fusion and deduplication, remove duplicate data and integrate cleaning results from different data sources to generate a deduplicated fusion data set; The deduplicated fused data set is stored in a designated storage medium and provided for subsequent data analysis and application through a data interface.
2. The method according to claim 1, characterized in that: The step of performing data quality assessment based on the preliminarily cleaned data set includes: The completeness of the data items is evaluated through the missing rate calculation formula, and the missing rate of each data is calculated ; Based on standard score Z and Kullback-Leibler divergence Assess the accuracy of the data; The weighted cosine similarity and temporal alignment error model are used to evaluate the consistency between data items and generate a consistency score.
3. The method according to claim 2, characterized in that The step of generating a data quality score sequence comprises: Based on the above completeness, accuracy and consistency assessment results, a quality score for each piece of data is generated based on a weighted calculation formula ; By weight coefficient , , Determine the contribution of completeness, accuracy, and consistency to the data quality score and generate a quality score sequence .
4. The big data processing method according to claim 1, characterized in that: The step of obtaining original data comprises: The raw data is acquired from a plurality of data sources, including sensor data, historical data, experimental data, and literature data.
5. The big data processing method according to claim 1, characterized in that: The step of automatically classifying based on the data type identification module includes: Through preset classification rules or machine learning algorithms, the type of raw data is automatically identified and labeled, and a label sequence corresponding to the data type is generated.
6. The big data processing method according to claim 1, characterized in that: The preliminary cleaning step comprises: Using a denoising algorithm to process the raw data to remove noise data, wherein the noise data is data with abnormal fluctuations and deviations from a normal pattern; Fill missing values in the data using an imputation algorithm, which includes interpolation or machine learning-based missing value prediction methods.
7. The big data processing method according to claim 1, characterized in that: The data quality assessment step includes: Based on statistical models or data distribution models, the data set after preliminary cleaning is evaluated to identify outliers, noise and inconsistent data in the data, and generate a data quality report.
8. The big data processing method according to claim 1, characterized in that: The deduplicated fused data set stored in the designated storage medium includes: The cleaned and deduplicated data includes sensor data, historical data, experimental data and other multi-source data, and stores metadata related to the data, including data source, timestamp and cleaning status.
9. The big data processing method according to claim 1, characterized in that: The efficient data output interface comprises: Provides a standardized API interface and supports multiple data format outputs, including CSV, JSON, and XML formats.
10. A big data processing system, applied to the big data processing method according to any one of claims 1 to 9, characterized in that: include: A data collection module (10) acquires data in the field of science and technology innovation from multiple data sources; A data cleaning module (20) pre-processes the collected raw data; A multi-source data fusion and deduplication module (30) merges cleaned data from different data sources and removes duplicate data items; A data storage module (40) stores the deduplicated fused data set in a designated storage medium; A data output and interface module (50) provides an efficient data output interface; A data analysis and decision support module (60) performs in-depth analysis based on the deduplicated fused data set and generates corresponding reports or decision support information; The system monitoring and adaptive module (70) continuously monitors the operating status of the entire system and optimizes system parameters through an adaptive algorithm.
Citation Information
Patent Citations
Multi-sensor fusion method and system based on multi-dimensional attribute correlation analysis
CN113761705A
Dam safety assessment method based on mass monitoring inspection information fusion model
CN117574321A
Comprehensive quality evaluation method and device for multi-mode forgery generated data
CN117591815A
Data conversion improvement method
CN118673006A
Urban and rural planning surveying and mapping data efficient processing method and surveying and mapping system
CN118734171A
Cited By
Processing method and system for collecting, cleaning and labeling culture big data and medium
CN120296078A
Multi-source heterogeneous bus data parallel acquisition and recording method and system
CN120561187A
Multi-source heterogeneous bus data parallel acquisition recording method and system thereof
CN120561187B
Cultural big data calculation cleaning method and system and medium
CN120804068A