Information acquisition system based on computer big data
By designing a computer big data information collection system with multiple modules, the problem of obtaining useful data from massive data and ensuring data security is solved, efficient and automated data collection and processing is achieved, and the accuracy, completeness and security of the data are ensured.
Patent Information
- Application Number
- CN202510140916.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology is difficult to effectively obtain data useful to enterprises from massive data, especially when indexing and collecting related information, and the security and traceability of data have not been effectively solved.
A computer-based big data information acquisition system is designed, including a data acquisition module, a data fusion module, a data storage and management module, a data processing and analysis module, a data analysis and mining module, and a data security and privacy protection module. The system uses random forest algorithm to clean and fusion data, uses encryption technology to ensure data security, and automatically performs data acquisition tasks through task scheduling tools.
It realizes efficient and automated data collection and processing, ensuring the accuracy and integrity of the data, and at the same time, through encryption and access control and other means, the security and privacy of the data are guaranteed.
Smart Images

Figure CN120086293A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data collection and processing, and specifically to a computer-based big data information collection system. Background Technique
[0002] Data is a form of expression of facts, concepts or instructions. The basic purpose of data processing is to extract and derive valuable and meaningful data for certain specific people from a large amount of data that may be chaotic and difficult to understand. Data collection and processing aims to collect, store, process and analyze large-scale data in order to extract valuable information and insights from it. It is widely used in various industries, including enterprise management, scientific research, marketing, healthcare, etc. Its core goal is to help people make informed decisions, optimize business processes and discover patterns and trends hidden in the data through effective data collection, processing and analysis.
[0003] Among them, a computer-based big data information collection system is a computer system specifically designed for collecting and processing large-scale data. Its purpose is to collect information from multiple data sources and integrate it into a centralized data warehouse for subsequent analysis and application. Its effect is to provide real-time, efficient and accurate data collection and processing to support data-driven decision-making and business optimization.
[0004] In the prior art, how to obtain useful data for enterprises and the like from a vast amount of data has become a problem to be solved. When indexing is required, how to index and collect relevant and required information from a vast amount of data is also a problem to be faced. In addition, the security and traceability of data are also a problem. Without effective protection measures, data is easily tampered with or illegally obtained.
[0005] Therefore, those skilled in the art have provided a computer-based big data information collection system to solve the problems raised in the above background technique. Summary of the Invention
[0006] (1) Technical Problems to be Solved
[0007] In view of the deficiencies of the prior art, the present invention provides a computer-based big data information collection system, which solves the problem of how to obtain useful data for enterprises and the like from a vast amount of data. When indexing is required, how to index and collect relevant and required information from a vast amount of data is also a problem to be faced.
[0008] (2) Technical Solutions
[0009] To achieve the above objectives, the present invention is realized through the following technical solutions:
[0010] A computer-based big data information collection system, including a data collection module, a data fusion module, a data storage and management module, a data processing and analysis module, a data analysis and mining module, and a data security and privacy protection module;
[0011] The data collection module is mainly responsible for collecting raw data from various data sources. It is the core component of the system and is responsible for ensuring that data can be collected and transmitted in an efficient and accurate manner;
[0012] The data fusion module uses the random forest algorithm to perform data cleaning and fusion processing to generate fused data information;
[0013] The data storage and management module is mainly responsible for storing, organizing, managing, and optimizing the data collected from various data sources to ensure that the data can be stored efficiently and securely and for subsequent processing;
[0014] The data processing and analysis module is mainly responsible for extracting, cleaning, transforming, and analyzing data from the storage system, ultimately providing valuable information for decision-making, and ensuring that the data analysis process runs efficiently and reliably;
[0015] The data analysis and mining module is mainly responsible for analyzing and processing the collected data in a big data environment, extracting valuable information and knowledge from it, and providing data-driven support for decision-makers by efficiently processing and analyzing massive amounts of data;
[0016] The data security and privacy protection module is mainly used to ensure the confidentiality, integrity, and availability of data during the processes of big data collection, storage, processing, analysis, and transmission, while protecting data security and user privacy. Through means such as encryption, access control, data masking, backup and recovery, etc., data security risks can be effectively reduced and compliance can be ensured.
[0017] Further, the data collection module includes a sensor and device unit, an Internet data unit, a database, a file and document unit, and a data interface;
[0018] The sensor and device unit includes Internet of Things devices, sensors, and smart hardware, which are used to collect various types of data, such as temperature, humidity, location, and motion data;
[0019] The Internet data unit includes behavioral data generated by social media, website access, e-commerce platforms, etc., such as logs, clickstreams, and comments;
[0020] The database is used to extract stored data from various databases, such as relational databases and NoSQL databases;
[0021] The file and document unit stores static data files in the formats of Excel, CSV files, PDF documents, etc.;
[0022] The data interface includes connection interfaces between different systems, such as API, database interface, and data stream, through which data from other systems can be collected.
[0023] Furthermore, the data fusion module includes a data cleaning submodule, a time series analysis submodule, and a machine learning submodule;
[0024] The data cleaning submodule uses a random forest algorithm to detect and correct outliers in the original data and generate a cleaned data set;
[0025] The time series analysis submodule uses time series analysis based on the cleaned data set to mine the trend and periodicity characteristics therein and generate time series analysis results;
[0026] The machine learning submodule uses a support vector machine based on the time series analysis results to refine the data fusion processing and generate fused data information.
[0027] Further, the data storage and management module includes a distributed storage system, a cloud storage unit and a database storage unit;
[0028] The distributed storage system is used to ensure the distribution, load balancing and high availability of data among multiple nodes;
[0029] The cloud storage unit is used to store data in the cloud to achieve elastic expansion and high availability. Common cloud storage services include Amazon S3, Google Cloud Storage, and Azure Blob Storage.
[0030] The database storage unit is used to store structured data, and usually uses a relational database (such as MySQL, PostgreSQL) or a non-relational database (such as MongoDB, Cassandra) to meet data storage requirements;
[0031] The data storage types of the data storage and management module include:
[0032] Batch storage: usually used to store large-scale historical data or batch imported data, such as file system storage and distributed file storage;
[0033] Real-time storage: used to store data generated in real time, suitable for the Internet of Things and sensor data, usually using streaming data storage systems such as Apache Kafka and Apache Pulsar;
[0034] Time-series data storage: For data that needs to be stored and queried in chronological order, it is particularly suitable for monitoring and IoT scenarios. Time-series databases such as InfluxDB and OpenTSDB.
[0035] Furthermore, the data processing and analysis module includes a data cleaning and preprocessing unit, a data transformation and feature engineering unit, a data analysis and modeling unit, a real-time data processing and stream data analysis unit, a data visualization unit, and a result output and report generation unit;
[0036] The data cleaning and preprocessing unit specifically includes:
[0037] Duplicate removal: Remove duplicate records to ensure data uniqueness;
[0038] Missing value handling: For missing data, common methods include deleting missing values, filling missing values (such as mean, median, interpolation, etc.), or using other machine learning algorithms for filling;
[0039] Outlier detection: Identify and process abnormal data through statistical methods or machine learning models;
[0040] Data standardization and normalization: Standardize (such as z-score standardization) or normalize (such as min-max normalization) data with different dimensions or ranges so that the data can be compared on the same scale;
[0041] Data format conversion: Unify data formats, such as date format conversion and text encoding standardization;
[0042] Data sampling: For massive data, perform data sampling to reduce computational overhead or meet specific analysis requirements.
[0043] The data transformation and feature engineering unit specifically includes:
[0044] Feature selection: Select the most relevant features from a large number of original features to reduce dimensionality and improve model performance;
[0045] Feature extraction: Transform the original data into new features, especially in text, image, or time-series data processing. For example, extract text features through TF-IDF and use PCA (Principal Component Analysis) for dimensionality reduction;
[0046] Feature construction: Construct new features based on domain knowledge. For example, extract information such as hours and weeks from timestamp data, or combine multiple fields into a new feature;
[0047] Data integration: Integrate data from different data sources to ensure data consistency and integrity.
[0048] The data analysis and modeling unit specifically includes:
[0049] Statistical analysis: Descriptive analysis, inferential analysis, and correlation analysis of data are carried out through statistical methods. Commonly used techniques include mean, variance, correlation coefficient, and regression analysis;
[0050] Reinforcement learning and deep learning: Reinforcement learning can be used to optimize the decision-making strategy of the model through reward and punishment mechanisms. For complex data types (such as images, speech, natural language processing), deep learning (such as convolutional neural network CNN, recurrent neural network RNN, transformer model, etc.) can effectively extract high-level features in the data and make predictions;
[0051] Data mining: Data mining techniques are used to discover potential patterns and rules from large amounts of data. Commonly used techniques include association rule mining, sequence pattern mining, and anomaly detection;
[0052] Time series analysis: For data with time series characteristics, time series analysis methods are used for trend prediction and periodic analysis.
[0053] Furthermore, the real-time data processing and stream data analysis unit specifically includes:
[0054] Streaming processing: Stream data analysis mainly relies on stream processing engines such as Apache Kafka, Apache Flink, Apache Storm, and Google Dataflow. These tools can process massive real-time data streams and provide low-latency data processing capabilities;
[0055] Real-time aggregation and window analysis: In the process of real-time data processing, it is often necessary to perform real-time aggregation and window analysis on the data for real-time reporting or early warning;
[0056] Event-driven architecture: By designing an event-driven system, the system can respond to specific events and trigger corresponding processing flows;
[0057] The data visualization unit helps to present complex data and analysis results in an intuitive way, facilitating business decision-makers to understand and analyze. Specifically, it includes:
[0058] Charts and dashboards: Such as bar charts, line charts, pie charts, heat maps, and tree maps;
[0059] Geographic Information System: For the analysis of geospatial data, the Geographic Information System can display and analyze the data through maps;
[0060] Interactive visualization: Provide an interactive visualization interface where users can explore the data through click, drag, and zoom operations.
[0061] After data processing and analysis, the result output and report generation unit outputs the final result and generates a report, specifically including:
[0062] Reports and dashboards: Automatically generate data reports and display them in real time through dashboards;
[0063] Export data: Export the analysis results in CSV, Excel, and PDF formats for use by other systems or users;
[0064] Automated reports: Use automated tools to regularly generate data analysis reports and notify relevant personnel via email or message push.
[0065] Furthermore, the data analysis and mining module specifically includes a data preprocessing sub-module, a data storage and management sub-module, a data analysis sub-module, and a data mining sub-module;
[0066] The data preprocessing sub-module is used to clean and transform the original data to ensure data quality. By removing redundant data, repairing missing values, handling outliers, and merging, matching, and integrating data from different sources, and then performing normalization, standardization, and normalization on the data for subsequent analysis;
[0067] The data storage and management sub-module is used to efficiently store and manage big data;
[0068] The data analysis sub-module extracts patterns, trends, and regularities in the data through various technical means, summarizes and statistically analyzes the data, such as mean, variance, and frequency distribution, to help understand the basic situation of the data, and explores the internal characteristics of the data through visualization means (such as histograms, box plots, and scatter plots);
[0069] The data mining sub-module is used to discover potential patterns, trends, and association rules from massive data and is realized through a series of algorithms and models. For example, use machine learning algorithms (such as decision trees, support vector machines, random forests, etc.) to classify the data and identify the characteristics of different categories; predict the continuous variables of the data through regression models; detect abnormal points or outliers in the data that do not conform to the expected patterns.
[0070] Furthermore, the data security and privacy protection module specifically includes a data encryption sub-module, an access control sub-module, a data desensitization and anonymization sub-module, a data segmentation and isolation sub-module, and a data backup and recovery sub-module;
[0071] The data encryption sub-module prevents the data from being illegally accessed during storage and transmission by encrypting the data, specifically including:
[0072] Transmission Encryption: The data transmission process is encrypted using protocols such as TLS / SSL to ensure that data is not stolen during network transmission;
[0073] Storage Encryption: Symmetric or asymmetric encryption technologies are used during data storage to ensure that even if the data storage medium is illegally accessed, the data content cannot be read;
[0074] End-to-End Encryption: Encryption is performed between the data generation end and the usage end to ensure that the data remains encrypted at any stage during transmission.
[0075] The access control sub-module restricts data access permissions by setting strict access control mechanisms to prevent unauthorized personnel from obtaining sensitive information, specifically including:
[0076] Identity Authentication: Visitors are authenticated using means such as username / password, biometric recognition, and two-factor authentication to ensure that only legitimate users can access the data;
[0077] Role Permission Management: Access permissions are assigned according to the roles and responsibilities of users, and only necessary personnel are allowed to access specific data to avoid abuse of permissions;
[0078] Access Log Recording: Detailed logs of each data access are recorded, including the identity of the visitor, time, and type of data accessed, for subsequent auditing and monitoring.
[0079] The data desensitization and anonymization sub-module makes the data not expose personal information by changing or masking sensitive data, removing or replacing the identity information in the data, ensuring that individual identities cannot be identified from the data, so as to perform data analysis without revealing privacy;
[0080] The data segmentation and isolation sub-module reduces the risk of data leakage in a big data environment by segmenting and isolating sensitive data. The sensitive data is split and stored in different storage media or databases. Even if a part is leaked, the sensitive information cannot be fully recovered, and data at different levels is isolated and managed to ensure that sensitive data and ordinary data are strictly distinguished in storage and access;
[0081] The data backup and recovery sub-module can perform regular backups according to the importance and update frequency of the data to ensure rapid recovery in case of data loss or damage. At the same time, the backup data also needs to be encrypted for storage to prevent unauthorized personnel from obtaining the backup data. It can also formulate a perfect disaster recovery strategy to ensure rapid recovery of data and services after the system crashes or is attacked.
[0082] (III) Beneficial Effects
[0083] The present invention provides a computer-based big data information collection system, which has the following beneficial effects:
[0084] 1. The present invention provides a computer-based big data information collection system. This system can integrate data from different types of data sources, such as databases, file systems, sensors, and web pages. It can converge multi-channel information in one system, avoiding the cumbersome process of collecting data separately in multiple different systems. Moreover, by using a task scheduling tool, it can automatically execute data collection tasks. The system can automatically start the collection process according to preset time, event trigger conditions, etc., greatly improving the efficiency of data collection.
[0085] 2. The present invention provides a computer-based big data information collection system. The data cleaning technology in the system can effectively handle data quality problems. Duplicate data removal can remove the duplicate data generated during the collection process, avoiding deviation in data analysis. Outlier handling can identify and correct or delete data that does not conform to business logic, ensuring that the data conforms to the actual situation. The data filling function can guarantee the integrity of the data, providing a complete data set for subsequent data analysis.
[0086] 3. The present invention provides a computer-based big data information collection system. It adopts data encryption technology to encrypt sensitive data during data transmission and storage processes, preventing the data from being stolen during transmission or illegally accessed during storage, and protecting the privacy information of enterprises and users. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] Figure 1 is the flowchart of the computer-based big data information collection system of the present invention;
[0088] Figure 2 is the flowchart of the data collection module of the present invention;
[0089] Figure 3 is the flowchart of the data fusion module of the present invention;
[0090] Figure 4 is the flowchart of the data storage and management module of the present invention;
[0091] Figure 5 is the flowchart of the data processing and analysis module of the present invention;
[0092] Figure 6 is the flowchart of the data analysis and mining module of the present invention;
[0093] Figure 7 is the flowchart of the data security and privacy protection module of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0094] The following will clearly and completely describe the technical solutions in the specific embodiments of the present invention in conjunction with the accompanying drawings in the specific embodiments of the present invention. Obviously, the described specific embodiments are only a part of the specific embodiments of the present invention, rather than all of the specific embodiments. All other specific embodiments obtained by those of ordinary skill in the art based on the specific embodiments of the present invention without creative efforts belong to the scope of protection of the present invention. Specific embodiments:
[0096] As Figures 1-7 shown, the specific embodiments of the present invention provide a computer big data information collection system,
[0097] including a data collection module, a data fusion module, a data storage and management module, a data processing and analysis module, a data analysis and mining module, and a data security and privacy protection module;
[0098] The data collection module is mainly responsible for collecting raw data from various data sources. It is the core component of the system and is responsible for ensuring that data can be collected and transmitted in an efficient and accurate manner;
[0099] The data collection module includes a sensor and device unit, an Internet data unit, a database, a file and document unit, and a data interface;
[0100] The sensor and device unit includes Internet of Things devices, sensors, and smart hardware, which are used to collect various types of data, such as temperature, humidity, location, and motion data;
[0101] The Internet data unit includes behavioral data generated by social media, website access, e-commerce platforms, etc., such as logs, clickstreams, and comments;
[0102] The database is used to extract stored data from various types of databases, such as relational databases and NoSQL databases;
[0103] The file and document unit stores static data files in formats such as Excel, CSV files, and PDF documents;
[0104] The data interface includes connection interfaces between different systems, such as APIs, database interfaces, and data streams. Through these interfaces, data from other systems can be collected.
[0105] The data fusion module uses the random forest algorithm to perform data cleaning and fusion processing to generate fused data information;
[0106] The data fusion module includes a data cleaning sub-module, a time series analysis sub-module, and a machine learning sub-module;
[0107] The data cleaning sub-module uses the random forest algorithm to detect and correct outliers in the original data, generating a cleaned data set;
[0108] The time series analysis sub-module, based on the cleaned data set, uses time series analysis to mine the trends and periodic characteristics therein, generating time series analysis results;
[0109] The machine learning sub-module, based on the time series analysis results, uses support vector machines for refined data fusion processing, generating fused data information.
[0110] The data storage and management module is mainly responsible for storing, organizing, managing, and optimizing the data collected from various data sources to ensure that the data can be stored efficiently and securely for subsequent processing;
[0111] The data storage and management module includes a distributed storage system, a cloud storage unit, and a database storage unit;
[0112] The distributed storage system is used to ensure the distribution, load balancing, and high availability of data among multiple nodes;
[0113] The cloud storage unit is used to store data in the cloud to achieve elastic expansion and high availability. Common cloud storage services include Amazon S3, Google Cloud Storage, and Azure Blob Storage;
[0114] The database storage unit is used to store structured data. Relational databases (such as MySQL and PostgreSQL) or non-relational databases (such as MongoDB and Cassandra) are usually used to meet the data storage requirements;
[0115] The data storage types of the data storage and management module include:
[0116] Batch storage: Usually used to store large-scale historical data or batch-imported data, such as file system storage and distributed file storage;
[0117] Real-time storage: Used to store real-time generated data, suitable for Internet of Things and sensor data. Stream data storage systems such as Apache Kafka and Apache Pulsar are usually adopted;
[0118] Time series data storage: For data that needs to be stored and queried in chronological order, especially suitable for monitoring and IoT scenarios. Time series databases such as InfluxDB and OpenTSDB.
[0119] The data processing and analysis module is mainly responsible for extracting, cleaning, transforming, and analyzing data from the storage system, ultimately providing valuable information for decision-making while ensuring the efficient and reliable operation of the data analysis process;
[0120] The data processing and analysis module includes a data cleaning and preprocessing unit, a data transformation and feature engineering unit, a data analysis and modeling unit, a real-time data processing and stream data analysis unit, a data visualization unit, and a result output and report generation unit;
[0121] The data cleaning and preprocessing unit specifically includes:
[0122] Duplicate removal: Removing duplicate records to ensure data uniqueness;
[0123] Missing value handling: For missing data, common methods include deleting missing values, filling missing values (such as mean, median, interpolation, etc.), or using other machine learning algorithms for filling;
[0124] Outlier detection: Identifying and handling abnormal data through statistical methods or machine learning models;
[0125] Data standardization and normalization: Standardizing (such as z-score standardization) or normalizing (such as min-max normalization) data with different dimensions or ranges so that the data can be compared on the same scale;
[0126] Data format conversion: Unifying data formats, such as date format conversion and text encoding standardization;
[0127] Data sampling: For massive data, data sampling is performed to reduce computational overhead or meet specific analysis requirements.
[0128] The data transformation and feature engineering unit specifically includes:
[0129] Feature selection: Selecting the most relevant features from a large number of original features to reduce dimensionality and improve model performance;
[0130] Feature extraction: Transforming the original data into new features, especially in text, image, or time series data processing. For example, text features are extracted through TF-IDF, and dimensionality reduction is performed using PCA (Principal Component Analysis);
[0131] Feature construction: Constructing new features based on domain knowledge. For example, information such as hours and weeks is extracted from timestamp data, or a new feature is combined from multiple fields;
[0132] Data integration: Integrating data from different data sources to ensure data consistency and integrity.
[0133] The data analysis and modeling unit specifically includes:
[0134] Statistical analysis: Conduct descriptive analysis, inferential analysis, and correlation analysis on data through statistical methods. Commonly used techniques include mean, variance, correlation coefficient, and regression analysis;
[0135] Reinforcement learning and deep learning: Reinforcement learning can be used to optimize the decision-making strategy of the model through reward and punishment mechanisms. For complex data types (such as images, speech, natural language processing), deep learning (such as convolutional neural network CNN, recurrent neural network RNN, transformer model, etc.) can effectively extract high-level features in the data and make predictions;
[0136] Data mining: Use data mining techniques to discover potential patterns and rules from a large amount of data. Commonly used techniques include association rule mining, sequence pattern mining, and anomaly detection;
[0137] Time series analysis: For data with time series characteristics, use time series analysis methods for trend prediction and cycle analysis;
[0138] The real-time data processing and stream data analysis unit specifically includes:
[0139] Stream processing: Stream data analysis mainly relies on stream processing engines such as Apache Kafka, Apache Flink, Apache Storm, and Google Dataflow. These tools can process massive real-time data streams and provide low-latency data processing capabilities;
[0140] Real-time aggregation and window analysis: In the process of real-time data processing, it is often necessary to perform real-time aggregation and window analysis on data for real-time reporting or warning;
[0141] Event-driven architecture: By designing an event-driven system, the system can respond to specific events and trigger corresponding processing flows;
[0142] The data visualization unit helps to present complex data and analysis results in an intuitive way, facilitating business decision-makers to understand and analyze. Specifically, it includes:
[0143] Charts and dashboards: Such as bar charts, line charts, pie charts, heat maps, and tree maps;
[0144] Geographic Information System: For the analysis of geospatial data, the Geographic Information System can display and analyze the data through maps;
[0145] Interactive visualization: Provide an interactive visualization interface, and users can explore the data through click, drag, and zoom operations.
[0146] After data processing and analysis, the result output and report generation unit outputs the final results and generates reports, specifically including:
[0147] Reports and dashboards: Automatically generate data reports and display them in real time through dashboards;
[0148] Export data: Export the analysis results in CSV, Excel, and PDF formats for use by other systems or users;
[0149] Automated reports: Use automated tools to regularly generate data analysis reports and notify relevant personnel via email or message push.
[0150] The data analysis and mining module is mainly responsible for analyzing and processing the collected data in a big data environment to extract valuable information and knowledge, and providing data-driven support for decision-makers by efficiently processing and analyzing massive data;
[0151] The data analysis and mining module specifically includes a data preprocessing sub-module, a data storage and management sub-module, a data analysis sub-module, and a data mining sub-module;
[0152] The data preprocessing sub-module is used to clean and transform the original data to ensure data quality. By removing redundant data, repairing missing values, handling outliers, and merging, matching, and integrating data from different sources, and then normalizing, standardizing, and normalizing the data for subsequent analysis;
[0153] The data storage and management sub-module is used to efficiently store and manage big data;
[0154] The data analysis sub-module extracts patterns, trends, and regularities in the data through various technical means, summarizes and statistically analyzes the data, such as mean, variance, frequency distribution, to help understand the basic situation of the data, and explores the internal characteristics of the data through visualization means (such as histograms, box plots, scatter plots);
[0155] The data mining sub-module is used to discover potential patterns, trends, and association rules from massive data and is achieved through a series of algorithms and models. For example, use machine learning algorithms (such as decision trees, support vector machines, random forests, etc.) to classify the data and identify the characteristics of different categories; predict the continuous variables of the data through regression models; detect abnormal points or outliers in the data that do not conform to the expected patterns.
[0156] The data security and privacy protection module is mainly used to ensure the confidentiality, integrity and availability of data during the processes of big data collection, storage, processing, analysis and transmission, while protecting data security and user privacy. By means of encryption, access control, data masking, backup and recovery, etc., it can effectively reduce data security risks and ensure compliance;
[0157] The data security and privacy protection module specifically includes a data encryption sub-module, an access control sub-module, a data masking and anonymization sub-module, a data segmentation and isolation sub-module, and a data backup and recovery sub-module;
[0158] The data encryption sub-module prevents data from being illegally accessed during storage and transmission by encrypting the data. Specifically, it includes:
[0159] Transmission encryption: Encrypt the data transmission process using protocols such as TLS / SSL to ensure that data is not stolen during network transmission;
[0160] Storage encryption: Use symmetric encryption or asymmetric encryption technology during data storage to ensure that even if the data storage medium is illegally accessed, the data content cannot be read;
[0161] End-to-end encryption: Encrypt between the data generation end and the usage end to ensure that the data remains encrypted at any link during transmission.
[0162] The access control sub-module restricts data access rights by setting strict access control mechanisms to prevent unauthorized personnel from obtaining sensitive information. Specifically, it includes:
[0163] Identity authentication: Authenticate visitors using means such as username / password, biometric recognition, and two-factor authentication to ensure that only legitimate users can access the data;
[0164] Role-based permission management: Allocate access permissions according to the roles and responsibilities of users, and only allow necessary personnel to access specific data to avoid permission abuse;
[0165] Access log recording: Record detailed logs of each data access, including the identity of the visitor, time, and type of data accessed, for subsequent auditing and monitoring.
[0166] The data masking and anonymization sub-module makes the data not expose personal information by changing or masking sensitive data, removes or replaces the identity information in the data, and ensures that individual identities cannot be identified from the data, so as to perform data analysis without leaking privacy;
[0167] The data segmentation and isolation sub-module reduces the risk of data leakage in the big data environment by segmenting and isolating sensitive data. It splits and stores sensitive data in different storage media or databases. Even if a part is leaked, the sensitive information cannot be fully recovered. It also manages the isolation of different levels of data to ensure that sensitive data and ordinary data are strictly distinguished in storage and access.
[0168] The data backup and recovery sub-module can perform regular backups according to the importance and update frequency of the data to ensure rapid recovery in case of data loss or damage. At the same time, the backup data also needs to be encrypted for storage to prevent unauthorized access to the backup data. It can also formulate a complete disaster recovery strategy to ensure rapid recovery of data and services after the system crashes or is attacked.
[0169] In the above embodiments of the present application, the system can integrate data from different types of data sources, such as databases, file systems, sensors, and web pages. It can converge multi-channel information in one system, avoiding the cumbersome process of collecting data separately in multiple different systems. And by using a task scheduling tool, it can automatically execute data collection tasks. The system can automatically start the collection process according to pre-set time, event trigger conditions, etc., greatly improving the efficiency of data collection.
[0170] The data cleaning technology in the system can effectively handle data quality problems. Duplicate data removal can remove duplicate data generated during the collection process to avoid biases in data analysis. Outlier handling can identify and correct or delete data that does not conform to business logic to ensure that the data conforms to the actual situation. The data filling function can ensure the integrity of the data, providing a complete data set for subsequent data analysis. And data encryption technology is adopted to encrypt sensitive data during data transmission and storage to prevent data from being stolen during transmission or illegally accessed during storage, protecting the privacy information of enterprises and users.
[0171] Although the specific embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these specific embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A computer-based big data information collection system, characterized by: It includes data acquisition module, data fusion module, data storage and management module, data processing and analysis module, data analysis and mining module, and data security and privacy protection module; The data acquisition module is mainly responsible for collecting raw data from various data sources. It is the core component of the system and is responsible for ensuring that data can be collected and transmitted in an efficient and accurate manner; The data fusion module uses a random forest algorithm to perform data cleaning and fusion processing to generate fused data information; The data storage and management module is mainly responsible for storing, organizing, managing and optimizing the data collected from various data sources to ensure that the data can be stored efficiently and securely and subsequently processed; The data processing and analysis module is mainly responsible for extracting, cleaning, converting and analyzing data from the storage system, ultimately providing valuable information for decision-making, while ensuring that the data analysis process runs efficiently and reliably; The data analysis and mining module is mainly responsible for extracting valuable information and knowledge from the collected data by analyzing and processing it in a big data environment, and providing data-driven support for decision makers by efficiently processing and analyzing massive data; The data security and privacy protection module is mainly used to ensure the confidentiality, integrity and availability of data during the collection, storage, processing, analysis and transmission of big data, while protecting data security and user privacy. Through encryption, access control, data desensitization, backup and recovery and other means, it can effectively reduce data security risks and ensure compliance.
2. The computer-based big data information collection system according to claim 1, characterized in that: The data acquisition module includes a sensor and device unit, an Internet data unit, a database, a file and document unit and a data interface; The sensor and device unit includes IoT devices, sensors, and smart hardware, which are used to collect various types of data, such as temperature, humidity, location, and motion data; The Internet data unit includes behavioral data generated by social media, website visits, e-commerce platforms, etc., such as logs, click streams, and comments; The database is used to extract stored data from various databases, such as relational databases and NoSQL databases; The file and document unit stores static data files in the formats of Excel, CSV files, PDF documents, etc.; The data interface includes connection interfaces between different systems, such as API, database interface, and data stream, through which data from other systems can be collected.
3. The computer-based big data information collection system according to claim 1, characterized in that: The data fusion module includes a data cleaning submodule, a time series analysis submodule, and a machine learning submodule; The data cleaning submodule uses a random forest algorithm to detect and correct outliers in the original data and generate a cleaned data set; The time series analysis submodule uses time series analysis based on the cleaned data set to mine the trend and periodicity characteristics therein and generate time series analysis results; The machine learning submodule uses a support vector machine based on the time series analysis results to refine the data fusion processing and generate fused data information.
4. The computer-based big data information collection system according to claim 1, characterized in that: The data storage and management module includes a distributed storage system, a cloud storage unit and a database storage unit; The distributed storage system is used to ensure the distribution, load balancing and high availability of data among multiple nodes; The cloud storage unit is used to store data in the cloud to achieve elastic expansion and high availability. Common cloud storage services include Amazon S3, Google Cloud Storage, and Azure Blob Storage. The database storage unit is used to store structured data, and usually uses a relational database or a non-relational database to meet the data storage requirements; The data storage types of the data storage and management module include: Batch storage: usually used to store large-scale historical data or batch imported data, such as file system storage and distributed file storage; Real-time storage: used to store data generated in real time, suitable for the Internet of Things and sensor data, usually using streaming data storage systems such as Apache Kafka and Apache Pulsar; Time series data storage: For data that needs to be stored and queried in chronological order, it is particularly suitable for monitoring and IoT scenarios. Time series databases such as InfluxDB and OpenTSDB.
5. The computer-based big data information collection system according to claim 1, characterized in that: The data processing and analysis module includes a data cleaning and preprocessing unit, a data conversion and feature engineering unit, a data analysis and modeling unit, a real-time data processing and streaming data analysis unit, a data visualization unit, and a result output and report generation unit; The data cleaning and preprocessing unit specifically includes: Deduplication: remove duplicate records to ensure data uniqueness; Missing value processing: For missing data, common methods include deleting missing values, filling missing values, or using other machine learning algorithms to fill them; Outlier detection: identifying and processing abnormal data through statistical methods or machine learning models; Data standardization and normalization: standardize or normalize data of different dimensions or ranges so that the data can be compared on the same scale; Data format conversion: unify data formats, such as date format conversion and text encoding standardization; Data sampling: For massive data, data sampling is performed to reduce computing overhead or meet specific analysis requirements. The data conversion and feature engineering unit specifically includes: Feature selection: Select the most relevant features from a large number of original features to reduce the dimension and improve the performance of the model; Feature extraction: converting raw data into new features, especially in text, image or time series data processing, for example, extracting text features through TF-IDF and using PCA principal component analysis for dimensionality reduction; Feature construction: construct new features based on domain knowledge, such as extracting information such as hours and days of the week from timestamp data, or combining multiple fields into a new feature; Data integration: Integrate data from different data sources to ensure data consistency and integrity. The data analysis and modeling unit specifically includes: Statistical analysis: Use statistical methods to conduct descriptive analysis, inferential analysis, and correlation analysis on data. Commonly used techniques include mean, variance, correlation coefficient, and regression analysis; Reinforcement learning and deep learning: Reinforcement learning can be used to optimize the decision-making strategy of the model through reward and penalty mechanisms. For complex data types, deep learning can effectively extract high-level features in the data and make predictions; Data mining: using data mining techniques to discover potential patterns and rules from large amounts of data. Commonly used techniques include association rule mining, sequence pattern mining, and anomaly detection; Time series analysis: For data with time series characteristics, use time series analysis methods to perform trend forecasting and cycle analysis.
6. The computer-based big data information collection system according to claim 1, characterized in that: The real-time data processing and stream data analysis unit specifically includes: Stream processing: Stream data analysis mainly relies on stream processing engines, such as Apache Kafka, Apache Flink, ApacheStorm, and Google Dataflow. These tools can process massive real-time data streams and provide low-latency data processing capabilities; Real-time aggregation and window analysis: In the process of real-time data processing, it is often necessary to perform real-time aggregation and window analysis on the data for real-time reporting or early warning; Event-driven architecture: By designing an event-driven system, the system can respond to specific events and trigger corresponding processing processes; The data visualization unit helps to present complex data and analysis results in an intuitive way, which is convenient for business decision makers to understand and analyze, including: Charts and dashboards: such as bar charts, line charts, pie charts, heat maps, and tree maps; Geographic Information System: For analysis involving geospatial data, GIS can display and analyze data through maps; Interactive visualization: Provides an interactive visualization interface where users can explore data by clicking, dragging, and zooming. After data processing and analysis, the result output and report generation unit outputs the final result and generates a report, which specifically includes: Reports and dashboards: Automatically generate data reports and display them in real time through dashboards; Export data: export analysis results to CSV, Excel, or PDF format for use by other systems or users; Automated reporting: Use automated tools to regularly generate data analysis reports and notify relevant personnel via email or push notifications.
7. The computer-based big data information collection system according to claim 1, characterized in that: The data analysis and mining module specifically includes a data preprocessing submodule, a data storage and management submodule, a data analysis submodule, and a data mining submodule; The data preprocessing submodule is used to clean and convert raw data to ensure data quality by removing redundant data, repairing missing values, processing outliers, merging, matching and integrating data from different sources, and then normalizing, standardizing and normalizing the data for subsequent analysis; The data storage and management submodule is used to efficiently store and manage big data; The data analysis submodule uses a variety of technical means to extract patterns, trends and rules in the data, summarizes the data, such as mean, variance, frequency distribution, to help understand the basic situation of the data, and explores the inherent characteristics of the data through visualization; The data mining submodule is used to discover potential patterns, trends, and association rules from massive data, and achieves this through a series of algorithms and models, such as using machine learning algorithms to classify data and identify features of different categories; predicting continuous variables of data through regression models; and detecting abnormal points or outliers in the data that do not conform to the expected pattern.
8. The computer-based big data information collection system according to claim 1, characterized in that: The data security and privacy protection module specifically includes a data encryption submodule, an access control submodule, a data desensitization and anonymization submodule, a data segmentation and isolation submodule, and a data backup and recovery submodule; The data encryption submodule encrypts data to prevent illegal access to data during storage and transmission, specifically including: Transmission encryption: Use protocols such as TLS / SSL to encrypt the data transmission process to ensure that the data cannot be stolen when it is transmitted over the network; Storage encryption: Symmetric or asymmetric encryption technology is used during data storage to ensure that the data content cannot be read even if the data storage medium is illegally accessed; End-to-end encryption: Encryption is performed between the data generator and the data user to ensure that the data remains encrypted at any stage during transmission. The access control submodule sets up a strict access control mechanism to limit access rights to data and prevent unauthorized personnel from obtaining sensitive information, including: Identity authentication: Use methods such as username / password, biometrics, and two-factor authentication to authenticate visitors and ensure that only legitimate users can access data; Role-based permission management: assign access rights based on user roles and responsibilities, allowing only necessary personnel to access specific data to avoid abuse of permissions; Access log recording: Record detailed logs of each data access, including visitor identity, time, and data type accessed, for subsequent auditing and monitoring. The data desensitization and anonymization submodule changes or masks sensitive data so that the data does not reveal personal information, removes or replaces identity information in the data, and ensures that individual identities cannot be identified from the data, so that data analysis can be performed without leaking privacy; The data segmentation and isolation submodule reduces the risk of data leakage in a big data environment by segmenting and isolating sensitive data. The sensitive data is split and stored in different storage media or databases. Even if a part is leaked, the sensitive information cannot be fully restored. Different levels of data are isolated and managed to ensure that sensitive data is strictly distinguished from ordinary data in terms of storage and access. The data backup and recovery submodule can perform regular backups based on the importance and update frequency of the data to ensure that it can be quickly restored when the data is lost or damaged. At the same time, the backup data also needs to be stored in encrypted form to prevent the backup data from being obtained by unauthorized personnel. A complete disaster recovery strategy can also be formulated to ensure that data and services can be quickly restored after a system crash or attack.
Citation Information
Cited By
Intelligent analysis method, system and equipment for performance data of terminal equipment and medium
CN120880945A
Big data information collecting and processing method and system
CN121217756A
Big data information collection processing method and system
CN121217756B