Data processing method of electronic information technology based on big data

By using distributed data acquisition and processing technology, combined with data preprocessing, storage and encryption mechanisms, the problems of data security and resource consumption in the era of big data have been solved, and secure and efficient data processing and value mining of IoT devices have been realized.

CN120951349AInactive Publication Date: 2025-11-14JIAN COLLEGE
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511055017.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-14
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing electronic information technology faces data security risks, privacy leaks, and resource consumption and environmental impact issues in the era of big data. In particular, data is vulnerable to attack and privacy can be stolen in Internet of Things (IoT) devices, and data centers consume a lot of energy, leading to resource shortages and increased environmental pressure.

Method used

Distributed data acquisition tools are used for real-time or batch data extraction. Combined with data preprocessing, distributed storage, encryption and access control, a distributed computing framework is used for feature extraction and analysis to achieve secure data storage and efficient processing. Automated decision-making is driven by a self-service analysis portal and a closed-loop feedback mechanism.

Benefits of technology

It achieves data security, scalability, and efficiency, ensures data integrity and privacy, reduces resource consumption, improves the real-time performance and accuracy of data processing, and supports end-to-end applications and value mining.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120951349A_ABST
    Figure CN120951349A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method of an electronic information technology based on big data, and relates to the technical field of Internet of Things equipment. According to the method, distributed streaming batch integrated processing is taken as a core, a data lake base is constructed through real-time acquisition, dynamic preprocessing and elastic storage, original data is stored in HDFS / HBase according to a partition strategy after standardized cleaning, and data security is guaranteed through encryption and fine-grained access control; on the basis of a storage layer, a Spark / Flink parallel computing framework is utilized to execute feature engineering and model training, and a stream processing module realizes millisecond-level real-time analysis through time window aggregation and complex event pattern recognition; the batch processing module adopts an incremental learning mechanism to optimize historical data mining efficiency; an analysis result is embedded into a service flow through an API or a self-service portal, and automatic decision making is driven; user interaction behavior data are returned to the feature library in real time, and the model precision is continuously iterated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Internet of Things (IoT) device technology, and more specifically to a data processing method based on big data electronic information technology. Background Technology

[0002] Numerous fields, such as finance, healthcare, the internet, and industrial manufacturing, have accumulated massive amounts of data resources. This data contains immense value, but due to its enormous scale, diverse data types, and rapid generation speed, traditional data processing technologies face many challenges. For example, in the era of big data, data is vulnerable to attack; for IoT devices, attackers may modify the data transmitted or stored by these devices.

[0003] Existing electronic information technology has the following main drawbacks in data processing:

[0004] 1) Data security risks: With the advent of the big data era, data has become a core resource, but data breaches and hacker attacks are frequent, threatening personal privacy and corporate security; 2) Privacy leakage risks: Personal information is easily stolen and misused, leading to privacy violations, especially in social media and online services; 3) Resource consumption and environmental impact: Infrastructure such as data centers consumes a large amount of energy, exacerbating resource shortages and ecological pressure. Therefore, we propose a data processing method based on big data electronic information technology. Summary of the Invention

[0005] The purpose of this invention is to provide a data processing method based on big data electronic information technology to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention specifically adopts the following technical solution:

[0007] A data processing method based on big data electronic information technology includes the following steps:

[0008] S1. Collect raw electronic data from multiple heterogeneous data sources and extract data in real time or in batches using distributed data acquisition tools;

[0009] S2. Preprocess the collected data to achieve data cleaning, format conversion, missing value handling, and data integration.

[0010] S3. Store the preprocessed data in a distributed storage system to achieve scalable and high-throughput storage of the data, while applying data partitioning and index optimization strategies to improve query performance.

[0011] S4. Perform real-time or batch analysis on stored data to extract information features;

[0012] S5. Transform the analysis results into structured output and integrate them into business systems or user interfaces to achieve end-to-end application and value mining of data.

[0013] Further, step S1 includes the following steps:

[0014] S11. Identify connection parameters for multiple heterogeneous data sources, including data source type, access protocol, and authentication information;

[0015] S12. Configure the extraction strategy of the distributed data acquisition tool and specify the real-time streaming acquisition or timed batch acquisition mode.

[0016] S13. Perform data extraction operations and handle high-concurrency data inflows through a load balancing mechanism;

[0017] S14. Conduct preliminary quality checks on the extracted data, including data integrity verification and error log recording;

[0018] S15. The extracted raw data is temporarily stored in a buffer, and subsequent preprocessing steps are triggered.

[0019] Further, step S2 includes the following steps:

[0020] S21. Data Cleaning: Identify and handle missing values, fill them with interpolation or rule-based default values, detect and correct outliers, and use statistical methods such as Z-score or interquartile range analysis to ensure data consistency and accuracy.

[0021] S22. Denoising: Apply filter technology to remove random noise, combine contextual information to smooth the data flow, remove outliers, and improve the data signal-to-noise ratio;

[0022] S23. Format Conversion: Unify the data structure, convert heterogeneous data sources into a standardized format, perform data type conversion and unit normalization to facilitate subsequent analysis.

[0023] Furthermore, S3 also includes data encryption and access control mechanisms, including the following steps:

[0024] S31. Static data encryption: Before data is written to the distributed storage node, AES-256 or Chinese cryptographic algorithm is used to transparently encrypt the data block, and the key is managed by an independent hardware security module.

[0025] S32. Dynamic authorization: Role-based access control or attribute-based access control policies, combined with data tag classification to implement fine-grained permission control, and audit logs to record all data access behaviors;

[0026] S33. Data lifecycle management: Automatically executes a tiered storage strategy for hot and cold data, and performs secure erasure or archived encrypted storage for expired or infrequently accessed data.

[0027] Further, step S4 includes the following steps:

[0028] S41. Feature Engineering: Deep feature extraction is performed on data stored in a distributed storage system to construct high-order feature combinations and time window statistics. Feature selection algorithms are applied to reduce dimensionality and eliminate redundancy.

[0029] S42, Model Training and Inference: A distributed computing framework is used to build and train the prediction model, supporting online model updates and A / B testing. For batch processing tasks, an incremental learning mechanism is applied to optimize the model iteration efficiency.

[0030] S43. Streaming Processing: For real-time data streams, deploy window-based aggregation computing and combine it with a complex event processing engine to identify specific patterns, achieving streaming analysis with millisecond-level latency.

[0031] Furthermore, S5 includes the following steps:

[0032] S51, Self-service Analysis Portal: Provides drag-and-drop report building tools, supports multi-dimensional drill-down analysis, cross-filtering and natural language query, and allows users to customize visualization components;

[0033] S52. Prediction Result Integration: Embed the prediction labels or scores generated by S4 into the business system workflow in real time to drive automated decision-making;

[0034] S53. Feedback closed-loop mechanism: Capture user interaction behavior with analysis results and automatically feed it back to the feature library for model optimization.

[0035] Furthermore, it also includes: S6, real-time monitoring of the data processing process, and triggering an alarm mechanism when an anomaly occurs.

[0036] Further, step S6 includes the following steps:

[0037] S61. Performance Metrics Collection: Continuously collect key performance metrics for each processing stage, including data throughput, processing latency, resource utilization, and queue backlog depth.

[0038] S62, Anomaly Detection Engine: It applies unsupervised learning algorithms to establish a baseline for system behavior and automatically identifies abnormal fluctuations that deviate from the baseline;

[0039] S63, Elastic Resource Scheduling: Dynamically scales up or down computing cluster resources based on monitoring indicators and predicted load, and triggers predefined alarm rules to notify the operation and maintenance system when resource bottlenecks or processing failures occur.

[0040] Furthermore, S7, data visualization, is used to display generated interactive reports and charts.

[0041] Further, step S7 includes the following steps:

[0042] S71. Data Mapping: Converting the processing results into the structured data format required for the visualization model, including time series, categorical variables, and numerical indicators;

[0043] S72. Chart Generation Engine: It adopts a front-end framework that integrates ECharts or D3.js libraries to dynamically render interactive charts and supports custom configurations for bar charts, heatmaps and scatter plots.

[0044] S73, User Interaction Module: Allows users to drill down into data, filter time ranges, and switch chart types via a web interface, and updates the displayed content in real time in response to user actions;

[0045] S74. Report Export: Provides a one-click export function to save the generated visualization results as PDF, PNG or CSV format for easy offline analysis and archiving.

[0046] The beneficial effects of this invention are as follows:

[0047] 1. Data-driven architecture: Based on distributed stream and batch processing, a data lake foundation is built through real-time acquisition, dynamic preprocessing and elastic storage. After standardized cleaning, the raw data is stored in HDFS / HBase according to the partitioning strategy, and data security is ensured through encryption and fine-grained access control.

[0048] 2. Intelligent Analysis Engine: Based on the storage layer, the Spark / Flink parallel computing framework is used to perform feature engineering and model training. The stream processing module achieves millisecond-level real-time analysis through time window aggregation and complex event pattern recognition, while the batch processing module adopts an incremental learning mechanism to optimize the efficiency of historical data mining.

[0049] 3. Closed-loop feedback system: Analysis results are embedded into business processes through APIs or self-service portals to drive automated decision-making; user interaction behavior data is fed back to the feature library in real time to continuously iterate the model accuracy. Attached Figure Description

[0050] Figure 1 This is a flowchart of the process of this invention;

[0051] Figure 2 This is a flowchart illustrating the workflow of collecting raw electronic data from multiple heterogeneous data sources in this invention.

[0052] Figure 3 This is a flowchart illustrating the preprocessing process for the collected data in this invention.

[0053] Figure 4 This is a flowchart illustrating the process of storing preprocessed data in a distributed storage system in this invention.

[0054] Figure 5 This is a flowchart of the batch processing analysis workflow in this invention;

[0055] Figure 6 This is a flowchart illustrating the process of converting analysis results into structured output in this invention.

[0056] Figure 7 This is a flowchart of the real-time monitoring data processing process in this invention;

[0057] Figure 8 This is a flowchart of the data visualization process in this invention. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0059] Please see Figure 1 - Figure 8 This invention provides a data processing method based on big data electronic information technology, comprising the following steps:

[0060] S1. Collect raw electronic data from multiple heterogeneous data sources and extract data in real time or in batches using distributed data acquisition tools. Specifically, this includes various data formats such as relational databases, NoSQL databases, API interfaces, and log files. Distributed stream processing is performed using Apache Kafka or Filter, supporting high throughput and low latency transmission to ensure data integrity and consistency and provide reliable input for subsequent processing.

[0061] S2. Preprocess the collected data to perform data cleaning (such as removing outliers, handling duplicate records, and correcting erroneous data), format conversion (converting heterogeneous data into standard JSON or Parquet formats to meet analysis needs), missing value handling (using interpolation, mean filling, or deletion of missing fields to ensure data integrity), and data integration (merging multi-source data for entity parsing, conflict detection, and resolution). Specifically, it utilizes Apache Spark or Flink batch processing frameworks to perform distributed computing, supporting parallel processing of large-scale datasets, optimizing memory usage and computational efficiency, ensuring data quality, consistency, and availability, and providing standardized input for subsequent data mining and modeling.

[0062] S3. The preprocessed data is stored in a distributed storage system to achieve scalable and high-throughput storage. Data partitioning and indexing optimization strategies are applied to improve query performance. Specifically, Hadoop HDFS or Apache HBase is used as the distributed storage system, supporting horizontal scaling to handle massive data growth. Data is distributed through strategies such as hash partitioning and range partitioning, and B-tree or LSM tree indexes are built to optimize query paths. Data compression technologies (such as Snappy or Gzip) are combined to reduce storage overhead, and replication mechanisms (such as HDFS's default three replicas) ensure high availability and fault tolerance, ensuring that the stored data efficiently serves subsequent real-time analysis, batch processing queries, and machine learning tasks.

[0063] S4. Based on the stored data, real-time analysis is performed using stream processing frameworks such as Apache Spark Streaming or Flink, and batch analysis is performed using Spark or MapReduce to extract key information features, including statistical indicators (such as mean and variance), time series features (such as trend and periodicity), association rules (such as frequent itemsets), and classification features (such as label distribution), thereby providing structured input for subsequent data mining, model training, and decision support.

[0064] S5. Transform the analysis results into structured output and integrate them into business systems or user interfaces to achieve end-to-end application and value mining of data.

[0065] In this embodiment, preferably, step S1 includes the following steps:

[0066] S11. Identify connection parameters for multiple heterogeneous data sources, including data source type, access protocol, and authentication information;

[0067] S12. Configure the extraction strategy of the distributed data acquisition tool and specify the real-time streaming acquisition or timed batch acquisition mode.

[0068] S13. Perform data extraction operations and handle high-concurrency data inflows through a load balancing mechanism;

[0069] S14. Conduct preliminary quality checks on the extracted data, including data integrity verification and error log recording;

[0070] S15. The extracted raw data is temporarily stored in a buffer, and subsequent preprocessing steps are triggered.

[0071] In this embodiment, preferably, step S2 includes the following steps:

[0072] S21. Data Cleaning: Identify and handle missing values, fill them with interpolation or rule-based default values, detect and correct outliers, and use statistical methods such as Z-score or interquartile range analysis to ensure data consistency and accuracy.

[0073] S22. Denoising: Apply filter technology to remove random noise, combine contextual information to smooth the data flow, remove outliers, and improve the data signal-to-noise ratio;

[0074] S23. Format Conversion: Unify the data structure, convert heterogeneous data sources into a standardized format, perform data type conversion and unit normalization to facilitate subsequent analysis.

[0075] In this embodiment, preferably, S3 further includes a data encryption and access control mechanism, including the following steps:

[0076] S31. Static data encryption: Before data is written to the distributed storage node, AES-256 or Chinese cryptographic algorithm is used to transparently encrypt the data block, and the key is managed by an independent hardware security module.

[0077] S32. Dynamic authorization: Role-based access control or attribute-based access control policies, combined with data tag classification to implement fine-grained permission control, and audit logs to record all data access behaviors;

[0078] S33. Data lifecycle management: Automatically executes a tiered storage strategy for hot and cold data, and performs secure erasure or archived encrypted storage for expired or infrequently accessed data.

[0079] In this embodiment, preferably, step S4 includes the following steps:

[0080] S41. Feature Engineering: Deep feature extraction is performed on data stored in a distributed storage system to construct high-order feature combinations and time window statistics. Feature selection algorithms are applied to reduce dimensionality and eliminate redundancy.

[0081] S42, Model Training and Inference: A distributed computing framework is used to build and train the prediction model, supporting online model updates and A / B testing. For batch processing tasks, an incremental learning mechanism is applied to optimize the model iteration efficiency.

[0082] S43. Streaming Processing: For real-time data streams, deploy window-based aggregation computing and combine it with a complex event processing engine to identify specific patterns, achieving streaming analysis with millisecond-level latency.

[0083] In this embodiment, preferably, S5 includes the following steps:

[0084] S51, Self-service Analysis Portal: Provides drag-and-drop report building tools, supports multi-dimensional drill-down analysis, cross-filtering and natural language query, and allows users to customize visualization components;

[0085] S52. Prediction Result Integration: Embed the prediction labels or scores generated by S4 into the business system workflow in real time to drive automated decision-making;

[0086] S53. Feedback closed-loop mechanism: Capture user interaction behavior with analysis results and automatically feed it back to the feature library for model optimization.

[0087] In this embodiment, preferably, it also includes: S6, real-time monitoring of the data processing process, and triggering an alarm mechanism when an anomaly occurs.

[0088] In this embodiment, preferably, step S6 includes the following steps:

[0089] S61. Performance Metrics Collection: Continuously collect key performance metrics for each processing stage, including data throughput, processing latency, resource utilization, and queue backlog depth.

[0090] S62, Anomaly Detection Engine: It applies unsupervised learning algorithms to establish a baseline for system behavior and automatically identifies abnormal fluctuations that deviate from the baseline;

[0091] S63, Elastic Resource Scheduling: Dynamically scales up or down computing cluster resources based on monitoring indicators and predicted load, and triggers predefined alarm rules to notify the operation and maintenance system when resource bottlenecks or processing failures occur.

[0092] In this embodiment, preferably, S7, data visualization, is used to display generated interactive reports and charts.

[0093] In this embodiment, preferably, step S7 includes the following steps:

[0094] S71. Data Mapping: Converting the processing results into the structured data format required for the visualization model, including time series, categorical variables, and numerical indicators;

[0095] S72. Chart Generation Engine: It adopts a front-end framework that integrates ECharts or D3.js libraries to dynamically render interactive charts and supports custom configurations for bar charts, heatmaps and scatter plots.

[0096] S73, User Interaction Module: Allows users to drill down into data, filter time ranges, and switch chart types via a web interface, and updates the displayed content in real time in response to user actions;

[0097] S74. Report Export: Provides a one-click export function to save the generated visualization results as PDF, PNG or CSV format for easy offline analysis and archiving.

[0098] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data processing method based on big data electronic information technology, characterized in that, Includes the following steps: S1. Collect raw electronic data from multiple heterogeneous data sources and extract data in real time or in batches using distributed data acquisition tools; S2. Preprocess the collected data to achieve data cleaning, format conversion, missing value handling, and data integration. S3. Store the preprocessed data in a distributed storage system to achieve scalable and high-throughput storage of the data, while applying data partitioning and index optimization strategies to improve query performance. S4. Perform real-time or batch analysis on stored data to extract information features; S5. Transform the analysis results into structured output and integrate them into business systems or user interfaces to achieve end-to-end application and value mining of data.

2. The data processing method based on big data electronic information technology according to claim 1, characterized in that, S1 includes the following steps: S11. Identify connection parameters for multiple heterogeneous data sources, including data source type, access protocol, and authentication information; S12. Configure the extraction strategy of the distributed data acquisition tool and specify the real-time streaming acquisition or timed batch acquisition mode. S13. Perform data extraction operations and handle high-concurrency data inflows through a load balancing mechanism; S14. Conduct preliminary quality checks on the extracted data, including data integrity verification and error log recording; S15. The extracted raw data is temporarily stored in a buffer, and subsequent preprocessing steps are triggered.

3. The data processing method based on big data electronic information technology according to claim 1, characterized in that, S2 includes the following steps: S21. Data cleaning: Identify and handle missing values, fill them with interpolation or rule-based default values, detect and correct outliers, and use statistical methods such as Z-score or interquartile range analysis to ensure data consistency and accuracy. S22. Denoising: Apply filter technology to remove random noise, combine contextual information to smooth the data flow, remove outliers, and improve the data signal-to-noise ratio; S23. Format Conversion: Unify the data structure, convert heterogeneous data sources into a standardized format, perform data type conversion and unit normalization to facilitate subsequent analysis.

4. The data processing method based on big data electronic information technology according to claim 1, characterized in that, The S3 also includes data encryption and access control mechanisms, including the following steps: S31. Static data encryption: Before data is written to the distributed storage node, AES-256 or Chinese cryptographic algorithm is used to transparently encrypt the data block, and the key is managed by an independent hardware security module. S32. Dynamic authorization: Role-based access control or attribute-based access control policies, combined with data tag classification to implement fine-grained permission control, and audit logs to record all data access behaviors; S33. Data lifecycle management: Automatically executes a tiered storage strategy for hot and cold data, and performs secure erasure or archived encrypted storage for expired or infrequently accessed data.

5. The data processing method based on big data electronic information technology according to claim 1, characterized in that, S4 includes the following steps: S41. Feature Engineering: Deep feature extraction is performed on data stored in a distributed storage system to construct high-order feature combinations and time window statistics. Feature selection algorithms are applied to reduce dimensionality and eliminate redundancy. S42, Model Training and Inference: A distributed computing framework is used to build and train the prediction model, supporting online model updates and A / B testing. For batch processing tasks, an incremental learning mechanism is applied to optimize the model iteration efficiency. S43. Streaming Processing: For real-time data streams, deploy window-based aggregation computing and combine it with a complex event processing engine to identify specific patterns, achieving streaming analysis with millisecond-level latency.

6. The data processing method based on big data electronic information technology according to claim 1, characterized in that, S5 includes the following steps: S51, Self-service Analysis Portal: Provides drag-and-drop report building tools, supports multi-dimensional drill-down analysis, cross-filtering and natural language query, and allows users to customize visualization components; S52. Prediction Result Integration: Embed the prediction labels or scores generated by S4 into the business system workflow in real time to drive automated decision-making; S53. Feedback closed-loop mechanism: Capture user interaction behavior with analysis results and automatically feed it back to the feature library for model optimization.

7. The data processing method based on big data electronic information technology according to claim 1, characterized in that, Also includes: S6. Monitor the data processing process in real time and trigger an alarm mechanism when an anomaly occurs.

8. The data processing method based on big data electronic information technology according to claim 7, characterized in that, S6 includes the following steps: S61. Performance Metrics Collection: Continuously collect key performance metrics for each processing stage, including data throughput, processing latency, resource utilization, and queue backlog depth. S62, Anomaly Detection Engine: It applies unsupervised learning algorithms to establish a baseline for system behavior and automatically identifies abnormal fluctuations that deviate from the baseline; S63, Elastic Resource Scheduling: Dynamically scales up or down computing cluster resources based on monitoring indicators and predicted load, and triggers predefined alarm rules to notify the operation and maintenance system when resource bottlenecks or processing failures occur.

9. A data processing method based on big data electronic information technology according to claim 1, characterized in that, Also includes: S7, Data Visualization, is used to display generated interactive reports and charts.

10. A data processing method based on big data electronic information technology according to claim 9, characterized in that, S7 includes the following steps: S71. Data Mapping: Converting the processing results into the structured data format required for the visualization model, including time series, categorical variables, and numerical indicators; S72. Chart Generation Engine: It adopts a front-end framework that integrates ECharts or D3.js libraries to dynamically render interactive charts and supports custom configurations for bar charts, heatmaps and scatter plots. S73, User Interaction Module: Allows users to drill down into data, filter time ranges, and switch chart types via a web interface, and updates the displayed content in real time in response to user actions; S74. Report Export: Provides a one-click export function to save the generated visualization results as PDF, PNG or CSV format for easy offline analysis and archiving.

Citation Information

Cited By

  • Engineering cost data processing method and system based on big data analysis

    CN122022208A