Computer data processing system based on big data analysis
Through diversified data acquisition, distributed storage, advanced data analysis and security protection methods, the challenges of traditional systems in data acquisition, storage, analysis and security are solved, and efficient, secure and visual data processing capabilities are achieved, improving user satisfaction and system practicality.
Patent Information
- Application Number
- CN202510255103.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional computer data processing systems have many challenges in data collection, storage, analysis, visualization and security, and are difficult to meet the needs of big data processing, including problems such as single data source, limited storage capacity, poor scalability, low analysis efficiency, insufficient security, single visualization methods and lack of interactivity.
It adopts diversified data acquisition methods, storage methods combining distributed file systems and cloud storage, data analysis of multiple machine learning and deep learning algorithms, diversified visual display, data security protection of symmetric and asymmetric encryption algorithms, a comprehensive data quality evaluation system and multi-level access control, as well as data mining and feedback mechanisms.
It realizes diversified, high-speed, safe and efficient data collection and storage, improves the accuracy of data analysis and visual interactivity, enhances data security and user satisfaction, and ensures data quality and real-time optimization capabilities of the system.
Smart Images

Figure CN120371887A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data computer processing, and in particular to a computer data processing system based on big data analysis. Background Art
[0002] In today's digital age, with the rapid development of information technology, the amount of data generated in various fields has shown an explosive growth, and the big data era has arrived. Computer data processing plays a crucial role in many aspects such as scientific research, business operation, and social management. However, traditional data processing systems face many challenges and are difficult to meet the growing demand for big data processing.
[0003] From the perspective of data collection, the data sources of traditional systems are relatively single, often limited to internal databases, and cannot make full use of diversified data resources such as the Internet and the Internet of Things. Moreover, during the collection process, there is a lack of an effective quality assessment mechanism, and a large amount of duplicate, incorrect, and incomplete information may exist in the collected data, which brings great difficulties to subsequent data processing and analysis. With the continuous enrichment of data types, such as the large emergence of unstructured data such as text, images, and videos, traditional collection methods are difficult to adapt to, resulting in the neglect of many valuable data.
[0004] In terms of data storage, traditional centralized storage systems have problems of limited capacity and poor scalability. Facing massive data, centralized storage is prone to performance bottlenecks and cannot meet the high-concurrency and high-availability requirements of big data storage. At the same time, the security and reliability of data also face severe challenges. Once the storage device fails or is attacked, it may lead to data loss or leakage, bringing huge losses to enterprises and organizations.
[0005] Data analysis is a key link in mining data value, but traditional data analysis methods are inefficient in processing large-scale and complex data. Traditional algorithms are often designed based on small-scale data sets, cannot make full use of the potential information in the data, and are difficult to discover the complex associations and patterns between data. Moreover, in the face of data analysis tasks with high real-time requirements, the response speed of traditional systems is slow and cannot provide support for decision-making in a timely manner.
[0006] Data visualization is an important means to present the results of data analysis to users in an intuitive way. However, traditional visualization methods are relatively single, lack interactivity and adaptability, and cannot meet the needs of different users. When displaying large-scale data, traditional charts and reports may appear chaotic, making it difficult to clearly convey the meaning of the data and affecting users' understanding and decision-making of the data.
[0007] Data security issues are also a major challenge faced by traditional data processing systems. As the importance of data becomes increasingly prominent, security incidents such as data leakage and malicious attacks occur frequently. Traditional security protection measures mainly focus on the network boundary, insufficiently protecting the data itself and unable to effectively cope with the irregular operations of internal personnel and increasingly complex network attack means.
[0008] Data quality directly affects the accuracy and reliability of data analysis results. Traditional data quality management methods mainly rely on manual inspection and rule matching, with low efficiency and prone to omissions. In the big data environment, the sources of data are extensive and the types are complex, making it difficult for traditional methods to comprehensively and real-time monitor and manage data quality.
[0009] In addition, traditional data processing systems lack data mining and feedback mechanisms, unable to fully exploit the potential value in data and difficult to adjust and optimize the system in a timely manner according to user feedback. Therefore, developing a computer data processing system based on big data analysis has important practical significance, which can effectively solve the problems existing in traditional systems, improve the efficiency, quality and security of data processing, and provide strong support for the development of various fields. Summary of the Invention
[0010] The computer data processing system based on big data analysis proposed by the present invention aims to solve the problems mentioned in the above prior art.
[0011] To achieve the above object, the present invention adopts the following technical solutions: A computer data processing system based on big data analysis, comprising:
[0012] Data acquisition module: Collect data through web crawlers, sensor networks, and API interfaces, covering structured, semi-structured, and unstructured data; The web crawler adopts a distributed architecture and intelligent scheduling algorithm to capture web data, and the sensor network collects data generated by Internet of Things devices. When collecting, use the data integrity formula to evaluate data quality, where I d is data integrity, N valid is the amount of valid data, N total is the total amount of collected data, mark and preprocess incomplete data;
[0013] Data storage module: Store data in a combined manner of the distributed file system HBase and cloud storage. HBase stores structured and semi-structured data, and cloud storage stores large-scale unstructured data; When storing data, use the data redundancy rate formula to set the redundancy, where R r is the data redundancy rate, N replicate is the amount of redundant data, N original is the amount of original data;
[0014] Data Analysis Module: Use machine learning algorithms and deep learning models for data analysis. Adopt the LSTM model to predict trends for time series data and use the random forest algorithm for classification problems; Screen key features through the feature importance formula where F i is the importance of the i-th feature, w ij is the weight of the i-th feature in the j-th model, and I ij is the importance score of the i-th feature in the j-th model;
[0015] Data Visualization Module: Display the analysis results in the form of charts, reports, and dashboards. Adopt responsive design technology and use the visualization clarity formula to evaluate the display effect, where C v is the visualization clarity, N clear is the amount of data clearly displayed, and N total-vis is the total amount of visualized data.
[0016] Furthermore, it also includes:
[0017] Data Security Module: Use a combination of the symmetric encryption algorithm AES and the asymmetric encryption algorithm RSA to encrypt data. Use the SSL / TLS protocol for encryption during data transmission. Evaluate the system security status through the data security risk formula where R s is the data security risk, P i is the probability of the i-th security threat occurring, and L i is the loss caused by the i-th security threat;
[0018] Data Quality Management Module: Establish a data quality evaluation index system and use the data accuracy formula to evaluate data quality, where A d is the data accuracy, N accurate is the amount of accurate data, and N total-eval is the total amount of data to be evaluated.
[0019] Furthermore, it also includes:
[0020] Data Mining Module: Use association rule mining and clustering analysis techniques to discover potential patterns and relationships in data. Screen valuable association rules through the association rule confidence formula where C rule is the association rule confidence, N support-both is the amount of data that supports both the antecedent and the consequent, and N support-antecedent is the amount of data that supports the antecedent;
[0021] Data feedback module: Collect user feedback on data analysis results and visualization, and use the feedback to optimize the system's analysis model and visualization scheme. Through the user satisfaction formula Evaluate the user's satisfaction with the system, where S u is the user satisfaction, R i is the satisfaction score of the i-th user, and n is the total number of users.
[0022] Furthermore, when the data security module performs multi-level access control, it combines behavioral analysis technology to monitor and analyze the user's operation behavior in real time, constructs a user behavior model, and through the behavior similarity formula Judge whether the user behavior is abnormal, where S b is the behavior similarity, w i is the weight of the i-th behavior feature, and M i is the matching degree of the i-th behavior feature.
[0023] Furthermore, when the data quality management module performs data cleaning, it uses an anomaly detection algorithm based on machine learning to identify and process abnormal data, and through the abnormal data ratio formula Evaluate the cleaning effect, where P ab is the abnormal data ratio, N abnormal is the amount of abnormal data, and N total-clean is the total amount of data cleaned.
[0024] Furthermore, when the data mining module performs clustering analysis, it uses the density peak clustering algorithm to cluster the data, and through the local density formula and the relative distance formula Determine the clustering center, where ρ i is the local density of the i-th data point, d ij is the distance between the i-th and the j-th data points, d c is the cut-off distance, and χ(x) is the step function.
[0025] Furthermore, after the data feedback module collects user feedback information, it uses sentiment analysis technology to analyze the feedback text, and through the sentiment polarity formula Judge the sentiment tendency of the user feedback, where P s is the sentiment polarity, w i is the weight of the i-th sentiment word, and S i is the sentiment score of the i-th sentiment word.
[0026] Furthermore, it also includes the following steps:
[0027] Data collection step: By collecting data, use the data integrity formula Evaluate the data quality and preprocess the incomplete data;
[0028] Data storage step: Store data by combining a distributed file system and cloud storage, and use the data redundancy rate formula Reasonably set the redundancy degree and perform data compression at the same time;
[0029] Data analysis step: Use machine learning and deep learning algorithms for analysis, and screen key features through the feature importance formula Screen key features;
[0030] Data visualization step: Display the analysis results and evaluate the display effect using the visualization clarity formula Evaluate the display effect.
[0031] Furthermore, it also includes:
[0032] Data security step: Encrypt data using encryption algorithms, and evaluate the security status and monitor and give early warnings in real time using the data security risk formula Evaluate the security status and monitor and give early warnings in real time;
[0033] Data quality management step: Establish an evaluation index system and evaluate data quality using the data accuracy formula Evaluate data quality.
[0034] Furthermore, it also includes:
[0035] Data mining step: Use association rule mining and clustering analysis techniques to discover data patterns, and screen rules through the association rule confidence formula Screen rules;
[0036] Data feedback step: Collect user feedback and judge the sentiment tendency using sentiment analysis techniques and the sentiment polarity formula Judge the sentiment tendency.
[0037] Compared with the existing technologies, the beneficial effects of the present invention are:
[0038] In terms of data collection, adopt diversified collection channels, and evaluate data quality in combination with the data integrity formula, which can comprehensively and accurately collect various types of data, avoid the limitations of traditional collection methods, and provide a rich and high-quality data basis for subsequent processing.
[0039] In terms of data storage, the combination of a distributed file system and cloud storage, combined with the data redundancy rate formula to reasonably set the redundancy degree, effectively solves the problems of limited capacity and poor scalability of traditional centralized storage, greatly improves the reliability and security of data storage, and reduces the risk of data loss and damage.
[0040] In the data analysis stage, a variety of advanced machine learning and deep learning algorithms are used to screen key features through the feature importance formula, significantly improving the efficiency and accuracy of analysis. It can extract valuable information and potential patterns from massive data, providing strong support for decision-making.
[0041] The data visualization module adopts diverse display forms and responsive design technologies, and optimizes the display effect using the visualization clarity formula, enhancing the intuitiveness and interactivity of data. This enables users to more easily understand and analyze data, improving the efficiency and quality of decision-making.
[0042] The data security module employs multiple encryption algorithms and multi-level access control mechanisms, and combines the data security risk formula to evaluate the security status in real time, effectively preventing data leakage and malicious attacks, and ensuring the security and privacy of data.
[0043] The data quality management module establishes a complete evaluation index system, evaluates data quality through the data accuracy formula, and performs cleaning and transformation operations to ensure the accuracy, consistency, and timeliness of data, improving the reliability of data analysis results.
[0044] The data mining module uses various mining technologies and related formulas to discover potential patterns and correlation relationships in data, providing valuable references for fields such as market prediction and customer segmentation. The data feedback module collects user feedback, uses sentiment analysis technology and sentiment polarity formula to judge sentiment tendency, and can optimize the system in a timely manner according to user needs, improving user satisfaction and the practicality of the system. Generally speaking, this system can comprehensively improve the ability and level of computer data processing, bringing significant benefits to the development of various industries. Brief Description of the Drawings
[0045] Figure 1 It is a schematic block diagram of a computer data processing system based on big data analysis proposed by the present invention;
[0046] Figure 2 It is a schematic block diagram of a computer data processing method based on big data analysis proposed by the present invention. Detailed Embodiments
[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0048] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. are based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation on the present invention.
[0049] In addition, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined. In addition, the terms "mounted", "connected" and "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection or an integral connection; it may be a mechanical connection or an electrical connection; it may be a direct connection or an indirect connection through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. The present invention will be further described in detail below with reference to the drawings.
[0050] Refer to Figure 1-2 : A computer data processing system based on big data analysis, comprising:
[0051] Data Acquisition Module: Data acquisition is the crucial starting link in the entire data processing flow. It collects data through a variety of advanced and efficient means, covering a wide range of structured, semi-structured, and unstructured data. In terms of web crawlers, a highly optimized distributed architecture is adopted. This architecture distributes the data crawling tasks to multiple nodes for parallel execution, greatly improving the speed of data acquisition. At the same time, it is paired with an intelligent scheduling algorithm, which can dynamically adjust the crawling priority and time interval according to various factors such as the update frequency and importance of web pages. For example, for news websites, due to the frequent content updates, the scheduling algorithm will shorten the crawling interval to ensure timely access to the latest information; for some relatively stable enterprise official websites, the crawling cycle will be appropriately extended to reasonably allocate resources. In this way, the web crawler can accurately crawl web page data according to the preset rules, improving the acquisition efficiency while ensuring the comprehensiveness of the acquisition and not missing any valuable data. Sensor networks play an indispensable role in data acquisition, mainly used to collect data generated by Internet of Things devices in real time. In the vast system of the Internet of Things, there are various types of sensors, including temperature sensors, pressure sensors, humidity sensors, etc. These sensors are widely deployed in various scenarios, such as smart home systems, industrial production environments, intelligent transportation networks, etc. Taking industrial production as an example, the sensor network can monitor the operating status of production equipment in real time, including key parameters such as the temperature, vibration frequency, and energy consumption of the equipment. Once an abnormality occurs in the equipment, the sensor can quickly capture the data change and transmit these real-time data to the data acquisition module for further processing. During the data acquisition process, the evaluation of data quality is crucial. For this reason, the data integrity formula (I d represents data integrity, N valid represents the amount of valid data, N total represents the total amount of collected data) is used for accurate evaluation. In actual operation, when a batch of data is collected, the system will automatically analyze the data and calculate the proportion of valid data in the total data volume, thereby obtaining the data integrity index. For those cases where the data integrity does not meet the standard, that is, there is incomplete data, the system will immediately mark it and start the preprocessing program. The preprocessing process includes operations such as filling in missing data and correcting incorrect data. For example, for the missing numerical data, methods such as mean filling and median filling can be used; for the errors in text data, natural language processing techniques can be used for error correction and normalization to ensure the accuracy and reliability of subsequent data processing and analysis.
[0052] Data Storage Module: As a key link in the data processing flow, it shoulders the important responsibility of storing data securely and efficiently. It innovatively adopts a storage method that combines the distributed file system HBase with cloud storage to meet the storage requirements of different types of data. HBase is a column-oriented distributed database that excels in storing structured and semi-structured data. Its high scalability benefits from its distributed architecture, enabling it to easily handle the rapid growth of data volume. When the data volume continues to climb, HBase can expand its storage and processing capabilities by adding more nodes without the need for large-scale adjustments to the system architecture. In terms of read and write performance, HBase adopts a unique storage structure and algorithm. It stores data by column family. This storage method allows it to skip unnecessary data blocks when reading data from a specific column family, greatly improving the read speed. At the same time, HBase's write operation is also very efficient. It uses a mechanism of in-memory writing and asynchronous flushing. First, it writes data into the MemStore in memory. When the MemStore reaches a certain threshold, it asynchronously flushes the data to the StoreFile on disk, thus reducing the number of disk I / O operations and improving the write performance. For example, when processing large-scale log data, HBase can quickly store the log information and retrieve relevant records promptly when needed. Cloud storage is mainly used to store large-scale unstructured data such as videos and audios. Cloud storage has a powerful storage capacity and high availability. It uses distributed storage technology to disperse data storage across multiple physical nodes, ensuring that data can still be accessed completely even if some nodes fail. At the same time, cloud storage also provides flexible access interfaces, allowing users to access and manage the data stored in the cloud anytime, anywhere through the network. Taking video data storage as an example, cloud storage can support the upload and storage of various formats of video files and provide efficient video streaming services to ensure smooth playback of videos for users without stuttering. During the data storage process, to ensure data reliability, the system uses the data redundancy rate formula (R r is the data redundancy rate, N replicate is the amount of redundant data, N originalSet redundancy reasonably according to the original data volume. By calculating the redundancy rate, the system can precisely control the amount of redundant data based on the importance of the data and the application scenario. For critical business data, the system will set a higher redundancy rate to ensure that the data can still be restored even if multiple nodes fail simultaneously. For some less important data, the redundancy rate can be appropriately reduced to save storage resources. In addition, to reduce storage costs, the system also uses data compression algorithms such as Snappy. Snappy is a fast lossless compression algorithm that has a high speed during both the compression and decompression processes and can effectively reduce the storage space of data without affecting data integrity. For example, when storing a large amount of text data, using the Snappy algorithm can compress the data volume to a fraction of the original, thus greatly reducing storage costs.
[0053] Data Analysis Module: The data analysis module is the core part of mining data value. It comprehensively applies a variety of advanced machine learning algorithms and deep learning models to deeply analyze the collected and stored data. In terms of machine learning algorithms, random forest and support vector machine are important components. Random forest is an ensemble learning algorithm based on decision trees. It constructs multiple decision trees and synthesizes the prediction results of these decision trees to obtain the final conclusion. In the process of constructing decision trees, random forest adopts the method of random sampling, extracting multiple different subsets from the original dataset to train each decision tree, which makes each tree have a certain difference, thus improving the generalization ability of the model. At the same time, random forest also adopts a random way in feature selection, further enhancing the stability of the model. For example, in customer credit assessment, random forest can accurately judge the credit level of customers based on multiple features such as age, income, and credit record. Support vector machine is a supervised learning algorithm for classification and regression. Its core idea is to find an optimal hyperplane in the feature space to separate data points of different classes as much as possible. For linearly separable data, support vector machine can directly find such a hyperplane; for linearly inseparable data, it maps the data to a high-dimensional space through a kernel function, and then finds the optimal hyperplane in the high-dimensional space. Support vector machine performs well in dealing with small-sample and high-dimensional data and is often used in fields such as image recognition and text classification. The long short-term memory network (LSTM) in deep learning models has unique advantages in processing time series data. LSTM is a special recurrent neural network (RNN). It effectively solves the problems of gradient disappearance and gradient explosion existing in traditional RNN when dealing with long sequence data by introducing memory units and gating mechanisms. The memory units of LSTM can save long-term information, while the input gate, forget gate, and output gate control the inflow, outflow, and retention of information. In time series prediction, such as stock price prediction and temperature change prediction, LSTM can capture the long-term dependence relationships in the data and accurately predict future trends. For example, when predicting the stock price trend, LSTM can make a relatively accurate prediction of the future stock price based on historical stock price data, trading volume and other information. In the data analysis process, the selection and importance evaluation of features are crucial. Through the feature importance formula (F i is the importance of the i-th feature, w ij is the weight of the i-th feature in the j-th model, I ij(i.e., the importance score of the i-th feature in the j-th model), the system can quantify the importance of each feature. In actual operation, first train the data in multiple different models to obtain the weights and importance scores of each feature in each model. Then, calculate the comprehensive importance score of each feature according to the above formula. By setting a certain threshold, filter out the key features with higher importance. This can not only reduce the dimensionality of the data, improve the calculation efficiency, but also avoid the interference of some unimportant features on the model performance, thereby improving the accuracy of the analysis. For example, in a customer purchase behavior prediction model, through feature importance evaluation, it is found that features such as the customer's purchase history and browsing records have higher importance, while some irrelevant demographic features have lower importance. Therefore, these key features can be focused on to improve the prediction accuracy of the model.
[0054] Data visualization module: As the terminal link of the data processing process, the data visualization module is dedicated to presenting the complex data analysis results to users in an intuitive and understandable form, greatly enhancing the readability and operability of the data. In terms of display forms, this module provides a variety of choices. The heat map is a commonly used visualization tool that represents the distribution and density of data through the depth of color. For example, in geographic information visualization, the heat map can clearly display information such as population density and business activity heat, and the darker the color area, the larger the relevant data volume. The scatter plot is suitable for showing the relationship between two variables. By plotting data points in a plane coordinate system, users can intuitively observe the correlation between variables, such as positive correlation, negative correlation, or no obvious correlation. In market research analysis, the scatter plot can be used to analyze the relationship between consumer age and purchase frequency. In addition, the report presents data in a structured table form, which is suitable for detailed display of specific numerical values and statistical information; the dashboard presents key data in an intuitive graph and indicators, and is often used for real-time monitoring and decision support, such as enterprise financial indicator monitoring, production progress monitoring, etc. To ensure that the visualization interface can be perfectly adapted to different devices, this module adopts responsive design technology. Responsive design is based on a flexible grid system and CSS media queries. In terms of the grid system, it divides the page into multiple columns of equal width, and different visualization elements can automatically adjust their layout in these columns according to the width of the device screen. For example, on a large-screen desktop computer, multiple charts and reports may be displayed at the same time; while on a small-screen mobile phone, these elements will be arranged in sequence to adapt to the screen space. CSS media queries allow developers to write different style rules for different device characteristics (such as screen width, height, resolution, etc.). When the screen size of the device changes, the browser will automatically apply the corresponding styles, thereby realizing the adaptive adjustment of the interface and ensuring that users can obtain a good visual experience on any device. In terms of evaluating the display effect, using the visualization clarity formula (Cv For visualization clarity, N clear For the amount of data clearly presented, N total-vis For the total amount of visualized data), a quantitative analysis is performed. In actual operation, the system scans and analyzes the data in the visualization interface to determine which data can be clearly recognized and understood by the user and which data may be fuzzy or overlapping. According to the calculated clarity index, the visualization layout is optimized. For example, if it is found that the data points in some charts are too dense to be distinguished, the size, color, or marker style of the chart will be adjusted to improve the clarity of the data. At the same time, the color combination will also be optimized, choosing colors with high contrast and easy to distinguish, and avoiding using overly similar colors that may cause data confusion, ensuring that the visualization effect can accurately convey information and has good visual aesthetics.
[0055] In the present invention, the following modules are further included:
[0056] Data Security Module: The data security module is a key component for ensuring the security of data throughout its entire life cycle. It comprehensively applies a variety of advanced technical means to safeguard the confidentiality, integrity, and availability of data in all aspects. In terms of data encryption, a combination of the symmetric encryption algorithm AES (Advanced Encryption Standard) and the asymmetric encryption algorithm RSA is adopted. AES is a symmetric key algorithm where the same key is used for both encryption and decryption. It features high efficiency and speed, enabling rapid encryption and decryption operations on large amounts of data. For example, in the internal database storage of an enterprise, AES can encrypt sensitive information of users, such as ID numbers and bank card numbers, effectively preventing data from being stolen during storage. The encryption process involves grouping the plaintext data according to a fixed block size (such as 128 bits, 192 bits, or 256 bits), and then using the key to perform encryption operations on each data block to generate ciphertext. RSA, on the other hand, is an asymmetric encryption algorithm that uses a pair of keys, namely the public key and the private key. The public key can be publicly distributed for encrypting data, while the private key is carefully kept by the user for decrypting data. RSA is commonly used in scenarios such as digital signatures and key exchanges. For example, before data transmission, the sender can use the recipient's public key to encrypt the data. After receiving the ciphertext, the recipient uses their own private key to decrypt it, ensuring that only the authorized recipient can obtain the data content. By combining AES and RSA, the advantages of the high efficiency of symmetric encryption and the security of asymmetric encryption are fully utilized, providing more reliable encryption protection for data. During the data transmission stage, the SSL (Secure Sockets Layer) / TLS (Transport Layer Security) protocol is used for encryption. The SSL / TLS protocol establishes a secure channel between the application layer and the transport layer to encrypt and authenticate the transmitted data. It negotiates the encryption algorithm and key through a handshake process to ensure the legitimacy of the identities of both communication parties. During the handshake process, the client and the server exchange certificates to verify each other's identities and negotiate the symmetric key used for encrypting data. Once the handshake is successful, all transmitted data will be encrypted, preventing data from being stolen and tampered with during transmission. For example, when a user accesses a website via the HTTPS protocol, the SSL / TLS protocol will be automatically activated to ensure the security of data transmission between the user and the website. To further ensure data security, a multi-level access control mechanism is set up. This mechanism assigns different data access levels based on the roles and permissions of users. First, users are classified into roles, such as administrators, ordinary users, and visitors. Administrators usually have the highest permissions and can perform operations such as accessing, modifying, and deleting all data; ordinary users are granted specific data access permissions according to their business needs, such as only being able to view certain specific reports or datasets; visitors may only have read-only permissions and can only browse public information. Through this detailed permission division, it is ensured that only authorized users can access the corresponding data, effectively preventing data leakage and abuse. In terms of evaluating the security status of the system, through the data security risk formula (R s is a data security risk, P i is the probability of the i-th security threat occurring, L i (is the loss caused by the i-th security threat) for quantitative analysis. The system will monitor various possible security threats in real time, such as cyberattacks, data breaches, malware, etc., and evaluate the probability of each threat occurring and the possible losses it may cause. For example, for common cyberattacks, by analyzing historical data and real-time network traffic, the probability of their occurrence is estimated; for the losses caused by data breaches, factors such as the value of the data, repair costs, and reputation losses are comprehensively considered for evaluation. According to the calculated risk value, the system can issue early warnings in a timely manner to remind the administrator to take corresponding measures to reduce risks, such as strengthening access control, updating security policies, performing data backups, etc., to ensure the secure and stable operation of the system.
[0057] Data Quality Management Module: The data quality management module is an important guarantee to ensure the effective play of data value. It comprehensively improves data quality by constructing a comprehensive evaluation system and implementing fine-grained data processing operations. First, the established data quality evaluation index system covers multiple key dimensions. The accuracy dimension focuses on the degree of conformity between the data and the actual situation. For example, in the financial field, the accuracy of data such as customer account balances and transaction records is crucial, and even minor errors may lead to serious consequences. When evaluating, the collected data is compared with authoritative data sources. For example, in the banking system, customer transaction data is compared with clearing center data to determine the accuracy of the data. The consistency dimension focuses on the logical consistency of data between different systems, at different times, or between different records. For example, in an enterprise's supply chain management system, the inventory quantity of products should be logically consistent in different links such as procurement, sales, and warehousing. If there is a situation where the procurement system shows that 100 pieces of a product are in stock, while the warehousing system records only 90 pieces, it indicates a data inconsistency problem. The timeliness dimension emphasizes the timeliness of data acquisition and update. In the e-commerce industry, the timely update of data such as real-time sales volume and inventory status of products can help merchants adjust their business strategies in a timely manner. If the data is delayed, it may lead to over-selling of products or inventory backlogs. When evaluating data quality, the data accuracy formula (A d is data accuracy, N accurate is the amount of accurate data, N total-evalThe total amount of data to be evaluated). In actual operation, first determine the evaluation scope and the total amount of data, then identify the accurate data volume through methods such as data verification and manual review, and then calculate the accuracy index. For example, in the evaluation of census data, from a large number of population information records, through cross-verification with authoritative data such as the public security household registration system, accurate records are found, and the accuracy of the population data in this area is calculated. To improve data quality, a series of data processing operations will be carried out. Data cleaning is to remove noise, error values, duplicate values, etc. from the data. For example, in customer information data, there may be situations such as incorrect phone number formats and incomplete address information, and cleaning is carried out through technical means such as regular expression matching and address standardization. Data transformation is to convert the data into the format and type that meet the analysis requirements. For example, unify the date format, convert the numerical values in string type to numerical type, etc., to facilitate subsequent analysis. Data integration is to integrate data from different data sources and solve the problems of data conflicts and redundancy. When an enterprise integrates data from multiple business systems, there may be a situation where the same customer has different identifiers in different systems. Through data matching and fusion technologies, the unification and integration of data are achieved. Through these operations, data quality is ensured, laying a solid foundation for reliable analysis results.
[0058] Data mining module: The core component that explores potential value from a large amount of data. It uses advanced algorithms and technologies to deeply mine the hidden patterns and correlation relationships in the data, providing strong support for decision-making. In terms of association rule mining, the Apriori algorithm is a classic and widely used method. Its basic principle is to find frequent item sets from the data set through an iterative method of layer-by-layer search, and then generate association rules. For example, in the analysis of supermarket sales data, the Apriori algorithm can help discover the associations between the products purchased by customers. It first sets a minimum support threshold. Support represents the frequency of a certain item set appearing in the data set. The algorithm starts from a single product, counts the support of each product, and finds the frequent 1-item sets that meet the minimum support. Then, based on these frequent 1-item sets, candidate 2-item sets (i.e., combinations of two products) are generated, and their support is counted again, and the frequent 2-item sets are screened out. This process is repeated continuously, generating higher-order candidate item sets and screening them until no higher-order item sets that meet the minimum support can be found. After obtaining the frequent item sets, association rules can be generated. The strength of the association rules is measured by confidence, through the association rule confidence formula (C rule is the confidence of the association rule, N support-both is the data volume that supports both the antecedent and the consequent, N support-antecedentCalculate based on the data volume supporting the antecedent. For example, the confidence of the association rule "beer → diaper" is the ratio of the number of customers who buy both beer and diapers to the number of customers who buy beer. By setting an appropriate confidence threshold, valuable association rules can be screened out to assist the supermarket in formulating product placement and promotion strategies. In cluster analysis, the K-Means algorithm is a commonly used unsupervised learning algorithm. Its goal is to divide the objects in the dataset into K different clusters, such that the objects within the same cluster have a high degree of similarity, while the objects between different clusters have a low degree of similarity. The specific steps of the algorithm are as follows: First, randomly select K objects as the initial cluster centers. Then, calculate the distance of each data point to these K cluster centers (usually using the Euclidean distance), and assign each data point to the cluster where the nearest cluster center is located. Next, recalculate the center of each cluster, which is the mean of all data points within the cluster. Repeat the above steps of assignment and center update until the cluster centers no longer change significantly or reach the preset number of iterations. For example, in customer segmentation, based on characteristics such as customers' purchase behavior, consumption amount, and purchase frequency, the K-Means algorithm can be used to divide customers into different groups, such as high-end customer groups, mid-end customer groups, and low-end customer groups, so that enterprises can formulate personalized marketing strategies for different groups. The results of data mining have a wide range of applications in multiple fields. In market prediction, by mining the patterns and trends in historical sales data, future market demand can be predicted to help enterprises reasonably arrange production and inventory. In the field of customer segmentation, the results of cluster analysis can help enterprises better understand the characteristics and needs of customer groups, thereby providing more accurate products and services and improving customer satisfaction and loyalty.
[0059] Data Feedback Module: The data feedback module is an important bridge connecting users and the data analysis system. It is dedicated to collecting user feedback information and optimizing the system based on this to improve the overall performance and user experience of the system. In terms of collecting feedback information, this module widely collects user feedback on data analysis results and visual displays through various channels. For satisfaction ratings, a dedicated rating entry is usually set in the data analysis report or visual interface. Users can choose the corresponding score from the pre-set rating levels (such as 1 - 5 or 1 - 10) according to their experience to express their satisfaction with the system. At the same time, to obtain more detailed improvement suggestions, a text input box is provided to encourage users to elaborate on the problems they think the system has and the expected improvement directions. For example, users may point out that some data analysis results are not accurate enough, or the color combination of visual charts is not reasonable enough. In addition, feedback information can also be collected through user research, online questionnaires, user forums, etc. to ensure a comprehensive understanding of user needs and opinions. In terms of using feedback to optimize the system, for feedback on data analysis results, the system will deeply analyze the problems and suggestions put forward by users. If users reflect that some analysis results are inaccurate, the development team will re-examine the analysis models and algorithms used, check whether the data input is correct, whether the feature selection is reasonable, and whether the model parameters need to be adjusted, etc. For example, in sales data analysis, if users think the prediction results deviate greatly from the actual situation, the development team may consider introducing more influencing factors as features or trying to use more complex machine learning models to improve the prediction accuracy. For feedback on visual displays, the visual layout, color combination, chart types, etc. will be adjusted according to user opinions. For example, if users feedback that the colors of a certain heat map are too similar to distinguish data differences, a color scheme with higher contrast will be re-selected; if users think the information in some reports is too complex, the report structure will be optimized to highlight key information. In terms of evaluating user satisfaction, it is quantitatively calculated through the user satisfaction formula (S u is the user satisfaction, R i is the satisfaction rating of the i-th user, and n is the total number of users). The system will regularly collect users' satisfaction ratings and calculate the average satisfaction. In addition to calculating the overall satisfaction, satisfaction evaluations can also be carried out separately for different types of users (such as ordinary users, professional users) or different usage scenarios (such as daily queries, special analysis) to make more targeted improvements. For example, if it is found that the satisfaction of professional users is low, more advanced analysis functions and customized visualization options may need to be provided; if the user satisfaction is not high in the special analysis scenario, the relevant analysis processes and display methods need to be optimized. According to the evaluation results, the system can be continuously adjusted and optimized to continuously improve users' satisfaction and usage experience with the system.
[0060] In the present invention, it further includes:
[0061] When the data security module performs multi-level access control, it combines behavior analysis technology to monitor and analyze the operation behavior of users in real time. A user behavior model is constructed, and through the behavior similarity formula (S b is the behavior similarity, w i is the weight of the i-th behavior feature, M i is the matching degree of the i-th behavior feature) to determine whether the user behavior is abnormal. Once an abnormal behavior is found, the access permission is immediately restricted and an alarm is issued.
[0062] When the data quality management module performs data cleaning, it uses an anomaly detection algorithm based on machine learning (such as the isolation forest algorithm) to identify and process abnormal data. Through the abnormal data ratio formula (P ab is the abnormal data ratio, N abnormal is the amount of abnormal data, N total-clean is the total amount of data to be cleaned) to evaluate the cleaning effect and ensure the stability of data quality.
[0063] In the present invention, it further includes:
[0064] When the data mining module performs clustering analysis, it uses the density peak clustering algorithm (DPC) to cluster the data. Through the local density formula (ρ i is the local density of the i-th data point, d ij is the distance between the i-th and the j-th data points, d c is the cut-off distance, χ(x) is the step function) and the relative distance formula to determine the clustering center and improve the accuracy and efficiency of clustering.
[0065] After the data feedback module collects user feedback information, it uses sentiment analysis technology to analyze the feedback text. Through the sentiment polarity formula (P s is the sentiment polarity, w i is the weight of the i-th sentiment word, s is the sentiment score of the i-th sentiment word) to judge the sentiment tendency of the user feedback, so as to optimize the system more accurately.
[0066] The present invention also discloses a computer data processing method based on big data analysis, including the following steps:
[0067] Data acquisition step: Collect data through various channels, and use the data integrity formula to evaluate the data quality and preprocess the incomplete data.
[0068] Data storage step: Store data by combining a distributed file system and cloud storage, and apply the data redundancy rate formula Reasonably set the redundancy degree and perform data compression simultaneously.
[0069] Data analysis step: Analyze using machine learning and deep learning algorithms, and screen key features through the feature importance formula to improve the analysis accuracy.
[0070] Data visualization step: Display the analysis results in various forms, and evaluate the display effect using the visualization clarity formula to optimize the visualization layout.
[0071] In the present invention, it further includes:
[0072] Data security step: Encrypt data using an encryption algorithm, set up a multi-level access control mechanism, and evaluate the security status using the data security risk formula for real-time monitoring and early warning.
[0073] Data quality management step: Establish an evaluation index system, evaluate data quality using the data accuracy formula and improve data quality through operations such as cleaning and transformation.
[0074] Data mining step: Discover data patterns using association rule mining and clustering analysis techniques, and screen rules through the association rule confidence formula to provide a basis for decision-making.
[0075] Data feedback step: Collect user feedback, judge the sentiment tendency using sentiment analysis technology and the sentiment polarity formula and optimize the system according to the feedback.
[0076] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A computer data processing system based on big data analysis, characterized in that, Including: Data collection module: Collects data through web crawlers, sensor networks, and API interfaces, covering structured, semi-structured, and unstructured data; The web crawler uses a distributed architecture and an intelligent scheduling algorithm to capture web page data. The sensor network collects the data generated by IoT devices. When collecting, the data integrity formula is used to evaluate the data quality, I d for data integrity, N υalid for the amount of valid data, N total for the total amount of collected data, mark and preprocess the incomplete data; Data storage module: It stores data by combining the distributed file system HBase with cloud storage. HBase stores structured and semi-structured data, and cloud storage stores large-scale unstructured data. When storing data, the data redundancy rate formula is used. Set the redundancy degree, R r Is the data redundancy rate, N replicate Is the amount of redundant data, N original Is the amount of original data; Data analysis module: Performs data analysis using machine learning algorithms and deep learning models. For time series data, the LSTM model is used to predict trends, and for classification problems, the random forest algorithm is used; Through the feature importance formula Screen key features, F i is the importance of the i-th feature, w ij is the weight of the i-th feature in the j-th model, I ij is the importance score of the i-th feature in the j-th model; Data Visualization Module: Presents the analysis results in the forms of charts, reports, and dashboards. Adopts responsive design technology and uses the visualization clarity formula to evaluate the display effect, where C υ is the visualization clarity, N clear is the amount of clearly displayed data, and N total-υis is the total amount of visualized data.
2. The computer data processing system based on big data analysis according to claim 1, wherein, It also includes: Data Security Module: The data is encrypted by combining the symmetric encryption algorithm AES and the asymmetric encryption algorithm RSA. The SSL / TLS protocol is used for encryption during the data transmission phase. The system security status is evaluated through the data security risk formula where R s is the data security risk, P i is the probability of the i-th type of security threat occurring, and L i is the loss caused by the i-th type of security threat; Data Quality Management Module: Establish a data quality assessment index system and use the data accuracy formula to evaluate data quality, A d is the data accuracy, N accurate is the amount of accurate data, N total-eval is the total amount of data to be evaluated.
3. A computer data processing system based on big data analysis according to claim 1, characterized in that, It also includes: Data mining module: Use association rule mining and clustering analysis techniques to discover potential patterns and correlation relationships in data. Through the confidence formula of association rules screen valuable association rules, where C rule is the confidence of the association rule, N support-both is the amount of data that supports both the antecedent and the consequent, and N support-antecedent is the amount of data that supports the antecedent; Data feedback module: Collect feedback information from users on data analysis results and visual displays, and use the feedback to optimize the system's analysis model and visualization scheme. Evaluate the user's satisfaction with the system through the user satisfaction formula where S u is the user satisfaction, and R i is the satisfaction score of the i-th user, and n is the total number of users.
4. A computer data processing system based on big data analysis according to claim 2, characterized in that, When the data security module performs multi-level access control, it combines behavioral analysis technology to monitor and analyze the user's operation behavior in real time, constructs a user behavior model, and passes the behavior similarity formula to determine whether the user behavior is abnormal. S b is the behavior similarity, w i is the weight of the i-th behavior feature, and M i is the matching degree of the i-th behavior feature.
5. A computer data processing system based on big data analysis according to claim 2, characterized in that, When the data quality management module performs data cleaning, it uses an anomaly detection algorithm based on machine learning to identify and process abnormal data, and evaluates the cleaning effect through the abnormal data proportion formula P ab is the proportion of abnormal data, N abnormal is the amount of abnormal data, and N total-clean is the total amount of data cleaned.
6. A computer data processing system based on big data analysis according to claim 3, characterized in that, When performing clustering analysis, the data mining module uses the density peak clustering algorithm to cluster the data, and determines the clustering center through the local density formula and the relative distance formula , where ρ i is the local density of the i-th data point, d ij is the distance between the i-th and j-th data points, d c is the cut-off distance, and χ(x) is the step function.
7. A computer data processing system based on big data analysis according to claim 3, characterized in that, After collecting user feedback information, the data feedback module analyzes the feedback text using sentiment analysis technology. Through the sentiment polarity formula to judge the sentiment tendency of the user feedback, where P s is the sentiment polarity, w i is the weight of the i-th sentiment word, and S i is the sentiment score of the i-th sentiment word.
8. A method for a computer data processing system based on big data analysis using the system according to any one of claims 1-7, characterized in that, Including the following steps: Data collection steps: By collecting data and using the data integrity formula Evaluate the data quality and preprocess the incomplete data; Data storage steps: Store data by combining a distributed file system and cloud storage, and use the data redundancy rate formula Reasonably set the redundancy and perform data compression at the same time; Data analysis steps: Use machine learning and deep learning algorithms for analysis, and screen key features through the feature importance formula Data visualization steps: Present the analysis results and use the visualization clarity formula Evaluate the presentation effect.
9. The computer data processing method based on big data analysis according to claim 8, wherein, It also includes: Data security steps: Encrypt data using encryption algorithms and apply the data security risk formula to evaluate the security status and monitor and give early warnings in real time; Data quality management steps: Establish an evaluation index system and use data accuracy formulas Evaluate data quality.
10. The computer data processing method based on big data analysis according to claim 8, characterized in that, It also includes: Data mining steps: Use association rule mining and clustering analysis techniques to discover data patterns, and filter rules through the confidence formula of association rules ; Data feedback steps: Collect user feedback and apply sentiment analysis techniques and sentiment polarity formulas Judge the sentiment tendency.