Computer data acquisition, processing and analysis system
By optimizing data collection through multi-source data verification and adaptive adjustment strategies, combined with machine learning and deep learning technologies, the problem of slow data processing speed is solved, and efficient data processing and system operation are achieved.
Patent Information
- Application Number
- CN202510905409.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-10-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, as the amount of data increases, the data screening and processing units may face the problem of slow processing speed, affecting the response time and efficiency of the system.
It adopts data collection and multi-source verification module, adaptive adjustment strategy module, intelligent data processing module, computer network status collection and diagnosis module, intelligent analysis module and computer network repair and big data management module. Through multi-source data verification, dynamic adjustment of data collection frequency and accuracy, machine learning algorithm to deal with noise and outliers, deep learning and knowledge graph technology to perform data mining and association analysis, real-time network diagnosis and repair.
It improves data accuracy and consistency, optimizes resource utilization, reduces data transmission and storage overhead, and ensures efficient system operation and data security.
Smart Images

Figure CN120804625A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data collection processing and analysis, and particularly relates to a computer data collection processing and analysis system. BACKGROUND
[0002] With the rise of the digital wave, more and more enterprises and organizations begin to attach importance to the value of data. As one of the important tools for digital transformation, data collection processing and analysis system is gradually penetrating into various industries. The rapid development of big data and artificial intelligence technology provides a broader application space for data collection processing and analysis system. Through the use of big data and artificial intelligence technology, the system can realize the rapid processing and analysis of massive data, and dig out more valuable information and insights.
[0003] After searching, the invention patent with Chinese patent number CN114826770A discloses a big data management platform for computer network intelligent analysis, belonging to the technical field of computer network intelligent analysis. It includes a data collection and processing module, a computer network state collection module, a computer network diagnosis module, an intelligent analysis module, a computer network repair module, and a big data management module. The data collection and processing module is used to collect and process data input into the big data management platform and transmit the collected and processed data to the intelligent analysis module. The computer network state collection module is used to collect load information, network traffic information, and network state information during computer network operation and transmit the collected information to the computer network diagnosis module. The computer network diagnosis module is used to receive the collected information transmitted by the computer network state collection module, diagnose the computer network based on the received information, and transmit the diagnosis results to the intelligent analysis module. Compared with the prior art, the invention patent with Chinese patent number CN114826770A marks the data that does not conform to the actual propagation path and the expected propagation path, determines the propagation path of the marked data of multi-path transmission based on the data increase, and facilitates the determination of specific fault points when the big data management platform or computer network fails. By managing the data or network before and after the fault point, the normal operation of the big data management platform can be ensured, and the data management effect of the platform is further improved.
[0004] However, in the above-mentioned use process, as the amount of data continues to increase, the data screening and processing unit may face the problem of slow processing speed. If a large amount of data cannot be screened and processed in time, the response time and efficiency of the entire system will be affected. Therefore, a computer data collection processing and analysis system is proposed. SUMMARY
[0005] The computer data acquisition processing and analysis system is provided to solve the problem of slow processing speed of data screening and processing units in the prior art, and to improve the response time and efficiency of the entire system.
[0006] To achieve the above object, the present application adopts the following technical scheme:
[0007] A computer data acquisition processing and analysis system comprises:
[0008] A data acquisition and multi-source verification module is responsible for acquiring data from multiple independent and reliable data sources and performing multi-source data verification to ensure the accuracy and consistency of the data.
[0009] An adaptive adjustment strategy module is responsible for dynamically adjusting the frequency and accuracy of data acquisition according to the network environment and device state to optimize resource utilization.
[0010] An intelligent data processing module is responsible for automatically processing noise, outliers and missing values in the data using machine learning algorithms, and performing data compression and encoding to reduce transmission and storage overhead.
[0011] A computer network state acquisition and diagnosis module is responsible for acquiring state information of the computer network during operation and performing preliminary diagnosis.
[0012] An intelligent analysis module is responsible for in-depth analysis of the causes of network anomalies and marking data anomalies, and uses deep learning and knowledge graph technology for data mining and correlation analysis.
[0013] A computer network repair and big data management module is responsible for repairing the computer network according to the analysis results of the intelligent analysis module and effectively managing the big data.
[0014] The data acquisition and multi-source verification module provides raw data, the adaptive adjustment strategy module adjusts the acquisition strategy according to the data quality and system resource state to form a closed-loop feedback, the acquisition data adjusted by the adaptive adjustment strategy module is input to the intelligent data processing module, the preprocessed data of the intelligent data processing module is transmitted to the computer network state acquisition and diagnosis module and the intelligent analysis module, the intelligent analysis module provides analysis results, and the computer network repair and big data management module performs repair and management operations according to the analysis results.
[0015] The above technical scheme further comprises:
[0016] Further, the data collection and multi-source verification module includes a data source interface unit, a data format analysis unit, a data cache and queue unit, and a multi-source data verification unit. The data source interface unit is responsible for establishing a connection with various data sources. The data format analysis unit performs format analysis on the received data and converts it into a unified data format within the system. During data collection, a data cache and queue mechanism is introduced. The data cache and queue unit is responsible for temporarily storing the parsed data in the cache or placing it in the message queue for subsequent processing. The multi-source data verification unit is responsible for cross-verification of the collected data. This step is crucial as it can identify and correct errors, inconsistencies, or redundant information in the data, thereby improving the accuracy and reliability of the data. The data source interface unit transmits the collected raw data to the data format analysis unit for format analysis. The parsed data is transmitted to the data cache and queue unit for temporary storage or queuing for processing. The data in the cache or queue is taken out one by one and enters the multi-source data verification unit for preliminary verification. For example, in network traffic data analysis, the system compares data from network devices, monitoring systems, and log files to find any inconsistencies and correct them accordingly.
[0017] Further, the adaptive adjustment strategy module monitors the network load and device resource usage in real time. When the network load is high or the device resources (such as CPU, memory, storage, etc.) are tight, the adaptive adjustment strategy module automatically reduces the data collection frequency to reduce the burden on system performance. In the case of good network conditions and sufficient device resources, the adaptive adjustment strategy module correspondingly increases the data collection accuracy and frequency. By dynamically adjusting the frequency and accuracy of data collection, the adaptive adjustment strategy module ensures data quality while minimizing unnecessary resource consumption.
[0018] Further, the intelligent data processing module includes a preprocessing unit, an intelligent processing unit, and a data compression and encoding technology unit. The preprocessing unit preprocesses raw data, including data cleaning (such as removing duplicates, formatting, etc.) and data conversion (such as converting text data to numerical data for machine learning model processing). Assuming we have a set of network traffic data, we first need to check if there are duplicate records in the data, and if so, perform deduplication processing. At the same time, we also need to convert text format data such as timestamps into time series data format suitable for analysis. The intelligent processing unit uses the pattern recognition capabilities of machine learning algorithms to automatically identify and process noise, outliers, and missing values in the data, and automatically monitor and correct data quality. The data compression and encoding technology unit optimizes the processed data, reduces redundant information in the data through algorithms while maintaining data integrity and core information, significantly reducing the bandwidth and storage space required for data transmission. At the same time, the data compression and encoding technology unit uses Huffman coding to encode data, improving data readability and maintainability.
[0019] Further, the computer network status acquisition and diagnosis module includes a computer network status acquisition unit and a network diagnosis unit. The computer network status acquisition unit is responsible for real-time acquisition of key parameters of computer network operation, including but not limited to network load information (such as CPU usage, memory occupancy), network traffic information (such as inbound and outbound data volume, packet size, transmission speed), and network status information (such as connection status, delay, packet loss rate, etc.). The computer network status acquisition unit transmits the collected raw data to the network diagnosis unit. The network diagnosis unit receives data from the network status acquisition unit, uses pre-set diagnostic logic to conduct in-depth analysis of network status to identify abnormal conditions in the network (such as congestion, failure, security threats, etc.). The network diagnosis unit outputs the diagnosis results in an easy-to-understand and operate form, typically including abnormal type, location, severity, and recommended repair measures, etc. The network diagnosis unit transmits the diagnosis results to the intelligent analysis module.
[0020] Further, the intelligent analysis module includes a network anomaly cause analysis unit and a marked data anomaly cause analysis unit. The computer network state acquisition and diagnosis module transmits the network state information collected and preliminarily diagnosed in real time to the network anomaly cause analysis unit. The network anomaly cause analysis unit performs in-depth analysis of the network anomaly cause based on the network state information. The intelligent data processing module transmits the processing result (including the data marked as abnormal) to the marked data anomaly cause analysis unit. The marked data anomaly cause analysis unit analyzes the cause of the processing result abnormality. After the intelligent analysis module completes the analysis of the abnormal cause, the analysis result is transmitted to the computer network repair and big data management module. The analysis result includes the network anomaly cause and the marked data anomaly cause.
[0021] Further, the computer network repair and big data management module includes a computer network repair unit and a big data management unit. According to the detailed analysis result transmitted by the intelligent analysis module, the computer network repair unit locates the fault point or potential security threat in the network. Subsequently, the computer network repair unit performs a series of targeted repair operations, including but not limited to adjusting the network configuration, isolating the infected device, optimizing the network traffic, etc., to ensure the stability and security of the network system. The computer network repair module has the ability to monitor the network state in real time, timely discovers and responds to the abnormal situation in the network, and prevents the fault from expanding or the security event from deteriorating through the preset emergency response mechanism. The big data management unit is responsible for data classification, storage optimization, and security protection. Based on the analysis result of the intelligent analysis module and the marked data anomaly cause analysis unit, the big data management unit classifies the collected big data, divides the data into different categories by identifying the type, source, purpose, etc. of the data, and implements storage optimization strategies according to the storage needs of the big data. Through compression, deduplication, distributed storage, etc., the storage space occupation is reduced, and the storage efficiency is improved. At the same time, a reasonable storage strategy is formulated according to the access frequency and importance of the data to ensure efficient access and long-term preservation of the data. The big data management unit is also responsible for data security protection. Through encryption, access control, audit, etc., data leakage, tampering, and illegal access, etc. are prevented. At the same time, regular backup and recovery drills are conducted to ensure quick recovery in case of data loss or damage.
[0022] Further, the multi-source data verification unit cross- verifies the collected data. The specific steps are as follows:
[0023] Data comparison and difference analysis: comparing the data from different data sources, putting the data of the same type or the same dimension together, and comparing their values, formats, timestamps, etc. one by one;
[0024] Assuming we are analyzing transaction data of an e-commerce website, the data comes from the website's backend database, payment system logs, and third-party logistics platforms. We may compare data items such as order numbers, transaction times, payment amounts, and product information for the same order to check for discrepancies between them;
[0025] Technology identification: After comparing the data discrepancies, train a machine learning model to learn the normal data pattern and identify abnormal data that does not conform to the normal data pattern. This can help us distinguish between normal differences (such as time differences due to time zone differences) and abnormalities (such as payment amounts not matching order amounts);
[0026] Problem positioning and correction: According to the identification results of the machine learning model, locate the specific data problems and take appropriate measures to correct them, including correcting incorrect data, deleting redundant data, merging duplicate data, etc.;
[0027] If the payment amount of an order does not match the order amount, it may be a payment system record error. In this case, contact the payment system provider for verification and adjust the data according to the actual situation;
[0028] Verification and feedback: Re-verify the corrected data to ensure that the problem is correctly solved. At the same time, feedback the experience and lessons learned during this process to the data source interface unit and data format parsing unit.
[0029] Further, the adaptive adjustment strategy module dynamically adjusts the data collection frequency, the specific steps are:
[0030] Monitoring mechanism: The adaptive adjustment strategy module uses network monitoring tools and system resource monitoring interfaces (such as CPU, memory, storage usage, etc.) to obtain real-time network load and device resource usage;
[0031] Threshold judgment: Set reasonable thresholds (such as network load rate, CPU usage, etc.) to determine whether the current system state is in a high load or resource shortage state;
[0032] Dynamic adjustment algorithm: When the system state exceeds the pre-set threshold, use a PID controller to calculate the control amount based on the deviation between the current state and the target state for stable control and adjustment;
[0033] Parameter setting: proportional gain Kp: according to the system response speed and the size of the steady-state error, select the appropriate proportional gain; integral time Ti: according to the system steady-state error requirements, select the appropriate integral time; differential time Td: according to the system oscillation characteristics and the response requirements of the rapid change, select the appropriate differential time; error: for the target to be controlled, collect the error of the feedback data;
[0034] The control amount is calculated by using the PID algorithm, and the calculation is performed according to the system state and the error;
[0035] According to the size of the current error, directly output the control amount proportional to the error, and the proportional gain Kp determines the speed of the control effect, and the output is: Output_P = Kp*Error;
[0036] According to the size of the accumulated error, output the control amount proportional to the accumulated error, and the integral time Ti determines the speed of the integral and the elimination ability of the steady-state error, and the output is: Output_I = Ki*∫Error dt;
[0037] According to the size of the error change rate, output the control amount proportional to the change rate. The differential time Td determines the sensitivity and smoothness of the error change rate, and the output is: Output_D = Kd*d(Error) / dt;
[0038] The output of the PID controller is the superposition of the three parts: ControlOutput = Output_P + Output_I + Output_D;
[0039] The calculated control amount is taken as the output, the frequency of data acquisition is dynamically adjusted, and the above steps are periodically repeated.
[0040] Further, the intelligent processing unit utilizes the pattern recognition capability of the machine learning algorithm to automatically monitor and correct the data quality, and the specific steps are:
[0041] Noise detection and processing: use support vector machine to detect noise in data, the support vector machine automatically finds abnormal values and outliers in data, which are often caused by noise, process the detected noise, and the processing method includes filtering, smoothing, etc. Filtering algorithm can remove high-frequency noise in data, and smoothing algorithm can make data smoother and reduce the influence of random fluctuations;
[0042] For the noise in network traffic data, we can use moving average filtering or low-pass filtering method. For example, when using moving average filtering, we can take the average value of the current data point and its previous and subsequent data points as the corrected value of the data point to reduce the influence of random fluctuations;
[0043] Outlier detection and processing: Outlier detection is the process of identifying and labeling points in a dataset that significantly deviate from other observations, which is achieved by Isolation Forest, outlier detection is performed using the Isolation Forest algorithm, one or more isolation trees are constructed to isolate data points, since outliers are usually isolated in data space, they will be isolated faster by isolation trees, by calculating the path length of each data point and comparing it with the threshold, outliers are identified;
[0044] First, the Isolation Forest model is trained using the training data. Then, the test data is input into the model for prediction, and the anomaly score of each data point is obtained. Finally, according to the anomaly score and the preset threshold, the outliers are labeled out;
[0045] Missing value detection and filling: The preprocessing unit has already performed preliminary missing value processing, but in this stage, the integrity of the data is checked again, by finding the K most similar samples to the missing value sample, and then predicting the missing value according to the corresponding values of these samples, to ensure that there are no missing values, for the detected missing values, select the appropriate filling method according to the characteristics and context of the data;
[0046] If the data of a certain time point in the network traffic data is missing, we can fill it according to the trend of the data before and after that time point using linear interpolation or polynomial interpolation.
[0047] The present application has the following advantages:
[0048] 1、In the present application, an intelligent data processing module is developed, which automatically identifies and processes noise, outliers and missing values in the data using machine learning algorithms, reducing the workload of subsequent processing, realizing data compression and encoding technology, and reducing the overhead of data transmission and storage under the premise of maintaining data integrity.
[0049] 2、In the present application, data is collected from multiple data sources and cross-validated to improve data accuracy, by comparing the data differences of different data sources, data errors can be discovered and corrected in time, according to the changes of network environment and device state, the frequency and accuracy of data collection are automatically adjusted, to ensure that high-quality data can be collected under limited resources.
[0050] 3、In the present application, a knowledge graph is constructed to associate and integrate data from different fields, forming a cross-field knowledge network, and improving the breadth and depth of data analysis. BRIEF DESCRIPTION OF DRAWINGS
[0051] Figure 1 A system block diagram of a computer data acquisition and processing analysis system is proposed in the present application. DETAILED DESCRIPTION
[0052] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work are within the protection scope of the present application.
[0053] Please refer to Figure 1 The present application is a computer data acquisition processing and analysis system, which includes:
[0054] Data acquisition and multi-source verification module: responsible for collecting data from multiple independent and reliable data sources, and performing multi-source data verification to ensure data accuracy and consistency;
[0055] Adaptive adjustment strategy module: responsible for dynamically adjusting the frequency and accuracy of data acquisition according to network environment and device status, and optimizing resource utilization;
[0056] Intelligent data processing module: responsible for automatically processing noise, outliers and missing values in data using machine learning algorithms, and performing data compression and encoding to reduce transmission and storage overhead;
[0057] Computer network state acquisition and diagnosis module: responsible for collecting state information of computer network during operation and performing preliminary diagnosis;
[0058] Intelligent analysis module: responsible for in-depth analysis of network anomaly causes and marking data anomaly causes, and performing data mining and correlation analysis using deep learning and knowledge graph technology;
[0059] Computer network repair and big data management module: responsible for repairing computer network according to analysis results of intelligent analysis module, and effectively managing big data;
[0060] The data acquisition and multi-source verification module provides raw data, the adaptive adjustment strategy module adjusts the acquisition strategy according to the data quality and system resource state, forms a closed-loop feedback, and the acquisition data adjusted by the adaptive adjustment strategy module is input to the intelligent data processing module, the preprocessed data of the intelligent data processing module is transmitted to the computer network state acquisition and diagnosis module and the intelligent analysis module, the intelligent analysis module provides analysis results, and the computer network repair and big data management module performs repair and management operations according to the analysis results.
[0061] The computer data acquisition, processing, and analysis system proposed in this invention works as follows: the data acquisition and multi-source verification module collects data from multiple registered data sources in parallel, performs multi-source comparison and verification on the collected data to ensure data accuracy and consistency. Inconsistent data is verified by attempting to obtain verification from other reliable sources or being marked as suspicious.
[0062] The adaptive adjustment strategy module continuously monitors the network environment and device status, such as bandwidth usage and CPU load, and dynamically adjusts the frequency and accuracy of data collection based on the monitoring results. For example, it reduces the collection frequency when the network is congested and improves the collection accuracy when the device load is low. The intelligent data processing module receives the adaptively adjusted data and uses machine learning algorithms to automatically process noise, outliers, and missing values. It then efficiently compresses and encodes the pre-processed data to reduce transmission and storage overhead.
[0063] The computer network status collection and diagnosis module collects real-time network status information during operation, such as latency, packet loss rate, and bandwidth. It performs preliminary diagnosis based on the collected status information and identifies possible network problems or anomalies. The intelligent analysis module receives pre-processed data and preliminary diagnosis results, and uses deep learning and knowledge graph technology to perform data mining and association analysis, identifying the root causes of network anomalies and data anomaly patterns, marking the identified anomalies and providing detailed analysis reports.
[0064] Based on the analysis results provided by the intelligent analysis module, the computer network repair module performs corresponding repair operations, such as reconfiguring routes, optimizing network settings, etc. At the same time, it effectively manages big data, including data archiving, backup, security auditing, etc., to ensure data integrity and availability.
[0065] In one embodiment, for the above-mentioned data collection and multi-source verification module, the data collection and multi-source verification module includes a data source interface unit, a data format analysis unit, a data cache and queue unit, and a multi-source data verification unit. The data source interface unit is responsible for establishing connections with various data sources. The data format analysis unit performs format analysis on the received data and converts it into a unified internal data format. During data collection, a data cache and queue mechanism is introduced. The data cache and queue unit is responsible for temporarily storing the parsed data in the cache or placing it in the message queue for subsequent processing. The multi-source data verification unit is responsible for cross-verification of the collected data. This step is crucial because it can identify and correct errors, inconsistencies, or redundant information in the data, thereby improving data accuracy and reliability. The data source interface unit transmits the collected raw data to the data format analysis unit for format analysis. The parsed data is transmitted to the data cache and queue unit for temporary storage or queuing for processing. The data in the cache or queue is taken out in turn and enters the multi-source data verification unit for preliminary verification. For example, in network traffic data analysis, the system compares data from network devices, monitoring systems, and log files to find any inconsistencies and correct them accordingly.
[0066] In one embodiment, for the above-mentioned adaptive adjustment strategy module, the adaptive adjustment strategy module monitors network load and device resource usage in real time. When the network load is high or the device resources (such as CPU, memory, storage, etc.) are strained, the adaptive adjustment strategy module automatically reduces the frequency of data collection to reduce the burden on system performance. In the case of good network conditions and sufficient device resources, the adaptive adjustment strategy module correspondingly increases the accuracy and frequency of data collection. By dynamically adjusting the frequency and accuracy of data collection, the adaptive adjustment strategy module ensures data quality while minimizing unnecessary resource consumption.
[0067] In one embodiment, for the above-mentioned intelligent data processing module, the intelligent data processing module includes a preprocessing unit, an intelligent processing unit, and a data compression and encoding technology unit. The preprocessing unit preprocesses the original data, including data cleaning (such as removing duplicates, formatting, etc.) and data conversion (such as converting text data into numerical data for machine learning model processing). Assuming we have a set of network traffic data, we first need to check if there are duplicate records in the data, and if so, perform deduplication processing. At the same time, we also need to convert text format data such as timestamps into time series data format suitable for analysis. The intelligent processing unit uses the pattern recognition capabilities of machine learning algorithms to automatically identify and process noise, outliers, and missing values in the data, and automatically monitor and correct data quality. The data compression and encoding technology unit optimizes the processed data, reduces redundant information in the data while maintaining data integrity and core information, significantly reduces the bandwidth and storage space required for data transmission, and improves data readability and maintainability through Huffman coding.
[0068] In one embodiment, for the above-mentioned computer network state collection and diagnosis module, the computer network state collection and diagnosis module includes a computer network state collection unit and a network diagnosis unit. The computer network state collection unit is responsible for real-time collection of key parameters of the computer network running, including but not limited to network load information (such as CPU usage, memory occupancy), network traffic information (such as inbound and outbound data volume, packet size, transmission speed), and network status information (such as connection status, delay, packet loss rate, etc.). The computer network state collection unit transmits the collected raw data to the network diagnosis unit. The network diagnosis unit receives data from the network state collection unit and uses pre-set diagnostic logic to analyze the network status in depth to identify abnormal conditions in the network (such as congestion, failure, security threats, etc.). The network diagnosis unit outputs the diagnosis results in an easy-to-understand and operate form, usually including abnormal type, location, severity, and recommended repair measures, etc. The network diagnosis unit transmits the diagnosis results to the intelligent analysis module.
[0069] In an embodiment, for the intelligent analysis module described above, the intelligent analysis module includes a network anomaly cause analysis unit and a marked data anomaly cause analysis unit, the computer network state acquisition and diagnosis module transmits the network state information collected and preliminarily diagnosed in real time to the network anomaly cause analysis unit, the network anomaly cause analysis unit performs in-depth analysis of the network anomaly cause based on the network state information, the intelligent data processing module transmits the processing result (including the data marked as abnormal) to the marked data anomaly cause analysis unit, the marked data anomaly cause analysis unit analyzes the cause of the processing result abnormality, and the intelligent analysis module transmits the analysis result to the computer network repair and big data management module after completing the analysis of the abnormal cause, the analysis result including the network anomaly cause and the marked data anomaly cause.
[0070] In an embodiment, for the computer network repair and big data management module described above, the computer network repair and big data management module includes a computer network repair unit and a big data management unit, according to the detailed analysis result transmitted by the intelligent analysis module, the computer network repair unit locates the fault point or potential security threat in the network, and then performs a series of targeted repair operations, including but not limited to adjusting the network configuration, isolating the infected device, optimizing the network traffic, etc., to ensure the stability and security of the network system, the computer network repair module has the ability to monitor the network state in real time, timely discovers and responds to abnormal situations in the network, and prevents the fault from expanding or the security event from deteriorating through the preset emergency response mechanism, the big data management unit is responsible for data classification, storage optimization and security protection, based on the analysis result of the intelligent analysis module and the marked data anomaly cause analysis unit, the big data management unit will classify and process the collected big data, divide the data into different categories by identifying the type, source, purpose, etc. of the data, and implement storage optimization strategies according to the storage needs of the big data. Through compression, deduplication, distributed storage and other technical means, the storage space occupation is reduced and the storage efficiency is improved, at the same time, reasonable storage strategies are formulated according to the access frequency and importance of the data to ensure efficient access and long-term preservation of the data, and the big data management unit is also responsible for the security protection of the data. Through encryption, access control, audit and other security measures, data leakage, tampering and illegal access, etc. are prevented. At the same time, regular backup and recovery drills are carried out to ensure quick recovery in case of data loss or damage.
[0071] In an embodiment, for the multi-source data verification unit described above, the multi-source data verification unit cross- verifies the collected data, and the specific steps are as follows:
[0072] Data comparison and difference analysis: Compare data from different data sources, put data of the same type or dimension together, and compare their values, formats, timestamps and other attributes one by one;
[0073] Imagine we're analyzing transaction data from an e-commerce website. The data comes from the website's backend database, payment system logs, and a third-party logistics platform. We might compare data items like order number, transaction time, payment amount, and product information for the same order to check for discrepancies.
[0074] Technical Identification: After identifying data discrepancies, we train a machine learning model to learn normal data patterns and identify abnormal data that does not conform to these patterns. This helps us distinguish between normal discrepancies (such as time differences due to different time zones) and abnormalities (such as a mismatch between the payment amount and the order amount).
[0075] Problem location and correction: Based on the identification results of the machine learning model, specific data problems are located and appropriate measures are taken to correct them. This includes correcting erroneous data, deleting redundant data, and merging duplicate data.
[0076] If the payment amount of an order is found to be inconsistent with the order amount, it may be due to an error in the payment system record. In this case, you need to contact the payment system provider to verify the situation and adjust the data according to the actual situation;
[0077] Verification and Feedback: Re-verify the corrected data to ensure the problem is correctly resolved. At the same time, feedback the experience and lessons learned from this process to the data source interface unit and the data format parsing unit.
[0078] In one embodiment, for the above-mentioned adaptive adjustment strategy module, the adaptive adjustment strategy module dynamically adjusts the data collection frequency, specifically in the following steps:
[0079] Monitoring mechanism: The adaptive adjustment strategy module obtains the current network load and device resource usage in real time through network monitoring tools and system resource monitoring interfaces (such as CPU, memory, storage, etc. usage);
[0080] Threshold judgment: Set reasonable thresholds (such as network load rate, CPU usage, etc.) to determine whether the current system status is in a high load or resource-constrained state;
[0081] Dynamic adjustment algorithm: When the system status exceeds the preset threshold, the PID controller is used to calculate the control quantity based on the deviation between the current state and the target state, and perform stable control and adjustment;
[0082] Parameter setting: proportional gain Kp: according to the system response speed and the size of the steady-state error, select the appropriate proportional gain; integral time Ti: according to the system steady-state error requirements, select the appropriate integral time; differential time Td: according to the system oscillation characteristics and the response requirements of the rapid change, select the appropriate differential time; error: for the target to be controlled, collect the error of the feedback data;
[0083] The control amount is calculated by using the PID algorithm, and the calculation is performed according to the system state and the error;
[0084] According to the size of the current error, directly output the control amount proportional to the error, and the proportional gain Kp determines the speed of the control effect, and the output is: Output_P=Kp*Error;
[0085] According to the size of the accumulated error, output the control amount proportional to the accumulated error, and the integral time Ti determines the speed of the integral and the elimination ability of the steady-state error, and the output is: Output_I=Ki*∫Error dt;
[0086] According to the size of the error change rate, output the control amount proportional to the change rate. The differential time Td determines the sensitivity and smoothness of the error change rate, and the output is: Output_D=Kd*d(Error) / dt;
[0087] The output of the PID controller is the superposition of the three parts: Control Output=Output_P+Output_I+Output_D;
[0088] The calculated control amount is taken as the output, and the frequency of data acquisition is dynamically adjusted, and the above steps are repeated periodically.
[0089] In one embodiment, for the above-mentioned intelligent processing unit, the intelligent processing unit uses the pattern recognition capability of the machine learning algorithm to automatically monitor and correct the data quality, and the specific steps are:
[0090] Noise detection and processing: use support vector machine to detect noise in data, support vector machine automatically finds abnormal values and outliers in data, abnormal values and outliers are often caused by noise, process the detected noise, processing methods include filtering, smoothing, etc. Filtering algorithm can remove high-frequency noise in data, and smoothing algorithm can make data smoother and reduce the influence of random fluctuations;
[0091] For the noise in network traffic data, we can use moving average filtering or low-pass filtering method. For example, when using moving average filtering, we can take the average value of the current data point and its previous and subsequent data points as the correction value of the data point to reduce the influence of random fluctuations;
[0092] Outlier detection and processing: Outlier detection is the process of identifying and labeling points in a dataset that significantly deviate from other observations. It is achieved by Isolation Forest, which uses the Isolation Forest algorithm to detect outliers, builds one or more Isolation Trees to isolate data points. Since outliers are usually more isolated in data space, they will be isolated faster by Isolation Trees. By calculating the path length of each data point and comparing it with the threshold, outliers are identified.
[0093] First, train the Isolation Forest model using the training data. Then, input the test data into the model for prediction to get the anomaly score of each data point. Finally, according to the anomaly score and the preset threshold, the outliers are marked out.
[0094] Missing value detection and filling: The pre-processing unit has already carried out preliminary missing value processing, but in this stage, the integrity of the data is checked again. By finding the K most similar samples to the missing value sample, the missing value is predicted according to the corresponding values of these samples, ensuring that there are no missing values. For the detected missing values, select the appropriate filling method according to the characteristics and context of the data.
[0095] If there is missing data at a certain time point in the network traffic data, we can fill it in by linear interpolation or polynomial interpolation according to the trend of the data before and after that time point.
[0096] Although the embodiments of the present application have been shown and described, it can be understood by those skilled in the art that various changes, modifications, replacements and deformations can be made to these embodiments without departing from the principles and spirits of the present application, and the scope of the present application is defined by the appended claims and their equivalents.
Claims
1. A computer data acquisition, processing and analysis system, characterized in that: include: Data collection and multi-source verification module: responsible for collecting data and performing multi-source data verification; Adaptive adjustment strategy module: responsible for dynamically adjusting the frequency and accuracy of data collection according to the network environment and device status; Intelligent data processing module: responsible for automatically processing noise, outliers and missing values in the data using machine learning algorithms, and performing data compression and encoding; Computer network status collection and diagnosis module: responsible for collecting status information of computer network during operation and performing preliminary diagnosis; Intelligent analysis module: responsible for in-depth analysis of the causes of network anomalies and labeled data anomalies, and using deep learning and knowledge graph technology for data mining and association analysis; Computer network repair and big data management module: responsible for repairing computer networks and effectively managing big data based on the analysis results of the intelligent analysis module; The data acquisition and multi-source verification module provides original data, the adaptive adjustment strategy module adjusts the acquisition strategy according to data quality and system resource status to form a closed-loop feedback, the acquired data adjusted by the adaptive adjustment strategy module serves as the input of the intelligent data processing module, the data pre-processed by the intelligent data processing module is passed to the computer network status acquisition and diagnosis module and the intelligent analysis module, the intelligent analysis module provides analysis results, and the computer network repair and big data management module performs repair and management operations according to the analysis results.
2. A computer data acquisition, processing and analysis system according to claim 1, characterized in that: The data acquisition and multi-source verification module includes a data source interface unit, a data format parsing unit, a data cache and queue unit, and a multi-source data verification unit. The data source interface unit is responsible for establishing connections with various data sources. The data format parsing unit performs format parsing on the received data and converts it into a unified data format within the system. During the data acquisition process, a data cache and queue mechanism is introduced. The data cache and queue unit is responsible for temporarily storing the parsed data in the cache or placing it in the message queue for subsequent processing. The multi-source data verification unit is responsible for cross-verifying the collected data. The data source interface unit transmits the collected original data to the data format parsing unit for format parsing. The parsed data is transmitted to the data cache and queue unit for temporary storage or queued for processing. The data in the cache or queue is taken out in sequence and enters the multi-source data verification unit for preliminary verification.
3. A computer data acquisition, processing and analysis system according to claim 1, characterized in that: The adaptive adjustment strategy module monitors the network load and device resource usage in real time. When the network condition is good and the device resources are sufficient, the adaptive adjustment strategy module will correspondingly improve the accuracy and frequency of data collection. By dynamically adjusting the frequency and accuracy of data collection, the adaptive adjustment strategy module can minimize unnecessary resource consumption while ensuring data quality.
4. A computer data acquisition, processing and analysis system according to claim 1, characterized in that: The intelligent data processing module includes a preprocessing unit, an intelligent processing unit and a data compression and encoding technology unit. The preprocessing unit preprocesses the original data, and the preprocessing includes data cleaning and data conversion. The intelligent processing unit uses the pattern recognition ability of the machine learning algorithm to automatically monitor and correct the data quality. The data compression and encoding technology unit optimizes the processed data and reduces redundant information in the data while maintaining data integrity and core information without loss. At the same time, the data compression and encoding technology unit uses Huffman coding to encode the data.
5. A computer data acquisition, processing and analysis system according to claim 1, characterized in that: The computer network status collection and diagnosis module includes a computer network status collection unit and a network diagnosis unit. The computer network status collection unit is responsible for real-time collection of key parameters during computer network operation. The computer network status collection unit transmits the collected raw data to the network diagnosis unit. The network diagnosis unit receives data from the network status collection unit and uses preset diagnostic logic to conduct in-depth analysis of the network status to identify abnormal conditions in the network. The network diagnosis unit transmits the diagnostic results to the intelligent analysis module.
6. A computer data acquisition, processing and analysis system according to claim 1, characterized in that: The intelligent analysis module includes a network anomaly cause analysis unit and a marked data anomaly cause analysis unit. The computer network status acquisition and diagnosis module transmits the network status information collected and preliminarily diagnosed in real time to the network anomaly cause analysis unit. The network anomaly cause analysis unit performs an in-depth network anomaly cause analysis based on the network status information. The intelligent data processing module transmits the processing result to the marked data anomaly cause analysis unit. The marked data anomaly cause analysis unit analyzes the cause of the abnormal processing result. After completing the abnormality cause analysis, the intelligent analysis module transmits the analysis result to the computer network repair and big data management module. The analysis result includes the network anomaly cause and the marked data anomaly cause.
7. A computer data acquisition, processing and analysis system according to claim 6, characterized in that: The computer network repair and big data management module includes a computer network repair unit and a big data management unit. Based on the detailed analysis results transmitted by the intelligent analysis module, the computer network repair unit locates the fault point or potential security threat in the network. Subsequently, the computer network repair unit performs a series of targeted repair operations. The computer network repair module has the ability to monitor the network status in real time, promptly discover and respond to abnormal situations in the network, and prevent the expansion of faults or the deterioration of security incidents through a preset emergency response mechanism. The big data management unit is responsible for data classification, storage optimization and security protection. Based on the analysis results of the intelligent analysis module and the data anomaly cause analysis unit, the big data management unit will classify and process the collected big data.
8. A computer data acquisition, processing and analysis system according to claim 2, characterized in that: The multi-source data verification unit performs cross-verification on the collected data, specifically the following steps: Data comparison and difference analysis: Compare data from different data sources, put data of the same type or dimension together, and compare their attributes one by one; Technical identification: After comparing data differences, a machine learning model is trained to learn normal data patterns and identify abnormal data that does not conform to normal data patterns; Problem location and correction: Based on the identification results of the machine learning model, specific data problems are located and corresponding measures are taken to correct them; Verification and feedback: Re-verify the corrected data and, at the same time, feed back the experience and lessons learned in this process to the data source interface unit and the data format parsing unit.
9. A computer data acquisition, processing and analysis system according to claim 3, characterized in that: The adaptive adjustment strategy module dynamically adjusts the data collection frequency, specifically the following steps: Monitoring mechanism: The adaptive adjustment strategy module obtains the current network load and device resource usage in real time through network monitoring tools and system resource monitoring interfaces; Threshold judgment: Set a reasonable threshold to determine whether the current system status is under high load or resource shortage; Dynamic adjustment algorithm: When the system status is detected to exceed the preset threshold, the PID controller is used to calculate the control quantity based on the deviation between the current state and the target state to perform stable control and adjustment.
10. A computer data acquisition, processing and analysis system according to claim 4, characterized in that: The intelligent processing unit uses the pattern recognition capability of machine learning algorithms to automatically monitor and correct data quality. The specific steps are as follows: Noise detection and processing: Use support vector machines to detect noise in data. The support vector machine automatically finds abnormal values and outliers in the data. These abnormal values and outliers are often caused by noise, and the detected noise is processed. Outlier detection and processing: Outlier detection is the process of identifying and marking points in a dataset that deviate significantly from other observations. This is achieved through the isolation forest algorithm, which constructs one or more isolation trees to isolate data points. Because outliers are generally isolated in the data space, they are more quickly isolated by the isolation trees. Outliers are identified by calculating the path length of each data point and comparing it with a threshold. Missing value detection and filling: Preliminary missing value processing has been performed in the preprocessing unit, but the integrity of the data is checked again at this stage. By finding the K samples most similar to the missing value samples, and then predicting the missing values based on the corresponding values of these samples, it is ensured that there are no omitted missing values. For the detected missing values, the appropriate filling method is selected according to the characteristics of the data and the context.
Citation Information
Patent Citations
Big data management platform for computer network intelligent analysis
CN114826770A