Data acquisition and analysis system for network big data information analysis

By designing a system that includes data acquisition, preprocessing, storage, analysis and visualization modules, the problems of insufficient real-time acquisition of multiple data sources, data quality assurance and analysis results in the prior art are solved, efficient data acquisition, processing and analysis are achieved, and accurate trend prediction results are generated.

CN120011713AInactive Publication Date: 2025-05-16HUAIAN XINGHE INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202411977047.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing network big data analysis systems are difficult to collect data from multiple heterogeneous data sources in real time, and the data quality assurance and accuracy of analysis results are insufficient.

Method used

A system including a data acquisition module, a data preprocessing module, a data storage module, a data analysis module and a data visualization module are designed. The system supports real-time acquisition of multiple heterogeneous data sources through the data acquisition module, uses the data preprocessing module to perform automated data cleaning, format conversion and denoising processing, and combines machine learning algorithms and data mining technology for in-depth analysis.

Benefits of technology

Real-time acquisition of multiple heterogeneous data sources is achieved, data quality is significantly improved, and a reliable data foundation is provided for subsequent analysis. In-depth analysis is used to generate efficient data models and trend prediction results, which improves the accuracy and reliability of the analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011713A_ABST
    Figure CN120011713A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of big data processing, in particular to a data acquisition and analysis system for network big data information analysis, which comprises a data acquisition module, a data preprocessing module, a data storage module, a data analysis module and a data visualization module, wherein the data acquisition module is used for acquiring original data from various data sources in real time through a network interface; the data preprocessing module is used for preprocessing the original data; the data storage module is used for carrying out storage and management by adopting a distributed database; and the data analysis module is used for performing deep analysis by applying a machine learning algorithm and a data mining technology to generate a trend prediction result. According to the method, real-time acquisition, automatic preprocessing, distributed storage and multi-algorithm fusion analysis of the multi-source data are realized, so that the problems of low data quality, poor processing efficiency, model isolation and the like in the prior art are solved, and the precision, efficiency and decision support capability of big data analysis are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data processing, and in particular to a data collection and analysis system for network big data information analysis. Background Art

[0002] With the rapid development of information technology and the popularization of the Internet, a large amount of network big data has been accumulated in all walks of life. These data not only contain information on business operations, user behavior, market trends, etc., but also contain potential value and can provide important decision-making support for decision makers. However, due to the wide variety, large scale and diverse formats of data, these data often present scattered, heterogeneous and complex characteristics, and traditional data processing and analysis methods are difficult to effectively extract useful information from them. How to efficiently collect, process, store and analyze these massive amounts of network big data has become a key issue that needs to be urgently solved in the current field of data science.

[0003] Existing network big data analysis systems usually face the following challenges: First, the data acquisition module often only supports a single type of data source and cannot process data from multiple heterogeneous data sources in real time, resulting in insufficient timeliness and integrity of the data; second, traditional data processing methods often rely on manual intervention and cannot achieve automated data cleaning, format conversion and denoising, resulting in difficulty in ensuring data quality, which in turn affects the accuracy of subsequent analysis results; third, although machine learning and data mining technologies can perform in-depth analysis of big data, existing technologies have information island problems when combining prediction models generated by different algorithms, making it difficult to achieve effective fusion and comprehensive analysis of multiple models, resulting in insufficient accuracy and reliability of prediction results. Summary of the invention

[0004] Based on the above objectives, the present invention provides a data collection and analysis system for network big data information analysis.

[0005] A data acquisition and analysis system for network big data information analysis includes a data acquisition module, a data preprocessing module, a data storage module, a data analysis module and a data visualization module; wherein: Data acquisition module: used to collect raw data from various data sources in real time through a network interface, and transmit the collected data to the data preprocessing module; Data preprocessing module: used to receive the raw data from the data acquisition module, perform data cleaning, format conversion and denoising, and transmit the preprocessed data to the data storage module; Data storage module: used to receive the cleaned data from the data preprocessing module, use a distributed database for storage and management, and provide storage data support for the data analysis module; Data analysis module: used to obtain stored data from the data storage module, apply machine learning algorithms and data mining technology to perform in-depth analysis, generate trend prediction results, and pass the analysis results to the data visualization module; Data visualization module: used to receive the analysis results from the data analysis module and display the analysis results to the user through visualization means.

[0006] Optionally, the data acquisition module includes a data source identification unit, a data connection management unit, a data acquisition unit, a data buffer unit and a data transmission unit; wherein: Data source identification unit: used to identify and classify data types and data formats from different types of network data sources, including social media platforms, IoT devices, and user behavior logs; Data connection management unit: used to establish network connections with various data sources; Data collection unit: used to collect raw data from various data sources in real time using corresponding protocols according to the classification information provided by the data source identification unit; Data buffer unit: used to temporarily store the collected raw data; Data transmission unit: used to transmit the original data in the buffer unit to the data preprocessing module through a secure network protocol.

[0007] Optionally, the data preprocessing module includes a data cleaning unit, a format conversion unit, a data denoising unit, a data verification unit and a data standardization unit; wherein: Data cleaning unit: used to identify and remove redundant information, erroneous data and inconsistent data in the original data. It uses a rule-based filtering algorithm to set data quality rules and thresholds, and uses a rule engine to filter the original data and remove erroneous data that does not meet the conditions. The format conversion unit is used to convert the data processed by the data cleaning unit into a unified standard format. It adopts a mapping conversion algorithm and defines data mapping rules to convert data in different formats into a unified structured format, including JSON, CSV or XML, so that data from different data sources can be uniformly represented; Data denoising unit: used to eliminate noise in data by applying wavelet transform denoising algorithm. Based on the multi-scale analysis principle of wavelet transform, the signal is decomposed into different frequency sub-bands and high-frequency noise components are removed, thereby retaining the main information of the data; Data verification unit: used to verify the data processed by the format conversion unit and the data denoising unit. It uses an integrity check algorithm to check the integrity of the data fields to verify whether each data record is missing the required field value. It also uses a consistency verification algorithm to check whether the associated data between different data sources is consistent, ensuring that the data meets the predetermined quality standards. Data normalization unit: used to normalize the verified data according to the predetermined standard. The Z-score normalization method is used to convert the data into standardized data with zero mean and unit variance by calculating the difference between each data point and the mean of the data set and dividing the difference by the standard deviation of the data set.

[0008] Optionally, the data storage module includes a data sharding unit, a data replication unit and an index management unit; wherein: A data sharding unit is used to divide the cleaned data into different data shards according to a predetermined sharding strategy, and allocate and store each shard to a corresponding database node; Data replication unit: used to replicate data between multiple database nodes to ensure that each data shard has a copy on at least two nodes; Index management unit: used to establish multi-dimensional indexes according to data types and query requirements, including primary key indexes, secondary indexes, and full-text indexes.

[0009] Optionally, the data slicing unit includes: Data partitioning: The cleaned data is divided into multiple data shards according to the predetermined sharding strategy. The sharding strategy includes sharding by time, by geographic location or by data type. The data is sharded using a hash algorithm. , using the following hash function To calculate its fragment identifier, the expression is: ,in, For application in data Multiple hash functions on is the total number of shards in the database, It’s data The shard identifier to which it belongs; Data allocation: based on data shard identifier , distribute data shards to different database nodes; specifically, the consistent hashing algorithm is used to map database nodes and data shards to a virtual hash ring, using shard identifiers and the hash value of the database node to determine the storage location of the data shard; Data storage: The divided and allocated data shards are stored in the specified database nodes through the network protocol. Once the data shards are allocated to the specified node, the data will be written to the storage medium of the node.

[0010] Optionally, the data analysis module includes a machine learning algorithm unit, a data mining technology unit, a prediction model generation unit, and a trend prediction result generation unit; wherein: Machine learning algorithm unit: used to apply the support vector machine algorithm to classify and regress the data obtained from the data storage module, and to achieve accurate data classification and trend prediction by constructing a hyperplane to maximize the class interval; Data Mining Technology Unit: Used to apply the Apriori association rule mining algorithm to mine frequent item sets in data. By calculating the support and confidence of the item set, strong association rules in the data are discovered to provide a basis for optimized decision-making. Prediction model generation unit: used to model time series data based on the long short-term memory network algorithm to capture long-term dependencies; Trend prediction result generation unit: used to combine the model generated by the support vector machine and the long short-term memory network to comprehensively calculate and output the prediction results of future trends.

[0011] Optionally, the machine learning algorithm unit includes: Data preprocessing: Standardize the raw data obtained from the data storage module to ensure that each feature is within the same scale range; Constructing a support vector machine model: Processing the standardized data through the selected kernel function, including linear kernel or radial basis kernel, and constructing a hyperplane for classification or regression analysis; Support vector machine model training: Using the Lagrange multiplier method in the support vector machine algorithm, the optimal hyperplane parameters are obtained by solving the following dual problem: ,in, is the Lagrange multiplier, represents the total number of training samples, is the tag value, and is the input feature; Model application: Use the trained support vector machine model to classify or regress new data. For classification tasks, the decision function is calculated as: ,in, is the predicted value, For input data, and are the hyperplane parameters obtained through training.

[0012] Optionally, the data mining technology unit includes: Frequent item set generation step: By using the Apriori algorithm to process the data, first generate a candidate item set from the data set and calculate the support of each item set; the support is defined as the frequency of the item set appearing in the database; Candidate item set expansion: Based on the current frequent item set, generate higher-order item sets of candidate two-item sets or three-item sets, calculate their support, and filter out the item sets that meet the minimum support threshold; Association rule generation: Based on frequent item sets, association rules are generated using confidence; Latent relationship disclosure: Combining the context of the data and the strength of the association rules to identify potential implicit connections.

[0013] Optionally, the prediction model generating unit includes: Data preparation: Split the data into training and test sets in chronological order and use standardization or normalization methods to process each feature to ensure that the input data is in the same numerical range; Long short-term memory network unit construction: By constructing a long short-term memory network structure, it is used to capture the long-term dependencies in time series data; the internal state is calculated and updated through the preset formula to calculate the current time step Output and update hidden state; Model training: Train the LSTM model through the back-propagation algorithm to optimize the weight matrix and bias terms to minimize the prediction error; Model validation and prediction: The trained LSTM model is used to validate and predict new time series data, and the output of each time step is calculated through forward propagation. , and generate the prediction results of future trends; the prediction formula is: ,in, is the predicted output value, is the hidden state of the current time step, is the weight matrix of the output layer, is the bias term.

[0014] Optionally, the trend prediction result generating unit includes: Support vector machine output processing: Receive the classification or regression output results of the support vector machine model in the machine learning algorithm unit, and set the output result to , and standardize the output result to form a standardized support vector machine model output ; Long short-term memory network output processing: Receive the time series prediction results from the LSTM model in the prediction model generation unit, and set the result to , and standardize the result to form the standardized LSTM prediction result ; Comprehensive prediction calculation: Take the weighted average of the standardized support vector machine model and LSTM model output results to obtain a comprehensive trend prediction result ; Trend forecast result output: Based on the comprehensive forecast results , generate forecast values ​​of future trends and pass them to the data visualization module.

[0015] Beneficial effects of the present invention: The present invention, through data collection, processing, storage and analysis technology, can effectively solve the problems of real-time collection of multiple data sources, data quality assurance and insufficient accuracy of analysis results in the prior art; firstly, the system supports the real-time collection of multiple heterogeneous data sources through the data collection module, ensuring the integrity and timeliness of the data, and avoiding the problems of data islands and lags in traditional methods; secondly, the data preprocessing module uses automated data cleaning, format conversion and denoising technology to significantly improve data quality, provide a reliable data basis for subsequent analysis, and avoid errors and inconsistencies caused by manual intervention; The present invention, by combining machine learning algorithms and data mining techniques, can perform in-depth analysis and generate efficient data models and trend prediction results; in particular, by integrating support vector machines and long short-term memory network algorithms, the system can effectively capture long-term dependencies in time series data and provide accurate trend predictions; and through the data visualization module, the analysis results can be intuitively displayed to users, facilitating quick understanding and decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only for the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0017] Figure 1 A schematic diagram of a data acquisition and analysis system according to an embodiment of the present invention; Figure 2 Schematic diagram of a data analysis module according to an embodiment of the present invention. DETAILED DESCRIPTION

[0018] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. At the same time, it is explained here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments, and those skilled in the art may also adopt other alternatives to implement some known technologies; and the accompanying drawings are only for more specific description of the embodiments, and are not intended to specifically limit the present invention.

[0019] like Figure 1-Figure 2 As shown, a data acquisition and analysis system for network big data information analysis includes a data acquisition module, a data preprocessing module, a data storage module, a data analysis module and a data visualization module; wherein: Data acquisition module: used to collect raw data from various data sources in real time through a network interface, and transmit the collected data to the data preprocessing module; Data preprocessing module: used to receive the raw data from the data acquisition module, perform data cleaning, format conversion and denoising to ensure data quality, and transmit the preprocessed data to the data storage module; Data storage module: used to receive cleaned data from the data preprocessing module, use a distributed database for storage and management, ensure data scalability and access speed, and provide storage data support for the data analysis module; Data analysis module: used to obtain stored data from the data storage module, apply machine learning algorithms and data mining technology to perform in-depth analysis, generate trend prediction results, and pass the analysis results to the data visualization module; Data visualization module: used to receive the analysis results from the data analysis module and display the analysis results to users through visualization to facilitate user understanding and decision-making.

[0020] The data acquisition module includes a data source identification unit, a data connection management unit, a data acquisition unit, a data buffer unit and a data transmission unit; wherein: Data source identification unit: used to identify and classify data types and data formats from different types of network data sources, including social media platforms, IoT devices, and user behavior logs; Data connection management unit: used to establish network connections with various data sources to ensure the stability and reliability of data transmission; Data collection unit: used to collect raw data from various data sources in real time using corresponding protocols according to the classification information provided by the data source identification unit; Data buffer unit: used to temporarily store the collected raw data to ensure the integrity and continuity of the data during data transmission; Data transmission unit: used to transmit the original data in the buffer unit to the data preprocessing module through a secure network protocol; by subdividing the data acquisition module into a data source identification unit, a data connection management unit, a data acquisition unit, a data buffer unit and a data transmission unit, the clear division of labor and collaborative work of each unit ensure the accuracy and reliability of the data acquisition process, and provide a solid data foundation for subsequent data preprocessing and analysis.

[0021] The data preprocessing module includes a data cleaning unit, a format conversion unit, a data denoising unit, a data verification unit and a data standardization unit; wherein: Data cleaning unit: used to identify and remove redundant information, erroneous data and inconsistent data in the original data. It uses a rule-based filtering algorithm to set data quality rules and thresholds, and uses a rule engine to filter the original data and remove erroneous data that does not meet the conditions. The rule engine determines the validity and consistency of the data based on preset rules (such as inconsistent data format, values ​​outside the normal range, duplicate data, etc.), thereby improving the accuracy and reliability of the data. The format conversion unit is used to convert the data processed by the data cleaning unit into a unified standard format. It adopts a mapping conversion algorithm and defines data mapping rules to convert data in different formats into a unified structured format, including JSON, CSV or XML, so that data from different data sources can be represented in a unified manner and facilitate subsequent data storage and analysis. Specifically, the algorithm automatically performs field mapping and type conversion by matching the name and data type of the data field to achieve consistency in data format; Data denoising unit: It is used to eliminate the noise in the data by applying the wavelet transform denoising algorithm. Based on the multi-scale analysis principle of wavelet transform, the signal (data) is decomposed into different frequency sub-bands and the high-frequency noise components are removed, thereby retaining the main information of the data. By performing wavelet transform on the data, the noise and useful signal are separated, and the denoised data has higher accuracy and stability. Data verification unit: used to verify the data processed by the format conversion unit and the data denoising unit. It uses an integrity check algorithm to check the integrity of the data fields to verify whether each data record is missing the required field value. It also uses a consistency verification algorithm to check whether the associated data between different data sources remain consistent, ensuring that the data meets the predetermined quality standards and avoiding the generation of inconsistent data. Data standardization unit: used to normalize the verified data according to the predetermined standards, using the Z-score standardization method, by calculating the difference between each data point and the mean of the data set, and dividing the difference by the standard deviation of the data set, the data is converted into standardized data with zero mean and unit variance, so that data from different data sources can be compared and analyzed on the same scale; through the collaborative work of the above units, the data preprocessing module can efficiently and accurately clean, convert and denoise the raw data from the data acquisition module. This modularization not only improves the efficiency and effect of data preprocessing, but also provides high-quality data support for subsequent data storage and analysis modules.

[0022] The data storage module includes a data sharding unit, a data replication unit, and an index management unit; wherein: The data sharding unit is used to divide the cleaned data into different data shards according to the predetermined sharding strategy, and allocate and store each shard to the corresponding database node to improve data access efficiency and system scalability; Data replication unit: used to replicate data between multiple database nodes, ensuring that each data shard has a copy on at least two nodes, thereby improving data high availability and fault tolerance; Index management unit: used to establish multi-dimensional indexes according to data types and query requirements, including primary key indexes, secondary indexes and full-text indexes, to optimize data retrieval speed and query performance; through the above units, data is stored in shards and replicated between multiple nodes, which can significantly improve the access efficiency, fault tolerance and scalability of the data storage module. At the same time, through precise index management, it can accelerate the data retrieval process and improve the overall system performance.

[0023] The data sharding unit includes: Data partitioning: The cleaned data is divided into multiple data shards according to the predetermined sharding strategy. The sharding strategy includes sharding by time, by geographic location or by data type. The data is sharded using a hash algorithm. , using the following hash function To calculate its fragment identifier, the expression is: ,in, For application in data Multiple hash functions on is the total number of shards in the database, It’s data The shard identifier to which it belongs; this step is used to divide the data into multiple non-overlapping shards according to different sharding strategies; Data allocation: based on data shard identifier , distribute data shards to different database nodes; specifically, the consistent hashing algorithm is used to map database nodes and data shards to a virtual hash ring, using shard identifiers The hash value of the database node is used to determine the storage location of the data shard; the specific calculation process is: ,in, For the The hash value of the database node, Sharding data The database nodes that should be stored. This step ensures that data is distributed to the most appropriate nodes through consistent hashing and minimizes data redistribution when nodes change. Data storage: The divided and allocated data shards are stored in the specified database nodes through the network protocol. Once the data shards are allocated to the specified node, the data will be written to the storage medium of the node, such as a distributed file system or database, to ensure data persistence and efficient access. By combining the hash algorithm with the consistent hash algorithm, the cleaned data can be effectively divided into multiple data shards according to the predetermined strategy, and each data shard can be efficiently allocated to different database nodes. This not only improves the scalability of data storage, but also enhances the system's load balancing capabilities, and can ensure the stability and efficiency of data access under high concurrency.

[0024] The data analysis module includes a machine learning algorithm unit, a data mining technology unit, a prediction model generation unit, and a trend prediction result generation unit; among which: Machine learning algorithm unit: used to apply the support vector machine algorithm to classify and regress the data obtained from the data storage module, and to achieve accurate data classification and trend prediction by constructing a hyperplane to maximize the class interval; Data Mining Technology Unit: Used to apply the Apriori association rule mining algorithm to mine frequent item sets in data. By calculating the support and confidence of the item set, strong association rules in the data are discovered to provide a basis for optimized decision-making. Prediction model generation unit: used to model time series data based on the long short-term memory network algorithm to capture long-term dependencies; Trend prediction result generation unit: used to combine the model generated by support vector machine and long short-term memory network, comprehensively calculate and output the prediction results of future trends; the above units can effectively improve the accuracy and timeliness of data analysis by combining the advanced algorithms of support vector machine and long short-term memory network, discover deep patterns and potential associations in the data, thereby providing users with more accurate prediction results and supporting more scientific decision-making.

[0025] The machine learning algorithm unit includes: Data preprocessing: Standardize the raw data obtained from the data storage module to ensure that each feature is within the same scale range; Constructing a support vector machine model: The standardized data is processed by a selected kernel function, including a linear kernel or a radial basis kernel, and a hyperplane is constructed for classification or regression analysis. For classification problems, the support vector machine constructs a hyperplane by maximizing the margin, and the optimization goal is: ,in, is the normal vector of the hyperplane, is the bias term; the constraints are: ,in, is a training sample, is the label value (classification value), and the formula aims to find the optimal hyperplane that maximizes the class interval; Support vector machine model training: Using the Lagrange multiplier method in the support vector machine algorithm, the optimal hyperplane parameters are obtained by solving the following dual problem: ,in, is the Lagrange multiplier, represents the total number of training samples, is the tag value, and is the input feature, and the model finds the optimal hyperplane parameters by solving the dual problem; Model application: Use the trained support vector machine model to classify or regress new data. For classification tasks, the decision function is calculated as: ,in, is the predicted value, For input data, and is the hyperplane parameter obtained through training; the above steps can effectively classify and regress the cleaned data obtained from the data storage module by applying the support vector machine algorithm, and build an accurate prediction model, thereby improving the prediction ability of data analysis. Especially in a complex data environment, the support vector machine can provide higher accuracy and good generalization ability by constructing a hyperplane with the maximum interval.

[0026] The data mining technology unit includes: Frequent item set generation step: By using the Apriori algorithm to process the data, first generate a candidate item set (a combination of single items) from the data set, and calculate the support of each item set; the support is defined as the frequency of the item set appearing in the database, and its calculation formula is: ,in, Representing Item Sets support level; Representing Item Sets The number of occurrences in the dataset; Indicates the total number of transactions in the data set; in the frequent item set generation step, filter out the item sets that meet the preset minimum support threshold and continue the next generation process; Candidate item set expansion: Based on the current frequent item set, generate higher-order item sets of candidate two-item sets or three-item sets, calculate their support, and filter out item sets that meet the minimum support threshold; this process removes item sets that do not meet the support threshold through pruning operations, reducing the amount of calculation and improving efficiency; Association rule generation: Based on frequent item sets, association rules are generated using confidence. The confidence calculation formula for association rules is: ,in, and Respectively represent the antecedent and consequent of the association rule; Representing Item Sets support level; Representing Item Sets support level; Uncovering potential relationships: Combining the context of the data and the strength of the association rules, potential implicit connections are identified to assist decision-making and optimize business processes. The above steps use the Apriori algorithm to mine frequent item sets and association rules, which can not only efficiently discover frequent patterns and strong correlations in the data, but also reveal potential business rules and optimization directions, thereby providing strong data support for subsequent decision-making.

[0027] The prediction model generation unit includes: Data preparation: First, process the time series data obtained from the data storage module and convert it into a format suitable for the long short-term memory network algorithm; specifically, divide the data into training sets and test sets in chronological order, and use standardization or normalization methods to process each feature to ensure that the input data is in the same numerical range, thereby improving the convergence speed and accuracy of the model; Long short-term memory network unit construction: By constructing a long short-term memory network structure, it is used to capture the long-term dependencies in time series data; the internal state is calculated and updated through the preset formula below to calculate the current time step Output and update hidden state; calculate forget gate , input gate , candidate memory and output gate The specific formula is: ; ; ; ,in, Represents the hidden state of the previous time step; Represents the input data of the current time step; are the weight matrices of the forget gate, input gate, candidate memory, and output gate respectively; are the bias terms of these gates respectively; is the Sigmoid activation function, is the hyperbolic tangent activation function; through the calculation of these gates, LSTM can effectively control the forgetting and memory of information, thereby capturing the long-term dependencies in time series data; Model training: Train the LSTM model through the back-propagation algorithm to optimize the weight matrix and bias terms to minimize the prediction error; Model validation and prediction: The trained LSTM model is used to validate and predict new time series data, and the output of each time step is calculated through forward propagation. , and generate the prediction results of future trends; the prediction formula is: ,in, is the predicted output value, is the hidden state of the current time step, is the weight matrix of the output layer, is the bias term; the above steps can effectively capture the long-term dependencies in time series data through the LSTM model, handle prediction problems with long-term trends or periodicity, and provide strong data support for subsequent decision-making.

[0028] The trend prediction result generating unit includes: Support vector machine output processing: Receive the classification or regression output results of the support vector machine model in the machine learning algorithm unit, and set the output result to , and standardize the output result to form a standardized support vector machine model output , to ensure comparability with other model outputs, the formula is: ,in, Represents the prediction result of the support vector machine model; represents the mean of the output of the support vector machine model; represents the standard deviation of the output of the support vector machine model; It is the standardized support vector machine model output; Long short-term memory network output processing: Receive the time series prediction results from the LSTM model in the prediction model generation unit, and set the result to , and standardize the result to form the standardized LSTM prediction result , the formula is: ,in, Represents the time series results predicted by the LSTM model; Represents the mean of the LSTM model output; Represents the standard deviation of the LSTM model output; is the standardized LSTM output; Comprehensive prediction calculation: Take the weighted average of the standardized support vector machine model and LSTM model output results to obtain a comprehensive trend prediction result , weighting coefficient and are the weighted proportions of the SVM and LSTM models respectively, and the calculation formula is: ,in, It is the trend result of comprehensive forecast; is the weight coefficient of the SVM model, and its value range is ; is the weight coefficient of the LSTM model; and They are the standardized SVM and LSTM output results respectively; Trend forecast result output: Based on the comprehensive forecast results , generate forecast values ​​of future trends and pass them to the data visualization module; by combining the prediction results of the support vector machine (SVM) and the long short-term memory network (LSTM) model and using the weighted average method for synthesis, the accuracy and robustness of trend prediction can be effectively improved. This method can simultaneously consider the classification characteristics (SVM) and time series characteristics (LSTM) of the data, thereby providing more comprehensive and accurate prediction results in more complex prediction scenarios.

[0029] The data visualization module includes: Chart generation unit: used to dynamically generate various types of charts (including bar charts, line charts and pie charts) based on the received analysis result data and apply the data-driven document technology based on D3.js to intuitively display the data analysis results; Dashboard configuration unit: used to configure and customize dashboards according to user needs. It adopts a drag-and-drop interface design method, allowing users to flexibly combine and layout various visualization elements by dragging predefined chart components to meet different display needs; Heat map drawing unit: used to visualize the spatial distribution of data related to geographic location through heat map algorithms (such as based on Gaussian kernel density estimation), showing the density and distribution trend of data in geographic space; Interactive screening unit: used to enable users to dynamically screen and filter visualization results. It uses a combination of front-end JavaScript frameworks (such as React or Vue.js) and back-end API interfaces to support real-time data updates and interactive operations of visualization content. Report generation unit: used to automatically generate and export visualization results and analysis reports into multiple formats (such as PDF, Excel).

[0030] The present invention covers any substitution, modification, equivalent method and scheme made on the essence and scope of the present invention. In order to make the public have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention, but those skilled in the art can fully understand the present invention without the description of these details. In addition, in order to avoid unnecessary confusion about the essence of the present invention, well-known methods, processes, procedures, components and circuits are not described in detail.

[0031] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A data collection and analysis system for network big data information analysis, characterized in that: It includes data acquisition module, data preprocessing module, data storage module, data analysis module and data visualization module; among which: Data acquisition module: used to collect raw data from various data sources in real time through a network interface, and transmit the collected data to the data preprocessing module; Data preprocessing module: used to receive the raw data from the data acquisition module, perform data cleaning, format conversion and denoising, and transmit the preprocessed data to the data storage module; Data storage module: used to receive the cleaned data from the data preprocessing module, use a distributed database for storage and management, and provide storage data support for the data analysis module; Data analysis module: used to obtain stored data from the data storage module, apply machine learning algorithms and data mining technology to perform in-depth analysis, generate trend prediction results, and pass the analysis results to the data visualization module; Data visualization module: used to receive the analysis results from the data analysis module and display the analysis results to the user through visualization means.

2. A data collection and analysis system for network big data information analysis according to claim 1, characterized in that: The data acquisition module includes a data source identification unit, a data connection management unit, a data acquisition unit, a data buffer unit and a data transmission unit; wherein: Data source identification unit: used to identify and classify data types and data formats from different types of network data sources, including social media platforms, IoT devices, and user behavior logs; Data connection management unit: used to establish network connections with various data sources; Data collection unit: used to collect raw data from various data sources in real time using corresponding protocols according to the classification information provided by the data source identification unit; Data buffer unit: used to temporarily store the collected raw data; Data transmission unit: used to transmit the original data in the buffer unit to the data preprocessing module through a secure network protocol.

3. A data collection and analysis system for network big data information analysis according to claim 1, characterized in that: The data preprocessing module includes a data cleaning unit, a format conversion unit, a data denoising unit, a data verification unit and a data standardization unit; wherein: Data cleaning unit: used to identify and remove redundant information, erroneous data and inconsistent data in the original data. It uses a rule-based filtering algorithm to set data quality rules and thresholds, and uses a rule engine to filter the original data and remove erroneous data that does not meet the conditions. The format conversion unit is used to convert the data processed by the data cleaning unit into a unified standard format. It adopts a mapping conversion algorithm and defines data mapping rules to convert data in different formats into a unified structured format, including JSON, CSV or XML, so that data from different data sources can be uniformly represented; Data denoising unit: used to eliminate noise in data by applying wavelet transform denoising algorithm. Based on the multi-scale analysis principle of wavelet transform, the signal is decomposed into different frequency sub-bands and high-frequency noise components are removed, thereby retaining the main information of the data; Data verification unit: used to verify the data processed by the format conversion unit and the data denoising unit. It uses an integrity check algorithm to check the integrity of the data fields to verify whether each data record is missing the required field value. It also uses a consistency verification algorithm to check whether the associated data between different data sources is consistent, ensuring that the data meets the predetermined quality standards. Data normalization unit: used to normalize the verified data according to the predetermined standard. The Z-score normalization method is used to convert the data into standardized data with zero mean and unit variance by calculating the difference between each data point and the mean of the data set and dividing the difference by the standard deviation of the data set.

4. A data collection and analysis system for network big data information analysis according to claim 1, characterized in that: The data storage module includes a data sharding unit, a data replication unit and an index management unit; wherein: A data sharding unit is used to divide the cleaned data into different data shards according to a predetermined sharding strategy, and allocate and store each shard to a corresponding database node; Data replication unit: used to replicate data between multiple database nodes to ensure that each data shard has a copy on at least two nodes; Index management unit: used to establish multi-dimensional indexes according to data types and query requirements, including primary key indexes, secondary indexes, and full-text indexes.

5. A data collection and analysis system for network big data information analysis according to claim 4, characterized in that: The data slicing unit comprises: Data partitioning: The cleaned data is divided into multiple data shards according to the predetermined sharding strategy. The sharding strategy includes sharding by time, by geographic location or by data type. The data is sharded using a hash algorithm. , using the following hash function To calculate its fragment identifier, the expression is: ,in, For application in data Multiple hash functions on is the total number of shards in the database, It’s data The shard identifier to which it belongs; Data allocation: based on data shard identifier , distribute data shards to different database nodes; specifically, the consistent hashing algorithm is used to map database nodes and data shards to a virtual hash ring, using shard identifiers and the hash value of the database node to determine the storage location of the data shard; Data storage: The divided and allocated data shards are stored in the specified database nodes through the network protocol. Once the data shards are allocated to the specified node, the data will be written to the storage medium of the node.

6. A data collection and analysis system for network big data information analysis according to claim 1, characterized in that: The data analysis module includes a machine learning algorithm unit, a data mining technology unit, a prediction model generation unit and a trend prediction result generation unit; wherein: Machine learning algorithm unit: used to apply the support vector machine algorithm to classify and regress the data obtained from the data storage module, and to achieve accurate data classification and trend prediction by constructing a hyperplane to maximize the class interval; Data Mining Technology Unit: Used to apply the Apriori association rule mining algorithm to mine frequent item sets in data. By calculating the support and confidence of the item set, strong association rules in the data are discovered to provide a basis for optimized decision-making. Prediction model generation unit: used to model time series data based on the long short-term memory network algorithm to capture long-term dependencies; Trend prediction result generation unit: used to combine the model generated by the support vector machine and the long short-term memory network to comprehensively calculate and output the prediction results of future trends.

7. A data collection and analysis system for network big data information analysis according to claim 6, characterized in that: The machine learning algorithm unit includes: Data preprocessing: Standardize the raw data obtained from the data storage module to ensure that each feature is within the same scale range; Constructing a support vector machine model: Processing the standardized data through the selected kernel function, including linear kernel or radial basis kernel, and constructing a hyperplane for classification or regression analysis; Support vector machine model training: Using the Lagrange multiplier method in the support vector machine algorithm, the optimal hyperplane parameters are obtained by solving the following dual problem: ,in, is the Lagrange multiplier, represents the total number of training samples, is the tag value, and is the input feature; Model application: Use the trained support vector machine model to classify or regress new data. For classification tasks, the decision function is calculated as: ,in, is the predicted value, For input data, and are the hyperplane parameters obtained through training.

8. A data collection and analysis system for network big data information analysis according to claim 7, characterized in that: The data mining technology unit includes: Frequent item set generation step: By using the Apriori algorithm to process the data, first generate a candidate item set from the data set and calculate the support of each item set; the support is defined as the frequency of the item set appearing in the database; Candidate item set expansion: Based on the current frequent item set, generate higher-order item sets of candidate two-item sets or three-item sets, calculate their support, and filter out the item sets that meet the minimum support threshold; Association rule generation: Based on frequent item sets, association rules are generated using confidence; Latent relationship disclosure: Combining the context of the data and the strength of the association rules to identify potential implicit connections.

9. A data collection and analysis system for network big data information analysis according to claim 8, characterized in that: The prediction model generating unit comprises: Data preparation: Split the data into training and test sets in chronological order and use standardization or normalization methods to process each feature to ensure that the input data is in the same numerical range; Long short-term memory network unit construction: By constructing a long short-term memory network structure, it is used to capture the long-term dependencies in time series data; the internal state is calculated and updated through the preset formula to calculate the current time step Output and update hidden state; Model training: Train the LSTM model through the back-propagation algorithm to optimize the weight matrix and bias terms to minimize the prediction error; Model validation and prediction: The trained LSTM model is used to validate and predict new time series data, and the output of each time step is calculated through forward propagation. , and generate the prediction results of future trends; the prediction formula is: ,in, is the predicted output value, is the hidden state of the current time step, is the weight matrix of the output layer, is the bias term.

10. A data collection and analysis system for network big data information analysis according to claim 9, characterized in that: The trend prediction result generating unit comprises: Support vector machine output processing: Receive the classification or regression output results of the support vector machine model in the machine learning algorithm unit, and set the output result to , and standardize the output result to form a standardized support vector machine model output ; Long short-term memory network output processing: Receive the time series prediction results from the LSTM model in the prediction model generation unit, and set the result to , and standardize the result to form the standardized LSTM prediction result ; Comprehensive prediction calculation: Take the weighted average of the standardized support vector machine model and LSTM model output results to obtain a comprehensive trend prediction result ; Trend forecast result output: Based on the comprehensive forecast results , generate forecast values ​​of future trends and pass them to the data visualization module.

Citation Information

Cited By

  • Media data analysis method and system based on artificial intelligence

    CN120632191A

  • Artificial intelligence-based media data analysis method and system

    CN120632191B

  • Financial intelligent analysis system based on cloud computing

    CN120765412A

  • Book information archive data processing method and system based on big data

    CN121071022A

  • Generative BI-based data analysis system

    CN121117283A