Internet big data information processing system
Through the Internet big data information processing system, redundant features are removed using Pearson correlation coefficient and principal component analysis, combined with distributed retrieval algorithm, the inefficiency problem caused by redundant features in industrial data processing is solved, and the speed and accuracy of data analysis are improved.
Patent Information
- Application Number
- CN202510358218.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
In industrial data processing, the presence of redundant features between features leads to prolonged analysis time and inefficiency, especially in the quality classification of injection molding products, highly relevant features such as injection pressure of injection molding machines and derived parameters occupying analysis resources, affecting work efficiency.
The Internet big data information processing system is adopted, including data acquisition, preprocessing, feature extraction, redundant feature removal and data analysis modules, and redundant features are removed through Pearson correlation coefficient calculation, variable clustering analysis and principal component analysis, and a distributed retrieval algorithm is used to improve data retrieval efficiency.
Effectively remove redundant features, optimize data structures, reduce analysis resource usage, shorten analysis time, improve industrial data processing efficiency, and ensure the integrity and accuracy of data analysis.
Smart Images

Figure CN120296005A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data information processing, and in particular to an Internet big data information processing system. Background Art
[0002] With the development of the Internet big data era, a huge amount of data has been generated in the industrial field. In the process of data processing, especially for industrial data, feature extraction is crucial to improve the classification accuracy.
[0003] However, in the feature engineering stage, there may be complex relationships between features, such as feature duplication or high correlation. Multiple features may actually provide the same information, resulting in the occupation of analysis resources and the extension of analysis time during data analysis. For example, in the quality classification of injection-molded products, the injection pressure of the injection molding machine and its directly related derived parameters may exist as features at the same time, and these highly correlated features are redundant features. The existence of such redundant features will cause the digital model to focus on duplicate information during the learning process and data analysis process, extending the analysis time and reducing work efficiency. Summary of the Invention
[0004] To solve the technical problems existing in the background art, the present invention proposes an Internet big data information processing system.
[0005] An Internet big data information processing system proposed by the present invention includes:
[0006] Data acquisition module: including sensors for acquiring industrial data;
[0007] Data preprocessing module: used for preprocessing the acquired industrial data
[0008] Data feature extraction module: used for extracting features from the data after preprocessing
[0009] Redundant feature removal module: used for judging redundant features of the extracted features, and if judged as redundant features, removing the redundant features;
[0010] Data analysis module: used for performing data analysis on the extracted features after removing redundant features;
[0011] Data retrieval module: used for data retrieval.
[0012] Preferably, in the redundant feature removal module, when judging redundant features, the Pearson correlation coefficient is first used to calculate the degree of correlation. If the Pearson correlation coefficient is greater than or equal to 0.4, variable clustering analysis is performed. If two features are classified into the same category, it is judged that redundant features appear, and the features are marked. At this time, for the marked features, the proportion of missing values is compared, and a marked feature with the lowest missing value is selected, and other marked features are deleted. If there are no missing values in the marked features, or there are two or more features with the lowest missing value, the noise levels are compared, and a marked feature with the lowest noise level is selected, and other marked features are deleted.
[0013] Preferably, the Pearson correlation coefficient is used to calculate the degree of feature correlation as follows: Suppose there are n features: X1, X2,..., X n , construct a correlation coefficient matrix, which is an n×n symmetric matrix. Calculate the degree of correlation for the features in the matrix. Suppose the two features in the matrix are X i and X j , and their sample data are x i1 , x i2 , …, x in , x j1 , x j2 , …, x jn ,
[0014] Calculate the means i of X j and and
[0015] That is, add up all the sample data of variable X i , and then divide by the number of observations n;
[0016] That is, add up all the sample data of variable X j , and then divide by the number of observations n;
[0017] Then, the calculation formula for the Pearson correlation coefficient r ij is:
[0018]
[0019] Preferably, when performing variable clustering analysis, use hierarchical clustering or K-means clustering algorithm to group the features, and observe whether two features with a Pearson correlation coefficient greater than or equal to 0.4 are clustered into the same category.
[0020] Preferably, in the redundant feature removal module, when the number of extracted features after redundant feature removal is greater than the set threshold, the principal component analysis dimensionality reduction technique is adopted to map the high-dimensional data to a low-dimensional space.
[0021] Preferably, in the data retrieval module, a distributed retrieval algorithm is adopted to distribute the retrieval tasks to multiple nodes for parallel execution.
[0022] Preferably, in the data retrieval module, for one of the retrieved marked features, the redundant features related to the marked feature appear together in the retrieval results.
[0023] An Internet big data information processing method includes the following steps:
[0024] Perform data preprocessing on the collected industrial data;
[0025] Extract features from the data after data preprocessing;
[0026] Judge redundant features for the extracted features. If judged as redundant features, remove the redundant features;
[0027] Perform data analysis on the data after feature extraction and removal of redundant features.
[0028] In the present invention, the proposed Internet big data information processing system has the following beneficial technical effects:
[0029] 1. In the redundant feature removal module, first calculate the degree of correlation using the Pearson correlation coefficient. When the Pearson correlation coefficient is greater than or equal to 0.4, perform variable cluster analysis. Group the features through hierarchical clustering or K-means clustering algorithms to judge whether the features are redundant, and mark the redundant features. For the marked features, further compare the missing value ratios. If there are missing values, select a marked feature with the lowest missing value and delete other marked features; if none of the marked features have missing values, or there are two or more features with the lowest missing values, compare the noise levels, select a marked feature with the lowest noise level, and delete other marked features. This processing method not only removes redundancy but also reasonably selects and retains the most valuable features, effectively dealing with the problem of redundant features existing between features in the process of industrial data processing. By removing redundant features and optimizing the data structure, the resources occupied during data analysis are reduced, the analysis time is shortened, and the problem of low efficiency caused by focusing on duplicate information during the analysis process due to redundant features is solved, thereby improving the overall efficiency of industrial data processing.
[0030] 2. In the data retrieval module, for one of the retrieved marked features, its related redundant features appear together in the retrieval results. This helps users comprehensively understand all the information related to a specific feature, avoid missing important associated content due to only obtaining a single feature, and provide more complete data support for further analysis and decision-making.
[0031] 3. When the number of extracted features after removing redundant features is greater than the set threshold, the principal component analysis dimensionality reduction technique is adopted. By mapping high-dimensional data to a low-dimensional space, the data structure is effectively simplified. This not only reduces the computational amount during data analysis but also improves the classification efficiency, enabling industrial data to be processed more efficiently in subsequent analysis and model training, and avoiding resource waste and low efficiency problems caused by too high data dimensions.
[0032] Additional aspects and advantages of the present invention will be given in part in the following description, will become apparent in part from the following description, or will be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a schematic block diagram of the system of the present invention;
[0034] Figure 2 is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0035] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar symbols represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.
[0036] As Figure 1 shown, an Internet big data information processing system includes:
[0037] Data acquisition module: including sensors for acquiring industrial data;
[0038] Data preprocessing module: used for preprocessing the acquired industrial data;
[0039] Data preprocessing includes noise data processing and data transformation;
[0040] In an optional embodiment,
[0041] Noise data processing includes:
[0042] Bin method: The data is divided into different intervals (bins), and noise is reduced through the smoothing process of the data within the bins. The data within each bin is processed, such as taking the average value of the data within the bin to replace the original data within the bin, thereby reducing the data fluctuation.
[0043] Regression method: A regression model is established to fit the data, and the predicted values of the model are used to smooth the noisy data.
[0044] Outlier handling: Identify and handle the outliers in the data. Outliers may be caused by data entry errors, measurement errors, or real extreme values. Statistical methods (such as the Z-score method) or distance-based methods (such as the K-nearest neighbor algorithm) can be used to identify outliers. For the identified outliers, options include deletion, correction, or retention.
[0045] Data transformations include:
[0046] Standardization and normalization:
[0047] Standardization: Convert the data to a distribution with a mean of 0 and a standard deviation of 1. It is applicable to the case where the data conforms to a normal distribution. In machine learning algorithms (such as support vector machines), the standardized input data can accelerate the convergence speed of the model.
[0048] Normalization: Map the data to a specific interval, such as [0, 1] or [-1, 1]. Common normalization methods include min-max normalization. In clustering algorithms in data mining, the normalized dataset can make features of different dimensions have the same weight, improving the clustering effect.
[0049] Data discretization: Convert continuous data into discrete data. Methods such as equal-width discretization (divide the data into equally spaced intervals) or equal-frequency discretization (make each interval contain the same number of data points) can be used.
[0050] Attribute construction: Construct new attributes based on existing attributes to better mine the information in the data.
[0051] In an optional embodiment, for records of industrial data containing missing values, if the missing values are greater than the set threshold, the data is directly deleted; otherwise, the missing value filling process is performed. The methods for filling missing values can include mean filling, median filling, and mode filling.
[0052] Data feature extraction module: Used to extract features from the data after preprocessing;
[0053] Data feature extraction is to extract valuable feature information for the model from the preprocessed data. Common industrial data feature extraction methods include time-domain feature extraction, frequency-domain feature extraction, signal morphology-based feature extraction, and model-based feature extraction.
[0054] Redundant feature removal module: used to judge redundant features for the extracted features. If a redundant feature is judged, the redundant feature is removed;
[0055] Data analysis module: used to perform data analysis on the extracted features after removing redundant features;
[0056] Data retrieval module: used for data retrieval.
[0057] In the redundant feature removal module, the Pearson correlation coefficient is first used to calculate the degree of correlation. When the Pearson correlation coefficient is greater than or equal to 0.4, variable cluster analysis is performed. The features are grouped through hierarchical clustering or K-means clustering algorithms to judge whether the features are redundant and mark the redundant features. For the marked features, the missing value ratio is further compared. If there are missing values, select a marked feature with the lowest missing value and delete other marked features; if there are no missing values in the marked features, or there are two or more features with the lowest missing value, then the noise level is compared, select a marked feature with the lowest noise level and delete other marked features. This processing method removes redundancy while reasonably selecting and retaining the most valuable features, effectively dealing with the redundant feature problem existing between features in the industrial data processing process. By removing redundant features and optimizing the data structure, the resources occupied during data analysis are reduced, the analysis time is shortened, and the problem of low efficiency caused by redundant features leading to repeated information attention during the analysis process is solved, thus improving the overall efficiency of industrial data processing.
[0058] For example, in the actual industrial application scenario of injection molding product quality classification, it can avoid the interference caused by highly correlated redundant features (such as the injection pressure of the injection molding machine and its directly related derived parameters) to the analysis and improve work efficiency.
[0059] The hierarchical clustering algorithm analyzes data at different levels based on the similarity between clusters to form a tree-shaped clustering structure. It is divided into two types: agglomerative and divisive. Agglomerative hierarchical clustering starts with each data point as a separate cluster and continuously merges similar clusters; divisive is the opposite, starting with all data points in one cluster and gradually splitting into smaller clusters;
[0060] Distance metrics are usually used to determine the similarity between clusters, such as Euclidean distance and Manhattan distance;
[0061] Using the hierarchical clustering algorithm does not require specifying the number of clusters in advance; the display form of the clustering results can intuitively show the hierarchical relationship between each cluster, which is very helpful for understanding the distribution structure of the data;
[0062] K-means clustering is a partitioning clustering algorithm whose goal is to partition a dataset into K clusters. Its basic idea is to minimize the sum of the distances from the data points within each cluster to the cluster center (centroid) through an iterative approach.
[0063] The algorithm is simple and efficient, and can also achieve good clustering results for large-scale datasets; the convergence speed is relatively fast.
[0064] In an optional embodiment, in the data retrieval module, a distributed retrieval algorithm is adopted to distribute the retrieval tasks to multiple nodes for parallel execution;
[0065] By adopting a distributed retrieval algorithm, the retrieval tasks are distributed to multiple nodes for parallel execution. This approach can effectively utilize system resources and greatly improve the efficiency of data retrieval, enabling the required information to be obtained more quickly when dealing with massive industrial data.
[0066] The distributed retrieval system includes:
[0067] Dataset servers: The distributed retrieval system contains multiple dataset servers, which store a large amount of data. Each dataset server is responsible for managing and maintaining a part of the data. Different dataset servers can be located at different geographical locations or network nodes, and the data has distribution and heterogeneity;
[0068] Proxy processors: Also known as coordinators or middleware, they are responsible for receiving users' retrieval requests and distributing the requests to each dataset server for retrieval. The proxy processors also need to integrate the retrieval results from different dataset servers and finally return them to the users. The performance and functions of the proxy processors play a key role in the efficiency and accuracy of the entire distributed retrieval system.
[0069] The main algorithm process of the distributed retrieval system is as follows:
[0070] Request distribution: After a user submits a retrieval request to the distributed retrieval system, the proxy processor first parses and analyzes the request. Then, according to a certain strategy, the request is distributed to different dataset servers. For example, appropriate dataset servers can be selected based on factors such as the data distribution and the server load conditions;
[0071] Parallel retrieval: After each dataset server receives the retrieval request, it performs parallel retrieval on the local data. Since the data is distributed on different servers, retrieval operations can be performed simultaneously on multiple servers, greatly improving the retrieval efficiency. During the parallel retrieval process, each dataset server will quickly find the data related to the retrieval request according to the local data index and retrieval algorithm;
[0072] Result integration: After the dataset server completes the retrieval, it returns the retrieval results to the proxy processor. The proxy processor needs to integrate the results from different servers, remove duplicate results, and sort the results according to certain sorting rules. For example, the results can be sorted according to factors such as the relevance score of the data, the timestamp, etc., and finally the integrated results are returned to the user.
[0073] In an optional embodiment, in the data retrieval module, for one of the retrieved marked features, the redundant features related to this marked feature appear together in the retrieval results.
[0074] In the data retrieval module, for one of the retrieved marked features, the redundant features related to it appear together in the retrieval results. This helps users comprehensively understand all information related to a specific feature, avoid missing important associated content due to only obtaining a single feature, and provide more complete data support for further analysis and decision-making.
[0075] In an optional embodiment, in the redundant feature removal module, when judging redundant features, first calculate the degree of correlation using the Pearson correlation coefficient. If the Pearson correlation coefficient is greater than or equal to the set threshold, then perform variable clustering analysis. If they are classified into the same class, it is determined that redundant features appear and the features are marked. At this time, for the marked features, compare the proportion of missing values, select a marked feature with the lowest missing value, and delete the other marked features. If all the marked features have no missing values, or there are two or more features with the lowest missing value, then compare the noise levels, select a marked feature with the lowest noise level, and delete the other marked features;
[0076] In an optional embodiment, the set threshold of the Pearson correlation coefficient is 0.4.
[0077] In an optional embodiment, the Pearson correlation coefficient calculates the degree of feature correlation as follows: Suppose there are n features: X1, X2,..., X n , construct a correlation coefficient matrix, which is an n×n symmetric matrix. Calculate the degree of correlation for the features in the matrix. Let the two features in the matrix be X i and X j , and their sample data are x i1 , x i2 ,..., x in , x j1 , x j2 ,…, x jn ,
[0078] Calculate the mean of X i and X j and
[0079] Add up all the sample data of variable X i and then divide by the number of observations n;
[0080] Add up all the sample data of variable X j and then divide by the number of observations n;
[0081] Then, the Pearson correlation coefficient r ij is calculated as follows:
[0082]
[0083] In an alternative embodiment, when performing variable clustering analysis, the features are grouped using hierarchical clustering or K-means clustering algorithms, and it is observed whether two features with a Pearson correlation coefficient greater than or equal to 0.4 are clustered in the same class;
[0084] In an alternative embodiment, in the redundant feature removal module, when the number of extracted features after redundant feature removal is greater than a set threshold, the principal component analysis dimensionality reduction technique is adopted to map the high-dimensional data to a low-dimensional space, simplify the data structure through dimensionality reduction, improve efficiency, and avoid problems such as algorithm resource waste and low efficiency caused by too high data dimensionality.
[0085] Such as Figure 2 shown in an Internet big data information processing method, including the following steps:
[0086] Perform data preprocessing on the collected industrial data;
[0087] Perform feature extraction on the data after data preprocessing;
[0088] Perform redundant feature judgment on the extracted features. If it is judged as a redundant feature, then remove the redundant feature;
[0089] Perform data analysis on the data after feature extraction and redundant feature removal.
[0090] In the redundant feature removal module, the Pearson correlation coefficient is first used to calculate the degree of correlation. When the Pearson correlation coefficient is greater than or equal to the set threshold, variable clustering analysis is performed. The features are grouped through hierarchical clustering or K-means clustering algorithms to determine whether the features are redundant, and the redundant features are marked. For the marked features, the missing value ratios are further compared. If there are missing values, a marked feature with the lowest missing value is selected, and other marked features are deleted. If there are no missing values in the marked features, or there are two or more features with the lowest missing value, the noise levels are compared, and a marked feature with the lowest noise level is selected, and other marked features are deleted. Such a processing method removes redundancy while reasonably selecting and retaining the most valuable features, effectively dealing with the problem of redundant features existing among features in the industrial data processing process. By removing redundant features and optimizing the data structure, the resources occupied during data analysis are reduced, the analysis time is shortened, and the problem of low efficiency caused by focusing on duplicate information during the analysis process due to redundant features is solved, thereby improving the overall efficiency of industrial data processing.
[0091] Meanwhile, the content not detailedly described in this specification belongs to the prior art well-known to those skilled in the art.
[0092] In the embodiments provided by the present invention, it should be understood that the disclosed system or method can be implemented in other ways. For example, the above-described invention embodiments are merely illustrative. For example, the division of modules is only a logical function division, and there may be other division methods in actual implementation.
[0093] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules. They can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0094] In addition, in each embodiment of the present invention, the functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above integrated modules can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.
[0095] For those operation and maintenance personnel in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and can be implemented in other specific forms without departing from the basic features of the present invention.
[0096] As described above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.
Claims
1. An Internet big data information processing system, characterized in that, It includes: Data acquisition module: including sensors for acquiring industrial data; Data preprocessing module: for preprocessing the acquired industrial data; Data feature extraction module: for extracting features from the data after data preprocessing; Redundant feature removal module: for judging redundant features of the extracted features. If a redundant feature is judged, the redundant feature is removed; Data analysis module: for analyzing the data after feature extraction and redundant feature removal; Data retrieval module: for data retrieval.
2. The Internet big data information processing system according to claim 1, characterized in that, In the redundant feature removal module, when judging redundant features, first calculate the correlation degree using the Pearson correlation coefficient. If the Pearson correlation coefficient is greater than or equal to the set threshold, perform variable clustering analysis. If they are classified into the same class, it is judged that redundant features appear and the features are marked. At this time, for the marked features, compare the proportion of missing values, select a marked feature with the lowest missing value, and delete other marked features. If there are no missing values in the marked features, or there are two or more features with the lowest missing value, compare the noise levels, select a marked feature with the lowest noise level, and delete other marked features.
3. The Internet big data information processing system according to claim 2, wherein The Pearson correlation coefficient is used to calculate the degree of correlation between features. Suppose there are n features: X1, X2,..., X n , a correlation coefficient matrix is constructed. This matrix is an n×n symmetric matrix. The degree of correlation between the features in the matrix is calculated. Let the two features in the matrix be X i and X j , and their sample data are x i1 , x i2 , …, x in , x j1 , x j2 , …, x jn ; Calculate X i and X j the mean value of and Sum all the sample data of variable X i and then divide by the number of observations n; Add up all the sample data of variable X j and then divide by the number of observations n; Then, the Pearson correlation coefficient r ij is calculated by the following formula:
4. The Internet big data information processing system according to claim 3, wherein When performing variable clustering analysis, use hierarchical clustering or K-means clustering algorithm to group the features and observe whether two features with Pearson correlation coefficient greater than or equal to the set threshold are clustered into the same class.
5. The Internet big data information processing system according to claim 1, characterized in that, In the redundant feature removal module, when the number of extracted features after redundant feature removal is greater than the set threshold, use the principal component analysis dimensionality reduction technique to map the high-dimensional data to a low-dimensional space.
6. The Internet big data information processing system according to claim 1, characterized in that In the data retrieval module, use a distributed retrieval algorithm to distribute the retrieval tasks to multiple nodes for parallel execution.
7. The Internet big data information processing system according to claim 1, characterized in that In the data retrieval module, for one of the retrieved marked features, the redundant features related to this marked feature appear together in the retrieval results.
8. The method for processing Internet big data information according to any one of claims 1-7, characterized in that, It includes the following steps: Preprocess the acquired industrial data; Extract features from the data after data preprocessing; Judge redundant features of the extracted features. If a redundant feature is judged, the redundant feature is removed; Analyze the data after feature extraction and redundant feature removal.
Citation Information
Patent Citations
Data processing method and system based on industrial Internet and intelligent manufacturing
CN112859788A
Big data storage system for preventing data redundancy based on data classification and peer comparison
CN116303404A
Internet-of-things system for quickly querying large-scale equipment data
CN117290405A
Method and system for screening redundant data of power system
CN119377205A
Intelligent industrial data clustering and grouping method, system and device and medium
CN119537980A