Big data intelligent analysis processing system based on information fusion
By introducing Bayesian fusion algorithm and random forest model optimized by electromagnetic interference variables in the big data intelligent analysis and processing system, problems such as inconsistent data formats and data islands are solved, efficient fusion and analysis of data are achieved, and the accuracy and response capabilities of the system are improved.
Patent Information
- Application Number
- CN202510481332.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-04-17
AI Technical Summary
In the existing big data intelligent analysis and processing system based on information fusion, since the data comes from different systems and platforms, problems such as inconsistent formats, missing data, errors or inconsistencies occur frequently, affecting the accuracy and reliability of the analysis results; privacy and security issues cannot be ignored, especially when processing personal sensitive information, how to maximize the value of data utilization while ensuring data security and personal privacy; the complexity and variability of technology bring challenges to the system, the phenomenon of data silos is serious, and data sharing and integration face obstacles, limiting the full utilization of data and the full use of system functions.
Data acquisition unit is used to collect data through a multi-source data acquisition module. The information fusion unit uses Bayesian fusion algorithm optimized by electromagnetic interference variables to fuse circuit performance data of different data sources. The data analysis unit performs prediction analysis through a random forest model and introduces the minimum impurity reduction variable optimization. The real-time processing unit evaluates and warnings based on the analysis results.
Improve the comprehensiveness and diversity of data, reduce data noise and error through Bayesian fusion algorithm, optimized random forest algorithm can extract deep-level patterns and laws, the system can make responses and decisions in real time, improving response speed and efficiency.
Smart Images

Figure CN120408505A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data technology, and in particular to a big data intelligent analysis and processing system based on information fusion. Background Art
[0002] The big data intelligent analysis and processing system based on information fusion is a highly integrated multidisciplinary technology solution designed to extract valuable information and knowledge from massive and diverse data sources. The system first uses advanced data acquisition technologies and IoT devices to collect data from various sources, including structured data (such as database records), semi-structured data (such as XML files), and unstructured data (such as text, images, and videos). Next, it leverages efficient data management and storage technologies, such as distributed file systems and NoSQL databases, to ensure secure data storage and fast access. Furthermore, it uses information fusion technology to cleanse, integrate, and standardize multi-source, heterogeneous data, eliminating redundancies and inconsistencies and improving data quality. Subsequently, it combines machine learning and deep learning algorithms to conduct in-depth analysis of the processed data, uncovering potential patterns and trends for accurate prediction and decision support. The system also incorporates advanced capabilities such as natural language processing and image recognition to enhance its understanding of complex data. To ensure security and privacy throughout the entire process, the system employs strict encryption and access control policies. Finally, advanced data visualization tools are used to present the analysis results to end users in an intuitive and easy-to-understand format, helping them make rapid, informed decisions. Overall, this system not only greatly improves the efficiency and accuracy of data processing, but also provides strong support for all walks of life and promotes the development of an intelligent society.
[0003] In existing big data intelligent analysis and processing systems based on information fusion, data originates from diverse systems and platforms, leading to frequent problems such as inconsistent formats, missing data, errors, and inconsistencies. If not properly addressed, these issues will directly impact the accuracy and reliability of analytical results. Furthermore, privacy and security issues cannot be ignored. Especially when processing sensitive personal information, maximizing the value of data while ensuring data security and privacy has become a pressing challenge. The complexity and variability of technology also pose challenges to the system. With the continuous expansion of data types and scale, the requirements for data processing and analysis technologies are also increasing, requiring continuous upgrades and optimization to maintain competitiveness. Data silos between different data sources remain severe, and data sharing and integration face numerous obstacles. This not only limits the full utilization of data but also hinders the full implementation of system functionality. Therefore, a big data intelligent analysis and processing system based on information fusion is designed. Summary of the Invention
[0004] The object of the present invention is to provide a big data intelligent analysis and processing system based on information fusion, so as to solve the problems raised in the above background technology. In the existing big data intelligent analysis and processing system based on information fusion, since the data comes from different systems and platforms, problems such as inconsistent formats, data missing, errors or inconsistencies frequently occur. If these problems are not properly handled, it will directly affect the accuracy and reliability of the analysis results. Secondly, privacy and security issues cannot be ignored. Especially when dealing with personal sensitive information, how to maximize the utilization value of data on the premise of ensuring data security and personal privacy has become a difficult problem to be solved urgently. The complexity and variability of technology also pose challenges to the system. With the continuous expansion of data types and scales, the requirements for data processing and analysis technologies are getting higher and higher, and the system needs to be continuously upgraded and optimized to maintain competitiveness. The phenomenon of data islands between different data sources is still serious, and there are many obstacles to data sharing and integration, which not only limits the comprehensive utilization of data, but also hinders the full play of system functions.
[0005] To achieve the above object, the present invention aims to provide a big data intelligent analysis and processing system based on information fusion, including a data acquisition unit, and the data acquisition unit collects circuit performance data of different data sources through a multi-source data acquisition module;
[0006] An information fusion unit, which is based on the circuit performance data of different data sources collected by the data acquisition unit, and fuses the circuit performance data of different data sources through a Bayesian fusion algorithm optimized by introducing electromagnetic interference variables;
[0007] A data analysis unit, which is based on the data fused by the information fusion unit, predicts and analyzes the circuit performance through a random forest model, and introduces a minimum impurity reduction variable to optimize the random forest model for improving the accuracy of the predicted value;
[0008] A real-time processing unit, which evaluates and gives early warnings according to the analysis results of the data analysis unit.
[0009] As a further improvement of this technical solution, the data acquisition unit includes a multi-source data acquisition module, and the multi-source data acquisition module includes a database module and a sensor module;
[0010] Among them, the database module is used to store the collected data;
[0011] The sensor module is used to collect circuit performance data.
[0012] As a further improvement of the technical solution, the information fusion unit fuses the circuit performance data of different data sources collected by the data acquisition unit by using a Bayesian fusion algorithm optimized by introducing an electromagnetic interference variable. The specific steps of fusing the circuit performance data of different data sources by using the Bayesian fusion algorithm optimized by introducing an electromagnetic interference variable are as follows;
[0013] S3.1. Perform data preprocessing on the data by removing noise, filling missing values, and processing outliers, and convert data in different formats into a unified format;
[0014] S3.2. Set the prior probability distribution P(x) of the system state;
[0015] S3.3. Use the Bayesian fusion algorithm to update the posterior probability distribution of the system state.
[0016] As a further improvement of the technical solution, in S3.3, the specific steps of using the Bayesian fusion algorithm to update the posterior probability distribution of the system state are as follows:
[0017] S3.31. Obtain the measurement value of the sensor from the sensor module;
[0018] S3.32. Calculate the likelihood function of the sensor through the Gaussian distribution;
[0019] S3.33. Calculate the joint likelihood function through the Gaussian distribution;
[0020] S3.34. Update the posterior probability distribution of the system state.
[0021] As a further improvement of the technical solution, in S3.34, the specific process of updating the posterior probability distribution of the system state is as follows:
[0022] Obtain the joint likelihood function from S3.33:
[0023]
[0024] Among them, P(z|x) represents the joint likelihood function; P(z i |x) represents the likelihood function of the i-th sensor; x represents the state vector; z represents the data; z i represents the data of the i-th sensor; n represents the number of z; i represents the index variable;
[0025] Obtain the updated posterior probability distribution of the system state:
[0026]
[0027] Among them, P(x|z) represents the posterior probability; P(z) represents the marginal probability of the data z;
[0028] Considering the accuracy of the posterior probability distribution for updating the system state, an electromagnetic interference variable is introduced to optimize the Bayesian fusion algorithm:
[0029]
[0030]
[0031]
[0032] Among them, represents the i-th sensor data after introducing the electromagnetic interference variable; ∈ i represents the electromagnetic interference variable of the i-th sensor; ∈ represents the inductive interference variable; P(z|x,∈) represents the joint likelihood function after introducing the electromagnetic interference variable; P(x|z)' represents the posterior probability after introducing the electromagnetic interference variable.
[0033] As a further improvement of this technical solution, the specific steps involved in predicting and analyzing the circuit performance through the random forest model and optimizing the random forest model by introducing the minimum impurity reduction variable are as follows:
[0034] S6.1. Obtain the fused data from the information fusion unit to form a data set;
[0035] S6.2. Divide the data set into a training set and a validation set;
[0036] S6.3. Construct a random forest model through the training set;
[0037] S6.4. Use the random forest model after introducing the minimum impurity reduction variable to predict and analyze the circuit performance;
[0038] S6.5. Use the validation set to evaluate the random forest model through cross-validation.
[0039] As a further improvement of this technical solution, in the above S6.3, the specific steps for constructing a random forest model through the training set are as follows:
[0040] S6.31. Randomly extract n samples from the training set as a new training set;
[0041] S6.32. Initialize the root node and use the new training set as the input data of the root node;
[0042] S6.33. Select the optimal feature and splitting point through the Gini impurity criterion and split the node into two child nodes;
[0043] S6.34. Repeat the operations of steps S6.32 - S6.33 for each child node to select the optimal feature and splitting point and split the node until the stopping condition is reached;
[0044] S6.35. Repeat the above steps T times to construct T decision trees and form a random forest.
[0045] As a further improvement of this technical solution, in S6.33, the specific process of selecting the optimal feature and splitting point through the Gini impurity criterion and splitting the node into two child nodes is as follows:
[0046] For a node N, its Gini impurity G(N) is defined as:
[0047]
[0048] where G(N) represents the Gini impurity of the current node N; K represents the number of classes; p k represents the proportion of samples belonging to class k in node N; k represents the index variable;
[0049] For each feature j and each candidate splitting point s:
[0050] Calculate the Gini impurity of the left child node:
[0051]
[0052] where G(N L ) represents the Gini impurity of the left child node; |N Lk | represents the number of samples belonging to class k in the left child node N L ; |N L | represents the number of samples in the left child node N L ;
[0053] Calculate the Gini impurity of the right child node:
[0054]
[0055] where G(N R ) represents the Gini impurity of the right child node; |N Rk | represents the number of samples belonging to class k in the right child node N R ; |N R | represents the number of samples in the left child node N R ;
[0056] Calculate the weighted Gini impurity after splitting:
[0057]
[0058] where G split(N) represents the weighted Gini impurity after splitting; |N| represents the total number of samples in node N;
[0059] Select the feature j with the minimum weighted Gini impurity * and the splitting point s * , as the optimal feature and splitting point:
[0060]
[0061] According to the optimal feature j * and the splitting point s * , split the current node N into two child nodes N L and N R ; j represents the feature index; s represents the splitting threshold.
[0062] As a further improvement of this technical solution, in S6.4, the process of using the random forest model with the minimum impurity reduction variable introduced to predict and analyze the circuit performance is as follows:
[0063] For the new sample x, let each tree in the random forest output a predicted class;
[0064] Determine the final predicted class through majority voting:
[0065]
[0066] Among them, represents the finally predicted class; T represents the total number of decision trees in the random forest; I represents the indicator function; y t represents the predicted class of the t-th tree for the sample; c represents one of the classes; represents selecting the class c that maximizes the above ratio as the final predicted class;
[0067] For the new sample x, let each tree in the random forest output a predicted value;
[0068] Determine the predicted value through the average value:
[0069]
[0070] Among them, represents the predicted value; y t ' represents the predicted value of the t-th tree for the sample;
[0071] Considering the accuracy of the predicted value, introduce the minimum impurity reduction variable to optimize the prediction process of the regression task of the random forest model;
[0072] When constructing each tree t, the splitting node satisfies:
[0073] G(N)-Gsplit (N) ≥ min impurity decrease ;
[0074] Among them, min impurity decrease represents the set minimum impurity reduction variable;
[0075] Then, the final predicted category is determined by majority voting after introducing the minimum impurity reduction variable:
[0076]
[0077] During the construction of each tree t, the splitting node satisfies:
[0078] MSE(N) - MSE split (N) ≥ min impurity decrease ;
[0079] Among them, MSE(N) represents the mean square error of the current node N; MSE split (N) represents the weighted mean square error after splitting;
[0080] Then, the final predicted value is determined by the average value after introducing the minimum impurity reduction variable:
[0081]
[0082] Among them, represents the final predicted value after introducing the minimum impurity reduction variable.
[0083] As a further improvement of this technical solution, the real-time processing unit evaluates and gives an early warning according to the analysis result of the data analysis unit. The specific steps for evaluation and early warning are as follows:
[0084] S10.1. Receive the prediction analysis result from the data analysis unit;
[0085] S10.2. Calculate the key performance indicators according to the analysis result;
[0086] S10.3. Compare the calculated performance indicators with the preset threshold;
[0087] S10.4. Based on the result of the status evaluation, assign a risk score to the system;
[0088] S10.5. Judge whether an early warning needs to be generated according to the result of the risk assessment.
[0089] Compared with the prior art, the beneficial effects of the present invention:
[0090] 1. In the big data intelligent analysis and processing system based on information fusion, the comprehensiveness and diversity of data are improved by collecting data from multiple data sources; through the Bayesian fusion algorithm, data from different data sources are made consistent and corrected, reducing data noise and errors. The optimized Bayesian fusion algorithm can handle the uncertainty and ambiguity of data, improving the credibility of data.
[0091] 2. In the big data intelligent analysis and processing system based on information fusion, by using advanced machine learning methods such as the optimized random forest algorithm, deep - level patterns and rules can be extracted from the fused data. According to the analysis results of the data analysis unit, real - time responses and decisions can be made, improving the response speed and efficiency of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0092] Figure 1 is the overall flow block diagram of the present invention;
[0093] The meanings of the various labels in the figure are as follows:
[0094] 1. Data acquisition unit; 2. Information fusion unit; 3. Data analysis unit; 4. Real - time processing unit. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0095] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0096] Please refer to Figure 1 As shown, a big data intelligent analysis and processing system based on information fusion is provided, including a data acquisition unit 1. The data acquisition unit 1 collects circuit performance data of different data sources through a multi - source data acquisition module;
[0097] In this example, the data acquisition unit 1 includes a multi - source data acquisition module, and the multi - source data acquisition module includes a database module and a sensor module;
[0098] Among them, the database module is used to store the collected data;
[0099] Specifically, the database module is one of the core components in the multi-source data acquisition system, responsible for storing and managing data collected from various data sources (including sensors, external systems, user inputs, etc.). It usually consists of one or more database management systems (DBMS), supporting Structured Query Language (SQL) or other query interfaces for efficient data insertion, query, update, and deletion operations. The database module can not only store a large amount of historical data but also optimize the data access speed through technologies such as indexing and partitioning to ensure that the real-time processing unit and other analysis modules can quickly obtain the required information. In addition, the database module also has functions such as data backup, recovery, security, and permission control to ensure the integrity and security of the data.
[0100] The sensor module is used to collect circuit performance data.
[0101] Specifically, the sensor module is the component in the multi-source data acquisition system that directly interacts with the physical world, responsible for real-time collection of circuit performance data. It usually consists of a series of high-precision sensors that can measure various parameters in the circuit, such as voltage, current, temperature, humidity, power, frequency, etc. The sensor module is connected to the circuit through analog or digital interfaces and can capture the transient behavior and long-term trends of the circuit with high frequency and high resolution. To ensure the accuracy and reliability of the data, the sensor module usually also includes signal conditioning circuits (such as filters, amplifiers), data acquisition cards (DAQ), and timestamp functions for preprocessing and synchronously marking the collected data. In addition, the sensor module may also have a self-calibration function to automatically adjust the measurement accuracy of the sensors and reduce drift and errors.
[0102] The big data intelligent analysis and processing system based on information fusion further includes an information fusion unit 2, which fuses the circuit performance data from different data sources collected by the data acquisition unit 1 through the Bayesian fusion algorithm.
[0103] In this example, the specific steps involved in fusing the circuit performance data from different data sources by the Bayesian fusion algorithm optimized by introducing electromagnetic interference variables are as follows:
[0104] Specifically, the Bayesian fusion algorithm is a method based on Bayes' theorem for fusing information from multiple data sources. This algorithm has wide applications in fields such as multi-sensor data fusion, multi-modal data processing, and multi-source information integration. The core idea of the Bayesian fusion algorithm is to integrate data from different sources through a probability model to improve the accuracy and reliability of the data.
[0105] S3.1. Preprocess the data by removing noise, filling missing values, and handling outliers, and convert data in different formats into a unified format;
[0106] Specifically, the specific steps for preprocessing the data are as follows:
[0107] S3.11. Remove noise using the moving average algorithm;
[0108]
[0109] where y t represents the smoothed value at time t; w t-i represents the original data at time t-l; n represents the window size; l represents the index variable;
[0110] S3.12. Fill in the missing values in the data using linear interpolation;
[0111]
[0112] where x t represents the interpolation result at time t; x t+1 and x t-1 represent the observed values at adjacent time points; t t+1 and t t-1 represent the corresponding timestamps;
[0113] S3.13. Identify and handle outliers in the data using the Z-score method;
[0114]
[0115] where Z represents the Z-score; w i represents the data point; μ represents the mean of the data; σ represents the standard deviation of the data;
[0116] S3.2. Set the prior probability distribution P(x) of the system state;
[0117] Specifically, the prior probability distribution P(x) is the initial estimate of the system state and can be a distribution based on historical data or other prior knowledge.
[0118] S3.3. Update the posterior probability distribution of the system state using the Bayesian fusion algorithm.
[0119] In this example, the specific steps for updating the posterior probability distribution of the system state using the Bayesian fusion algorithm are as follows:
[0120] S3.31. Obtain the measurement values of the sensor from the sensor module;
[0121] S3.32. Calculate the likelihood function of the sensor through the Gaussian distribution;
[0122] Specifically, the process of calculating the likelihood function of the sensor through the Gaussian distribution is as follows:
[0123] Assume that the measurement value z of the i-th sensor i follows a Gaussian distribution at the given position x:
[0124]
[0125] where represents the variance of the i-th sensor.
[0126] S3.33. Calculate the joint likelihood function through the Gaussian distribution;
[0127] Specifically, the expression of the joint likelihood function calculated through the Gaussian distribution is:
[0128]
[0129] Through the joint likelihood function, the information of multi-sensor data can be more comprehensively reflected, improving the accuracy of system state estimation.
[0130] S3.34. Update the posterior probability distribution of the system state.
[0131] In this example, the specific process of updating the posterior probability distribution of the system state is as follows:
[0132] Obtain the joint likelihood function from S3.33:
[0133]
[0134] where P(x|x) represents the joint likelihood function; P(z i |x) represents the likelihood function of the i-th sensor; x represents the state vector; z represents the data; z i represents the data of the i-th sensor; n represents the number of z; i represents the index variable;
[0135] Obtain the posterior probability distribution of updating the system state:
[0136]
[0137] where P(x|z) represents the posterior probability; P(z) represents the marginal probability of the data z;
[0138] Considering the accuracy of updating the posterior probability distribution of the system state, an electromagnetic interference variable is introduced to optimize the Bayesian fusion algorithm:
[0139]
[0140]
[0141]
[0142] wherein, represents the i-th sensor data after introducing the electromagnetic interference variable; ∈ i represents the electromagnetic interference variable of the i-th sensor; ∈ represents the inductance interference variable; P(z|x,∈) represents the joint likelihood function after introducing the electromagnetic interference variable; P(x|z)' represents the posterior probability after introducing the electromagnetic interference variable.
[0143] Specifically, by explicitly considering the influence of electromagnetic interference in the sensor data, the model can more accurately capture the state of the actual system, reduce the errors caused by interference, and thus ensure that the posterior probability distribution is closer to the real situation.
[0144] Data preprocessing ensures the quality and consistency of the input data; setting the prior probability distribution provides an initial estimate for the system state; the Bayesian fusion algorithm updates and optimizes the estimate of the system state by combining multi-source observation data, improving the prediction accuracy. Finally, the posterior probability distribution output by the information fusion unit 2 provides reliable data support for subsequent data analysis and decision-making.
[0145] The big data intelligent analysis and processing system based on information fusion further includes a data analysis unit 3. The data analysis unit 3 analyzes the data through the random forest algorithm based on the data fused by the information fusion unit 2;
[0146] In this example, the specific steps involved in predicting and analyzing the circuit performance through the random forest model and optimizing the random forest model by introducing the minimum impurity reduction variable are as follows:
[0147] Specifically, Random Forest is an ensemble learning method mainly used for classification and regression tasks. It improves the accuracy and robustness of the model by constructing multiple decision trees and summarizing their prediction results. The main feature of the random forest algorithm is to reduce the overfitting problem of a single decision tree by introducing randomness and diversity, and to improve the generalization ability of the model by integrating multiple trees.
[0148] S6.1. Obtain the fused data from the information fusion unit 2 to form a data set;
[0149] S6.2. Divide the data set into a training set and a validation set;
[0150] S6.3. Construct a random forest model through the training set;
[0151] In this example, the specific steps of constructing a random forest model through the training set are as follows:
[0152] S6.31. Randomly select n samples from the training set as the new training set;
[0153] Specifically, the implementation process of S6.31 is as follows:
[0154]
[0155] S6.32. Initialize the root node and use the new training set as the input data of the root node;
[0156] Specifically, the implementation process of S6.32 is as follows:
[0157] def initialize_root_node(X,y):
[0158] root_node = {'X':X,'y':y}
[0159] return root_node;
[0160] S6.33. Select the optimal feature and splitting point through the Gini impurity criterion and split the node into two child nodes;
[0161] Specifically, Gini Impurity is a metric used to measure the purity of a dataset and is commonly used for feature selection and node splitting in decision tree algorithms. It reflects the probability that two randomly selected samples from the dataset have different class labels. The lower the value of Gini Impurity, the higher the purity of the dataset, meaning that most samples in the dataset belong to the same class; conversely, the higher the Gini Impurity, the greater the degree of mixing in the dataset and the more evenly distributed the class distribution. When constructing a decision tree, the Gini impurity criterion is used to select the optimal feature and splitting point to maximize the purity of the child nodes, thereby improving the classification performance of the model.
[0162] In this example, the specific process of selecting the optimal feature and splitting point through the Gini impurity criterion and splitting the node into two child nodes is as follows:
[0163] For a node N, its Gini impurity G(N) is defined as:
[0164]
[0165] where G(N) represents the Gini impurity of the current node N; K represents the number of classes; p k represents the proportion of samples in node N that belong to class k; k represents the index variable;
[0166] For each feature j and each candidate splitting point s:
[0167] Calculate the Gini impurity of the left child node:
[0168]
[0169] where G(N L ) represents the Gini impurity of the left child node; |N Lk | represents the number of samples belonging to class k in the left child node N L ; |N L | represents the number of samples in the left child node N L .
[0170] Calculate the Gini impurity of the right child node:
[0171]
[0172] where G(N R ) represents the Gini impurity of the right child node; |N Rk | represents the number of samples belonging to class k in the right child node N R ; |N R | represents the number of samples in the left child node N R .
[0173] Calculate the weighted Gini impurity after splitting:
[0174]
[0175] where G split (N) represents the weighted Gini impurity after splitting; |N| represents the total number of samples in node N;
[0176] Select the feature j * with the minimum weighted Gini impurity and the split point s * as the optimal feature and split point:
[0177]
[0178] According to the optimal feature j * and the split point s * , split the current node Z into two child nodes N L and N R ; j represents the feature index; s represents the split threshold.
[0179] Specifically, the process of selecting the optimal feature and splitting point through the Gini impurity criterion and splitting the node into two child nodes ensures that the model can efficiently extract key features from multi-source data and construct a decision tree with a reasonable structure and strong generalization ability. This process not only improves the accuracy of the model but also enhances the robustness and stability of the system. Especially when dealing with complex, high-dimensional, and noisy large datasets, it can effectively reduce overfitting and improve the prediction performance of the model. By optimizing the splitting conditions, the system can effectively fuse information between different data sources, capture potential patterns in the data, and thus provide reliable support for subsequent intelligent analysis and decision-making.
[0180] S6.34. Repeat the operations in S6.32 - S6.33 to select the optimal feature and splitting point and split the node for each child node until the stopping condition is met.
[0181] Specifically, the implementation process of S6.34 is as follows:
[0182]
[0183]
[0184] S6.35. Repeat the above steps T times to construct T decision trees and form a random forest. Specifically, the implementation process of S6.35 is as follows:
[0185]
[0186]
[0187] By randomly sampling from the training set to increase the diversity and robustness of the model, and constructing decision trees by recursively selecting the optimal feature and splitting point to capture complex patterns in the data. Finally, by constructing multiple decision trees and integrating their prediction results, the random forest model can not only reduce the risk of overfitting but also improve the generalization ability and prediction accuracy of the model, thus providing reliable support for the intelligent analysis and real-time processing of the system.
[0188] S6.4. Use the random forest model with the variable of introducing the minimum impurity reduction to predict and analyze the circuit performance.
[0189] Specifically, the variable of the minimum impurity reduction splits the node by selecting the feature that can minimize the node impurity to the greatest extent. Specifically, it measures the reduction in the overall impurity of the child nodes relative to the impurity of the parent node after splitting on a given feature. By selecting the feature that causes the most reduction in impurity as the splitting basis, the model can construct a purer and better-classifying decision tree, thereby improving the prediction accuracy and generalization ability.
[0190] In this example, the process of using the random forest model after introducing the minimum impurity reduction variable to predict and analyze the circuit performance is as follows:
[0191] For a new sample x, let each tree in the random forest output a predicted class;
[0192] Determine the final predicted class through majority voting:
[0193]
[0194] Among them, represents the finally predicted class; T represents the total number of decision trees in the random forest; I represents the indicator function; y t represents the predicted class of the t-th tree for the sample; c represents one of the classes; represents selecting the class c that maximizes the above ratio as the final predicted class;
[0195] For a new sample x, let each tree in the random forest output a predicted value;
[0196] Determine the predicted value through the average value:
[0197]
[0198] Among them, represents the predicted value; y t ' represents the predicted value of the t-th tree for the sample;
[0199] Considering the accuracy of the predicted value, introduce the minimum impurity reduction variable to optimize the prediction process of the regression task of the random forest model;
[0200] When constructing each tree t, the splitting node satisfies:
[0201] G(N)-G split (N)≥min impurity decrease ;
[0202] Among them, min impurity decrease represents the set minimum impurity reduction variable;
[0203] Then determine the final predicted class through majority voting after introducing the minimum impurity reduction variable:
[0204]
[0205] When constructing each tree t, the splitting node satisfies:
[0206] MSE(N)-MSE split (N)≥min impuritydecrease ;
[0207] Among them, MSE(N) represents the mean squared error of the current node N; MSE split (N) represents the weighted mean squared error after splitting;
[0208] Then, the final prediction value is determined by the average value after introducing the minimum impurity reduction variable:
[0209]
[0210] Among them, represents the final prediction value after introducing the minimum impurity reduction variable.
[0211] Specifically, using the random forest model after introducing the minimum impurity reduction variable, the specific steps for predicting and analyzing the circuit performance play a key role, ensuring the accuracy and reliability of the system. These steps make each tree in the random forest predict new samples, and determine the final prediction result through majority voting (classification task) or average value (regression task), thus making full use of the integrated advantages of multiple trees and improving the generalization ability and prediction accuracy of the model. Especially after introducing the weighted average algorithm, by calculating the mean squared error of each tree on the validation set and assigning corresponding weights, the prediction process of the regression task is further optimized. This method not only considers the prediction performance of each tree, but also ensures the rationality of the weighted average by normalizing the weights, thus significantly improving the accuracy and stability of the prediction value and providing stronger support for the intelligent analysis and real-time processing of the system.
[0212] By restricting unnecessary splits, the model can more accurately capture the true patterns in the data while maintaining a simple structure, reducing the influence of noise and redundant features, thereby improving the prediction accuracy. The introduction of the minimum impurity reduction variable makes the model more stable. Especially when dealing with multi-source data, it can effectively handle the differences and noises between different data sources, avoid overfitting, and ensure the reliable performance of the model in complex environments. By reducing unnecessary splits, the training and prediction speeds of the model are improved, reducing the consumption of computing resources, and it is especially suitable for large-scale data sets and real-time processing scenarios. In the information fusion process, the optimized random forest model can better integrate information from different sensors or data sources, provide more reliable prediction results, and support more informed decision-making.
[0213] S6.5. Evaluate the random forest model using the validation set through cross-validation.
[0214] Specifically, the implementation process is as follows:
[0215]
[0216]
[0217]
[0218] Through cross-validation, the system can evaluate the performance of the model on multiple different data subsets, thus providing a more comprehensive understanding of how the model performs under different data distributions. This not only helps to identify potential overfitting or underfitting issues in the model, but also provides more reliable performance metrics, offering a scientific basis for model selection and tuning. Ultimately, this step improves the prediction accuracy and robustness of the system, ensuring its effectiveness in dealing with complex and changing data environments in practical applications.
[0219] The big data intelligent analysis and processing system based on information fusion further includes a real-time processing unit 4. The real-time processing unit 4 conducts evaluation and early warning based on the analysis results of the data analysis unit 3.
[0220] In this example, when the real-time processing unit 4 conducts evaluation and early warning based on the analysis results of the data analysis unit 3, the specific steps for evaluation and early warning are as follows:
[0221] S10.1: Receive the predictive analysis results from the data analysis unit 3;
[0222] S10.2: Calculate the key performance indicators based on the analysis results;
[0223] Specifically, the process of calculating the key performance indicators based on the analysis results is as follows:
[0224] Select indicators: Determine the key performance indicators (KPIs) to be calculated according to the characteristics of the system and the application scenario. Common KPIs include power consumption, temperature, current, voltage, signal quality, probability of failure, etc.
[0225] Indicator calculation: Based on the received analysis results, use predefined formulas or algorithms to calculate each KPI; for example:
[0226] Power consumption = voltage × current
[0227] Temperature change rate = (current temperature - previous temperature) / time interval
[0228] Probability of failure = probability value output by the classification model
[0229] Multi-source data fusion: If there are multiple sensors or data sources in the system, data from different sources can be fused to obtain more accurate performance indicators.
[0230] S10.3: Compare the calculated performance indicators with the preset thresholds;
[0231] Specifically, the process of comparing the calculated performance metrics with the preset thresholds is as follows:
[0232] Threshold setting: Set a reasonable threshold range for each KPI. The thresholds can be based on historical data, empirical rules, expert knowledge, or adjusted dynamically. For example:
[0233] Normal operating temperature range: 20°C - 60°C
[0234] Maximum allowable power consumption: 50W
[0235] Lower limit of signal quality: 80%
[0236] Threshold comparison: Compare the calculated performance metrics with the preset thresholds item by item to determine whether they exceed the normal range. For example:
[0237] If the current temperature > 60°C, trigger a high - temperature warning.
[0238] If the power consumption > 50W, trigger a warning for excessive power consumption.
[0239] If the signal quality < 80%, trigger a warning for signal quality problems.
[0240] S10.4. Assign a risk score to the system based on the results of the status assessment;
[0241] Specifically, the process of assigning a risk score to the system based on the results of the status assessment is as follows:
[0242] Risk scoring model: Use a predefined risk scoring model to comprehensively consider multiple KPIs and their weights to assign a risk score to the system. The risk score can be a numerical value representing the current risk level of the system, or a multi - dimensional score covering different risk factors (such as security, reliability, efficiency, etc.).
[0243] Risk score calculation: Calculate the risk score according to the results of the threshold comparison. For example:
[0244] If a certain KPI exceeds the threshold, increase the corresponding risk score value.
[0245] If multiple KPIs exceed the threshold simultaneously, further increase the risk score.
[0246] Methods such as weighted summation, logistic regression, decision tree, etc. can be used to calculate the risk score.
[0247] Uncertainty analysis: Considering the uncertainty and noise in the data, use probability models (such as Bayesian networks, Monte Carlo simulations) to evaluate the risk probabilities in different scenarios and further optimize the risk score.
[0248] S10.5. Determine whether an early warning needs to be generated based on the results of the risk assessment.
[0249] Specifically, the process of determining whether an early warning needs to be generated based on the results of the risk assessment is as follows:
[0250] Trigger condition judgment: Determine whether an early warning needs to be generated based on the risk score. Common trigger conditions include:
[0251] The risk score reaches or exceeds a certain critical value.
[0252] A certain performance indicator exceeds the preset threshold.
[0253] Trend analysis shows a potential deterioration trend.
[0254] Early warning level classification: Classify the early warning into different levels (such as low, medium, high) according to the severity of the risk. Different levels of early warnings can trigger different response measures.
[0255] Early warning content generation: Generate specific early warning information, including:
[0256] The type of early warning (such as performance anomaly, fault warning, safety warning, etc.).
[0257] Specific performance indicators and their current values.
[0258] Analysis of possible causes.
[0259] Suggested countermeasures.
[0260] Notification and response: Send the early warning information to relevant personnel or systems through various channels. Common notification methods include:
[0261] Email or text message: Sent to maintenance personnel or management personnel.
[0262] Alarm system: Trigger the audible and visual alarm device to remind on-site staff.
[0263] Visual interface: Display the early warning information on the monitoring system or dashboard for convenient real-time viewing. Automated response: For certain emergency situations, the system can automatically execute predefined operations, such as shutting down equipment, switching to a backup system, adjusting operating parameters, etc., to prevent further damage or accidents.
[0264] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.
Claims
1. A big data intelligent analysis and processing system based on information fusion, characterized in that, including a data acquisition unit (1), which acquires circuit performance data of different data sources through a multi-source data acquisition module; an information fusion unit (2), which fuses the circuit performance data of different data sources based on the circuit performance data of different data sources acquired by the data acquisition unit (1) through a Bayesian fusion algorithm optimized by introducing an electromagnetic interference variable; a data analysis unit (3), which predicts and analyzes the circuit performance through a random forest model based on the data fused by the information fusion unit (2), and introduces a minimum impurity reduction variable to optimize the random forest model for improving the accuracy of the prediction value; a real-time processing unit (4), which evaluates and gives an early warning according to the analysis result of the data analysis unit (3).
2. The big data intelligent analysis and processing system based on information fusion according to claim 1, wherein: The data acquisition unit (1) includes a multi-source data acquisition module, and the multi-source data acquisition module includes a database module and a sensor module; wherein, the database module is used for storing the collected data; the sensor module is used for acquiring circuit performance data.
3. The big data intelligent analysis and processing system based on information fusion according to claim 1, characterized in that: The specific steps involved in the information fusion unit (2) fusing the circuit performance data of different data sources through a Bayesian fusion algorithm optimized by introducing an electromagnetic interference variable are as follows; S3.1, preprocess the data by removing noise, filling missing values, and processing outliers, and convert data in different formats into a unified format; S3.2, set the prior probability distribution P(x) of the system state; S3.3, update the posterior probability distribution of the system state using the Bayesian fusion algorithm.
4. The big data intelligent analysis and processing system based on information fusion according to claim 3, characterized in that: In the S3.3, the specific steps of updating the posterior probability distribution of the system state using the Bayesian fusion algorithm are: S3.31, obtain the measurement value of the sensor from the sensor module; S3.32, calculate the likelihood function of the sensor through a Gaussian distribution; S3.33, calculate the joint likelihood function through a Gaussian distribution; S3.34, update the posterior probability distribution of the system state.
5. The big data intelligent analysis and processing system based on information fusion according to claim 4, characterized in that: In the S3.34, the specific process of updating the posterior probability distribution of the system state is: obtain the joint likelihood function from S3.33: Among them, P(z|x) represents the joint likelihood function; P(z i |x) represents the likelihood function of the i-th sensor; x represents the state vector; z represents the data; z i represents the data of the i-th sensor; n represents the number of z; i represents the index variable; obtain the updated posterior probability distribution of the system state: wherein, P(x|z) represents the posterior probability; P(z) represents the marginal probability of the data z; considering the accuracy of updating the posterior probability distribution of the system state, introduce an electromagnetic interference variable to optimize the Bayesian fusion algorithm: Among them, represents the i-th sensor data after introducing the electromagnetic interference variable; ∈ i represents the electromagnetic interference variable of the i-th sensor; ∈ represents the inductance interference variable; P(z|x,∈) represents the joint likelihood function after introducing the electromagnetic interference variable; P(x|z)' represents the posterior probability after introducing the electromagnetic interference variable.
6. The big data intelligent analysis and processing system based on information fusion according to claim 1, characterized in that: The specific steps involved in the data analysis unit (3) predicting and analyzing the circuit performance through a random forest model and introducing a minimum impurity reduction variable to optimize the random forest model are as follows: S6.1, obtain the fused data from the information fusion unit (2) to form a data set; S6.2, divide the data set into a training set and a validation set; S6.3, construct a random forest model through the training set; S6.4, use the random forest model after introducing the minimum impurity reduction variable to predict and analyze the circuit performance; S6.5, use the validation set to evaluate the random forest model through cross-validation.
7. The big data intelligent analysis and processing system based on information fusion according to claim 6, characterized in that: In the S6.3, the specific steps of constructing a random forest model through the training set are: S6.
31. Randomly select n samples from the training set as the new training set; S6.
32. Initialize the root node and use the new training set as the input data of the root node; S6.
33. Select the optimal feature and splitting point according to the Gini impurity criterion and split the node into two child nodes; S6.
34. Repeat the operations of S6.32 - S6.33, i.e., select the optimal feature and splitting point and split the node, for each child node until the stopping condition is met; S6.
35. Repeat the above steps T times to construct T decision trees to form a random forest.
8. The big data intelligent analysis and processing system based on information fusion according to claim 7, characterized in that: In S6.33, the specific process of selecting the optimal feature and splitting point according to the Gini impurity criterion and splitting the node into two child nodes is as follows: For a node N, its Gini impurity G(N) is defined as: Among them, G(N) represents the Gini impurity of the current node N; K represents the number of categories; p k represents the proportion of samples belonging to category k in node N; k represents the index variable; For each feature j and each candidate splitting point s: Calculate the Gini impurity of the left child node: Among them, G(N L ) represents the Gini impurity of the left child node; |N Lk | represents the number of samples belonging to class k in the left child node N L ; |N L | represents the number of samples of the left child node N L ; Calculate the Gini impurity of the right child node: Among them, G(N R ) represents the Gini impurity of the right child node; |N Rk | represents the number of samples belonging to class k in the right child node N R ; |N R | represents the number of samples in the left child node N R ; Calculate the weighted Gini impurity after splitting: Among them, G split (N) represents the weighted Gini impurity after splitting; |N| represents the total number of samples in node N; Select the feature j with the minimum weighted Gini impurity * and the splitting point s * , as the optimal feature and splitting point: According to the optimal feature j * and the splitting point s * , split the current node N into two child nodes N L and N R ; j represents the feature index; s represents the splitting threshold.
9. The big data intelligent analysis and processing system based on information fusion according to claim 6, wherein: In S6.4, the process of using the random forest model with the minimum impurity reduction variable introduced to predict and analyze the circuit performance is as follows: For a new sample x, let each tree in the random forest output a predicted class; Determine the final predicted class through majority voting: Among them, represents the finally predicted class; T represents the total number of decision trees in the random forest; I represents the indicator function; y t represents the predicted class of the t-th tree for the sample; c represents one of the classes; means to select the class c that maximizes the above ratio as the finally predicted class; For a new sample x, let each tree in the random forest output a predicted value; Determine the predicted value through the average value: Among them, represents the predicted value; y t ' represents the predicted value of the t-th tree for the sample; Considering the accuracy of the predicted value, introduce the minimum impurity reduction variable to optimize the prediction process of the regression task of the random forest model; When in the construction process of each tree t, the splitting node satisfies: G(N)-G split (N)≥min impurity decrease ; where min impurity decrease represents the set minimum impurity reduction variable; Then determine the final predicted class through majority voting with the minimum impurity reduction variable introduced: When in the construction process of each tree t, the splitting node satisfies: MSE(N)-MSE split (N)≥min impurity decrease ; Among them, MSE(N) represents the mean square error of the current node N; MSE split (N) represents the weighted mean square error after splitting; Then determine the final predicted value through the average value with the minimum impurity reduction variable introduced: Among them, represents the final predicted value after introducing the minimum impurity reduction variable.
10. The big data intelligent analysis and processing system based on information fusion according to claim 1, wherein: The real-time processing unit (4) conducts evaluation and early warning according to the analysis result of the data analysis unit (3). The specific steps for evaluation and early warning are as follows: S10.
1. Receive the prediction and analysis result from the data analysis unit (3); S10.
2. Calculate the key performance indicators according to the analysis result; S10.
3. Compare the calculated performance indicators with the preset thresholds; S10.
4. Assign a risk score to the system based on the result of the status evaluation; S10.
5. Judge whether it is necessary to generate an early warning according to the result of the risk assessment.
Citation Information
Patent Citations
An electric energy meter failure rate assessment method and device
CN109767061A
Aero-engine on-wing reliability evaluation method based on monitoring information fusion
CN116776454A
Comprehensive electronic system electromagnetic compatibility analysis method based on knowledge graph
CN117150038A
Transmission monitoring system for relay protection overhaul test of intelligent substation
CN118171195A
Main transformer fault prediction method based on multivariable data fusion
CN118779831A