Big data intelligent analysis processing system based on information fusion

By combining multi-source data acquisition, a Bayesian fusion algorithm for electromagnetic interference variable optimization, and a random forest model, the problems of inconsistent data formats and privacy security were solved, enabling efficient data fusion and real-time analysis, and improving the system's accuracy and responsiveness.

CN120408505BActive Publication Date: 2025-12-26GUANGZHOU JIAOJIREN TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510481332.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-12-26
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

In existing big data intelligent analysis and processing systems based on information fusion, problems such as inconsistent formats, missing data, errors, or inconsistencies frequently occur due to data originating from different systems and platforms, affecting the accuracy and reliability of analysis results. Privacy and security issues cannot be ignored, especially when processing sensitive personal information. How to maximize the value of data utilization while ensuring data security and personal privacy is a key challenge. The complexity and variability of technology pose challenges to the system, with severe data silos and obstacles to data sharing and integration, limiting the comprehensive utilization of data and the full realization of system functions.

Method used

The data acquisition unit collects data through a multi-source data acquisition module. The information fusion unit uses a Bayesian fusion algorithm with electromagnetic interference variable optimization to fuse circuit performance data from different data sources. The data analysis unit performs predictive analysis using a random forest model and introduces minimum impurity to reduce variable optimization. The real-time processing unit evaluates and issues warnings based on the analysis results.

Benefits of technology

It improves the comprehensiveness and diversity of data, reduces data noise and errors through Bayesian fusion algorithm, and the optimized random forest algorithm can extract deep-level patterns and rules. The system can respond and make decisions in real time, improving response speed and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408505B_ABST
    Figure CN120408505B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of big data, in particular to a big data intelligent analysis processing system based on information fusion. The system comprises the following steps: a data acquisition unit collects circuit performance data of different data sources through a multi-source data acquisition module; an information fusion unit fuses the circuit performance data of different data sources by introducing a Bayesian fusion algorithm optimized by an electromagnetic interference variable based on the circuit performance data of different data sources collected by the data acquisition unit; a data analysis unit predicts and analyzes the circuit performance through a random forest model based on the fused data of the information fusion unit, and introduces a minimum impurity reduction variable to optimize the random forest model, so that the accuracy of the predicted value is improved; and a real-time processing unit evaluates and gives an early warning according to the analysis result of the data analysis unit. The application fuses data from different data sources through the Bayesian fusion algorithm, and the consistency and correction of the data are realized, so that data noise and errors are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of big data, in particular to a big data intelligent analysis processing system based on information fusion. BACKGROUND

[0002] The big data intelligent analysis processing system based on information fusion is a highly integrated multidisciplinary technical solution, aiming to extract valuable information and knowledge from massive and diversified data sources. The system first collects data from different channels through advanced data acquisition technology and Internet of Things devices, including structured data (such as database records), semi-structured data (such as XML files), and unstructured data (such as text, pictures, and videos). Then, using efficient data management and storage technologies such as distributed file systems and NoSQL databases, it ensures safe storage and fast access to data. On this basis, information fusion technology is used to clean, integrate, and standardize multi-source heterogeneous data, eliminating redundancy and contradictions and improving data quality. Subsequently, combined with machine learning and deep learning algorithms, the processed data is analyzed in depth to mine potential patterns and trends, enabling accurate prediction and decision support. At the same time, the system also incorporates advanced functions such as natural language processing and image recognition to enhance its ability to understand complex data. To ensure the security and privacy of the entire process, the system uses strict encryption technology and access control strategies. Finally, with the help of advanced data visualization tools, the analysis results are presented to the end user in an intuitive and easy-to-understand form, helping them make scientific and reasonable judgments quickly. Overall, this system not only greatly improves the efficiency and accuracy of data processing, but also provides strong support for various industries and promotes the development of an intelligent society.

[0003] In the existing big data intelligent analysis processing system based on information fusion, since the data comes from different systems and platforms, there are frequent problems such as non-uniform format, data missing, errors, or inconsistencies, which will directly affect the accuracy and reliability of the analysis results if not properly handled. Secondly, privacy and security issues cannot be ignored, especially when dealing with personal sensitive information, how to maximize the value of data while ensuring data security and personal privacy has become a difficult problem to be solved. The complexity and variability of technology also pose challenges to the system, as the types and scale of data continue to expand, the requirements for data processing and analysis technology are also increasing, and the system needs to be continuously upgraded and optimized to maintain competitiveness. The phenomenon of data silos between different data sources is still serious, and data sharing and integration face many obstacles, which not only limit the comprehensive use of data, but also hinder the full play of system functions. Therefore, the big data intelligent analysis processing system based on information fusion is designed. SUMMARY

[0004] The present application aims to provide a big data intelligent analysis processing system based on information fusion to solve the problems in the prior art, such as the non-uniform format, data loss, errors or inconsistencies of data from different systems and platforms, which directly affect the accuracy and reliability of the analysis results. In addition, privacy and security issues cannot be ignored, especially when dealing with personal sensitive information. How to maximize the value of data while ensuring data security and personal privacy has become a difficult problem to be solved. The complexity and variability of technology also bring challenges to the system. With the continuous expansion of data types and scales, the requirements for data processing and analysis technology are also increasing. The system needs to be upgraded and optimized to maintain competitiveness. The data island phenomenon between different data sources is still serious, and data sharing and integration face many obstacles, which not only limit the comprehensive use of data, but also hinder the full play of system functions.

[0005] To achieve the above-mentioned purpose, the present application aims to provide a big data intelligent analysis processing system based on information fusion, comprising a data acquisition unit, wherein the data acquisition unit acquires circuit performance data of different data sources through a multi-source data acquisition module;

[0006] an information fusion unit, wherein the information fusion unit fuses the circuit performance data of different data sources based on the circuit performance data of different data sources acquired by the data acquisition unit through a Bayesian fusion algorithm optimized by introducing an electromagnetic interference variable;

[0007] a data analysis unit, wherein the data analysis unit performs predictive analysis on the circuit performance based on the fused data of the information fusion unit through a random forest model, and optimizes the random forest model by introducing a minimum impurity reduction variable to improve the accuracy of the predicted value;

[0008] a real-time processing unit, wherein the real-time processing unit evaluates and warns according to the analysis result of the data analysis unit.

[0009] As a further improvement of the present technical solution, the data acquisition unit comprises a multi-source data acquisition module, and the multi-source data acquisition module comprises a database module and a sensor module.

[0010] The database module is used for storing the collected data.

[0011] The sensor module is used for acquiring circuit performance data.

[0012] As a further improvement of the technical solution, the information fusion unit fuses the circuit performance data of different data sources based on the circuit performance data collected by the data acquisition unit, and the specific steps of fusing the circuit performance data of different data sources through the Bayesian fusion algorithm optimized by introducing electromagnetic interference variables are:

[0013] S3.1, data preprocessing is performed on the data by removing noise, missing value filling, and abnormal value processing, and different formats of data are converted into a unified format;

[0014] S3.2, set the prior probability distribution P(x) of the system state;

[0015] S3.3, update the posterior probability distribution of the system state using the Bayesian fusion algorithm.

[0016] As a further improvement of the technical solution, in S3.3, the specific steps of updating the posterior probability distribution of the system state using the Bayesian fusion algorithm are:

[0017] S3.31, obtain the measurement value of the sensor from the sensor module;

[0018] S3.32, calculate the likelihood function of the sensor through Gaussian distribution;

[0019] S3.33, calculate the joint likelihood function through Gaussian distribution;

[0020] S3.34, update the posterior probability distribution of the system state.

[0021] As a further improvement of the technical solution, in S3.34, the specific process of updating the posterior probability distribution of the system state is:

[0022] The joint likelihood function is obtained from S3.33:

[0023]

[0024] Where P(z|x) represents the joint likelihood function; P(z i |x) represents the likelihood function of the i-th sensor; x represents the state vector; z represents the data; z i represents the data of the i-th sensor; n represents the number of z; i represents the index variable;

[0025] The posterior probability distribution of the system state is obtained:

[0026]

[0027] Where P(x|z) represents the posterior probability; P(z) represents the marginal probability of data z.

[0028] Considering the accuracy of updating the posterior probability distribution of the system state, an electromagnetic interference variable is introduced to optimize the Bayesian fusion algorithm:

[0029]

[0030]

[0031]

[0032] wherein, represents the i-th sensor data after introducing the electromagnetic interference variable; ∈ i represents the electromagnetic interference variable of the i-th sensor; ∈ represents the electromagnetic interference variable; P(z|x, ∈) represents the joint likelihood function after introducing the electromagnetic interference variable; P(x|z)' represents the posterior probability after introducing the electromagnetic interference variable.

[0033] As a further improvement of the technical solution, the specific steps involved in predicting and analyzing the circuit performance by the random forest model and introducing the minimum impurity reduction variable to optimize the random forest model are:

[0034] S6.1, obtaining the fused data from the information fusion unit to form a data set;

[0035] S6.2, dividing the data set into a training set and a validation set;

[0036] S6.3, constructing a random forest model through the training set;

[0037] S6.4, using the random forest model after introducing the minimum impurity reduction variable to predict and analyze the circuit performance;

[0038] S6.5, using the validation set to evaluate the random forest model through cross-validation.

[0039] As a further improvement of the technical solution, in S6.3, the specific steps for constructing a random forest model through the training set are:

[0040] S6.31, randomly selecting n samples from the training set as a new training set;

[0041] S6.32, initializing the root node and taking the new training set as the input data of the root node;

[0042] S6.33, selecting the optimal feature and split point through the Gini impurity criterion, and splitting the node into two child nodes;

[0043] S6.34, repeat the steps of selecting optimal feature and split point, splitting node for each child node until the stopping condition is reached;

[0044] S6.35, repeat the above steps T times to build T decision trees to form a random forest.

[0045] As a further improvement of the technical solution, in S6.33, the specific process of selecting the optimal feature and split point by the Gini impurity criterion and splitting the node into two child nodes is as follows:

[0046] For a node N, its Gini impurity G(N) is defined as:

[0047]

[0048] Where G(N) represents the Gini impurity of the current node N; K represents the number of categories; p k represents the proportion of samples belonging to category k in node N; k represents the index variable;

[0049] For each feature j and each candidate split point s:

[0050] Calculate the Gini impurity of the left child node:

[0051]

[0052] Where G(N L ) represents the Gini impurity of the left child node; |N Lk | represents the number of samples belonging to category k in the left child node N L ; |N L | represents the number of samples in the left child node N L ;

[0053] Calculate the Gini impurity of the right child node:

[0054]

[0055] Where G(N R ) represents the Gini impurity of the right child node; |N Rk | represents the number of samples belonging to category k in the right child node N R ; |N R | represents the number of samples in the left child node N R ;

[0056] Calculate the weighted Gini impurity after splitting:

[0057]

[0058] Where G split(N) represents the weighted Gini impurity after splitting; |N| represents the total number of samples in node N;

[0059] selecting the feature j with the minimum weighted Gini impurity * and the split point s * , as the optimal feature and split point:

[0060]

[0061] According to the optimal feature j * and the split point s * , the current node N is split into two child nodes N L and N R ; j represents the feature index; s represents the split threshold.

[0062] As a further improvement of the technical solution, in S6.4, the process of predicting and analyzing the circuit performance using the random forest model after introducing the minimum impurity reduction variable is:

[0063] For a new sample x, let each tree in the random forest output a predicted class;

[0064] Determine the final predicted class by majority vote:

[0065]

[0066] wherein, represents the final predicted class; T represents the total number of decision trees in the random forest; I represents the indicator function; y t represents the predicted class of the sample by the tth tree; c represents one of the classes; represents selecting the class c with the maximum ratio as the final predicted class;

[0067] For a new sample x, let each tree in the random forest output a predicted value;

[0068] Determine the predicted value by averaging:

[0069]

[0070] wherein, represents the predicted value; y t ' represents the predicted value of the sample by the tth tree;

[0071] Considering the accuracy of the predicted value, introduce the minimum impurity reduction variable to optimize the regression task prediction process of the random forest model;

[0072] When the splitting node satisfies:

[0073] G(N)-Gsplit (N)≥min impurity decrease ;

[0074] Among them, min impurity decrease This indicates the reduction of the variable to represent the minimum impurity setting;

[0075] The final predicted category is determined by majority voting after reducing the number of variables by introducing minimum impurity.

[0076]

[0077] During the construction of each tree t, the splitting nodes satisfy:

[0078] MSE(N)-MSE split (N)≥min impurity decrease ;

[0079] Where MSE(N) represents the mean square error of the current node N; MSE split (N) represents the weighted mean square error after splitting;

[0080] The final predicted value is determined by averaging the values ​​after reducing the variables by introducing minimum impurity:

[0081]

[0082] in, This represents the final predicted value after introducing minimum impurity to reduce the number of variables.

[0083] As a further improvement to this technical solution, the real-time processing unit evaluates and issues warnings based on the analysis results of the data analysis unit. The specific steps for evaluation and warning are as follows:

[0084] S10.1 Receive predictive analysis results from the data analysis unit;

[0085] S10.2. Based on the analysis results, calculate the key performance indicators;

[0086] S10.3 Compare the calculated performance indicators with the preset thresholds;

[0087] S10.4. Based on the results of the status assessment, assign a risk score to the system;

[0088] S10.5. Based on the results of the risk assessment, determine whether an early warning needs to be generated.

[0089] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0090] 1、The big data intelligent analysis processing system based on information fusion, through collecting data from multiple data sources, improves the comprehensiveness and diversity of data; through the Bayesian fusion algorithm, the data from different data sources are consistent and corrected, reducing data noise and error. The optimized Bayesian fusion algorithm can handle the uncertainty and fuzziness of data, and improve the credibility of data.

[0091] 2、The big data intelligent analysis processing system based on information fusion, through the use of advanced machine learning methods such as optimized random forest algorithm, can extract deep patterns and rules from the fused data. According to the analysis result of the data analysis unit, it can make response and decision in real time, improve the response speed and efficiency of the system. BRIEF DESCRIPTION OF DRAWINGS

[0092] Figure 1 The overall flowchart of the present application is shown in the figure;

[0093] The meanings of various labels in the figure are as follows:

[0094] 1, data acquisition unit; 2, information fusion unit; 3, data analysis unit; 4, real-time processing unit. DETAILED DESCRIPTION

[0095] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0096] Please refer to Figure 1 As shown in the figure, the big data intelligent analysis processing system based on information fusion is provided, which includes a data acquisition unit 1, the data acquisition unit 1 collects circuit performance data of different data sources through a multi-source data acquisition module;

[0097] In this example, the data acquisition unit 1 includes a multi-source data acquisition module, and the multi-source data acquisition module includes a database module and a sensor module;

[0098] The database module is used to store the collected data.

[0099] Specifically, the database module is one of the core components in the multi-source data acquisition system, responsible for storing and managing data collected from various data sources, including sensors, external systems, user inputs, etc. It is usually composed of one or more database management systems (DBMS), supporting structured query language (SQL) or other query interfaces, to efficiently perform data insertion, query, update, and deletion operations. The database module not only can store a large amount of historical data, but also can optimize data access speed through indexing, partitioning, and other techniques to ensure that the real-time processing unit and other analysis modules can quickly obtain the required information. In addition, the database module also has functions such as data backup, recovery, security, and permission control to ensure the integrity and security of data.

[0100] The sensor module is used to collect circuit performance data.

[0101] Specifically, the sensor module is a component in the multi-source data acquisition system that directly interacts with the physical world, responsible for real-time collection of circuit performance data. It is usually composed of a series of high-precision sensors that can measure various parameters in the circuit, such as voltage, current, temperature, humidity, power, frequency, etc. The sensor module connects with the circuit through analog or digital interfaces, and can capture the transient behavior and long-term trends of the circuit at high frequency and high resolution. In order to ensure the accuracy and reliability of the data, the sensor module usually includes signal conditioning circuits (such as filters, amplifiers), data acquisition cards (DAQ), and timestamp functions to pre-process and synchronize the collected data. In addition, the sensor module may also have a self-calibration function to automatically adjust the measurement accuracy of the sensor, reducing drift and error.

[0102] The big data intelligent analysis and processing system based on information fusion further comprises an information fusion unit 2, which fuses the circuit performance data of different data sources collected by the data acquisition unit 1 through a Bayesian fusion algorithm;

[0103] In this example, the specific steps involved in the Bayesian fusion algorithm optimized by introducing electromagnetic interference variables to fuse the circuit performance data of different data sources are as follows:

[0104] Specifically, the Bayesian fusion algorithm is a method based on Bayes' theorem, used to fuse information from multiple data sources. This algorithm has a wide range of applications in multi-sensor data fusion, multi-modal data processing, multi-source information integration, etc. The core idea of the Bayesian fusion algorithm is to integrate data from different sources through a probability model, thereby improving the accuracy and reliability of the data.

[0105] S3.1, data preprocessing is performed on the data by removing noise, missing value filling, and outlier processing, and data in different formats are converted into a unified format;

[0106] Specifically, the specific steps of data preprocessing on the data are:

[0107] S3.11, using a moving average algorithm to remove noise;

[0108]

[0109] where y t represents the smoothed value at time t; w t-i represents the original data at time t-l; n represents the window size; and l represents the index variable;

[0110] S3.12, using linear interpolation to fill in missing values in the data;

[0111]

[0112] where x t represents the interpolation result at time t; x t+1 and x t-1 represent the observation values of adjacent time points; t t+1 and t t-1 represent the corresponding time stamps;

[0113] S3.13, using Z-score method to identify and process outliers in the data;

[0114]

[0115] where Z represents the Z-score; w i represents the data point; μ represents the mean of the data; and σ represents the standard deviation of the data;

[0116] S3.2, setting the prior probability distribution P(x) of the system state;

[0117] Specifically, the prior probability distribution P(x) is an initial estimate of the system state, which can be a distribution based on historical data or other prior knowledge.

[0118] S3.3, using Bayesian fusion algorithm to update the posterior probability distribution of the system state.

[0119] In this example, the specific steps of using Bayesian fusion algorithm to update the posterior probability distribution of the system state are:

[0120] S3.31, obtaining the measurement value of the sensor from the sensor module;

[0121] S3.32, calculating the likelihood function of the sensor by Gaussian distribution;

[0122] Specifically, the process of calculating the likelihood function of the sensor through Gaussian distribution is as follows:

[0123] Assume that the measurement value z i obeys Gaussian distribution at a given position x:

[0124]

[0125] wherein, represents the variance of the i-th sensor.

[0126] S3.33, calculate the joint likelihood function through Gaussian distribution;

[0127] Specifically, the expression of calculating the joint likelihood function through Gaussian distribution is as follows:

[0128]

[0129] Through the joint likelihood function, the information of multi-sensor data can be more comprehensively reflected, and the accuracy of system state estimation can be improved.

[0130] S3.34, update the posterior probability distribution of the system state.

[0131] In this example, the specific process of updating the posterior probability distribution of the system state is as follows:

[0132] The joint likelihood function is obtained from S3.33:

[0133]

[0134] wherein, P(x|x) represents the joint likelihood function; P(z i |x) represents the likelihood function of the i-th sensor; x represents the state vector; z represents the data; z i represents the data of the i-th sensor; n represents the number of z; i represents the index variable;

[0135] The posterior probability distribution of the system state is obtained after updating:

[0136]

[0137] wherein, P(x|z) represents the posterior probability; P(z) represents the marginal probability of data z;

[0138] Considering the accuracy of updating the posterior probability distribution of the system state, an electromagnetic interference variable is introduced to optimize the Bayesian fusion algorithm:

[0139]

[0140]

[0141]

[0142] wherein, represents the i-th sensor data after introducing the electromagnetic interference variable; ∈ i represents the electromagnetic interference variable of the i-th sensor; ∈ represents the electromagnetic interference variable; P(z|x, ∈) represents the joint likelihood function after introducing the electromagnetic interference variable; P(x|z)' represents the posterior probability after introducing the electromagnetic interference variable.

[0143] Specifically, by explicitly considering the influence of electromagnetic interference in the sensor data, the model can more accurately capture the state of the actual system, reduce the error caused by interference, and thus ensure that the posterior probability distribution is closer to the true situation.

[0144] Data preprocessing ensures the quality and consistency of the input data; setting the prior probability distribution provides an initial estimate of the system state; the Bayesian fusion algorithm updates and optimizes the estimate of the system state by combining multi-source observation data, thereby improving the prediction accuracy. Finally, the posterior probability distribution output by the information fusion unit 2 provides reliable data support for subsequent data analysis and decision-making.

[0145] The big data intelligent analysis and processing system based on information fusion further comprises a data analysis unit 3, which analyzes the data based on the fused data from the information fusion unit 2 through a random forest algorithm;

[0146] In this example, the circuit performance is predicted and analyzed through a random forest model, and the minimum impurity reduction variable is introduced to optimize the random forest model. The specific steps involved are as follows:

[0147] Specifically, random forest (Random Forest) is an ensemble learning method mainly used for classification and regression tasks. It improves the accuracy and robustness of the model by constructing multiple decision trees and aggregating their prediction results. The main feature of the random forest algorithm is to reduce the overfitting problem of individual decision trees by introducing randomness and diversity, and to improve the generalization ability of the model by integrating multiple trees.

[0148] S6.1, obtaining the fused data from the information fusion unit 2 to form a data set;

[0149] S6.2, dividing the data set into a training set and a validation set;

[0150] S6.3, constructing a random forest model through the training set;

[0151] In this example, the specific steps for constructing a random forest model through the training set are as follows:

[0152] S6.31, randomly select n samples from the training set as a new training set;

[0153] Specifically, the implementation process of S6.31 is as follows:

[0154]

[0155] S6.32, initialize the root node, and take the new training set as the input data of the root node;

[0156] Specifically, the implementation process of S6.32 is as follows:

[0157] def initialize_root_node(X, y):

[0158] root_node = {'X': X, 'y': y}

[0159] return root_node

[0160] S6.33, select the optimal feature and split point by the Gini Impurity criterion, and split the node into two child nodes;

[0161] Specifically, Gini Impurity is an index used to measure the purity of a dataset, commonly used in feature selection and node splitting in decision tree algorithms. It reflects the probability that two randomly selected samples from the dataset have different class labels. The lower the Gini Impurity value, the higher the purity of the dataset, i.e., most samples in the dataset belong to the same class; on the contrary, the higher the Gini Impurity, the greater the degree of mixture of the dataset, and the more uniform the class distribution. In building a decision tree, the Gini Impurity criterion is used to select the optimal feature and split point to maximize the purity of the child nodes, thereby improving the classification performance of the model.

[0162] In this example, the specific process of selecting the optimal feature and split point by the Gini Impurity criterion and splitting the node into two child nodes is as follows:

[0163] For a node N, its Gini Impurity G(N) is defined as:

[0164]

[0165] where G(N) represents the Gini Impurity of the current node N; K represents the number of classes; p k represents the proportion of samples belonging to class k in node N; k represents the index variable;

[0166] For each feature j and each candidate split point s:

[0167] Calculate the Gini impurity of the left child node:

[0168]

[0169] where G(N L ) represents the Gini impurity of the left child node; |N Lk | represents the number of samples in the left child node N L that belong to class k; |N L | represents the number of samples in the left child node N L .

[0170] Calculate the Gini impurity of the right child node:

[0171]

[0172] where G(N R ) represents the Gini impurity of the right child node; |N Rk | represents the number of samples in the right child node N R that belong to class k; |N R | represents the number of samples in the left child node N R .

[0173] Calculate the weighted Gini impurity after splitting:

[0174]

[0175] where G split (N) represents the weighted Gini impurity after splitting; |N| represents the total number of samples in node N.

[0176] Select the feature j * and the split point s * with the smallest weighted Gini impurity as the optimal feature and split point:

[0177]

[0178] Split the current node Z into two child nodes N * and N * according to the optimal feature j L and the split point s R ; j represents the feature index; s represents the split threshold.

[0179] Specifically, the process of selecting the optimal feature and split point through the Gini impurity criterion and splitting the node into two child nodes ensures that the model can efficiently extract key features from multi-source data and construct a decision tree with a reasonable structure and strong generalization ability. This process not only improves the accuracy of the model but also enhances the robustness and stability of the system, especially when dealing with complex, high-dimensional, and noisy large data sets. It effectively reduces overfitting and improves the predictive performance of the model. By optimizing the splitting condition, the system can effectively integrate information between different data sources and capture potential patterns in the data, providing reliable support for subsequent intelligent analysis and decision-making.

[0180] S6.34, repeat the steps of selecting the optimal feature and split point and splitting the node for each child node until the stopping condition is reached;

[0181] Specifically, the implementation process of S6.34 is as follows:

[0182]

[0183]

[0184] S6.35, repeat the above steps T times to build T decision trees and form a random forest. Specifically, the implementation process of S6.35 is as follows:

[0185]

[0186]

[0187] By randomly sampling from the training set to increase the diversity and robustness of the model, and by recursively selecting the optimal feature and split point to build decision trees, the complex patterns in the data can be captured. Finally, by building multiple decision trees and integrating their prediction results, the random forest model not only reduces the risk of overfitting but also improves the generalization ability and prediction accuracy of the model, providing reliable support for the intelligent analysis and real-time processing of the system.

[0188] S6.4, use the random forest model with the minimum impurity reduction variable introduced to predict and analyze the performance of the circuit;

[0189] Specifically, the minimum impurity reduction variable selects the feature that can reduce the impurity of the node to the greatest extent for node splitting. Specifically, it measures the reduction of the overall impurity of the child nodes relative to the impurity of the parent node after splitting on a given feature. By selecting the feature that reduces impurity the most as the basis for splitting, the model can build a more pure and better classified decision tree, thereby improving the accuracy and generalization ability of the prediction.

[0190] In this example, the process of predicting the performance of the circuit using the random forest model with the minimum impurity reduction variable introduced is as follows:

[0191] For a new sample x, let each tree in the random forest output a predicted class;

[0192] Determine the final predicted class by majority vote:

[0193]

[0194] Wherein, represents the final predicted class; T represents the total number of decision trees in the random forest; I represents the indicator function; y t represents the predicted class of the sample by the tth tree; c represents one of the classes; represents selecting the class c with the maximum ratio as the final predicted class;

[0195] For a new sample x, let each tree in the random forest output a predicted value;

[0196] Determine the predicted value by averaging:

[0197]

[0198] Wherein, represents the predicted value; y t represents the predicted value of the sample by the tth tree;

[0199] Considering the accuracy of the predicted value, introduce the minimum impurity reduction variable to optimize the regression task prediction process of the random forest model;

[0200] When each tree t is constructed, the split node satisfies:

[0201] G(N)-G split (N)≥min impurity decrease ;

[0202] Wherein, min impurity decrease represents the set minimum impurity reduction variable;

[0203] Then determine the final predicted class by majority vote after introducing the minimum impurity reduction variable:

[0204]

[0205] When each tree t is constructed, the split node satisfies:

[0206] MSE(N)-MSE split (N)≥min impuritydecrease ;

[0207] where MSE(N) represents the mean squared error of the current node N; MSE split (N) represents the weighted mean squared error after splitting;

[0208] The final prediction value is determined by introducing the average value of the minimum impurity reduction variable:

[0209]

[0210] where, represents the final prediction value after introducing the minimum impurity reduction variable.

[0211] Specifically, the specific steps of using the random forest model with the minimum impurity reduction variable play a key role in the prediction and analysis of circuit performance, ensuring the accuracy and reliability of the system. These steps allow each tree in the random forest to make predictions on new samples, and determine the final prediction result through majority voting (classification task) or average value (regression task), thereby fully utilizing the integration advantage of multiple trees, improving the generalization ability and prediction accuracy of the model. Especially after introducing the weighted average algorithm, by calculating the mean squared error of each tree on the validation set and assigning corresponding weights, the prediction process of the regression task is further optimized. This method not only considers the prediction performance of each tree, but also ensures the rationality of the weighted average by normalizing the weights, thereby significantly improving the accuracy and stability of the prediction value, providing stronger support for intelligent analysis and real-time processing of the system.

[0212] By limiting unnecessary splits, the model can accurately capture the true patterns in the data while maintaining a simple structure, reducing the impact of noise and redundant features, and improving the accuracy of the prediction. The introduction of the minimum impurity reduction variable makes the model more stable, especially when dealing with multi-source data, it can effectively deal with the differences and noise between different data sources, avoid overfitting, and ensure the reliable performance of the model in complex environments. By reducing unnecessary splits, the training and prediction speed of the model is improved, reducing the consumption of computing resources, especially suitable for large-scale data sets and real-time processing scenarios. In the process of information fusion, the optimized random forest model can better integrate information from different sensors or data sources, provide more reliable prediction results, and support more intelligent decision-making.

[0213] S6.5, use the validation set to evaluate the random forest model through cross-validation.

[0214] Specifically, the implementation process is as follows:

[0215]

[0216]

[0217]

[0218] Through cross-validation, the system can evaluate the performance of the model on multiple different subsets of data, providing a more comprehensive understanding of the model's performance under different data distributions. This not only helps to discover potential overfitting or underfitting problems of the model, but also provides more reliable performance indicators for the selection and optimization of the model. Ultimately, this step improves the prediction accuracy and robustness of the system, ensuring that it can effectively cope with complex and variable data environments in practical applications.

[0219] The big data intelligent analysis processing system based on information fusion further comprises a real-time processing unit 4, which evaluates and warns according to the analysis result of the data analysis unit 3.

[0220] In this example, the real-time processing unit 4 evaluates and warns according to the analysis result of the data analysis unit 3, and the specific steps for evaluation and warning are as follows:

[0221] S10.1, receiving the prediction analysis result from the data analysis unit 3;

[0222] S10.2, calculating the key performance indicators according to the analysis result;

[0223] Specifically, the process of calculating the key performance indicators according to the analysis result is as follows:

[0224] Select indicators: According to the characteristics and application scenarios of the system, determine the key performance indicators (KPIs) that need to be calculated. Common KPIs include power consumption, temperature, current, voltage, signal quality, failure probability, etc.

[0225] Indicator calculation: Based on the received analysis result, use predefined formulas or algorithms to calculate each KPI; for example:

[0226] Power consumption = voltage × current

[0227] Temperature change rate = (current temperature - previous time temperature) / time interval

[0228] Failure probability = probability value output by the classification model

[0229] Multi-source data fusion: If there are multiple sensors or data sources in the system, the data from different sources can be fused to obtain more accurate performance indicators.

[0230] S10.3, compare the calculated performance indicators with the preset threshold;

[0231] Specifically, the process of comparing the calculated performance indicators with the preset threshold values is as follows:

[0232] Threshold setting: Set reasonable threshold ranges for each KPI. Threshold values can be based on historical data, empirical rules, expert knowledge, or dynamic adjustments. For example:

[0233] Normal operating temperature range: 20-60°C

[0234] Maximum allowed power consumption: 50W

[0235] Lower limit of signal quality: 80%

[0236] Threshold comparison: Compare the calculated performance indicators with the preset threshold values item by item to determine whether they are outside the normal range. For example:

[0237] If the current temperature > 60°C, trigger a high temperature warning.

[0238] If the power consumption > 50W, trigger a power consumption warning.

[0239] If the signal quality < 80%, trigger a signal quality problem warning.

[0240] S10.4, based on the results of state evaluation, assign a risk score to the system;

[0241] Specifically, the process of assigning a risk score to the system based on the results of state evaluation is as follows:

[0242] Risk score model: Use a predefined risk score model to consider multiple KPIs and their weights to assign a risk score to the system. The risk score can be a numerical value representing the current risk level of the system, or a multi-dimensional score covering different risk factors (such as safety, reliability, efficiency, etc.).

[0243] Risk score calculation: Calculate the risk score based on the results of threshold comparison. For example:

[0244] If a KPI exceeds the threshold, increase the corresponding risk score.

[0245] If multiple KPIs exceed the threshold at the same time, further increase the risk score.

[0246] Weighted summation, logistic regression, decision tree, etc. can be used to calculate the risk score.

[0247] Uncertainty analysis: Considering the uncertainty and noise in the data, use probabilistic models (such as Bayesian networks, Monte Carlo simulation) to evaluate the risk probability under different scenarios, and further optimize the risk score.

[0248] S10.5, Determine whether an alert needs to be generated based on the results of the risk assessment.

[0249] Specifically, the process of determining whether an alert needs to be generated based on the results of the risk assessment is as follows:

[0250] Trigger condition determination: Based on the risk score, determine whether an alert needs to be generated. Common trigger conditions include:

[0251] The risk score reaches or exceeds a certain threshold.

[0252] A certain performance indicator exceeds a preset threshold.

[0253] Trend analysis shows potential deterioration trend.

[0254] Alert level classification: According to the severity of the risk, the alert is divided into different levels (such as low, medium and high). Different levels of alerts can trigger different response measures.

[0255] Alert content generation: Generate specific alert information, including:

[0256] Alert type (such as performance anomaly, failure warning, safety warning, etc.).

[0257] Specific performance indicators and their current values.

[0258] Possible cause analysis.

[0259] Suggested countermeasures.

[0260] Notification and response: Send alert information to relevant personnel or systems through various channels. Common notification methods include:

[0261] Email or SMS: Send to maintenance personnel or managers.

[0262] Alarm system: Trigger sound and light alarm devices to remind on-site workers.

[0263] Visual interface: Display alert information on the monitoring system or dashboard for easy real-time viewing. Automatic response: For some emergency situations, the system can automatically perform predefined operations such as shutting down equipment, switching to backup systems, adjusting operating parameters, etc. to prevent further damage or accidents.

[0264] The above shows and describes the basic principles, main features and advantages of the present application. Those skilled in the art should understand that the present application is not limited by the above examples, and the above examples and descriptions in the specification are only preferred examples of the present application and are not intended to limit the present application. Without departing from the spirit and scope of the present application, various changes and improvements can be made to the present application, and these changes and improvements all fall within the scope of the claimed present application.

Claims

1. A big data intelligent analysis processing system based on information fusion, characterized in that, The utility model relates to a kind of circuit performance prediction method and system based on multi-source data fusion. Data acquisition unit (1), the data acquisition unit (1) is collected by multi-source data acquisition module circuit performance data of different data sources; Information fusion unit (2), the information fusion unit (2) is fused by introducing electromagnetic interference variable optimization bayesian fusion algorithm to the circuit performance data of different data sources based on the circuit performance data of different data sources collected by data acquisition unit (1); Electromagnetic interference variable optimization bayesian fusion algorithm: ; ; ; wherein represents the i-th sensor data after introducing the electromagnetic interference variable; represents the i-th sensor data after introducing the electromagnetic interference variable; represents the electromagnetic interference variable of the i-th sensor; represents the electromagnetic interference variable; represents the joint likelihood function after introducing the electromagnetic interference variable; represents the posterior probability after introducing the electromagnetic interference variable, represents the i-th sensor data, represents the i-th sensor data, represents the marginal probability of the data represents the number of represents represents the number of is a prior probability distribution of setting the system state; Data analysis unit (3), the data analysis unit (3) is based on the data fused by information fusion unit (2), and the circuit performance is predicted and analyzed by random forest model, and the random forest model is optimized by introducing minimum impurity reduction variable, to improve the accuracy of predicted value; Real-time processing unit (4), the real-time processing unit (4) carries out evaluation and early warning according to the analysis result of data analysis unit (3). 2.The big data intelligent analysis processing system based on information fusion of claim 1, wherein: The data acquisition unit (1) includes multi-source data acquisition module, and the multi-source data acquisition module includes database module and sensor module. The database module is used for storing collected data. The sensor module is used for collecting circuit performance data. 3.The big data intelligent analysis processing system based on information fusion of claim 1, wherein: The specific steps involved in the information fusion unit (2) fusing the circuit performance data of different data sources by introducing electromagnetic interference variable optimization bayesian fusion algorithm are as follows: S3.1, data preprocessing is carried out by removing noise, missing value filling and abnormal value processing, and different formats of data are converted into a unified format; S3.2, Set the prior probability distribution of the system state ; S3.3, the posterior probability distribution of system state is updated using bayesian fusion algorithm. 4.The big data intelligent analysis processing system based on information fusion of claim 3, wherein: In S3.3, the specific steps of updating the posterior probability distribution of system state using bayesian fusion algorithm are as follows: S3.31, the measurement value of sensor is obtained from sensor module; S3.32, the likelihood function of sensor is calculated by Gaussian distribution; S3.33, the joint likelihood function is calculated by Gaussian distribution; S3.34, the posterior probability distribution of system state is updated. 5.The big data intelligent analysis processing system based on information fusion of claim 4, wherein: In S3.34, the specific process of updating the posterior probability distribution of system state is as follows: The joint likelihood function is obtained from S3.33: ; wherein, represents a joint likelihood function; represents a likelihood function of the th sensor; represents a state vector; represents data; represents data of the th sensor; represents a number of represents an index variable; The posterior probability distribution of system state is obtained by updating: ; wherein, denotes the posterior probability; denotes the data marginal probability; Considering the accuracy of updating the posterior probability distribution of system state, electromagnetic interference variable optimization bayesian fusion algorithm is introduced. 6.The big data intelligent analysis processing system based on information fusion of claim 1, wherein: The specific steps involved in the data analysis unit (3) predicting and analyzing the circuit performance by random forest model and optimizing the random forest model by introducing minimum impurity reduction variable are as follows: S6.1, the data set is formed by obtaining the fused data from information fusion unit (2); S6.2, the data set is divided into training set and validation set; S6.3, the random forest model is constructed by training set; S6.4, the circuit performance is predicted and analyzed by using the random forest model with minimum impurity reduction variable introduced; S6.5, the random forest model is evaluated by cross-validation using validation set. 7.The big data intelligent analysis processing system based on information fusion of claim 6, wherein: In S6.3, the specific steps of constructing random forest model by training set are as follows: S6.31, a number of samples are randomly extracted from training set as new training set. S6.32, initialize the root node, and take the new training set as the input data of the root node; S6.33, select the optimal feature and split point by the Gini impurity criterion, and split the node into two child nodes; S6.34, repeat the steps of S6.32-S6.33 to select the optimal feature and split point and split the node until the stopping condition is reached; S6.35, repeat the above steps subsequently, construct a decision tree, forming a random forest. 8.The big data intelligent analysis processing system based on information fusion of claim 7, wherein: In S6.33, the specific process of selecting the optimal feature and split point by the Gini impurity criterion and splitting the node into two child nodes is as follows: For a node its Gini impurity is defined as: ; wherein, represents the Gini impurity of the current node ; represents the number of classes; represents the proportion of samples belonging to class in node ; represents the index variable; For each feature and each candidate split point : Calculate the Gini impurity of the left child node: ; wherein, Gini impurity of the left child node; Gini impurity of the left child node number of samples in the left child node belonging to the class number of samples in the left child node number of samples in the left child node number of samples in the left child node Calculate the Gini impurity of the right child node: ; wherein, Gini impurity of the right child node; Gini impurity of the right child node number of samples in the right child node belonging to class number of samples in the right child node number of samples in the right child node number of samples in the right child node Calculate the weighted Gini impurity after splitting: ; wherein, represents the weighted Gini impurity after split; represents the total number of samples in the node . selecting the feature with the minimum weighted gini impurity and split point as the optimal feature and split point ; According to the optimal feature and the split point , the current node is split into two child nodes and ; represents the feature index; represents the split threshold. 9.The big data intelligent analysis processing system based on information fusion of claim 6, wherein: In S6.4, the process of using the random forest model with the minimum impurity reduction variable introduced to predict and analyze the performance of the circuit is as follows: For a new sample, let each tree in the random forest output a predicted class; Determine the final predicted class by majority voting: ; wherein, represents the final predicted class; represents the total number of decision trees in the random forest; represents the indicator function; represents the prediction class of the sample by the th tree; represents one of the classes; represents the class that is selected to be the largest as the final predicted class; For a new sample, let each tree in the random forest output a predicted value; Determine the predicted value by averaging: ; wherein, denotes a predicted value; denotes the predicted value of the tree for the sample; Considering the accuracy of the predicted value, introduce the minimum impurity reduction variable to optimize the random forest model's regression task prediction process; When each tree is constructed, the split nodes satisfy: ; wherein, represents the set minimum impurity reduction variable; Then determine the final predicted class by majority voting after introducing the minimum impurity reduction variable: ; When each tree is constructed, the split nodes satisfy: ; wherein, represents the mean squared error of the current node ; represents the weighted mean squared error after splitting; Then determine the final predicted value by averaging after introducing the minimum impurity reduction variable: ; wherein, represents the final prediction value after introduction of the minimum impurity reduction variable. 10.The big data intelligent analysis processing system based on information fusion of claim 1, wherein: The real-time processing unit (4) evaluates and warns based on the analysis results of the data analysis unit (3), and the specific steps for evaluation and warning are as follows: S10.1, receive the prediction analysis result from the data analysis unit (3); S10.2, calculate the key performance indicators based on the analysis results; S10.3, compare the calculated performance indicators with the preset threshold; S10.4, based on the results of the state evaluation, assign a risk score to the system; S10.5, determine whether an early warning is needed based on the results of the risk evaluation.

Citation Information

Patent Citations

  • Comprehensive electronic system electromagnetic compatibility analysis method based on knowledge graph

    CN117150038A

  • Main transformer fault prediction method based on multivariable data fusion

    CN118779831A