Data visualization system based on big data analysis

By combining the deep forest model with the pigeon flock optimization algorithm, the hyperparameter configuration of big data analysis is optimized, which solves the efficiency and accuracy problems in big data processing, and realizes efficient, intelligent and interactive data analysis and decision support.

CN120508592AInactive Publication Date: 2025-08-19HAIKOU QISHUN BIG DATA TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510620409.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to process and analyze big data efficiently and accurately, especially in hyperparameter configuration and data visualization, which has problems such as low efficiency and easy to fall into local optimal solutions. Traditional methods are difficult to meet the multi-dimensional, high complexity and real-time requirements of big data.

Method used

The deep forest model is used combined with the pigeon flock optimization algorithm to automatically optimize hyperparameter configuration, and classification, regression and prediction are performed through big data analysis. At the same time, interactive analysis is introduced to update visual results in real time to provide decision support and risk control.

Benefits of technology

It improves the accuracy and efficiency of big data analysis, enhances the flexibility and interactivity of data visualization, ensures the stability and efficiency of the system in the big data environment, and supports real-time decision-making and risk management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508592A_ABST
    Figure CN120508592A_ABST
Patent Text Reader

Abstract

The invention discloses a data visualization system based on big data analysis, and the system comprises a reading module which is used for reading massive big data sets; the deep forest model module is used for constructing a deep forest model and inputting the standardized data set into the model for training and analysis; the pigeon inspired optimization algorithm module is used for configuring a pigeon inspired optimization algorithm according to the initial output of the deep forest model and optimizing hyper-parameter configuration of the deep forest model; the big data analysis module is used for carrying out classification, regression and prediction analysis on big data according to the optimal hyper-parameter configuration; the interactive analysis module is used for updating an analysis result in real time according to an interaction demand input by a user; the decision support module is used for providing big data driven decision support; and the risk control module is used for analyzing the business operation condition and carrying out risk assessment and control. According to the invention, accurate data visualization is provided for the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data analysis, and in particular to a data visualization system based on big data analysis. Background Art

[0002] With the rapid development of information technology, the application of big data technology has become increasingly widespread across various industries, particularly in the fields of data analysis and decision support. The ability to process and analyze big data has become a key driver for industrial upgrading and intelligent decision-making. However, big data analysis faces many technical challenges, the most prominent of which is how to efficiently and accurately process and analyze massive amounts of data and extract valuable information from it. Currently, traditional data analysis methods often struggle to cope with the multidimensional, complex, and highly nonlinear characteristics of big data. Therefore, optimizing data analysis models and improving their predictive accuracy and computational efficiency has become a hot topic in current technical research.

[0003] In the process of big data analysis, the application of machine learning and deep learning technologies has gradually become a core means of solving problems. Deep learning models, particularly the deep forest model, have achieved remarkable results in tasks such as data classification, regression, and prediction due to their powerful feature extraction and data representation capabilities. Deep forest models, through their multi-layered decision tree networks, are able to learn complex patterns and nonlinear relationships in data, enabling relatively effective data analysis and prediction. Despite their widespread application in various fields, deep forest models still face numerous challenges and limitations in practical applications.

[0004] First, the hyperparameter configuration of deep forest models remains a major challenge in their application. The performance of deep forest models is highly dependent on the selection of hyperparameters, including the number of trees, the maximum tree depth, the minimum number of sample splits per tree, and the minimum number of leaf node samples per tree. Traditionally, hyperparameter selection relies on empirical experience or manual tuning methods such as grid search. However, these methods are often time-consuming and difficult to guarantee the optimal hyperparameter configuration. Especially when processing large-scale data, manual hyperparameter tuning is not only inefficient but also prone to local optimal solutions. Therefore, how to optimize the hyperparameter configuration of deep forest models through automated methods has become a critical issue facing current technology.

[0005] Secondly, while traditional optimization methods such as particle swarm optimization and genetic algorithms can address hyperparameter optimization to a certain extent, these methods often suffer from high computational overhead, slow convergence, and a tendency to get stuck in local optimal solutions. Furthermore, when faced with large data feature spaces, the efficiency and accuracy of traditional optimization algorithms cannot meet the demands of real-time analysis and decision-making. Therefore, more efficient and precise optimization algorithms are needed to improve the efficiency and accuracy of hyperparameter searches and reduce computational time and resource consumption.

[0006] Furthermore, the application of data visualization technology in big data analysis faces certain bottlenecks. Big data is characterized by high dimensionality, large volume, and high complexity. Traditional data visualization methods are often limited to processing low-dimensional, relatively simple datasets, and struggle to meet the real-time visualization requirements for complex data in big data environments. Especially during real-time data updates and interactive analysis, the responsiveness and presentation quality of traditional visualization methods often fall short of user needs. Efficiently presenting big data analysis results while ensuring real-time and interactivity has become a pressing challenge for data visualization technology.

[0007] To address these technical deficiencies, in recent years, deep forest model hyperparameter optimization methods based on the pigeon swarm optimization algorithm have gradually become a research hotspot. The pigeon swarm optimization algorithm is a new type of swarm intelligence optimization algorithm that simulates the collective behavior and information exchange mechanisms of a flock of pigeons, effectively searching for the global optimal solution in a multi-dimensional feature space. Compared to traditional optimization algorithms, the pigeon swarm optimization algorithm has stronger global search capabilities and can effectively avoid the problem of local optimal solutions. In the hyperparameter optimization of the deep forest model, the pigeon swarm optimization algorithm can quickly find the optimal hyperparameter configuration in a vast feature space by simulating the interaction and collaborative behavior between pigeons, thereby improving the performance and computational efficiency of the deep forest model.

[0008] However, despite the great potential of the pigeon swarm optimization algorithm for hyperparameter optimization of deep forest models, its application still faces several challenges. First, the search efficiency and accuracy of the pigeon swarm optimization algorithm still need to be further improved. When processing large-scale data, the pigeon swarm optimization algorithm has a high computational overhead, especially when the number of pigeons is large, the algorithm's time complexity and resource consumption increase exponentially. Second, the convergence speed and global search capability of the pigeon swarm optimization algorithm are significantly affected by the initial parameter settings and the position of the pigeons. Therefore, how to further improve the efficiency and stability of the optimization algorithm by dynamically adjusting the search strategy and position of the pigeons is an urgent problem to be solved.

[0009] Therefore, how to provide a data visualization system based on big data analysis is an urgent problem that those skilled in the art need to solve. Summary of the Invention

[0010] One objective of the present invention is to provide a data visualization system based on big data analysis. This system automatically optimizes hyperparameter configurations using a deep forest model combined with a pigeon flock optimization algorithm. It also uses big data analysis to perform classification, regression, and predictive analysis. Furthermore, the system incorporates interactive analysis, enabling real-time updates of visualization results based on user needs, while also providing decision support and risk control. By visualizing big data, the system displays analysis results under different parameter configurations, providing users with precise decision-making and risk control strategies.

[0011] A data visualization system based on big data analysis according to an embodiment of the present invention includes:

[0012] Reading module, used to read massive data sets;

[0013] Deep Forest Model module, used to build a deep forest model and input standardized data sets into the model for training and analysis;

[0014] Pigeon Flock Optimization algorithm module, used to configure the Pigeon Flock Optimization algorithm based on the preliminary output of the deep forest model and optimize the hyperparameter configuration of the deep forest model;

[0015] Big data analysis module, used to perform classification, regression and predictive analysis of big data based on optimal hyperparameter configuration;

[0016] Interactive analysis module, used to update analysis results in real time based on user input interaction requirements;

[0017] Decision support module, used to provide big data-driven decision support;

[0018] Risk control module, used to analyze business operations and conduct risk assessment and control;

[0019] Optionally, modules can be connected using the following methods:

[0020] S1. Collect massive data sets, preprocess them, and generate standardized data sets;

[0021] S2. Build a deep forest model. Input the standardized dataset into the multi-level deep forest model. By mapping the features of the standardized dataset into a multi-level tree structure, the model learns the patterns and nonlinear relationships in the data layer by layer, generating the initial output of the deep forest model.

[0022] S3. Based on the initial output of the deep forest model, the pigeon swarm optimization algorithm configures a location exploration unit for each pigeon. Based on the migration paths and collective behavior of all pigeons, it simulates the collaboration and information exchange mechanism of the pigeon swarm to optimize the hyperparameter configuration of the deep forest model.

[0023] S4. Based on the optimized hyperparameter configuration, the pigeon flock optimization algorithm is used to simulate the collective behavior and migration path of the pigeon flock, traverse the big data feature space, and find the optimal parameter configuration of the deep forest model;

[0024] S5. Use the optimal parameter configuration to perform classification, regression, and prediction analysis on big data, output the analysis results, and display the analysis results through the big data visualization unit;

[0025] S6. Update visualization results in real time based on user interaction needs, support interactive analysis of big data, and display analysis results under different parameter configurations;

[0026] S7. Based on the analysis results, provide big data-driven decision support for further system optimization, business decision-making and risk control.

[0027] Optionally, the hyperparameters include the number of trees, the maximum depth of the tree, the minimum number of sample splits for each tree, the minimum number of leaf node samples for each tree, the learning rate, and the subsample ratio.

[0028] Optionally, the nonlinear relationship includes an exponential relationship, a logarithmic relationship, a quadratic relationship, and a periodic relationship.

[0029] Optionally, S2 includes the following specific steps:

[0030] S21. Based on the standardized data set, initialize the deep forest model and configure the multi-level structure of the deep forest model. Each level contains multiple decision trees. Initially, the parameters of each level are set to the default values, that is, 100 trees, the maximum tree depth is 5, the minimum number of sample splits per tree is 2, and the minimum number of leaf node samples per tree is 1.

[0031] S22. Input the standardized data set into the first layer of the deep forest model, perform preliminary feature extraction on the input data through the decision tree of the first layer, and generate the output of the first layer;

[0032] S23, passing the output data of the first layer to the second layer of the deep forest model, and using the decision tree of the second layer to further process and learn the patterns and nonlinear relationships in the output data to generate the output of the second layer;

[0033] S24, repeat S23, pass the output data to the next layer layer by layer, and gradually extract the features and nonlinear relationships in the output data through the combination of each layer of trees until the final layer;

[0034] S25. Output the preliminary output of the deep forest model.

[0035] Optionally, S3 includes the following specific steps:

[0036] S31. Initialize the pigeon swarm optimization algorithm based on the preliminary output of the deep forest model, set each pigeon as a candidate solution, and the candidate solution is the hyperparameter configuration of the deep forest model;

[0037] S32, configuring a position exploration unit for each pigeon body to simulate the position of the pigeon body in the search space, where the position of the pigeon body in the search space is the current value of each hyperparameter in the deep forest model;

[0038] S33. Based on the pigeon swarm algorithm, the pigeon swarm optimization algorithm dynamically adjusts the weight of each pigeon according to the initial output of the deep forest model through a dynamic adaptive adjustment mechanism:

[0039]

[0040] in, The new position of the i-th pigeon after the current position is updated, P i represents the current position information of the i-th pigeon at the current position, ω represents the pigeon inertia weight, α represents the adjustment factor that controls the pigeon's exploration range, represents the distance between the i-th pigeon and the j-th pigeon, β represents the factor that controls the speed of convergence between the pigeon and the global optimal solution, P j The position of the j-th pigeon, N represents the total number of pigeons in the pigeon group, i represents the i-th pigeon in the pigeon group, j represents the reference pigeon of the i-th pigeon, P global represents the location of the global optimal solution, P j Indicates the position of the j-th pigeon;

[0041] S34. Calculate the fitness value of each pigeon based on its current position and the performance evaluation of the deep forest model. Adopt an adaptive adjustment mechanism to weight the fitness and design a multi-level filtering mechanism to sort and screen the fitness of the pigeons, retaining the best performing pigeons. Further optimize the hyperparameter configuration through multiple rounds of refined screening.

[0042] S35. By simulating the individual behavior of the pigeon flock and continuously adjusting the hyperparameter configuration, the pigeons adjust themselves according to their own performance in each round of iteration.

[0043] Optionally, S4 includes the following specific steps:

[0044] S41, receiving the hyperparameter configuration optimized by the pigeon flock optimization algorithm and initializing the big data feature space;

[0045] S42. Set the initial position of the pigeon group according to the hyperparameter configuration. The initial position of each pigeon represents a hyperparameter combination in the deep forest model, and set the initial exploration range of the pigeon.

[0046] S43. Calculate the performance index of the deep forest model corresponding to the current pigeon body based on the pigeon body position and hyperparameter configuration:

[0047]

[0048] Among them, E i represents the fitness value of the i-th pigeon, i represents the i-th pigeon in the pigeon group, The new position of the i-th pigeon after the current position is updated, Indicates the best performing pigeon body position in history, represents the distance between the i-th pigeon body and the best pigeon body in history, δ represents the attenuation factor, γ represents the adjustment factor, error i represents the error value of the deep forest model obtained by the hyperparameter configuration corresponding to the i-th pigeon, ∈ represents the tolerance parameter, η represents the factor of weighted error, N represents the total number of pigeons in the pigeon group, and k represents the pigeon with the best performance in history;

[0049] S44. Introducing an adaptive gradient weighting mechanism during the pigeon group update process, by dynamically adjusting the search step size of each pigeon, so that the search step size of each pigeon is not only based on fitness, but also combined with the relative position in the feature space. When the distance between the pigeon and the global optimal solution is less than or equal to 0.1, the search step size is reduced. When the distance between the pigeon and the global optimal solution is greater than 0.1, the search step size is increased.

[0050] S45. Introducing a multi-objective adaptive optimization strategy, through which multi-dimensional fitness calculations are performed during the pigeon body update process, balancing the accuracy of the deep forest model and computing resources;

[0051] S46, through the feedback loop of S44 and S45, adjust the exploration range and direction of the pigeon body in real time, and eventually make the pigeon body converge toward the optimized hyperparameter solution in the feature space and output the optimal hyperparameter configuration.

[0052] Optionally, S5 includes the following specific steps:

[0053] S51. Initialize the deep forest model based on the optimal hyperparameter configuration and apply it to large data sets for classification, regression, and predictive analysis. At the same time, introduce an adaptive training set selection mechanism to automatically adjust the ratio of training and test sets based on data distribution, allowing the model to dynamically adjust the training and test splits based on different data characteristics.

[0054] S52. Dynamically expand the training set during the training process, gradually introduce new data samples during the training of the deep forest model, and continuously adjust the learning weights of the deep forest model;

[0055] S53, applying the trained deep forest model to the test set, predicting the test set, and optimizing the deep forest model in real time based on the prediction results;

[0056] S54. Based on the prediction results, calculate the performance indicators of the model, including accuracy, error rate, and mean square error, and weight different performance indicators according to the importance of each type of data to generate analysis results;

[0057] S55. Organize the analysis results into a data structure to form the data format required for visualization, including classification results, regression values, and performance evaluation indicators;

[0058] S56. Through the big data visualization unit, the analysis results are presented to the user in a graphical and interactive manner, supporting multiple visualization forms and providing dynamic analysis functions.

[0059] Optionally, S6 includes the following specific steps:

[0060] S61, receiving the interaction requirements input by the user, and determining the data content and visualization method to be displayed based on the specific parameter configuration selected by the user;

[0061] S62. Reload the corresponding data set according to the parameter configuration specified by the user, and perform data preprocessing to adapt to the new parameter configuration;

[0062] S63. Recalculate the analysis results of the deep forest model based on the new data set and parameter configuration:

[0063]

[0064] Among them, R i represents the analysis result of the hyperparameter configuration corresponding to the i-th pigeon body, j represents the reference pigeon body of the i-th pigeon body, Indicates the new position of the i-th pigeon after the current position is updated. Indicates the position of the best pigeon in history, represents the distance between pigeon body i and pigeon body j, α' represents the distance weighting factor, β' represents the error weighting factor, error i represents the error value of the deep forest model obtained by the hyperparameter configuration corresponding to the i-th pigeon, γ represents the adjustment factor, N represents the total number of pigeons in the pigeon group, i represents the i-th pigeon in the pigeon group, ∈ represents the tolerance parameter, and δ represents the error square weighting factor;

[0065] S64. After the real-time calculation is completed, the visualization graphics, including heat maps, scatter plots, and time series graphs, are updated through the big data visualization unit, and the display method of the graphics is adjusted according to the dynamic changes of the new data set;

[0066] S65. Based on the user's operations on the visualization interface, including filtering, zooming in and out, the new analysis results are updated and presented in real time, and the display content is adjusted;

[0067] S66. Provide user-defined view function, allowing users to select and save visualization results under different parameter configurations.

[0068] Optionally, S7 includes the following specific steps:

[0069] S71. Based on the new analysis results, identify and extract the key factors that affect performance and business decisions;

[0070] S72. Based on new analysis results and key factors, use forecast data and historical data to conduct multi-dimensional analysis to identify potential optimization opportunities, business growth points, and potential risks;

[0071] S73. Generate optimization suggestions and formulate specific action plans based on the identified potential optimization opportunities, business growth points, and potential risks;

[0072] S74. Apply optimization suggestions to real-time monitoring and tracking of optimization results, and make dynamic adjustments based on real-time data feedback;

[0073] S75. Support senior decision makers' strategic decisions through analysis results and optimization suggestions, and provide comprehensive business data reports and actionable decision-making basis;

[0074] S76. Conduct risk assessment and control based on business and operational conditions, predict potential risks based on analysis results, propose preventive measures and implement risk control strategies.

[0075] The beneficial effects of the present invention are:

[0076] By implementing this invention, we can effectively address several challenges in the prior art by using a deep forest model combined with a pigeon flock optimization algorithm for automated analysis and optimization of big data. First, the deep forest model extracts complex patterns and nonlinear relationships from data layer by layer through a multi-layered tree structure, significantly improving the accuracy and efficiency of data analysis. The introduction of the pigeon flock optimization algorithm, by simulating the collective behavior and information exchange mechanisms of pigeon flocks, optimizes the hyperparameter configuration of the deep forest model, avoiding the inefficiency and tendency to fall into local optimal solutions encountered in traditional manual hyperparameter adjustment.

[0077] Furthermore, the present invention ensures efficient search and optimization in a big data environment by dynamically adjusting the search step size of the pigeon body and introducing a multi-objective adaptive optimization strategy. This optimization not only improves the efficiency of hyperparameter search but also takes into account the computational resource balance of the deep forest model, making the system more stable and efficient when processing big data.

[0078] This invention also addresses user interaction needs, enhancing the flexibility and interactivity of data visualization by displaying analysis results under different parameter configurations in real time through a big data visualization unit. Users can view and update analysis results in real time based on their needs, enabling more accurate data analysis and decision support. Through these innovations, this invention provides an efficient, intelligent, and interactive solution for big data analysis and decision-making, effectively enhancing the practical value of big data analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0080] Figure 1 This is a method flow chart of a data visualization system based on big data analysis proposed by the present invention;

[0081] Figure 2 This is a system flow chart of a data visualization system based on big data analysis proposed by the present invention;

[0082] Figure 3 This is a data flow diagram of a data visualization system based on big data analysis proposed by the present invention. DETAILED DESCRIPTION

[0083] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0084] refer to Figure 1-3 , a data visualization system based on big data analysis, including:

[0085] Reading module, used to read massive data sets;

[0086] Deep Forest Model module, used to build a deep forest model and input standardized data sets into the model for training and analysis;

[0087] Pigeon Flock Optimization algorithm module, used to configure the Pigeon Flock Optimization algorithm based on the preliminary output of the deep forest model and optimize the hyperparameter configuration of the deep forest model;

[0088] Big data analysis module, used to perform classification, regression and predictive analysis of big data based on optimal hyperparameter configuration;

[0089] Interactive analysis module, used to update analysis results in real time based on user input interaction requirements;

[0090] Decision support module, used to provide big data-driven decision support;

[0091] Risk control module, used to analyze business operations and conduct risk assessment and control;

[0092] This paper combines the deep forest model with the pigeon flock optimization algorithm to improve the accuracy and efficiency of big data analysis by optimizing hyperparameter configuration. By adjusting hyperparameter configuration in real time, model performance is optimized, ensuring the accuracy of classification, regression, and predictive analysis. The introduction of an interactive analysis module allows users to update analysis results in real time based on their needs, enhancing the flexibility and adaptability of decision support. The risk control module enables real-time analysis of business operations, prediction of potential risks, and automatic adjustment of strategies, ensuring the stability and security of the system in practical applications.

[0093] In this embodiment, the modules are connected through the following methods:

[0094] S1. Collect massive data sets, preprocess them, and generate standardized data sets;

[0095] S2. Build a deep forest model. Input the standardized dataset into the multi-level deep forest model. By mapping the features of the standardized dataset into a multi-level tree structure, the model learns the patterns and nonlinear relationships in the data layer by layer, generating the initial output of the deep forest model.

[0096] S3. Based on the initial output of the deep forest model, the pigeon swarm optimization algorithm configures a location exploration unit for each pigeon. Based on the migration paths and collective behavior of all pigeons, it simulates the collaboration and information exchange mechanism of the pigeon swarm to optimize the hyperparameter configuration of the deep forest model.

[0097] S4. Based on the optimized hyperparameter configuration, the pigeon flock optimization algorithm is used to simulate the collective behavior and migration path of the pigeon flock, traverse the big data feature space, and find the optimal parameter configuration of the deep forest model;

[0098] S5. Use the optimal parameter configuration to perform classification, regression, and prediction analysis on big data, output the analysis results, and display the analysis results through the big data visualization unit;

[0099] S6. Update visualization results in real time based on user interaction needs, support interactive analysis of big data, and display analysis results under different parameter configurations;

[0100] S7. Based on the analysis results, provide big data-driven decision support for further system optimization, business decision-making and risk control.

[0101] This paper combines the deep forest model with the pigeon flock optimization algorithm to optimize the classification, regression, and prediction accuracy of big data analysis by dynamically adjusting hyperparameter configurations. By simulating the collective behavior and collaboration mechanisms of pigeon flocks, it ensures efficient traversal of the big data feature space, automatically finds the optimal hyperparameter configuration, and improves model performance. The introduction of an interactive visualization module enables users to adjust analysis parameters and update results in real time, enhancing the system's flexibility and adaptability. Through big data-driven decision support, it provides real-time system optimization, business decision-making, and risk control capabilities, ensuring the efficiency and accuracy of the data analysis process.

[0102] In this embodiment, the hyperparameters include the number of trees, the maximum depth of the tree, the minimum number of sample splits per tree, the minimum number of leaf node samples per tree, the learning rate, and the subsample ratio.

[0103] The present invention improves the prediction accuracy and generalization ability of the model by optimizing the hyperparameters of the deep forest model, including the number of trees, maximum depth, minimum number of sample splits, minimum number of leaf node samples, learning rate, and subsample ratio. By dynamically adjusting these hyperparameters, the model can be adapted to different data characteristics in big data analysis, avoiding overfitting and underfitting problems. This optimization process combines multi-dimensional parameter adjustment to ensure that the model maintains a low computational cost while efficiently processing data, thereby improving the operating efficiency and stability of the system.

[0104] In this embodiment, the nonlinear relationship includes an exponential relationship, a logarithmic relationship, a quadratic relationship, and a periodic relationship.

[0105] This paper enhances the deep forest model's ability to learn complex patterns in data by introducing a variety of nonlinear relationships, such as exponential, logarithmic, quadratic, and periodic relationships. These nonlinear relationships ensure that the model can adapt to diverse data characteristics, accurately capture nonlinear changes between data, and improve the accuracy of predictive analysis. By combining these relationships, the model can more comprehensively understand the underlying structure of the data in different data sets and application scenarios, avoiding the limitations of traditional linear models, thereby improving the system's performance and stability in complex data analysis.

[0106] In this embodiment, S2 includes the following specific steps:

[0107] S21. Based on the standardized dataset, initialize the Deep Forest 7 model and configure a multi-layer structure of the Deep Forest model. Each layer contains multiple decision trees. Initially, the parameters of each layer are set to the default values, that is, 100 trees, the maximum tree depth is 5, the minimum number of sample splits per tree is 2, and the minimum number of leaf node samples per tree is 1.

[0108] S22. Input the standardized data set into the first layer of the deep forest model, perform preliminary feature extraction on the input data through the decision tree of the first layer, and generate the output of the first layer;

[0109] S23, passing the output data of the first layer to the second layer of the deep forest model, and using the decision tree of the second layer to further process and learn the patterns and nonlinear relationships in the output data to generate the output of the second layer;

[0110] S24, repeat S23, pass the output data to the next layer layer by layer, and gradually extract the features and nonlinear relationships in the output data through the combination of each layer of trees until the final layer;

[0111] S25. Output the preliminary output of the deep forest model.

[0112] This paper achieves efficient data feature extraction and pattern learning by initializing a deep forest model and configuring a multi-level decision tree structure. Each layer of the decision tree extracts complex features and nonlinear relationships from the data by layer-by-layer transmission and combination, ensuring that the deep forest model can efficiently learn complex patterns in big data. The hierarchical feature extraction of the initial output enhances the model's understanding of the data and improves prediction accuracy. By adjusting the default parameter settings of the tree, this paper ensures rapid adaptation and optimization of results on different data sets while maintaining efficient utilization of computing resources.

[0113] In this embodiment, S3 includes the following specific steps:

[0114] S31. Initialize the pigeon swarm optimization algorithm based on the preliminary output of the deep forest model, set each pigeon as a candidate solution, and the candidate solution is the hyperparameter configuration of the deep forest model;

[0115] S32, configuring a position exploration unit for each pigeon body to simulate the position of the pigeon body in the search space, where the position of the pigeon body in the search space is the current value of each hyperparameter in the deep forest model;

[0116] S33. Based on the pigeon swarm algorithm, the pigeon swarm optimization algorithm dynamically adjusts the weight of each pigeon according to the initial output of the deep forest model through a dynamic adaptive adjustment mechanism:

[0117]

[0118] in, The new position of the i-th pigeon after the current position is updated, P i represents the current position information of the i-th pigeon at the current position, ω represents the pigeon inertia weight, α represents the adjustment factor that controls the pigeon's exploration range, represents the distance between the i-th pigeon and the j-th pigeon, β represents the factor that controls the speed of convergence between the pigeon and the global optimal solution, P j The position of the j-th pigeon, N represents the total number of pigeons in the pigeon group, i represents the i-th pigeon in the pigeon group, j represents the reference pigeon of the i-th pigeon, P global represents the location of the global optimal solution, P j Indicates the position of the j-th pigeon;

[0119] S34. Calculate the fitness value of each pigeon based on its current position and the performance evaluation of the deep forest model. Adopt an adaptive adjustment mechanism to weight the fitness and design a multi-level filtering mechanism to sort and screen the fitness of the pigeons, retaining the best performing pigeons. Further optimize the hyperparameter configuration through multiple rounds of refined screening.

[0120] S35. By simulating the individual behavior of the pigeon flock and continuously adjusting the hyperparameter configuration, the pigeons adjust themselves according to their own performance in each round of iteration.

[0121] The present invention dynamically adjusts hyperparameter configurations through a pigeon swarm optimization algorithm, simulates the behavior of pigeons in feature space, and achieves precise optimization of deep forest model hyperparameters. Each pigeon represents a candidate solution. With each round of iteration, the pigeons self-adjust based on performance feedback, gradually approaching the global optimal solution. The introduction of an adaptive adjustment mechanism and a multi-level filtering mechanism effectively improves the accuracy and efficiency of the optimization process. By weighting and sorting the fitness of the pigeons, the global search capability of hyperparameter optimization is ensured, while avoiding local optimal solutions, improving the accuracy and stability of the system in big data analysis.

[0122] In this embodiment, S4 includes the following specific steps:

[0123] S41, receiving the hyperparameter configuration optimized by the pigeon flock optimization algorithm and initializing the big data feature space;

[0124] S42. Set the initial position of the pigeon group according to the hyperparameter configuration. The initial position of each pigeon represents a hyperparameter combination in the deep forest model, and set the initial exploration range of the pigeon.

[0125] S43. Calculate the performance index of the deep forest model corresponding to the current pigeon body based on the pigeon body position and hyperparameter configuration:

[0126]

[0127] Among them, E i represents the fitness value of the i-th pigeon, i represents the i-th pigeon in the pigeon group, The new position of the i-th pigeon after the current position is updated, Indicates the best performing pigeon body position in history, represents the distance between the i-th pigeon body and the best pigeon body in history, δ represents the attenuation factor, γ represents the adjustment factor, error i represents the error value of the deep forest model obtained by the hyperparameter configuration corresponding to the i-th pigeon, ∈ represents the tolerance parameter, η represents the factor of weighted error, N represents the total number of pigeons in the pigeon group, and k represents the pigeon with the best performance in history;

[0128] S44. Introducing an adaptive gradient weighting mechanism during the pigeon group update process, by dynamically adjusting the search step size of each pigeon, so that the search step size of each pigeon is not only based on fitness, but also combined with the relative position in the feature space. When the distance between the pigeon and the global optimal solution is less than or equal to 0.1, the search step size is reduced. When the distance between the pigeon and the global optimal solution is greater than 0.1, the search step size is increased.

[0129] S45. Introducing a multi-objective adaptive optimization strategy, through which multi-dimensional fitness calculations are performed during the pigeon body update process, balancing the accuracy of the deep forest model and computing resources;

[0130] S46, through the feedback loop of S44 and S45, adjust the exploration range and direction of the pigeon body in real time, and eventually make the pigeon body converge toward the optimized hyperparameter solution in the feature space and output the optimal hyperparameter configuration.

[0131] The present invention optimizes the hyperparameter configuration through a pigeon flock optimization algorithm, dynamically adjusting the exploration path of the pigeon body in the big data feature space. Combining an adaptive gradient weighting mechanism and a multi-objective optimization strategy, it is possible to efficiently adjust the search step size of each pigeon body, balance accuracy and computing resources, and avoid the trap of local optimal solutions. Through real-time feedback and adjustment, it is ensured that the pigeon body converges toward the optimal hyperparameter solution, improving the performance and computational efficiency of the deep forest model. This method not only improves the accuracy of hyperparameter optimization, but also significantly enhances the processing power and flexibility of the system in a big data environment.

[0132] In this embodiment, S5 includes the following specific steps:

[0133] S51. Initialize the deep forest model based on the optimal hyperparameter configuration and apply it to large data sets for classification, regression, and predictive analysis. At the same time, introduce an adaptive training set selection mechanism to automatically adjust the ratio of training and test sets based on data distribution, allowing the model to dynamically adjust the training and test splits based on different data characteristics.

[0134] S52. Dynamically expand the training set during the training process, gradually introduce new data samples during the training of the deep forest model, and continuously adjust the learning weights of the deep forest model;

[0135] S53, applying the trained deep forest model to the test set, predicting the test set, and optimizing the deep forest model in real time based on the prediction results;

[0136] S54. Based on the prediction results, calculate the performance indicators of the model, including accuracy, error rate, and mean square error, and weight different performance indicators according to the importance of each type of data to generate analysis results;

[0137] S55. Organize the analysis results into a data structure to form the data format required for visualization, including classification results, regression values, and performance evaluation indicators;

[0138] S56. Through the big data visualization unit, the analysis results are presented to the user in a graphical and interactive manner, supporting multiple visualization forms and providing dynamic analysis functions.

[0139] This invention uses an adaptive training set selection mechanism and dynamically expanded training sets to ensure that the deep forest model can automatically adjust the ratio of training and test sets based on different data characteristics, thereby improving the model's learning ability in a big data environment. The real-time optimized deep forest model can continuously adjust learning weights based on prediction results, thereby improving the accuracy of classification, regression, and predictive analysis. At the same time, by weighting different performance indicators and displaying the analysis results in a graphical and interactive manner, the flexibility and operability of data visualization are enhanced, helping users to monitor and optimize model effects in real time.

[0140] In this embodiment, S6 includes the following specific steps:

[0141] S61, receiving the interaction requirements input by the user, and determining the data content and visualization method to be displayed based on the specific parameter configuration selected by the user;

[0142] S62. Reload the corresponding data set according to the parameter configuration specified by the user, and perform data preprocessing to adapt to the new parameter configuration;

[0143] S63. Recalculate the analysis results of the deep forest model based on the new data set and parameter configuration:

[0144]

[0145] Among them, R i represents the analysis result of the hyperparameter configuration corresponding to the i-th pigeon body, j represents the reference pigeon body of the i-th pigeon body, Indicates the new position of the i-th pigeon after the current position is updated. Indicates the position of the best pigeon in history, represents the distance between pigeon body i and pigeon body j, α' represents the distance weighting factor, β' represents the error weighting factor, error irepresents the error value of the deep forest model obtained by the hyperparameter configuration corresponding to the i-th pigeon, γ represents the adjustment factor, N represents the total number of pigeons in the pigeon group, i represents the i-th pigeon in the pigeon group, ∈ represents the tolerance parameter, and δ represents the error square weighting factor;

[0146] S64. After the real-time calculation is completed, the visualization graphics, including heat maps, scatter plots, and time series graphs, are updated through the big data visualization unit, and the display method of the graphics is adjusted according to the dynamic changes of the new data set;

[0147] S65. Based on the user's operations on the visualization interface, including filtering, zooming in and out, the new analysis results are updated and presented in real time, and the display content is adjusted;

[0148] S66. Provide user-defined view function, allowing users to select and save visualization results under different parameter configurations.

[0149] This method uses interactive analysis, combined with user-entered parameter configurations, to update the analysis results of the Deep Forest model in real time. It dynamically adjusts the dataset and model configuration to ensure the timeliness and accuracy of the visualization content. By supporting multiple visualization formats such as heat maps, scatter plots, and time series graphs, users can filter, zoom in, and zoom out data views in real time as needed, and view analysis results under different parameter configurations. Furthermore, a custom view function is provided, allowing users to save and query visualization results under different configurations, enhancing the flexibility and interactivity of data analysis.

[0150] In this embodiment, S7 includes the following specific steps:

[0151] S71. Based on the new analysis results, identify and extract the key factors that affect performance and business decisions;

[0152] S72. Based on new analysis results and key factors, use forecast data and historical data to conduct multi-dimensional analysis to identify potential optimization opportunities, business growth points, and potential risks;

[0153] S73. Generate optimization suggestions and formulate specific action plans based on the identified potential optimization opportunities, business growth points, and potential risks;

[0154] S74. Apply optimization suggestions to real-time monitoring and tracking of optimization results, and make dynamic adjustments based on real-time data feedback;

[0155] S75. Support senior decision makers' strategic decisions through analysis results and optimization suggestions, and provide comprehensive business data reports and actionable decision-making basis;

[0156] S76. Conduct risk assessment and control based on business and operational conditions, predict potential risks based on analysis results, propose preventive measures and implement risk control strategies.

[0157] This invention extracts key factors based on analysis results and conducts multi-dimensional analysis combining forecasted and historical data to accurately identify business optimization opportunities and potential risks. The generated optimization recommendations can be monitored and dynamically adjusted in real time to ensure that optimization results are maintained in a constantly changing business environment. By supporting the strategic decisions of senior decision makers and providing actionable business data reports and decision-making basis, it effectively ensures the continued growth of the business. Furthermore, the risk assessment and control mechanism can predict potential risks based on real-time data feedback, taking proactive preventative measures and improving the stability and security of the overall business.

[0158] Example 1:

[0159] To verify the feasibility of this invention, it was applied to a large-scale power equipment monitoring system at a power company. The company uses this system to monitor the status of power equipment in real time, specifically monitoring and analyzing partial discharge (PD). Partial discharge in power equipment is a common fault phenomenon that can cause damage to the equipment and even lead to serious accidents. Therefore, timely identification of PD signals and accurate prediction of equipment status can effectively improve the safety and reliability of equipment operation.

[0160] In this power company, the big data analysis system is used to monitor the operating status of equipment and provide fault warnings. Equipment sensors collect equipment operating data at all times, including voltage, current, and partial discharge signals. These data are uploaded to the central server for further data processing and analysis. The original system used traditional machine learning models for data classification and fault prediction, but due to the wide variety of equipment, complex failure modes, and the traditional model's hyperparameter adjustment method that relies too much on manual experience, the model's prediction accuracy fluctuates greatly. In order to improve the accuracy of the analysis model and the precision of the prediction, the deep forest model of the present invention is introduced in combination with the pigeon flock optimization algorithm, and combined with interactive data visualization technology to display various types of data and analysis results in real time.

[0161] During implementation, the equipment's operating data is first preprocessed, including denoising and standardization, to ensure data quality. A deep forest model is then used for feature extraction and pattern recognition to construct a preliminary equipment failure prediction model. Next, a pigeon flock optimization algorithm automatically optimizes the deep forest model's hyperparameter configuration based on the model's preliminary output. This optimized model achieves higher accuracy and greater generalization capabilities for equipment classification, regression, and predictive analysis, accurately identifying partial discharge signals and potential failures across different equipment. This process presents analysis results through a visualization module, allowing users to interact with the visualization interface, adjust parameter configurations in real time, and update analysis results.

[0162] To verify the effectiveness of this invention, experiments were conducted on a dataset of actual equipment at a power company. The dataset contained partial discharge signals from multiple pieces of power equipment, including transformers, circuit breakers, and cables. The equipment's operating status and partial discharge signals were standardized and then fed into a deep forest model. Hyperparameters were optimized using the pigeon flock optimization algorithm, and the data was displayed using a visualization module.

[0163] Table 1 Comparison of optimization effects of a data visualization system based on big data analysis

[0164]

[0165] Table 1 shows that the prediction accuracy of the traditional model was 85%, while the accuracy reached 93% after optimizing the Deep Forest model with the Pigeon Swarm Optimization algorithm. Specifically, the optimized model significantly improves the accuracy of identifying partial discharge signals in transformers and cables under low voltage conditions, effectively reducing both false alarm and missed alarm rates. This system enables power companies to predict equipment failures in a timely manner and take proactive measures, avoiding potential equipment damage and safety incidents.

[0166] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A data visualization system based on big data analysis, characterized in that: include: Reading module, used to read massive data sets; Deep Forest Model module, used to build a deep forest model and input standardized data sets into the model for training and analysis; Pigeon Flock Optimization algorithm module, used to configure the Pigeon Flock Optimization algorithm based on the preliminary output of the deep forest model and optimize the hyperparameter configuration of the deep forest model; Big data analysis module, used to perform classification, regression and predictive analysis of big data based on optimal hyperparameter configuration; Interactive analysis module, used to update analysis results in real time based on user input interaction requirements; Decision support module, used to provide big data-driven decision support; The risk control module is used to analyze business operations and conduct risk assessment and control.

2. A data visualization system based on big data analysis according to claim 1, characterized in that: The modules are implemented as follows: S1. Collect massive data sets, preprocess them, and generate standardized data sets; S2. Build a deep forest model. Input the standardized dataset into the multi-level deep forest model. By mapping the features of the standardized dataset into a multi-level tree structure, the model learns the patterns and nonlinear relationships in the data layer by layer, generating the initial output of the deep forest model. S3. Based on the initial output of the deep forest model, the pigeon swarm optimization algorithm configures a location exploration unit for each pigeon. Based on the migration paths and collective behavior of all pigeons, it simulates the collaboration and information exchange mechanism of the pigeon swarm to optimize the hyperparameter configuration of the deep forest model. S4. Based on the optimized hyperparameter configuration, the pigeon flock optimization algorithm is used to simulate the collective behavior and migration path of the pigeon flock, traverse the big data feature space, and find the optimal parameter configuration of the deep forest model; S5. Use the optimal parameter configuration to perform classification, regression, and prediction analysis on big data, output the analysis results, and display the analysis results through the big data visualization unit; S6. Update visualization results in real time based on user interaction needs, support interactive analysis of big data, and display analysis results under different parameter configurations; S7. Based on the analysis results, provide big data-driven decision support for further system optimization, business decision-making and risk control.

3. The data visualization system based on big data analysis according to claim 2, characterized in that: The hyperparameters include the number of trees, the maximum depth of the tree, the minimum number of sample splits per tree, the minimum number of leaf node samples per tree, the learning rate, and the subsample ratio.

4. The data visualization system based on big data analysis according to claim 2, characterized in that: The nonlinear relationship includes an exponential relationship, a logarithmic relationship, a quadratic relationship, and a periodic relationship.

5. The data visualization system based on big data analysis according to claim 2, characterized in that: The S2 includes the following specific steps: S21. Based on the standardized data set, initialize the deep forest model and configure the multi-level structure of the deep forest model. Each level contains multiple decision trees. Initially, the parameters of each level are set to the default values, that is, 100 trees, the maximum tree depth is 5, the minimum number of sample splits per tree is 2, and the minimum number of leaf node samples per tree is 1. S22. Input the standardized data set into the first layer of the deep forest model, perform preliminary feature extraction on the input data through the decision tree of the first layer, and generate the output of the first layer; S23, passing the output data of the first layer to the second layer of the deep forest model, and using the decision tree of the second layer to further process and learn the patterns and nonlinear relationships in the output data to generate the output of the second layer; S24, repeat S23, pass the output data to the next layer layer by layer, and gradually extract the features and nonlinear relationships in the output data through the combination of each layer of trees until the final layer; S25. Output the preliminary output of the deep forest model.

6. The data visualization system based on big data analysis according to claim 2, characterized in that: The S3 includes the following specific steps: S31. Initialize the pigeon swarm optimization algorithm based on the preliminary output of the deep forest model, set each pigeon as a candidate solution, and the candidate solution is the hyperparameter configuration of the deep forest model; S32, configuring a position exploration unit for each pigeon body to simulate the position of the pigeon body in the search space, where the position of the pigeon body in the search space is the current value of each hyperparameter in the deep forest model; S33. Based on the pigeon swarm algorithm, the pigeon swarm optimization algorithm dynamically adjusts the weight of each pigeon according to the initial output of the deep forest model through a dynamic adaptive adjustment mechanism: in, The new position of the i-th pigeon after the current position is updated, P i represents the current position information of the i-th pigeon at the current position, ω represents the pigeon inertia weight, α represents the adjustment factor that controls the pigeon's exploration range, represents the distance between the i-th pigeon and the j-th pigeon, β represents the factor that controls the speed of convergence between the pigeon and the global optimal solution, P j The position of the j-th pigeon, N represents the total number of pigeons in the pigeon group, i represents the i-th pigeon in the pigeon group, j represents the reference pigeon of the i-th pigeon, P global represents the location of the global optimal solution, P j Indicates the position of the j-th pigeon; S34. Calculate the fitness value of each pigeon based on its current position and the performance evaluation of the deep forest model. Adopt an adaptive adjustment mechanism to weight the fitness and design a multi-level filtering mechanism to sort and screen the fitness of the pigeons, retaining the best performing pigeons. Further optimize the hyperparameter configuration through multiple rounds of refined screening. S35. By simulating the individual behavior of the pigeon flock and continuously adjusting the hyperparameter configuration, the pigeons adjust themselves according to their own performance in each round of iteration.

7. The data visualization system based on big data analysis according to claim 2, characterized in that: The S4 includes the following specific steps: S41, receiving the hyperparameter configuration optimized by the pigeon flock optimization algorithm and initializing the big data feature space; S42. Set the initial position of the pigeon group according to the hyperparameter configuration. The initial position of each pigeon represents a hyperparameter combination in the deep forest model, and set the initial exploration range of the pigeon. S43. Calculate the performance index of the deep forest model corresponding to the current pigeon body based on the pigeon body position and hyperparameter configuration: Among them, E i represents the fitness value of the i-th pigeon, i represents the i-th pigeon in the pigeon group, The new position of the i-th pigeon after the current position is updated, Indicates the best performing pigeon body position in history, represents the distance between the i-th pigeon body and the best pigeon body in history, δ represents the attenuation factor, γ represents the adjustment factor, error i represents the error value of the deep forest model obtained by the hyperparameter configuration corresponding to the i-th pigeon, ∈ represents the tolerance parameter, η represents the factor of weighted error, N represents the total number of pigeons in the pigeon group, and k represents the pigeon with the best performance in history; S44. Introducing an adaptive gradient weighting mechanism during the pigeon group update process, by dynamically adjusting the search step size of each pigeon, so that the search step size of each pigeon is not only based on fitness, but also combined with the relative position in the feature space. When the distance between the pigeon and the global optimal solution is less than or equal to 0.1, the search step size is reduced. When the distance between the pigeon and the global optimal solution is greater than 0.1, the search step size is increased. S45. Introducing a multi-objective adaptive optimization strategy, through which multi-dimensional fitness calculations are performed during the pigeon body update process, balancing the accuracy of the deep forest model and computing resources; S46, through the feedback loop of S44 and S45, adjust the exploration range and direction of the pigeon body in real time, and eventually make the pigeon body converge toward the optimized hyperparameter solution in the feature space and output the optimal hyperparameter configuration.

8. The data visualization system based on big data analysis according to claim 2, characterized in that: The S5 includes the following specific steps: S51. Initialize the deep forest model based on the optimal hyperparameter configuration and apply it to large data sets for classification, regression, and predictive analysis. At the same time, introduce an adaptive training set selection mechanism to automatically adjust the ratio of training and test sets based on data distribution, allowing the model to dynamically adjust the training and test splits based on different data characteristics. S52. Dynamically expand the training set during the training process, gradually introduce new data samples during the training of the deep forest model, and continuously adjust the learning weights of the deep forest model; S53, applying the trained deep forest model to the test set, predicting the test set, and optimizing the deep forest model in real time based on the prediction results; S54. Based on the prediction results, calculate the performance indicators of the model, including accuracy, error rate, and mean square error, and weight different performance indicators according to the importance of each type of data to generate analysis results; S55. Organize the analysis results into a data structure to form the data format required for visualization, including classification results, regression values, and performance evaluation indicators; S56. Through the big data visualization unit, the analysis results are presented to users in a graphical and interactive manner, supporting multiple visualization forms and providing dynamic analysis functions.

9. The data visualization system based on big data analysis according to claim 2, characterized in that: The S6 includes the following specific steps: S61, receiving the interaction requirements input by the user, and determining the data content and visualization method to be displayed based on the specific parameter configuration selected by the user; S62. Reload the corresponding data set according to the parameter configuration specified by the user, and perform data preprocessing to adapt to the new parameter configuration; S63. Recalculate the analysis results of the deep forest model based on the new data set and parameter configuration: Among them, R i represents the analysis result of the hyperparameter configuration corresponding to the i-th pigeon body, j represents the reference pigeon body of the i-th pigeon body, Indicates the new position of the i-th pigeon after the current position is updated. Indicates the position of the best pigeon in history, represents the distance between pigeon body i and pigeon body j, α' represents the distance weighting factor, β' represents the error weighting factor, error i represents the error value of the deep forest model obtained by the hyperparameter configuration corresponding to the i-th pigeon, γ represents the adjustment factor, N represents the total number of pigeons in the pigeon group, i represents the i-th pigeon in the pigeon group, ∈ represents the tolerance parameter, and δ represents the error square weighting factor; S64. After the real-time calculation is completed, the visualization graphics, including heat maps, scatter plots, and time series graphs, are updated through the big data visualization unit, and the display method of the graphics is adjusted according to the dynamic changes of the new data set; S65. Based on the user's operations on the visualization interface, including filtering, zooming in and out, the new analysis results are updated and presented in real time, and the display content is adjusted; S66. Provide user-defined view function, allowing users to select and save visualization results under different parameter configurations.

10. The data visualization system based on big data analysis according to claim 2, characterized in that: The S7 includes the following specific steps: S71. Based on the new analysis results, identify and extract the key factors that affect performance and business decisions; S72. Based on new analysis results and key factors, use forecast data and historical data to conduct multi-dimensional analysis to identify potential optimization opportunities, business growth points, and potential risks. S73. Generate optimization suggestions and formulate specific action plans based on the identified potential optimization opportunities, business growth points, and potential risks; S74. Apply optimization suggestions to real-time monitoring and tracking of optimization results, and make dynamic adjustments based on real-time data feedback; S75. Support senior decision makers' strategic decisions through analysis results and optimization suggestions, and provide comprehensive business data reports and actionable decision-making basis; S76. Conduct risk assessment and control based on business and operational conditions, predict potential risks based on analysis results, propose preventive measures and implement risk control strategies.