Visual big data processing equipment
By building visual big data processing equipment and combining hardware and software, we solved the unified management and compatibility issues of big data components, achieved efficient and stable big data processing services, and reduced costs.
Patent Information
- Application Number
- CN202510709660.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-09
AI Technical Summary
Existing big data components lack a unified management and scheduling mechanism, resources are scattered, the construction process is complex, and there are many compatibility issues, resulting in low processing efficiency and high costs.
Provides a visual big data processing device that combines hardware and software, including disk arrays, memory bars, central processing units, network interfaces, operating systems, as well as user interaction interfaces, My SQL databases, data processing frameworks, Kafka distributed caches, data visualization tools, etc., to build a unified management and analysis platform that supports both stream-batch and lake-warehouse integrated data management and has security and backup mechanisms.
It realizes the centralized management of big data components, simplifies the server construction process, improves processing efficiency and stability, reduces implementation costs, and provides highly reliable and high-performance services.
Smart Images

Figure CN120610992A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of big data processing technology, and specifically refers to a visual big data processing device. Background Art
[0002] With the rapid development of information technology, big data has become a vital resource for business decision-making, scientific research, and everyday life. However, big data processing and analysis face numerous challenges. First, big data components are numerous, including but not limited to Hadoop, Spark, and Kafka. These components operate independently and lack a unified management and scheduling mechanism, resulting in fragmented resources and difficulty in efficient utilization. Second, the process of setting up a big data server is complex and varied, involving multiple steps such as hardware selection, software installation, configuration and debugging. This presents a high level of technical difficulty and places high demands on operations and maintenance personnel. Furthermore, compatibility issues may exist between different big data components, further increasing the difficulty and cost of big data applications. Therefore, the market urgently needs a technical solution that can address these issues, enable centralized management of big data components, simplify the server setup process, and improve the efficiency and stability of big data processing.
[0003] Currently, there are several technical solutions for big data processing on the market, but most of these solutions can only solve partial problems and cannot fully meet the needs of big data processing. For example, some solutions provide management tools for big data components, but these tools are often limited to managing specific components and cannot achieve unified management across components. Other solutions provide big data server setup services, but these services generally only provide standardized setup processes and cannot be flexibly adjusted to meet users' actual needs. In addition, some solutions attempt to solve the compatibility issues of big data components through virtualization technology, but these solutions often require extensive customization and development work at the hardware and software levels, resulting in high implementation costs and difficulty in ensuring system stability and performance. Therefore, existing technical solutions have obvious limitations in solving the problems faced by big data processing. Summary of the Invention
[0004] In order to solve the above-mentioned existing problems, the present invention provides a visual big data processing device that serves as an intermediate device for small companies, Internet of Things companies or various Internet of Things data collection devices, and mainly performs data collection, processing and analysis.
[0005] The technical solution adopted by the present invention is as follows: The visualization big data processing device of the present invention includes hardware equipment and a big data processing platform. The hardware part includes a disk array, a memory bar, a central processing unit, a network interface and an operating system. The big data processing platform includes a user interaction interface, a My SQL database, a data processing framework, a Kafka distributed cache, data visualization, a data lake, a data middle platform, a data API, data reports, data analysis, a data warehouse and data mining.
[0006] Furthermore, the user interaction interface design provides functions such as big data component management, task scheduling, and data monitoring, which facilitates users to perform daily management and maintenance of the big data environment.
[0007] The data processing framework uses Flink data collection and Spark in-memory computing. The Spark in-memory computing is used for large-scale data processing and analysis, providing a fast and versatile data processing engine.
[0008] The Kafka distributed cache is used to build real-time data pipelines and stream applications, and supports a high-throughput distributed publish-subscribe messaging system.
[0009] The tools used for data visualization and data analysis include Tableau data analysis software, Power BI data analysis software, and Apache Superset open source data exploration and visualization platform, which facilitate users to intuitively analyze and display big data.
[0010] The data lake adopts the Huid data lake platform and HDFS file system, and the data middle platform adopts the JavaWeb-Spring Boot framework.
[0011] The data warehouse uses the Hive data warehouse tool and the Hadoop system framework. The Hadoop system framework provides distributed storage and processing capabilities and supports efficient management of large-scale data sets.
[0012] Furthermore, the user interaction interface, My SQL database, data processing framework and data API constitute a stream-batch integrated data management architecture, and the data visualization, data lake, data middle platform, data reports, data analysis, data warehouse and data mining constitute a lake-warehouse integrated data management architecture.
[0013] Furthermore, the big data processing platform also includes a Post Gre SQL database management system, an Hbase database, a Flume massive log collection system, and a Zookeeper application coordination service. The Zookeeper application coordination service provides consistency services for distributed applications, such as configuration management, naming services, and distributed synchronization.
[0014] Furthermore, the big data processing platform also includes a security and backup mechanism, which has a built-in firewall and intrusion detection system to ensure the security of the big data environment; at the same time, it provides regular automatic backup and disaster recovery functions to ensure the security and integrity of the data.
[0015] Furthermore, the disk array consists of four 4TB hard drives, the memory stick has a capacity of 32G, the central processing unit model is Intel Xeon E-21243.3GHz, the network interface data transmission rate is Gigabit Ethernet, and the operating system is Ubuntu Server 18.04LTS.
[0016] The beneficial effects achieved by the present invention using the above structure are as follows:
[0017] 1. Provide a big data solution by combining hardware and software.
[0018] 2. Provide enterprises with highly reliable and high-performance big data services in a simple and efficient manner.
[0019] 3. Provide a simple, efficient and low-cost big data server. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 Schematic diagram of the big data processing platform structure of the visual big data processing equipment proposed in this solution. DETAILED DESCRIPTION
[0021] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0022] like Figure 1 As shown in the figure, the visualization big data processing device proposed in this solution includes hardware equipment and a big data processing platform. The hardware part includes a disk array, a memory bar, a central processing unit, a network interface and an operating system. The big data processing platform includes a user interaction interface, a My SQL database, a data processing framework, a Kafka distributed cache, data visualization, a data lake, a data middle platform, a data API, data reports, data analysis, a data warehouse and data mining.
[0023] The user interaction interface design provides functions such as big data component management, task scheduling, and data monitoring, which facilitates users to perform daily management and maintenance of the big data environment.
[0024] The data processing framework uses Flink data collection and Spark in-memory computing. The Spark in-memory computing is used for large-scale data processing and analysis, providing a fast and versatile data processing engine.
[0025] The Kafka distributed cache is used to build real-time data pipelines and stream applications, and supports a high-throughput distributed publish-subscribe messaging system.
[0026] The tools used for data visualization and data analysis include Tableau data analysis software, Power BI data analysis software, and Apache Superset open source data exploration and visualization platform, which facilitate users to intuitively analyze and display big data.
[0027] The data lake adopts the Huid data lake platform and HDFS file system, and the data middle platform adopts the JavaWeb-Spring Boot framework.
[0028] The data warehouse uses the Hive data warehouse tool and the Hadoop system framework. The Hadoop system framework provides distributed storage and processing capabilities and supports efficient management of large-scale data sets.
[0029] The user interaction interface, My SQL database, data processing framework and data API constitute a stream-batch integrated data management architecture, and the data visualization, data lake, data middle platform, data reports, data analysis, data warehouse and data mining constitute a lake-warehouse integrated data management architecture.
[0030] The big data processing platform also includes a PostGre SQL database management system, an Hbase database, a Flume massive log collection system, and a Zookeeper application coordination service. The Zookeeper application coordination service provides consistency services for distributed applications, such as configuration management, naming services, distributed synchronization, etc. The big data processing platform also includes a security and backup mechanism. The security and backup mechanism has a built-in firewall and intrusion detection system to ensure the security of the big data environment; at the same time, it provides regular automatic backup and disaster recovery functions to ensure the security and integrity of the data.
[0031] Among them, the disk array consists of 4 4TB hard drives, the memory stick capacity is 32G, the central processing unit model is Intel Xeon E-21243.3GHz, the network interface data transmission rate is Gigabit Ethernet, and the operating system is Ubuntu Server 18.04LTS.
[0032] During specific use, the user interaction interface design provides functions such as big data component management, task scheduling, and data monitoring, which facilitate users to carry out daily management and maintenance of the big data environment. Spark in-memory computing is used for large-scale data processing and analysis, providing a fast and general data processing engine. Kafka distributed cache is used to build real-time data pipelines and streaming applications, supporting high-throughput distributed publish-subscribe messaging systems. The tools used for data visualization and data analysis include Tableau data analysis software, PowerBI data analysis software, and Apache Superset open source data exploration and visualization platform, which facilitate users to intuitively analyze and display big data. The data lake uses the Huid data lake platform and HDFS file system, the data middle platform uses the JavaWeb-Spring Boot framework, and the data warehouse uses the Hive data warehouse tool and the Hadoop system framework. The Hadoop system framework provides distributed storage and processing capabilities, and supports the efficient management of large-scale data sets.
[0033] It provides a unified user interface and management tools, enabling centralized management and configuration of big data components. Users can install, configure, and monitor big data components through a simple graphical interface, eliminating the need for in-depth knowledge of each component. Furthermore, the big data processing equipment offers automated deployment and expansion capabilities, automatically adjusting resource allocation based on user needs to ensure efficient and stable big data processing. At the hardware level, the big data processing equipment utilizes high-performance disk arrays and memory configurations, as well as stable network interfaces and operating systems, providing a solid hardware foundation for big data processing.
[0034] The big data processing equipment supports multiple access methods, including but not limited to LAN access, internet remote access, and VPN access, to meet the needs of different users in different scenarios. LAN access is suitable for enterprise internal network environments, allowing users to directly access the services provided by the big data box through the internal network. Internet remote access allows users to remotely manage the big data box via the internet, facilitating big data processing and analysis from different locations. VPN access provides a secure and reliable remote access method. By establishing an encrypted VPN tunnel, users can securely access the big data processing equipment and protect the security of data transmission.
[0035] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0036] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. Visual big data processing equipment, characterized by: It includes hardware equipment and big data processing platform. The hardware part includes disk array, memory bar, central processing unit, network interface and operating system. The big data processing platform includes user interaction interface, My SQL database, data processing framework, Kafka distributed cache, data visualization, data lake, data middle platform, data API, data report, data analysis, data warehouse and data mining.
2. The visualization big data processing device according to claim 1, characterized in that: The user interface design provides big data component management, task scheduling, and data monitoring functions, making it easier for users to perform daily management and maintenance of the big data environment; The data processing framework uses Flink data collection and Spark in-memory computing. Spark in-memory computing is used for large-scale data processing and analysis, providing a fast and versatile data processing engine. The Kafka distributed cache is used to build real-time data pipelines and streaming applications, supporting high-throughput distributed publish-subscribe messaging systems; The data visualization and data analysis tools used include Tableau data analysis software, Power BI data analysis software, and Apache Superset open source data exploration and visualization platform, which facilitate users to intuitively analyze and display big data; The data lake uses the Huid data lake platform and HDFS file system, and the data middle platform uses the JavaWeb-SpringBoot framework; The data warehouse uses the Hive data warehouse tool and the Hadoop system framework. The Hadoop system framework provides distributed storage and processing capabilities and supports efficient management of large-scale data sets.
3. The visualization big data processing device according to claim 1, characterized in that: The user interaction interface, MySQL database, data processing framework and data API constitute a stream-batch integrated data management architecture, and the data visualization, data lake, data middle platform, data reports, data analysis, data warehouse and data mining constitute a lake-warehouse integrated data management architecture.
4. The visualization big data processing device according to claim 1, characterized in that: The big data processing platform also includes a PostGre SQL database management system, an Hbase database, a Flume massive log collection system, and a Zookeeper application coordination service. The Zookeeper application coordination service provides consistency services for distributed applications.
5. The visualization big data processing device according to claim 1, characterized in that: The big data processing platform also includes a security and backup mechanism with built-in firewalls and intrusion detection systems to ensure the security of the big data environment; it also provides regular automatic backup and disaster recovery functions to ensure the security and integrity of the data.
6. The visual big data processing device according to claim 1, characterized in that: The disk array consists of four 4TB hard drives, the memory stick has a capacity of 32G, the central processing unit model is Intel Xeon E-2124 3.3GHz, the network interface data transmission rate is Gigabit Ethernet, and the operating system is Ubuntu Server 18.04LTS.