Big data center based on database structure and construction method thereof

Through a big data center based on database structure, distributed storage architecture and data cleaning and encryption technology are used to solve the scalability and security problems of traditional centralized architectures, and efficient data management and security are achieved.

CN120029994APending Publication Date: 2025-05-23HUAINAN UNITED UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510016499.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Big data centers with traditional centralized architectures have problems such as poor scalability, high risk of single point failure, complex data management, and difficult to ensure data consistency and security.

Method used

It adopts a big data center based on database structure, adopts distributed storage architecture, data layered storage, data cleaning and standardized processing, machine learning and deep learning algorithm analysis, encryption technology and access control, and combines regular backup strategies to ensure the security and availability of data.

Benefits of technology

It improves the stability and availability of data storage, reduces the additional overhead of data processing, shortens the data analysis cycle, ensures the confidentiality and integrity of data, and reduces the risk of single point of failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029994A_ABST
    Figure CN120029994A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of big data centers, in particular to a big data center based on a database structure and a construction method thereof. According to the technical scheme, the system comprises a data storage module, a data processing module, a data management module and a data security module. Through redundancy design, the risk of data loss caused by hardware faults is reduced, the data are divided into different logic modules according to functions through the logic layer, all parts of data are clear in order and convenient to manage, the stability of data storage and use is further guaranteed, meanwhile, the data are subjected to standardized processing, and the data storage and use efficiency is improved. Extra processing expenses caused by inconsistent and nonstandard data formats in the subsequent mining and analysis process are reduced, the period of data from collection to value output is shortened by utilizing machine learning and deep learning algorithms, and support is quickly provided for services; in addition, the data is encrypted through an encryption technology, so that the confidentiality of the data is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data centers, and in particular to a big data center based on a database structure and a construction method thereof. Background Art

[0002] With the rapid development of information technology and the acceleration of digitalization, the data generated by various industries has exploded, and data has become an important asset for enterprises and organizations. In this context, big data centers have emerged to efficiently store, process and manage massive and diverse data to meet the needs of enterprises for data analysis, decision support, business optimization, etc.

[0003] Traditional data storage methods often use centralized architectures. When faced with large-scale data, this architecture gradually exposes problems such as poor scalability and high risk of single point failures. In addition, data integration, classification, storage, and retrieval become extremely complex. Data in different formats and with different semantics are difficult to manage uniformly, which easily leads to data inconsistency and data redundancy, affecting data quality and the accuracy of analysis results. Therefore, we propose a big data center based on a database structure and its construction method. Summary of the invention

[0004] The purpose of the present invention is to propose a big data center based on a database structure and a construction method thereof in view of the shortcomings of the centralized architecture existing in the background technology, such as poor scalability and high risk of single point failure, as well as the problem of complex data management.

[0005] First aspect: The present application provides a big data center based on a database structure, including a data storage module, a data processing module, a data management module and a data security module:

[0006] The data storage module is used to store data. The data storage module is based on a distributed storage architecture and has a data hierarchical storage function, including a physical layer, a logical layer and an application layer;

[0007] The data processing module includes a data preprocessing unit and a data mining and analysis unit, wherein the data preprocessing unit is used to perform data cleaning, data type standardization and data integration on the data, and the data mining and analysis unit is used to process and analyze the data;

[0008] The data management module is used to manage the metadata of the data to ensure the consistency and integrity of the data;

[0009] The data security module ensures data security through encryption technology, access control technology and data backup and recovery mechanism.

[0010] Optionally, the physical layer uses disk array technology to ensure the reliability of data storage. The disk array technology includes redundant disk array technology. The redundant disk array technology performs striping processing on multiple disks in the disk array and stores data in a dispersed manner on multiple disks. The redundant disk array algorithm formula is as follows:

[0011]

[0012] Among them, d i Represents the data of the i-th disk in the disk array, where n is the number of disks in the disk array;

[0013] The logic layer divides data into different logic modules according to functions, including user data, business data, and system configuration data. Each logic module has independent functions and interfaces. Data interaction is achieved between modules through data interfaces, and the data interfaces comply with communication protocols.

[0014] The application layer is used to provide data access and processing interfaces for users.

[0015] Optionally, the data cleaning removes noise, impurities and outliers in the data through regular expressions, and the formula is as follows:

[0016]

[0017] Among them, α i is the character set in the regular expression, β i The number of times the corresponding character is repeated;

[0018] The data type standardization includes the unification of data formats and a data dictionary. The unified data format is used to ensure the consistency of data, and the data dictionary is used to define data elements.

[0019] Optionally, the data format includes a date type, a numeric type, and a string type. For the date type, the ISO standard format is adopted. The algorithm formula of the date format is as follows:

[0020] DATE=YYYY-MM-dd

[0021] Among them, YYYY represents year, MM represents month, and dd represents day;

[0022] For the numeric type, define the value range and precision of the data type, where the value range of the integer type is [-2 31 ,2 31 -1], the precision of floating point types is ±10 -7 ;

[0023] For the string type, the length of the string shall not exceed 255 characters, and the character set is ASCII encoding;

[0024] The data dictionary is used to define and explain data elements. The data elements include data name, data type, value range, and description. The structure of the data dictionary is as follows:

[0025]

[0026] Optionally, the data mining and analysis unit uses a machine learning algorithm and a deep learning algorithm to process and analyze data. The machine learning algorithm uses a decision tree algorithm. The decision tree algorithm classifies and predicts data by constructing a tree structure. The decision tree algorithm formula is as follows:

[0027]

[0028] Among them, Entropy(S) is the information entropy of the data set S, c is the number of category labels in the data set S, and p i is the probability of category i appearing in data set S, Gain(S,A) is the information gain of feature A for data set S, Values(A) is the set of all possible values ​​of feature A, S v It is a subset of feature A in the data set S with a value of v. By calculating the information gain of each feature, the feature with the largest information gain is selected as the splitting node of the decision tree, so as to construct a decision tree to classify and predict the data;

[0029] The deep learning algorithm adopts a neural network model, and the neural network model formula is as follows:

[0030] z l+1 =W l+1 a l +b l+1

[0031] a l+1 =f(z l+1 )

[0032] Among them, l represents the number of layers of the neural network, z l+1 is the weighted sum of the inputs to layer l+1, including the bias b l+1 , W l+1 is the weight matrix of the l+1th layer, a l is the output of the lth layer, f(·) is the activation function, including the Sigmoid function, through the weighted and connected activation operations of multiple layers of neurons, the neural network model learns and analyzes the data, and adjusts the weight W and bias b to fit the characteristics and patterns of the data.

[0033] Optionally, the data management module is used to manage metadata of the data to ensure consistency and integrity of the data. The metadata includes data name, data type, value range, and description. The data management module includes index and directory management of the data.

[0034] Optionally, the encryption technology uses an encryption algorithm formula to encrypt data, and the encryption algorithm formula is as follows:

[0035] E=E key (P)

[0036] Among them, E is the encrypted ciphertext, E key is the encryption key, and P is the original data;

[0037] The access control technology manages user rights and limits user access to data through user identity authentication and authorization mechanisms;

[0038] The data backup and recovery mechanism adopts a regular backup strategy, and the storage format of the backup data is JSON format.

[0039] Second aspect: This application provides a method for constructing a big data center based on a database structure, comprising the following steps:

[0040] Optional: Hardware facilities construction: The hardware facilities include servers, storage devices and network devices, and the hardware environment is built;

[0041] Software system development: The software system includes an operating system, a database management system and an application program. The operating system is Linux. The database management system includes database table structure, index and storage strategy.

[0042] Data migration and integration: collect, clean, convert and integrate data. The data sources of the collected data include business systems, sensors, and networks. The data cleaning includes removing noise, impurities and outliers from the data. The data conversion converts the data into a format suitable for storage and analysis. The integrated data should be stored in the data storage module.

[0043] System testing and optimization: Perform functional, performance and security tests on the system, and make optimization adjustments based on the test results. The functional test is used to ensure the normal operation of the system functions. The performance test evaluates the response time, throughput and resource utilization of the system. The security test is used to check the security and vulnerabilities of the system. Based on the test results, the system is optimized to improve the performance and stability of the system.

[0044] System deployment and maintenance: Deploy the system to the production environment and perform routine maintenance and data updates. The deployment process includes installing the operating system, database management system and applications, configuring the server and network environment. Routine maintenance includes monitoring the system operation status, handling faults and abnormal situations, and data updates to ensure data accuracy and timeliness.

[0045] Optionally, the testing methods adopted by the security test include vulnerability scanning, penetration testing and security auditing. The scope of the security test covers the system's network, operating system, database, and application program, and the security vulnerabilities and risks found are repaired and reinforced.

[0046] In summary, the present application includes at least one of the following beneficial technical effects:

[0047] The present invention adopts a distributed storage architecture and has a data layered storage function. The physical layer stores data in multiple disks in a dispersed manner. Through redundant design, even if some disks fail, the redundant disk array algorithm can be used to restore data, reducing the risk of data loss due to hardware failure. The data is divided into different logical modules according to function through the logical layer, and interaction is achieved through the data interface, so that each part of the data is clear and easy to manage, further ensuring the stability of data storage and use; at the same time, the application layer provides users with a unified data access and processing interface, which is convenient for users to easily obtain the required data, and enhances the availability of data from the user's perspective;

[0048] The present invention removes noise, impurities and outliers in the data through a data preprocessing unit, and performs data type standardization and data integration at the same time, integrating data from different sources, thereby achieving standardized processing of the data, reducing the additional processing overhead caused by inconsistent and non-standard data formats in the subsequent mining and analysis process, and using machine learning and deep learning algorithms to complete the processing and analysis tasks of massive data within a reasonable time, shortening the cycle from data collection to output value, and quickly providing support for business;

[0049] The data security module in the present invention uses encryption technology to encrypt data, so that the data exists in ciphertext form during storage and transmission, which ensures the confidentiality of the data and prevents the leakage of sensitive information. It cooperates with access control technology based on user identity authentication and authorization mechanism to limit user access rights to data, distinguish the data scope accessible to different user roles, eliminate illegal access from the source, and reduce the risk of data being maliciously tampered with or improperly used; at the same time, a regular backup strategy is adopted to store backup data to ensure data security and business continuity, and minimize losses. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1A structural block diagram of a big data center based on a database structure is given in the present invention. DETAILED DESCRIPTION

[0051] The technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments.

[0052] The components of the embodiments of the present invention generally described and shown in the drawings herein may be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention.

[0053] Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in the field without making any creative work shall fall within the scope of protection of the present invention.

[0054] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.

[0055] It should be noted that the terms "include", "comprises" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0056] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0057] Example

[0058] On the one hand, the present application provides a big data center based on a database structure, such as Figure 1 As shown, it includes a data storage module, a data processing module, a data management module and a data security module;

[0059] The data storage module is used to store data. The data storage module is based on a distributed storage architecture and has a data layered storage function, including a physical layer, a logical layer, and an application layer;

[0060] Among them, the physical layer uses disk array technology to ensure the reliability of data storage. Disk array technology includes redundant disk array technology. Redundant disk array technology stripes multiple disks in the disk array and stores data on multiple disks in a dispersed manner. The redundant disk array algorithm formula is as follows:

[0061]

[0062] Among them, d i Represents the data of the i-th disk in the disk array, and n is the number of disks in the disk array. Through the redundant design, even if some disks fail, the redundant disk array algorithm can be used to recover the data, reducing the risk of data loss due to hardware failure;

[0063] In addition, the logic layer divides the data into different logic modules according to function, including user data, business data, and system configuration data. Each logic module has independent functions and interfaces. Data interaction is achieved between modules through data interfaces. The data interfaces follow the communication protocol. The logic layer divides the data into different logic modules according to function, and interaction is achieved through data interfaces, so that each part of the data is clearly organized and easy to manage, further ensuring the stability of data storage and use; the application layer is used to provide users with data access and processing interfaces, so that users can easily obtain the required data, thereby enhancing the availability of data from the user's perspective.

[0064] The data processing module includes a data preprocessing unit and a data mining and analysis unit. The data preprocessing unit is used to clean the data, standardize the data type, and integrate the data, thereby achieving standardized processing of the data and reducing the additional processing overhead caused by inconsistent and non-standard data formats in the subsequent mining and analysis process. The data mining and analysis unit is used to process and analyze the data. It uses machine learning and deep learning algorithms to complete the processing and analysis tasks of massive data within a reasonable time, shorten the cycle from data collection to output value, and quickly provide support for the business;

[0065] Among them, data cleaning uses regular expressions to remove noise, impurities and outliers in the data. The formula is as follows:

[0066]

[0067] Among them, α i is the character set in the regular expression, β i The number of times the corresponding character is repeated;

[0068] In addition, data type standardization includes the unification of data formats and data dictionaries. The unified data format is used to ensure data consistency. The data dictionary is used to define data elements. The data formats include date type, numeric type, and string type. For the date type, the ISO standard format is used. The algorithm formula for the date format is as follows:

[0069] DATE=YYYY-MM-dd

[0070] Among them, YYYY represents year, MM represents month, and dd represents day;

[0071] For numeric types, define the value range and precision of the data type. The value range of integer types is [-2 31 ,2 31 -1], the precision of floating point types is ±10 -7 For string types, the string length must not exceed 255 characters, and the character set is ASCII encoding;

[0072] Finally, the data dictionary is used to define and explain data elements. Data elements include data name, data type, value range, and description. The structure of the data dictionary is:

[0073]

[0074] The data mining and analysis unit uses machine learning algorithms and deep learning algorithms to process and analyze data. The machine learning algorithm uses a decision tree algorithm. The decision tree algorithm classifies and predicts data by building a tree structure. The decision tree algorithm formula is as follows:

[0075]

[0076] Among them, Entropy(S) is the information entropy of the data set S, c is the number of category labels in the data set S, and p i is the probability of category i appearing in data set S, Gain(S,A) is the information gain of feature A for data set S, Values(A) is the set of all possible values ​​of feature A, S v It is a subset of feature A in the data set S with a value of v. By calculating the information gain of each feature, the feature with the largest information gain is selected as the splitting node of the decision tree, so as to construct a decision tree to classify and predict the data;

[0077] The deep learning algorithm uses a neural network model. The neural network model formula is as follows:

[0078] z l+1 =W l+1 a l +b l+1

[0079] a l+1 =f(z l+1 )

[0080] Among them, l represents the number of layers of the neural network, z l+1 is the weighted sum of the inputs to layer l+1, including the bias b l+1 , W l+1 is the weight matrix of the l+1th layer, a l is the output of the lth layer, f(·) is the activation function, including the Sigma ID function. Through the weighted sum of multiple layers of neurons and the connection with the activation operation, the neural network model learns and analyzes the data, and adjusts the weight W and bias b to fit the characteristics and patterns of the data.

[0081] The data management module is used to manage the metadata of the data to ensure the consistency and integrity of the data. The metadata includes the data name, data type, value range, description, and the data management module includes data index and directory management.

[0082] The data security module ensures data security through encryption technology, access control technology, and data backup and recovery mechanism. The encryption technology uses encryption algorithm formula to encrypt data. The encryption algorithm formula is as follows:

[0083] E=E key (P)

[0084] Among them, E is the encrypted ciphertext, E key is the encryption key, P is the original data, and encryption technology is used to encrypt the data so that the data exists in ciphertext during storage and transmission, ensuring the confidentiality of the data and preventing the leakage of sensitive information;

[0085] Access control technology manages user rights. Through user identity authentication and authorization mechanisms, it limits user access to data, distinguishes the data scope accessible to different user roles, and eliminates illegal access at the source.

[0086] The data backup and recovery mechanism adopts a regular backup strategy. The storage format of the backup data is JSON format. A regular backup strategy is used to store backup data to ensure data security and business continuity and minimize losses.

[0087] On the other hand, the present application provides a method for constructing a big data center based on a database structure, comprising the following steps:

[0088] (I) Hardware facilities construction

[0089] Server selection and configuration: Select servers, each server is equipped with 2 multi-core processors, 256GB memory, and the server uses redundant power supply modules. When one power supply fails, the other power supply can continue to power the server to ensure the continuous operation of the server and avoid data processing interruptions caused by power failures; and is equipped with a high-speed network interface card that supports a 10Gbps network transmission rate to meet the needs of high-speed data transmission within the big data center and reduce data transmission delays.

[0090] Storage device deployment: A storage system based on a distributed storage architecture is used, preferably a Ceph storage cluster; at the physical layer, redundant disk array technology is used. The disk array consists of five 4TB enterprise-level hard disks. Through striping, data is evenly distributed on each disk to improve data storage reliability and read / write performance. For a certain financial transaction data file, its data blocks are stored on these five disks according to the RAID 5 algorithm, and parity information is calculated and stored on one of the disks. When a disk fails, the data on the failed disk can be restored using the data and parity information on the remaining disks to ensure data integrity and availability;

[0091] In addition, the total capacity of the storage system is planned to be 100TB, and storage nodes can be easily expanded to increase storage capacity according to data growth.

[0092] Network equipment construction: The core switch adopts a three-layer switch with a backplane bandwidth of 100Gbps and low-latency forwarding capability to ensure smooth data exchange between a large number of servers and storage devices within the big data center; and adopts a redundant link design to establish multiple physical links between the core switch and the server and between the core switches, and bundle the links into logical links through link aggregation technology, increasing the link bandwidth while providing link redundancy backup. When a link fails, the data traffic can automatically switch to other normal links to ensure uninterrupted operation of the network.

[0093] (II) Software system development

[0094] Operating system customization: Customized development based on the Linux operating system, optimized kernel parameters, adjusted file system cache parameters to improve the read and write performance of large data files, and optimized network parameters, including adjusting the TCP window size and maximum number of connections, to meet the high-concurrency data transmission requirements of large data centers;

[0095] Install and configure relevant system management tools, including the monitoring tool Zabbix, which is used to monitor key performance indicators such as CPU, memory, disk I / O, and network traffic of the server in real time;

[0096] Database management system configuration: Select the open-source database management system MySQL. When creating the database table structure, design it reasonably according to the characteristics of financial data and business requirements. Adopt the partition table technology to store transaction data by time range (one partition per month); Configure the storage engine of the database as InnoDB, and utilize its transaction support and row-level locking mechanism to ensure data consistency and integrity;

[0097] Application development: Develop a data collection program. Use the Java language combined with the open-source Sqoop tool to collect data from business systems, sensors, and the network. Among them, business systems include the core transaction system and the customer relationship management system. Sensor data sources include market quotation data collection sensors, and network data sources are the API interfaces of third-party financial data providers. The collection program collects data at a frequency of once every 15 minutes, and conducts preliminary verification and filtering on the collected data to ensure the legality and integrity of the data;

[0098] (III) Data migration and integration

[0099] Data collection: According to the predetermined collection plan, collect transaction data for the past year from the core transaction system, including transaction records, account information, and product information, and extract this data to the temporary storage area of the big data center through the Sqoop tool; At the same time, collect market data from the API interface of the market quotation data provider. Market data includes real-time stock prices, exchange rates, and bond yields, and transmit this data to the Kafka message queue in the big data center in real time through the Flume tool for subsequent real-time processing and analysis;

[0100] Extract data such as customer basic information, credit records, and investment preferences from the customer relationship management system, and use ETL tools to clean, transform, and load this data, and integrate it into the customer information table in the relational database of the big data center to ensure data consistency and integrity.

[0101] Data cleaning and transformation: Use the data preprocessing module to clean the collected transaction data, removing duplicate records and invalid data. Among them, duplicate records are removed by checking the transaction serial number, and invalid data refers to data with a transaction amount of 0 or negative;

[0102] For date data, convert it from various original formats to ISO standard formats. For numeric data, adjust the data precision and handle abnormal values. Unify the precision of transaction amounts to 6 decimal places. Mark and review amounts that are obviously beyond the normal transaction range to confirm whether they are data errors or abnormal transactions. Data greater than 100 million yuan is judged as abnormal data based on business experience.

[0103] Convert the cleaned data to make it suitable for storage and analysis, perform word segmentation on the text description fields in the transaction data, extract keywords for subsequent text analysis and data mining, and associate and integrate customer information data with the data of credit rating agencies to supplement customer credit scores and other information, enrich customer portrait data, and provide more comprehensive data support for risk assessment and precision marketing;

[0104] Data integration and storage: Integrate the cleaned and converted transaction data, market data and customer information data, and classify and store them according to the data subject and business logic. Specifically, store the transaction data in the Hive table and partition them according to the transaction date to facilitate quick query and analysis of transaction conditions in different time periods; store the market data in the time series database to facilitate efficient time series analysis of the changing trends of market data; store the customer information data in the relational database, and establish appropriate indexes and table associations to facilitate business operations such as customer relationship management and precision marketing; at the same time, store the integrated data metadata information in the data management module of the big data center, including the data name, type, source, storage location, and update time, so as to manage and query the data in a unified manner.

[0105] (IV) System testing and optimization

[0106] Functional testing: Conduct comprehensive functional testing on each functional module of the big data center. For the data collection function, test the accuracy and completeness of data collected from different data sources, check whether the collected data is consistent with the data in the data source, and whether there is data loss or repeated collection. For the data mining and analysis function, use the test data set with known results to conduct model training and prediction, and verify whether the output results of the decision tree algorithm and neural network model meet expectations.

[0107] Test the system's user interface to ensure that it is user-friendly and easy to operate. Users can easily perform operations such as data query, analysis task submission, and result viewing. Check whether the layout of the interface is reasonable, whether the menus and buttons function normally, whether the data display is clear and intuitive, and whether there are friendly prompts for incorrect user input operations, etc., to improve the user experience;

[0108] Performance testing: Use performance testing tools such as JMeter and LoadRunner to perform performance testing on the big data center, simulating a large number of concurrent user access and data processing scenarios, and evaluating the system's response time, throughput, and resource utilization. The number of concurrent users is set to gradually increase from 100 to 1,000, and the response time of the system under different concurrent pressures is observed to ensure that the average response time of the system does not exceed 5 seconds under high concurrency to meet the real-time requirements of the business. At the same time, monitor the utilization of resources such as the server's CPU, memory, disk I / O, and network bandwidth to ensure that resource utilization is within a reasonable range, CPU utilization does not exceed 80%, and memory utilization does not exceed 70%, to avoid resource bottlenecks that lead to system performance degradation;

[0109] Conduct special tests on data storage and query performance, testing the write speed and query efficiency of large-scale data storage; insert 100 million transaction data into the Hi ve table in batches, record the time required for insertion, and test the response time for complex queries on these data. By optimizing the database table structure, index design, and query statements, improve data storage and query performance to ensure that the system can quickly respond to user data requests;

[0110] Security testing: Use vulnerability scanning tools to conduct comprehensive vulnerability scanning on the big data center's network, operating system, database, and application programs to check whether the system has known security vulnerabilities. Vulnerability scanning tools include Nessus. Timely install system updates and security patches to repair discovered vulnerabilities and improve system security.

[0111] Conduct penetration testing to simulate hacker attacks, try to break through the system's security defenses, obtain sensitive data or perform illegal operations; try to obtain user information in the database through SQL injection attacks, find unsafe ports open in the system through network port scanning, try to log in to the system through brute force password cracking, etc. Based on the results of the penetration test, strengthen the system's access control, input verification, encryption mechanism and other security measures to prevent potential security risks;

[0112] Conduct security audits to check whether the system's security log records are complete and accurate, and whether they can trace back user operations and system security events; audit data access logs to check whether there are abnormal data access records, such as unauthorized users accessing sensitive data; audit system login logs to check whether there are multiple failed logins followed by successful logins, which may indicate a password cracking risk. Through security audits, potential security issues can be discovered and addressed in a timely manner to ensure system security and compliance;

[0113] Optimization and adjustment: According to the test results, conduct targeted optimization and adjustment on the big data center. For the vulnerabilities and risks discovered in the security test, take corresponding security reinforcement measures.

[0114] (V) System Deployment and Maintenance

[0115] System deployment: Install a customized Linux operating system, database management system, and application programs in the production environment. The application programs include data collection programs, preprocessing programs, and mining and analysis programs; Install and perform basic configurations of the operating system on the server, including network settings, user permission management, etc. Install the database management system, and import the tested and optimized database table structures, indexes, and initial data; Deploy the application programs to ensure that the dependencies between various programs are correctly configured and can be started and run normally;

[0116] Configure the server and network environment. According to the design in the hardware facility construction stage, perform cluster configuration on the server. Use the Hadoop cluster management tool to configure multiple servers into a Hadoop cluster to achieve distributed computing and storage functions; Configure the network environment, including setting network parameters such as IP addresses, subnet masks, gateways, etc., to ensure normal communication between servers, between servers and storage devices, and with the external network; At the same time, configure firewall rules to restrict illegal access from the external network to the internal resources of the big data center, and only open the necessary service ports to ensure the security of the system;

[0117] Daily maintenance: Establish a 7×24-hour monitoring mechanism, and use monitoring tools to monitor the running status of the system in real time, including the hardware health status of the server, system performance indicators, the running status of application programs, and data storage conditions; Once a system failure or abnormal situation is detected, send an alarm in a timely manner to notify the system administrator for handling;

[0118] Regularly conduct inspections and maintenance on the system, including checking whether the hardware devices of the server are running normally, checking the update status of the operating system and application programs, installing security patches and upgraded versions in a timely manner, fixing known vulnerabilities and problems, and improving the stability and security of the system; Optimize and maintain the database, such as performing database backup operations, checking whether the database indexes need to be rebuilt, optimizing query statements, etc., to ensure the efficient operation of the database.

[0119] The above specific embodiments are merely several alternative embodiments of the present invention. Based on the technical solution of the present invention and the relevant inspirations of the above embodiments, those skilled in the art can make various alternative improvements and combinations to the above specific embodiments.

Claims

1. A big data center based on a database structure, characterized in that: Including data storage module, data processing module, data management module and data security module: The data storage module is used to store data. The data storage module is based on a distributed storage architecture and has a data hierarchical storage function, including a physical layer, a logical layer and an application layer; The data processing module includes a data preprocessing unit and a data mining and analysis unit, wherein the data preprocessing unit is used to perform data cleaning, data type standardization and data integration on the data, and the data mining and analysis unit is used to process and analyze the data; The data management module is used to manage the metadata of the data to ensure the consistency and integrity of the data; The data security module ensures data security through encryption technology, access control technology and data backup and recovery mechanism.

2. A big data center based on a database structure according to claim 1, characterized in that: The physical layer adopts disk array technology, which includes redundant disk array technology. The redundant disk array technology performs striping on multiple disks in the disk array and stores data in multiple disks in a dispersed manner. The redundant disk array algorithm formula is as follows: Among them, d i Represents the data of the i-th disk in the disk array, where n is the number of disks in the disk array; The logic layer divides data into different logic modules according to functions, including user data, business data, and system configuration data. Each logic module has independent functions and interfaces. Data interaction is achieved between modules through data interfaces, and the data interfaces comply with communication protocols. The application layer is used to provide data access and processing interfaces for users.

3. The big data center based on the database structure according to claim 1, characterized in that: The data cleaning removes noise, impurities and outliers in the data through regular expressions. The formula is as follows: Among them, α i is the character set in the regular expression, β i The number of times the corresponding character is repeated; The data type standardization includes the unification of data formats and a data dictionary. The unified data format is used to ensure the consistency of data, and the data dictionary is used to define data elements.

4. The big data center based on the database structure according to claim 3, characterized in that: The data formats include date type, numeric type, and string type. For the date type, the ISO standard format is adopted. The algorithm formula of the date format is as follows: DATE=YYYY-MM-dd Among them, YYYY represents year, MM represents month, and dd represents day; For the numeric type, define the value range and precision of the data type, where the value range of the integer type is [-2 31 ,2 31 -1], the precision of floating point types is ±10 -7 ; For the string type, the length of the string shall not exceed 255 characters, and the character set is ASCII encoding; The data dictionary is used to define and explain data elements. The data elements include data name, data type, value range, and description. The structure of the data dictionary is as follows:

5. The big data center based on database structure according to claim 1, characterized in that: The data mining and analysis unit uses a machine learning algorithm and a deep learning algorithm to process and analyze data. The machine learning algorithm uses a decision tree algorithm. The decision tree algorithm classifies and predicts data by constructing a tree structure. The decision tree algorithm formula is as follows: Among them, Entropy(S) is the information entropy of the data set S, c is the number of category labels in the data set S, and p i is the probability of category i appearing in data set S, Gain(S,A) is the information gain of feature A for data set S, Values(A) is the set of all possible values ​​of feature A, S v It is a subset of feature A in the data set S with a value of v. By calculating the information gain of each feature, the feature with the largest information gain is selected as the splitting node of the decision tree, so as to construct a decision tree to classify and predict the data; The deep learning algorithm adopts a neural network model, and the neural network model formula is as follows: z l+1 =W l+1 a l +b l+1 a l+1 =f(z l+1 ) Among them, l represents the number of layers of the neural network, z l+1 is the weighted sum of the inputs to layer l+1, including the bias b l+1 , W l+1 is the weight matrix of the l+1th layer, a l is the output of the lth layer, f(·) is the activation function, including the Sigmoid function, through the weighted and connected activation operations of multiple layers of neurons, the neural network model learns and analyzes the data, and adjusts the weight W and bias b to fit the characteristics and patterns of the data.

6. The big data center based on database structure according to claim 1, characterized in that: The data management module is used to manage the metadata of the data to ensure the consistency and integrity of the data. The metadata includes the data name, data type, value range, and description. The data management module includes index and directory management of the data.

7. The big data center based on database structure according to claim 1, characterized in that: The encryption technology uses an encryption algorithm formula to encrypt data. The encryption algorithm formula is as follows: E=E key (P) Among them, E is the encrypted ciphertext, E key is the encryption key, and P is the original data; The access control technology manages user rights and limits user access to data through user identity authentication and authorization mechanisms; The data backup and recovery mechanism adopts a regular backup strategy, and the storage format of the backup data is JSON format.

8. A method for constructing a big data center based on a database structure, characterized in that: The following steps are involved: Hardware facilities construction: The hardware facilities include servers, storage devices and network equipment to build the hardware environment; Software system development: The software system includes an operating system, a database management system and an application program. The operating system is Linux. The database management system includes database table structure, index and storage strategy. Data migration and integration: collect, clean, convert and integrate data. The data sources of the collected data include business systems, sensors, and networks. The data cleaning includes removing noise, impurities and outliers from the data. The data conversion converts the data into a format suitable for storage and analysis. The integrated data should be stored in the data storage module. System testing and optimization: Perform functional, performance and security tests on the system, and make optimization adjustments based on the test results. The functional test is used to ensure the normal operation of the system functions. The performance test evaluates the response time, throughput and resource utilization of the system. The security test is used to check the security and vulnerabilities of the system. The system is optimized based on the test results. System deployment and maintenance: Deploy the system to the production environment and perform routine maintenance and data updates. The deployment process includes installing the operating system, database management system and application programs, configuring the server and network environment, and routine maintenance includes monitoring the system operation status, handling faults and abnormal situations.

9. The method for constructing a big data center based on a database structure according to claim 8, characterized in that: The testing methods used in the security test include vulnerability scanning, penetration testing and security auditing. The scope of the security test covers the system's network, operating system, database, and application programs, and the security vulnerabilities and risks found are repaired and reinforced.