Log data warehouse establishment method and system, electronic equipment and storage medium
By establishing a log data warehouse, collecting and correlating DNS and AAA log data, generating integrated log data and determining dimension tables, the problem of diverse data formats and lack of a unified platform in traditional systems is solved, efficient log data management and real-time analysis are achieved, and network security and performance monitoring capabilities are improved.
Patent Information
- Application Number
- CN202510556584.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-15
AI Technical Summary
Traditional log management systems are difficult to efficiently store, manage and analyze DNS and AAA log data, resulting in diverse data formats, lack of a unified platform and insufficient real-time performance, affecting network security and performance monitoring.
Establish a log data warehouse, collect DNS and AAA log data, perform preprocessing and association analysis, generate integrated log data, and determine dimension tables based on business needs, establish a distributed data warehouse architecture, support diversified data formats, and provide real-time analysis and encryption processing.
It improves data storage and management efficiency, realizes unified log data management and analysis, improves network operation and maintenance efficiency and security, and supports real-time data query and abnormal detection.
Smart Images

Figure CN120492547A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a method, system, electronic device, and storage medium for establishing a log data warehouse. Background Art
[0002] In modern network management and security, DNS (Domain Name System) and AAA (Authentication, Authorization, Accounting) log data are critical monitoring and analysis resources. DNS log data records domain name resolution requests and responses, helping network administrators understand network traffic, detect malicious activity, and optimize network performance. AAA log data records user authentication, authorization, and accounting information, fundamental to ensuring network security and resource utilization.
[0003] As network size and complexity increase, traditional log management systems are struggling to meet the demands for efficient data storage, management, and analysis. Specifically, the diverse data formats used in related technologies hinder comprehensive data analysis. Furthermore, DNS and AAA log data are distributed across different platforms, requiring cross-platform data processing, which impacts processing efficiency. Summary of the Invention
[0004] The main purpose of the embodiments of the present application is to provide an efficient method, system, electronic device and storage medium for establishing a log data warehouse.
[0005] To achieve the above objectives, one aspect of an embodiment of the present application proposes a method for establishing a log data warehouse, the method comprising: collecting DNS log data and AAA log data; preprocessing the DNS log data and the AAA log data, and performing correlation analysis on the DNS log data and the AAA log data to generate integrated log data; determining a dimension table for the data warehouse based on business needs and the integrated log data, and storing the integrated log data in the data warehouse. By performing correlation analysis on the DNS log data and the AAA log data to generate unified integrated log data, and then performing data storage and analysis, the present application is conducive to improving the efficiency of data analysis and processing.
[0006] In some embodiments, the method provided by the embodiments of the present application, wherein the correlation analysis of the DNS log data and the AAA log data to generate integrated log data includes:
[0007] Correlating the DNS log data with the AAA log data according to the timestamp of the DNS log data and the session time of the AAA log data to obtain a correlation result;
[0008] Alternatively, correlating the DNS log data and the AAA log data according to the port number of the DNS log data and the port range of the AAA log data to obtain a correlation result;
[0009] Alternatively, according to the IP address of the DNS log data and the IP address of the AAA log data, the DNS log data and the AAA log data are associated to obtain an association result;
[0010] Based on the association result, integrated log data including preset log fields is generated.
[0011] In some embodiments, the method provided by the embodiments of the present application, wherein determining the dimension table of the data warehouse based on the business requirements and the integrated log data, includes:
[0012] Determining a time dimension table according to timestamp information associated with the DNS log data or the AAA log data;
[0013] Determine a geographic location dimension table based on the IP addresses associated with the DNS log data and in association with the geographic locations;
[0014] Classify users and user behaviors to determine a user type dimension table; the users are related to the DNS log data or the AAA log data;
[0015] Classify and process the query request of the DNS log data or the AAA log data to determine a query type dimension table;
[0016] The time dimension table, the geographic location dimension table, the user type dimension table, and the query type dimension table are aggregated to obtain a dimension table of the data warehouse.
[0017] In some embodiments, the method provided by the embodiments of the present application further includes:
[0018] Determine the hash value of the integrated log data and create a hash table;
[0019] Alternatively, frequently queried fields are determined based on historical query records, and dedicated indexes corresponding to the frequently queried fields are established;
[0020] Alternatively, new data and updated data in the data warehouse are recorded, and incremental indexes of the new data and updated data are created.
[0021] In some embodiments, the method provided by the embodiments of the present application further includes:
[0022] encrypting the integrated log data;
[0023] Alternatively, the permissions of the access user are determined based on the role of the access user.
[0024] In some embodiments, the method provided by the embodiments of the present application further includes:
[0025] Compressing the integrated log data using a Huffman coding algorithm;
[0026] Establish a distributed data warehouse architecture.
[0027] In some embodiments, the method provided by the embodiments of the present application further includes:
[0028] Establish the data architecture of the log data warehouse;
[0029] The data architecture includes:
[0030] A basic layer, used to store the originally collected DNS log data and the AAA log data;
[0031] An intermediate layer, configured to pre-process the DNS log data and the AAA log data;
[0032] The subject layer is used to perform correlation analysis on the DNS log data and the AAA log data to generate integrated log data;
[0033] The application layer is used to provide data query and data analysis.
[0034] To achieve the above objectives, another aspect of the present application provides a system for establishing a log data warehouse, the system comprising:
[0035] The first module is used to collect DNS log data and AAA log data;
[0036] The second module is used to pre-process the DNS log data and the AAA log data, and perform correlation analysis on the DNS log data and the AAA log data to generate integrated log data;
[0037] The third module is used to determine the dimension table of the data warehouse according to business requirements and the integrated log data, and store the integrated log data in the data warehouse.
[0038] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned method when executing the computer program.
[0039] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program implements the above-mentioned method when executed by a processor.
[0040] The embodiments of the present application include at least the following beneficial effects: The method provided by the embodiments of the present application includes: collecting DNS log data and AAA log data; preprocessing the DNS log data and the AAA log data, and performing correlation analysis on the DNS log data and the AAA log data to generate integrated log data; determining a dimension table of a data warehouse based on business needs and the integrated log data, and storing the integrated log data in the data warehouse. By performing correlation analysis on the DNS log data and the AAA log data to generate unified integrated log data, and then performing data storage and analysis, the present application is conducive to improving the efficiency of data analysis and processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 This is a flowchart of an embodiment of a method for establishing a log data warehouse provided by the present application;
[0042] Figure 2 This is a flowchart of an embodiment of the association analysis process provided by this application;
[0043] Figure 3 This is a flowchart of an embodiment of a dimension table determination process provided by this application;
[0044] Figure 4 is a flowchart of an embodiment of the indexing process provided by this application;
[0045] Figure 5 This is a structural diagram of an embodiment of the log data warehouse architecture provided by this application;
[0046] Figure 6 This is a flowchart of another embodiment of the method for establishing a log data warehouse provided by the present application;
[0047] Figure 7 This is a flowchart of an embodiment of the collection process provided by this application;
[0048] Figure 8 This is a flow chart of an embodiment of the flow process provided by this application;
[0049] Figure 9 This is a structural diagram of a system for establishing a log data warehouse provided by an embodiment of the present application;
[0050] Figure 10 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are merely examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.
[0052] It will be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0053] The terms "at least one", "plurality", "each", "any", etc. used in this application include "at least one", "two" or more, "plurality" or "each", "any" or "any one", "each" or "any one" as used herein.
[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0055] Before explaining the embodiments of the present application in detail, some of the nouns and terms involved in the embodiments of the present application are first explained. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.
[0056] DNS (Domain Name System): The Domain Name System (DNS) is the internet's phone book. It converts user-friendly domain names (such as www.example.com) into computer-readable IP addresses (such as 192.0.2.1), allowing users to access websites by domain name without having to remember complex numeric addresses. DNS's main functions include domain name resolution, domain name resolution caching, and reverse resolution.
[0057] AAA: Authentication, Authorization, and Accounting: In computer network and information security, AAA is a framework for controlling access to computer resources, tracking user activity, and ensuring the authenticity of user identities.
[0058] CNAME (Canonical Name Record): A Canonical Name Record (CNAME) is a type of DNS record that points a domain alias to the official name of another domain. For example, if you have a domain example.com and its alias www.example.com, you can use a CNAME record to point www.example.com to example.com. This way, users visiting www.example.com will be automatically redirected to example.com's IP address.
[0059] Data warehouse: The full name is data warehouse (Data Warehouse), which is a database system used for reporting and analysis. It provides decision support for enterprises by integrating data from one or more data sources.
[0060] Classification algorithms are algorithms used in the field of machine learning to solve classification problems. Classification is the problem of assigning each instance in a dataset to one or more categories. These algorithms are typically used to predict the category of a data point, making decisions based on input features.
[0061] AES encryption: The full name is Advanced Encryption Standard, which is a widely used symmetric key encryption algorithm.
[0062] Huffman Coding: A widely used data compression algorithm. It is a greedy algorithm for lossless data compression that encodes common patterns in the data using a variable-length coding table, thereby reducing the overall data size.
[0063] In modern network management and security, DNS (Domain Name System) and AAA (Authentication, Authorization, Accounting) log data are critical monitoring and analysis resources. DNS log data records domain name resolution requests and responses, helping network administrators understand network traffic, detect malicious activity, and optimize network performance. AAA log data records user authentication, authorization, and accounting information, fundamental to ensuring network security and resource utilization.
[0064] As network scale and complexity increase, traditional log management systems have become unable to meet the needs of efficient data storage, management, and analysis. Existing log management systems often have the following problems:
[0065] Huge data volume: DNS and AAA log data is huge, and traditional storage systems have difficulty processing and storing this data efficiently.
[0066] Diverse data formats: Log data generated by different devices and services comes in various formats, which increases the complexity of data integration and analysis.
[0067] High real-time requirements: Network security and performance monitoring require real-time analysis of log data to promptly detect and respond to abnormal activities.
[0068] Lack of a unified platform: Existing operators' DNS and AAA data storage is usually scattered across different platforms and tools, lacking a unified management and analysis platform, leading to data silos.
[0069] To address these challenges, establishing a comprehensive data warehouse based on DNS and AAA log data is crucial. This approach should offer efficient data storage and management, support the integration of various log data formats, and provide real-time data analysis. This solution allows network administrators to better monitor network status, promptly identify and respond to security threats, and optimize network resource usage, thereby improving overall network security and performance.
[0070] In view of this, the present invention provides a method for establishing and analyzing a DNS and AAA log data warehouse in an embodiment to address the various challenges currently faced in log data management and analysis, and to provide an efficient, unified, and powerful log data management and analysis platform. Specifically, the present invention has the following objectives:
[0071] Improve data storage and management efficiency: By designing an efficient data warehouse structure, it can process and store large amounts of DNS and AAA log data, solving the performance bottleneck problem of traditional storage systems when facing massive data.
[0072] Support diverse data formats: Develop a unified data format conversion and integration mechanism that is compatible with multiple log data formats generated by different devices and services, simplifying the data processing process.
[0073] Implement offline data analysis: Provide offline data analysis functions, support rapid query and processing of log data, and help network administrators promptly detect and respond to network anomalies and security threats.
[0074] Build a unified management platform: Establish an integrated log data management and analysis platform to eliminate data silos, provide centralized log data storage, management, and visualization tools, and improve network operation and maintenance efficiency.
[0075] Enhanced data storage security: By encrypting sensitive DNS and AAA log data with ASE, data can be prevented from being stolen or tampered with during transmission, thus avoiding potential security threats.
[0076] By providing these features, this invention aims to provide network administrators with a comprehensive and efficient log data management and analysis solution, improving data security, reliability, and operational efficiency. By establishing an efficient DNS and AAA log data warehouse, it enables rapid collection, cleaning, storage, and analysis of massive amounts of DNS and AAA log data, improving overall data processing efficiency.
[0077] The method for establishing a log data warehouse provided in the embodiment of the present application relates to the field of data processing and network security technology. The method for establishing a log data warehouse provided in the embodiment of the present application can be applied to a terminal, can also be applied to a server, and can also be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or can be configured as a server cluster or distributed system composed of multiple physical servers, and can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements the method for establishing a log data warehouse, etc., but is not limited to the above forms.
[0078] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0079] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0080] Figure 1 This is an optional flowchart of the method for establishing a log data warehouse provided in an embodiment of the present application; Figure 1 The method may include but is not limited to steps S100 to S300.
[0081] Step S100: collecting DNS log data and AAA log data;
[0082] Step S200: pre-processing the DNS log data and the AAA log data, and performing correlation analysis on the DNS log data and the AAA log data to generate integrated log data;
[0083] Step S300: Determine the dimension table of the data warehouse according to business requirements and integrated log data, and store the integrated log data in the data warehouse.
[0084] This application collects DNS log data and AAA log data in real time. Data preprocessing includes data cleaning and format conversion to generate standardized data. This application performs correlation analysis on DNS log data and AAA log data to generate integrated log data that combines the DNS and AAA log data. Based on business needs, the types and dimensions of the data to be stored are determined. Dimension tables are designed based on the characteristics of the integrated log data. Dimension tables are used for data analysis and processing.
[0085] This paper proposes a method and system for establishing and analyzing a DNS and AAA log data warehouse. This system aims to enhance network management and security monitoring capabilities by efficiently storing, managing, and analyzing large amounts of DNS and AAA log data. The system combines DNS and AAA log data and provides a unified data warehouse structure and offline analysis capabilities to meet the needs of network operations and security monitoring.
[0086] In some embodiments, reference Figure 2 As shown, the method provided in the embodiment of the present application performs correlation analysis on DNS log data and AAA log data to generate integrated log data, including:
[0087] Step S210, correlating the DNS log data and the AAA log data according to the timestamp of the DNS log data and the session time of the AAA log data to obtain a correlation result;
[0088] Alternatively, step S220, correlating the DNS log data and the AAA log data according to the port number of the DNS log data and the port range of the AAA log data to obtain a correlation result;
[0089] Alternatively, step S230, correlating the DNS log data and the AAA log data according to the IP address of the DNS log data and the IP address of the AAA log data to obtain a correlation result;
[0090] Step S240 : generating integrated log data including preset log fields based on the correlation result.
[0091] It can be understood that the association of time information can be determined by whether the timestamp is within the session time; the association of port information can be determined by whether the port number is within the port range; and the relationship of IP addresses can be determined by the consistency of the two IP addresses.
[0092] In some embodiments, reference Figure 3 The method provided in the embodiment of the present application determines the dimension table of the data warehouse according to business needs and integrated log data, including:
[0093] Step S310: determining a time dimension table based on timestamp information related to the DNS log data or the AAA log data;
[0094] Step S320: determining a geographic location dimension table based on the IP addresses associated with the DNS log data and the associated geographic locations;
[0095] Step S330: Classify the users and their behaviors to determine a user type dimension table; the users are related to the DNS log data or the AAA log data;
[0096] Step S340: Classify the query request of the DNS log data or the AAA log data and determine the query type dimension table;
[0097] Step S350 , summarizing the time dimension table, the geographic location dimension table, the user type dimension table, and the query type dimension table to obtain a dimension table of the data warehouse.
[0098] It should be noted that the design of the dimension table can be adjusted according to actual conditions, and this application does not limit the specific type of the dimension table. In other embodiments, this application can also set the weight of the dimension table and update it from time to time to improve data query and analysis efficiency.
[0099] In some embodiments, reference Figure 4 , the method provided in the embodiment of the present application further includes:
[0100] Step S400, determining the hash value of the integrated log data and creating a hash table;
[0101] Alternatively, in step S500, frequently queried fields are determined based on historical query records, and dedicated indexes corresponding to the frequently queried fields are established;
[0102] Alternatively, in step S600, new data and updated data in the data warehouse are recorded, and incremental indexes of the new data and updated data are established.
[0103] This application can also create an index through other methods.
[0104] In some embodiments, the method provided by the embodiments of the present application further includes:
[0105] Encrypt the integrated log data;
[0106] Alternatively, the permissions of the access user are determined based on the access user's role.
[0107] This application sets access permissions for different user roles to improve the security of the data warehouse.
[0108] In some embodiments, the method provided by the embodiments of the present application further includes:
[0109] The integrated log data is compressed using the Huffman coding algorithm;
[0110] Establish a distributed data warehouse architecture.
[0111] The present application may also compress the data in other ways.
[0112] In some embodiments, the method provided by the embodiments of the present application further includes:
[0113] Establish the data architecture of the log data warehouse;
[0114] Reference Figure 5 , the data architecture includes:
[0115] The basic layer is used to store the original collected DNS log data and AAA log data;
[0116] The middle layer is used to pre-process DNS log data and AAA log data;
[0117] The subject layer is used to perform correlation analysis on DNS log data and AAA log data to generate integrated log data;
[0118] The application layer is used to provide data query and data analysis.
[0119] The following is a detailed introduction and description of the solution of the embodiment of the present invention in conjunction with specific application examples. Figure 6 As shown, the data warehouse establishment method provided by this application includes:
[0120] 1.1. Multidimensional model design of data warehouse.
[0121] S11. Data Warehouse Construction: A data warehouse is a subject-oriented, integrated, non-volatile, time-varying data collection that supports management decision-making. It extracts data from distributed, heterogeneous data sources and, through data cleansing, transformation, and integration, creates a high-quality data collection for analysis and decision support.
[0122] S12. Select the star schema for multi-dimensional modeling: This solution uses the star schema to build a data warehouse and supports multi-dimensional analysis of DNS and AAA log data. A new dynamic weighting algorithm is added to analyze historical query frequencies through machine learning, automatically adjust the association priority of dimension tables, and improve query performance. Materialized views are established for high-frequency combined queries (such as "region + protocol + time period"), and an incremental update mechanism is supported, which only recalculates partitions with changes exceeding 5%, reducing computing resource consumption, shortening query response time, and ensuring data freshness.
[0123] W=0.7*Q+0.3*D
[0124] Where W represents the weight of the dimension, Q represents the number of queries, and D represents the data cardinality.
[0125] The comparative advantages are shown in Table 1:
[0126]
[0127] Table 1
[0128] 1.2. Dimensional design.
[0129] S13. Time dimension design: The time dimension divides time into days and hours, serving as a partition key to optimize query performance. Each DNS and AAA log record is associated with a specific timestamp, enabling analysis of network traffic and behavior patterns by time period. Applying time series can be used to identify traffic patterns and trends:
[0130]
[0131] Where T t is the total flow at time t, X i is the flow rate in the i-th time interval, w i is the corresponding weight.
[0132] S14. Geographic Location Dimension: The geographic location dimension records the mapping between IP addresses and geographic locations. By associating IP addresses with geographic locations, the solution supports regional analysis of DNS query and traffic patterns, identifying regional network issues or threats.
[0133] S15. Design of user type dimension: The user type dimension is used to distinguish different types of users, such as dedicated line users and broadband users. By analyzing the behavior patterns of different user types, abnormal activities or network behavior characteristics of specific user groups can be detected. User behavior analysis can distinguish different types of users through classification algorithms:
[0134] U c =classify(u,D)
[0135] Among them U c is the classification of user u, and D is the user behavior dataset.
[0136] S16. Design of query type dimension: The query type dimension categorizes DNS query requests, such as A record query, AAAA record query, etc. By identifying different types of query patterns and trends, we can better understand the composition of network traffic and optimize DNS services. Query pattern identification can be achieved through statistical analysis:
[0137]
[0138] Among them, P q is the proportion of specific query type q, N q is the number of queries of this type, N t is the total number of queries.
[0139] 1.3. Efficient data indexing mechanism.
[0140] S17. Develop an efficient data indexing mechanism: To accelerate DNS queries and log data retrieval, the system has developed an efficient data indexing mechanism using hash tables and B-tree index structures. During data storage, the system automatically creates a hash value for each log record, allowing for rapid location of data records using the hash table. The hash function design can be expressed as follows:
[0141] h(k)=HashFunction(m)modm
[0142] Where h(k) is the hash value of keyword k, HashFunction is the hash function, and m is the size of the hash table.
[0143] S18. Create dedicated indexes for frequently queried fields: The system uses a consistent hashing algorithm to build scalable hash indexes for frequently queried fields (such as account numbers, IP addresses, domain names, CNAMEs, and customer names), and supports dynamic expansion of index nodes.
[0144] A built-in hash collision optimization mechanism (chain address method + secondary detection method) achieves a collision rate of <0.3%. A single node can process 120,000 precise queries per second (average latency ≤ 1ms). These indexes are automatically generated when log data is imported and updated regularly, ensuring that relevant data can be quickly located during query execution, reducing query response time.
[0145] S19. Range search B+ tree index: Build a hierarchical B+ tree index for continuous fields such as timestamps and response duration, and automatically shard by time window (30 days / layer by default) to achieve constant efficiency in cross-year data retrieval. Using SSD to optimize the storage structure, the node fan-out coefficient is increased to 512, and the tree height is compressed by 40%. A billion-level data range query (such as "202x-01 to 202x-12") takes ≤0.8 seconds. To improve indexing efficiency, the system adopts an incremental indexing strategy, which only indexes new or updated data, rather than rebuilding the entire index structure each time. This not only saves system resources, but also ensures that the index is always up to date.
[0146] 1.4. Security Enhancements.
[0147] S110, Data Encryption: To ensure the confidentiality of DNS and AAA log data, the system encrypts all stored and transmitted data. It uses the Advanced Encryption Standard (AES) to encrypt sensitive data to prevent theft or tampering during transmission. The AES encryption process can be expressed as the following mathematical formula:
[0148] C=E(K,P)
[0149] Among them, C is the ciphertext, E is the encryption function, K is the key, and P is the plaintext.
[0150] S111. Access Control: This system implements a role-based access control (RBAC) mechanism for access control. Unlike traditional RBAC solutions, the innovative solution of this invention implements fine-grained control at the field level and supports advanced features such as IP end segment hiding and regular expression matching. In addition, the permission verification mechanism uses zero-knowledge proof (ZKP) technology, which can be verified without transmitting ciphertext, greatly enhancing security. The dynamic policy engine reduces the delay of hot policy loading to no more than 500ms, greatly improving flexibility and real-time performance. At the same time, the operation fingerprint is stored on the blockchain to ensure the integrity and non-tamperability of the audit trail.
[0151] 1.5. Data compression and processing optimization.
[0152] S112. Data compression strategy: Due to the huge amount of DNS and AAA log data, the system introduces a data compression mechanism to reduce storage space usage and improve data reading efficiency.
[0153] In some embodiments, a Huffman Coding algorithm is used to compress data. The compression algorithm is as follows:
[0154]
[0155] where p iThe construction of the Huffman tree is based on the probability of each character, and Cost is the total expected length of the code, l i Is the encoded length of the character.
[0156] A multi-mode compression algorithm has been introduced, employing dictionary compression, Delta-RLE encoding, and bit-slicing compression for different data types, achieving high compression rates of 95%, 80%, and 67%, respectively, effectively reducing storage space. Furthermore, a new direct computation capability in compressed state has been added, allowing aggregation operations to be performed directly in the compressed state without decompressing the data, significantly reducing I / O throughput and improving data processing efficiency. These upgrades not only improve storage efficiency but also significantly enhance the system's performance and practicality.
[0157] Adopting adaptive compression algorithm selector, the comparison is shown in Table 2:
[0158] Data Type Compression algorithm Compression ratio Text log Zstandard 8:1 Digital indicators Delta+Snappy 12:1 Timestamp Gorilla 10:1
[0159] Table 2
[0160] Design SIMD optimized parsing pipeline:
[0161]
[0162] The log parsing throughput reaches 1.2GB / s / core, which is 6 times higher than traditional solutions.
[0163] The technical advantages are summarized in Table 3:
[0164]
[0165]
[0166] Table 3
[0167] S113. Distributed Data Warehouse Architecture: The solution adopts a distributed data warehouse architecture to address the storage and processing needs of massive data volumes. Through this distributed storage and computing framework, the solution can scale horizontally to support larger-scale data processing tasks. A comparison is shown in Table 4.
[0168]
[0169] Table 4
[0170] S114, stream-batch fusion analysis engine: The stream computing DAG and offline MapReduce tasks are compiled into a unified intermediate code (IR); state data is stored in SSD-RAID, and offline data is cached in a distributed memory pool.
[0171] Streaming and batch SQL syntax
[0172]
[0173] Quantified advantages: Resource utilization increased from 30% to 78%, and real-time alarm latency was reduced from 2 seconds to 0.5 seconds.
[0174] 1.6. Data integration and correlation analysis.
[0175] S115. Setting of association conditions: To associate DNS logs with AAA logs, the system performs association based on the following conditions:
[0176] Time dimension: The timestamp of the DNS log must be between the session start and end time of the AAA log.
[0177] Port matching: The port number of the DNS request should be within the start port and end port range of the AAA log.
[0178] IP address matching: The IP address in the DNS request must be consistent with the IP address in the AAA log.
[0179] S116. Generate unified log records: After association and integration, the system generates unified log records containing the following fields: date, time, account, request IP, request port, destination IP, destination port, ID, domain name, domain name customer affiliation, request type, resolution result, CNAME, CNAME customer affiliation, resolution type, IPv4 upstream traffic, IPv4 downstream traffic, IPv6 upstream traffic, IPv6 downstream traffic.
[0180] 1.7. Business domain and purpose hierarchical structure design.
[0181] S117, Operational Data Store (ODS): The base layer is the underlying layer of the data warehouse and is primarily responsible for storing raw DNS and AAA log data. Raw data is collected from various distributed systems and, after simple data cleaning and preprocessing, is directly stored in the ODS.
[0182] S118, Data Warehouse Detail (DWD): The middle layer performs data cleaning, preprocessing, and format conversion to generate standardized log data. The purpose of this layer is to convert raw data into a structured, standardized form to facilitate subsequent analysis and querying.
[0183] S119, Data Warehouse Summary (DWS): At the topic layer, the solution correlates and integrates DNS and AAA log data to generate themed log records. These records are further processed and analyzed to provide data support for specific topics (such as security monitoring and performance analysis).
[0184] S120, Application Data Store (ADS): The application layer provides efficient data access services for real-time query and analysis. Data in this layer is optimized for real-time or near-real-time analysis scenarios, such as network traffic monitoring and anomaly detection.
[0185] The following describes a specific implementation plan based on the DNS and AAA log data warehouse. It should be understood that this patent application is intended to enable those skilled in the art to more easily and efficiently process services, particularly to gain a deeper understanding of, and subsequently implement and even optimize, this solution. This patent application is by no means limited by its textual content. In fact, this implementation plan is intended to provide a method and approach to solving problems for those skilled in the field.
[0186] Relevant technical personnel can know that the present invention is an implementable computer analysis solution based on DNS, AAA and other data using big data processing as a technical means.
[0187] The principle of the present invention will be explained in detail below with reference to another representative embodiment of the present invention.
[0188] 21. Data collection and preprocessing:
[0189] Collect DNS and AAA log data and pre-process them separately to ensure the uniformity and integrity of the data format.
[0190] The DNS log data format is as follows:
[0191] Destination IP|Destination Port|Requesting IP|Requesting Port|ID|Domain Name|Request Type|Resolution Result|CNAME|Resolution Type|Date|Resolution Log.
[0192] The format of AAA log data is as follows:
[0193] Account|Status|IPv4|IPv6|IP|Start Port|End Port|Start Time|End Time|V4 Upstream Traffic|V4 Downstream Traffic|V6 Upstream Traffic|V6 Downstream Traffic|Brasip|NAS-Port.
[0194] Specific implementation steps:
[0195] Determine business needs:
[0196] Determine the data types and dimensions that need to be stored and analyzed based on the data analysis requirements of each operator's business scenario.
[0197] Design dimension table:
[0198] Based on the characteristics of log data, design the time dimension table, IP dimension table, domain name dimension table, and user account dimension table. The following example:
[0199]
[0200]
[0201] To meet different business needs, data is organized into multiple hierarchical structures:
[0202] Basic layer (ODS, Operational Data Store): stores raw DNS and AAA log data.
[0203]
[0204]
[0205] Middle layer (DWD, Data Warehouse Detail): performs data cleaning, preprocessing, and format conversion to generate standardized log data.
[0206]
[0207] Theme layer (DWS, Data Warehouse Summary): associates and integrates DNS and AAA log data to generate themed log records;
[0208]
[0209]
[0210] Application layer (ADS, Application Data Store): provides efficient data access services for real-time query and analysis.
[0211] CREATE TABLE ads_combined_logs AS SELECT*FROM dws_combined_logs.
[0212] Through the above-mentioned data warehouse model design, the present invention realizes the efficient storage, management and analysis of DNS and AAA log data, and improves the data storage and management efficiency.
[0213] 22. Data loading and query:
[0214] Periodically load data from raw log files into the ODS layer.
[0215] Perform data cleaning and format conversion (methods, formulas) at the DWD layer.
[0216] Perform data association and integration at the DWS layer to generate themed log records.
[0217] Provides an efficient query interface at the ADS layer to support real-time data analysis and visualization needs.
[0218] In another embodiment, the present application proposes a method for establishing and analyzing a DNS and AAA log data warehouse to address the various challenges currently faced in log data management and analysis, including the following steps:
[0219] Step S31: data collection.
[0220] Reference Figure 7 As shown, this invention builds a multi-dimensional, fine-grained analysis system for DNS and AAA log data in complex operator business scenarios. The data acquisition module acquires DNS resolution logs and AAA authentication logs in real time, is compatible with mainstream network equipment and operating system log output formats, and uses an adaptive parsing engine to automatically identify and convert disparate source data into a unified structure.
[0221] Collection range and frequency:
[0222] DNS logs: Covers request and response logs of recursive resolvers, authoritative servers, and corporate intranet DNS. The collection frequency can be dynamically adjusted based on network traffic to ensure real-time data.
[0223] AAA logs: Comprehensively collect logs of AAA protocols such as RADIUS and TACACS+, supporting the collection of data from the entire authentication, authorization, and billing process. The frequency can be flexibly set based on business needs.
[0224] Technical Implementation: Multi-threading and asynchronous I / O technologies are used to enhance the concurrent processing capabilities of log collection, enabling a single node to collect 100,000 log entries per second. Data compression algorithms are also employed to reduce the amount of data transmitted and network bandwidth usage. A breakpoint-resume mechanism ensures that data collection can resume from the last interruption point after a network outage or system failure, ensuring data integrity.
[0225] Step S32: data cleaning and analysis.
[0226] Data cleaning:
[0227] Deduplication: Identify and remove duplicate log records through hashing algorithms, reducing the waste of storage and computing resources caused by redundant data.
[0228] Fixes: Corrected obvious errors (such as incorrect IP address format), supplemented missing values (such as filling missing port numbers based on context).
[0229] Filtering: Accurately filter data based on preset rules (such as retaining logs of specific domain names and IP segments) to improve data quality.
[0230] Data analysis:
[0231] Structured parsing: Parse unstructured log text into structured data to facilitate subsequent processing and analysis.
[0232] Field extraction: Use regular expressions and other techniques to extract key fields from logs, such as domain names and IP addresses in DNS queries.
[0233] Semantic parsing: Perform semantic analysis on specific fields, such as converting DNS query type codes into readable text, to enhance data understandability.
[0234] Step S33: dimensional modeling.
[0235] In terms of dimensional modeling, the present invention designs a fine-grained, hierarchical dimensional table system.
[0236] Time dimension table: accurate to the second level, supports time series analysis, and includes fields such as year, quarter, month, day, and hour, facilitating data aggregation and analysis at different time granularities.
[0237] IP dimension table: Integrates GeoIP data to provide precise geographic location information, covering fields such as IP address, country, province, city, district, latitude and longitude, and supports geographic distribution analysis.
[0238] Domain name dimension table: associates customer information to track business activities. It contains fields such as domain name, customer, and business type to facilitate business-level analysis.
[0239] User account dimension table: Integrates multi-source identity data to build a complete user profile, covering fields such as account number, name, user type, department, and permission level to serve user behavior analysis.
[0240] Step S34: Data warehouse model design.
[0241] To meet different business needs, data is organized into multiple hierarchical structures:
[0242] Step S341, basic layer (ODS, Operational Data Store).
[0243] The foundational layer is the underlying layer of the data warehouse, primarily responsible for storing raw DNS and AAA log data. Raw data is collected from various distributed systems and, after simple data cleaning and preprocessing, is directly stored in the ODS. The data structure at this layer remains consistent with the source system, ensuring data integrity and originality.
[0244] Technical Details: A distributed file system (HDFS) is used to store raw logs, ensuring high data availability and durability. Tools such as Flume and Kafka are used for real-time log collection and transmission, ensuring data timeliness. Hive external table technology is used to provide lightweight metadata management for raw logs, facilitating subsequent data processing.
[0245] Step S342: middle layer (DWD, Data Warehouse Detail).
[0246] The middle layer performs data cleaning, preprocessing, and format conversion to generate standardized log data. The middle layer further cleans and converts the data from the ODS layer, unifying the data format, handling missing values and outliers, and performing data standardization. The purpose of this layer is to convert raw data into a structured, standardized form to facilitate subsequent analysis and querying.
[0247] Technical details: see Figure 8 As shown, the Flink distributed computing framework is used to efficiently process large amounts of data. Machine learning algorithms (such as clustering and classification) are applied to identify and correct anomalies and errors in the data. Data lineage tools are used to track changes in data during the cleaning and transformation processes, facilitating troubleshooting and data governance.
[0248] Step S343, subject layer (DWS, Data Warehouse Summary).
[0249] The topic layer associates and integrates the data from the DWD layer based on business needs and analysis objectives, generating themed log records. These records are further processed and analyzed to provide data support for specific topics (such as security monitoring, performance analysis, and user behavior analysis).
[0250] Technical Details: Utilizes the HiveQL query language to perform complex joins and aggregations on multi-source data. Data cube technology is used to pre-calculate common analytical metrics and accelerate multidimensional data analysis. Data encryption and desensitization technologies are used to protect sensitive information and ensure data security.
[0251] Step S344: Application layer (ADS, Application Data Store).
[0252] The application layer optimizes and indexes the data in the DWS layer to improve the efficiency of data query and analysis. It is suitable for real-time or near-real-time analysis scenarios, such as network traffic monitoring, anomaly detection, and real-time display of business indicators.
[0253] Technical Details: Build a data caching mechanism (e.g., Redis) to reduce repetitive computations and improve response speed. Develop a RESTful API to facilitate data access by front-end applications and data analysis tools.
[0254] Step S35: Data loading and query.
[0255] Data is regularly loaded from the original log files into the ODS layer. The data loading process uses an incremental loading mechanism, transferring and loading only new or updated data, reducing data transmission volume and loading time.
[0256] Technical Details: CDC (Change Data Capture) technology is used to capture data changes in source systems in real time, enabling incremental loading. Data compression and encoding technologies are used to optimize data storage and reduce storage space usage. Data partitioning strategies are implemented to divide data by time, business area, and other dimensions, improving query performance.
[0257] Data cleaning and format conversion are performed at the DWD layer. Data cleaning involves removing duplicate records, addressing missing values, and correcting outliers to ensure data accuracy and completeness. Format conversion standardizes data into a standardized format for subsequent processing and analysis.
[0258] Technical Details: Utilizes the Flink distributed computing model to achieve parallel cleaning and transformation of large-scale data. Data quality monitoring tools monitor data cleaning results in real time, enabling timely identification and resolution of issues. A data version control mechanism is employed to save different versions of cleaning results, facilitating data backtracking and comparative analysis.
[0259] Data association and integration are performed at the DWS layer to generate themed log records. By associating different data tables and fields, business-meaningful themed datasets are generated to support multi-dimensional analysis and decision-making.
[0260] Technical Details: Utilize graph database technologies (such as Neo4j) to build data relationship maps, visually displaying complex connections between data. Leverage data aggregation and summarization techniques to generate multi-granular data views to meet analytical needs at different levels. Develop data lineage and impact analysis tools to track the flow and impact of data during the integration process.
[0261] The ADS layer provides an efficient query interface to support real-time data analysis and visualization. Data at this layer is optimized for real-time or near-real-time analysis scenarios, such as network traffic monitoring and anomaly detection. The query interface supports standard SQL queries as well as custom analysis functions, facilitating complex data analysis and report generation.
[0262] Technical Details: Utilizes an OLAP engine to accelerate complex query responses and support rapid analysis of large-scale data. Develops an intelligent query optimizer to automatically select the optimal query path and improve query efficiency. Integrates a data visualization tool (Tableau) to enable intuitive display and interactive exploration of analysis results.
[0263] In summary, this application provides the following ideas for building a data warehouse:
[0264] Data warehouse construction: A subject-oriented data warehouse was built, and a star schema was used for multi-dimensional data analysis.
[0265] Dimension design: Time, geography, user, and query type dimensions are designed to analyze log data from multiple perspectives.
[0266] Efficient data indexing mechanism: Implements efficient data indexing, including hash and B-tree indexing, to optimize query efficiency.
[0267] Security enhancement measures: Data encryption and access control are implemented to enhance data security.
[0268] Data compression and processing optimization: The application of data compression and distributed architecture improves data processing and storage efficiency.
[0269] Data integration and correlation analysis: By setting correlation conditions, DNS and AAA log data can be integrated for easy analysis.
[0270] Business domain and usage hierarchical structure design: A multi-layer data architecture from ODS to ADS was designed to meet data processing requirements at different levels.
[0271] Specifically:
[0272] This application proposes a multi-dimensional model design for data warehouses and multi-dimensional modeling to optimize query performance and data aggregation efficiency.
[0273] This application divides and labels data in multiple dimensions, designing multiple dimensions such as time, geographic location, user type, and query type to support analysis of DNS and AAA log data from different perspectives, and identify user behavior patterns and query trends through classification algorithms and statistical analysis.
[0274] This application develops an efficient data indexing mechanism, including hash tables and B-tree index structures, as well as creating dedicated indexes for frequently queried fields and adopting incremental indexing strategies to speed up data retrieval and reduce query response time.
[0275] This application implements data encryption and role-based access control (RBAC) mechanisms to ensure data confidentiality and prevent unauthorized access.
[0276] This application introduces a data compression mechanism, uses the Huffman coding algorithm to reduce storage space occupancy, and adopts a distributed data warehouse architecture to meet the storage and processing needs of massive data.
[0277] This application sets association conditions to achieve the association of DNS logs and AAA logs through time dimension, port and IP address matching, and generates unified log records to support in-depth analysis.
[0278] This application designs a hierarchical structure including operational data store (ODS), data warehouse detail layer (DWD), data warehouse subject layer (DWS) and application data store (ADS) to support the whole process of data management and optimization from raw data to real-time analysis.
[0279] This application improves the real-time and accuracy of data.
[0280] Real-time data processing: By using real-time stream processing technology, this application can collect and process data as it is generated, greatly improving the real-time nature of the data.
[0281] Data cleaning and analysis: During the data collection process, an efficient analysis and cleaning mechanism is adopted to ensure the accuracy and consistency of the data.
[0282] This application achieves data integration and enhanced analysis capabilities.
[0283] Multi-source data integration: This application achieves seamless integration of DNS logs and AAA logs. By correlating key fields such as IP addresses and time intervals, it unifies information from different data sources and provides a comprehensive perspective.
[0284] Rich geographic information: By introducing IP address ownership information, it provides more detailed geographic dimension support for analysis and decision-making.
[0285] This application achieves automation and improves scheduling efficiency.
[0286] Scheduled task scheduling: Using scheduling tools, we automated daily data integration tasks, reducing manual intervention (daily manual tasks decreased by 92%) and ensuring the timeliness and stability of data processing. Resource optimization: Dynamic scheduling algorithms improve cluster resource utilization, supporting the processing of an average of 580TB of logs per day.
[0287] This application implements data storage and query optimization.
[0288] Data Warehouse: Utilizing Hive for data storage and management fully leverages its strengths in large-scale data processing and querying, supporting efficient data storage, management, and analysis, enabling organized storage of petabyte-level data. A well-designed Hive table structure ensures efficient data storage and fast querying, improving analytical processing performance. Query Acceleration: A rational partitioning strategy reduces typical query response time to 1 / 15th of the original solution.
[0289] This application achieves scalability and maintainability. It enhances system scalability. It achieves module decoupling: acquisition, processing, and storage modules can be independently scaled, supporting a smooth evolution from a single cluster to a hybrid cloud architecture. It also achieves elastic capacity expansion: the native scalability of big data components ensures linear throughput growth (120TB → 580TB daily processing).
[0290] Modular design: The data warehouse adopts a modular design, and each link of data collection, analysis, processing, storage and display is relatively independent, which facilitates system maintenance and expansion.
[0291] Scalable architecture: The system architecture based on the big data technology stack, such as Flink, Hive, and HDFS, supports large-scale data processing and storage, has good scalability, and can adapt to the rapid growth of data volume.
[0292] See also Figure 9 The present application also provides a system for establishing a log data warehouse, which can implement the above-mentioned method for establishing a log data warehouse. The system includes:
[0293] The first module 910 is used to collect DNS log data and AAA log data;
[0294] The second module 920 is used to pre-process the DNS log data and the AAA log data, and perform correlation analysis on the DNS log data and the AAA log data to generate integrated log data;
[0295] The third module 930 is used to determine the dimension table of the data warehouse according to business requirements and integrated log data, and store the integrated log data in the data warehouse.
[0296] In some embodiments, the system provided by the embodiments of the present application further includes a fourth module for:
[0297] Determine the hash value of the integrated log data and create a hash table;
[0298] Alternatively, based on historical query records, frequently queried fields are determined, and dedicated indexes corresponding to the frequently queried fields are created;
[0299] Alternatively, new and updated data in the data warehouse is recorded and incremental indexes of the new and updated data are created.
[0300] In some embodiments, the system provided by the embodiments of the present application further includes a fifth module for:
[0301] Encrypt the integrated log data;
[0302] Alternatively, the permissions of the access user are determined based on the access user's role.
[0303] In some embodiments, the system provided by the embodiments of the present application further includes a sixth module for:
[0304] The integrated log data is compressed using the Huffman coding algorithm;
[0305] Establish a distributed data warehouse architecture.
[0306] In some embodiments, the system provided by the embodiments of the present application further includes a seventh module for:
[0307] Establish the data architecture of the log data warehouse;
[0308] The data architecture includes:
[0309] The basic layer is used to store the original collected DNS log data and AAA log data;
[0310] The middle layer is used to pre-process DNS log data and AAA log data;
[0311] The subject layer is used to perform correlation analysis on DNS log data and AAA log data to generate integrated log data;
[0312] The application layer is used to provide data query and data analysis.
[0313] It can be understood that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0314] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-mentioned method for establishing a log data warehouse. The electronic device can be any smart terminal including a tablet computer, an in-vehicle computer, or the like.
[0315] It can be understood that the contents of the above method embodiments are applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0316] See also Figure 10 , Figure 10 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:
[0317] The processor 901 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0318] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902, and the processor 901 calls and executes the method for establishing a log data warehouse in the embodiments of this application;
[0319] Input / output interface 903, used to implement information input and output;
[0320] Communication interface 904, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0321] Bus 905 , which transmits information between various components of the device (e.g., processor 901 , memory 902 , input / output interface 903 , and communication interface 904 );
[0322] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .
[0323] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned method for establishing a log data warehouse.
[0324] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiment, the functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0325] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0326] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0327] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0328] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0329] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0330] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0331] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0332] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0333] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0334] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0335] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0336] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A method for establishing a log data warehouse, characterized in that: The method comprises: Collect DNS log data and AAA log data; Preprocessing the DNS log data and the AAA log data, and performing correlation analysis on the DNS log data and the AAA log data to generate integrated log data; According to business requirements and the integrated log data, a dimension table of a data warehouse is determined, and the integrated log data is stored in the data warehouse.
2. The method according to claim 1, characterized in that The performing correlation analysis on the DNS log data and the AAA log data to generate integrated log data includes: Correlating the DNS log data with the AAA log data according to the timestamp of the DNS log data and the session time of the AAA log data to obtain a correlation result; Alternatively, correlating the DNS log data and the AAA log data according to the port number of the DNS log data and the port range of the AAA log data to obtain a correlation result; Alternatively, according to the IP address of the DNS log data and the IP address of the AAA log data, the DNS log data and the AAA log data are associated to obtain an association result; Based on the association result, integrated log data including preset log fields is generated.
3. The method according to claim 1, characterized in that Determining the dimension table of the data warehouse based on business requirements and the integrated log data includes: Determining a time dimension table according to timestamp information associated with the DNS log data or the AAA log data; Determine a geographic location dimension table based on the IP addresses associated with the DNS log data and in association with the geographic locations; Classify users and user behaviors to determine a user type dimension table; the users are related to the DNS log data or the AAA log data; Classify and process the query request of the DNS log data or the AAA log data to determine a query type dimension table; The time dimension table, the geographic location dimension table, the user type dimension table, and the query type dimension table are aggregated to obtain a dimension table of the data warehouse.
4. The method according to claim 1, wherein The method further comprises: Determine the hash value of the integrated log data and create a hash table; Alternatively, frequently queried fields are determined based on historical query records, and dedicated indexes corresponding to the frequently queried fields are established; Alternatively, new data and updated data in the data warehouse are recorded, and incremental indexes of the new data and updated data are created.
5. The method according to claim 1, wherein The method further comprises: encrypting the integrated log data; Alternatively, the permissions of the access user are determined based on the role of the access user.
6. The method according to claim 1, characterized in that The method further comprises: Compressing the integrated log data using a Huffman coding algorithm; Establish a distributed data warehouse architecture.
7. The method according to claim 1, characterized in that The method further comprises: Establish the data architecture of the log data warehouse; The data architecture includes: A basic layer, used to store the originally collected DNS log data and the AAA log data; An intermediate layer, configured to pre-process the DNS log data and the AAA log data; The subject layer is used to perform correlation analysis on the DNS log data and the AAA log data to generate integrated log data; The application layer is used to provide data query and data analysis.
8. A system for establishing a log data warehouse, characterized in that: The system comprises: The first module is used to collect DNS log data and AAA log data; The second module is used to pre-process the DNS log data and the AAA log data, and perform correlation analysis on the DNS log data and the AAA log data to generate integrated log data; The third module is used to determine the dimension table of the data warehouse according to business requirements and the integrated log data, and store the integrated log data in the data warehouse.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.