Log analysis method and system under object storage item

By calling object storage services and Hadoop clusters in object storage projects and combining Spark engine to calculate the log data, the problem of log analysis under object storage projects is solved, and data support and security monitoring of the business platform is achieved.

CN120045534APending Publication Date: 2025-05-27AISINO CORPORATION
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411928218.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively perform log analysis under object storage projects, and cannot meet the business system's mining of log commercial value and solving the metric tasks.

Method used

By calling the object storage service, the business log data is stored in the Hadoop cluster and the target business log data is uploaded to the distributed file system HDFS. Then, the Spark engine is used to calculate the target business log data in the Hive data table, obtain the indicator data and store it in the Hive warehouse, and finally generate a log analysis report.

Benefits of technology

It realizes log analysis under object storage projects, can provide data support for the business platform, realizes security and compliance audits, protects business-sensitive data, and grasps various business status offline, and conducts security monitoring of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045534A_ABST
    Figure CN120045534A_ABST
Patent Text Reader

Abstract

The invention discloses a log analysis method and system under an object storage project, and the method comprises the steps: calling an object storage service to store business data in a Hadoop cluster, obtaining business log data generated by a calling operation, and storing the business log data in a Tomcat container; checking the service log data in the Tomcat container to obtain target service log data, and uploading the target service log data to an HDFS (Hadoop Distributed File System); loading the HDFS data into a Hive data table, carrying out index calculation on target business log data in the Hive data table by utilizing a Spark engine to obtain index data, and storing the index data into a Hive warehouse; and exporting the index data in the Hive warehouse to other databases, and generating a log analysis report based on the index data. According to the method, data support can be provided for marketing, sales strategies and the like of a service platform, service sensitive data are protected, each service state is mastered offline, and safety monitoring is carried out on the data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of log analysis, and more particularly, to a log analysis method and system under an object storage project. Background Art

[0002] Log analysis is the process of extracting valuable information from log data generated by computer systems, network devices, application programs, etc. Log analysis plays an important role in troubleshooting, security monitoring, performance optimization, and compliance requirements.

[0003] The types of logs include: system logs, application logs, and network logs. System logs record the activities of the operating system, including system startup and shutdown times, kernel messages, activities of device drivers, etc. In the Linux system, a common system log file is / var / log / messages, which contains various important information during the system operation, such as hardware failure notifications, startup and stop records of system services. Application logs are generated by various application programs and are used to record the running status of the application programs and user operations. For example, an enterprise resource planning (ERP) software will record operations such as user logins, module accesses, and data modifications. Taking a customer relationship management (CRM) system as an example, its logs may record the update operations of salespersons on customer information, including modifying customer contact information, adding sales opportunities, etc. These logs help to track the activity trajectories of users within the application. Network logs mainly involve the activities of network devices (such as routers, firewalls). Firewall logs will record information such as the source IP address, destination IP address, port number, protocol type of network connections, and whether the connections are allowed. For network administrators, by analyzing these logs, they can monitor the direction of network traffic and promptly discover abnormal network connection behaviors.

[0004] Through log analysis, text search, statistical analysis, correlation analysis, and visualization analysis can be performed. Prior art: A device for batch collecting logs based on a Web application, including: a log container instantiation unit for instantiating a log container to serve as a bridge for communication between front-end code and local storage; a log collection unit for collecting logs during the operation of the Web application, collecting and organizing the pre-determined status information to be recorded, and writing it into the log container; a log container writing unit for using local storage as a temporary log storage space to temporarily write the serialized log text into local storage; a log export unit for traversing the log records in the log container and exporting them as text files or exporting them to a log analysis server.

[0005] For object storage projects, common open-source technologies are mainly used on the distributed file system (Ceph) to implement data storage and query functions. The purpose of the log analysis service in this context is to explore the commercial value of logs and solve the metric tasks concerned by the business system.

[0006] Therefore, a log analysis method for object storage projects is needed. Summary of the Invention

[0007] The present invention proposes a log analysis method and system for object storage projects to solve the problem of how to implement log analysis for object storage projects.

[0008] To solve the above problems, according to one aspect of the present invention, a log analysis method for object storage projects is provided. The method includes:

[0009] Call the object storage service to store business data in the Hadoop cluster, obtain the business log data generated by the call operation, and store the business log data in the Tomcat container;

[0010] Check the business log data in the Tomcat container to obtain the target business log data, and upload the target business log data to the distributed file system HDFS;

[0011] Load the HDFS data into the Hive data table, and use the Spark engine to calculate metrics for the target business log data in the Hive data table, obtain the metric data, and store it in the Hive warehouse;

[0012] Export the metric data in the Hive warehouse to other databases, and generate a log analysis report based on the metric data.

[0013] Preferably, the business log data is stored in the Tomcat container in text form.

[0014] Preferably, the checking the business log data in the Tomcat container to obtain the target business log data includes:

[0015] Use a shell script program to send SSH commands regularly or at preset time intervals to check the Tomcat container based on the SSH commands, and filter out the business log data corresponding to the call of the object storage service to obtain the target business log data.

[0016] Preferably, the loading the HDFS data into the Hive data table includes:

[0017] Create the corresponding table structure in Hive according to the data format and business requirements to receive the data loaded from HDFS;

[0018] Use the LOAD DATA statement to load the HDFS data into the Hive data table.

[0019] Preferably, the method further includes:

[0020] Before calculating the metrics for the target business log data in the Hive data table, perform an ELT operation on the data in the Hive data table.

[0021] Preferably, calculating the metrics for the target business log data in the Hive data table using the Spark engine to obtain the metric data includes:

[0022] Create a SparkSession object and use the spark.sql() method through the SparkSession object to execute SQL statements to read the target business log data from the Hive table;

[0023] Perform data cleaning and transformation operations on the target business log data;

[0024] Based on the target business log data that has undergone data cleaning and transformation operations, calculate the metrics according to the requirements to obtain the metric data.

[0025] Preferably, the method further includes:

[0026] Generate an analysis report based on the log analysis report according to the template and send the analysis report to the requester according to the preset file sending method.

[0027] According to another aspect of the present invention, a log analysis system under an object storage project is provided. The system includes:

[0028] A business log data acquisition unit, configured to call an object storage service to store business data in a Hadoop cluster, acquire the business log data generated by the call operation, and store the business log data in a Tomcat container;

[0029] A target business log data acquisition unit, configured to check the business log data in the Tomcat container to obtain the target business log data, and upload the target business log data to the distributed file system HDFS;

[0030] An index data calculation unit, configured to load HDFS data into a Hive data table, and use the Spark engine to perform index calculations on the target business log data in the Hive data table, obtain index data, and store it in the Hive warehouse;

[0031] A log analysis report determination unit, configured to export the index data in the Hive warehouse to other databases, and generate a log analysis report based on the index data.

[0032] Preferably, the business log data is stored in the Tomcat container in text form.

[0033] Preferably, the target business log data acquisition unit checks the business log data in the Tomcat container to obtain target business log data, including:

[0034] Using a shell script program to send SSH commands regularly or at preset time intervals to check the Tomcat container based on the SSH commands, and filtering out the business log data corresponding to the call object storage service to obtain the target business log data.

[0035] Preferably, the index data calculation unit loads HDFS data into a Hive data table, including:

[0036] According to the data format and business requirements, create a corresponding table structure in Hive to receive the data loaded from HDFS;

[0037] Use the LOAD DATA statement to load the HDFS data into the Hive data table.

[0038] Preferably, the index data calculation unit further includes:

[0039] Before performing index calculations on the target business log data in the Hive data table, perform an ELT operation on the data in the Hive data table.

[0040] Preferably, the index data calculation unit uses the Spark engine to perform index calculations on the target business log data in the Hive data table to obtain index data, including:

[0041] Create a SparkSession object, and use the spark.sql() system through the SparkSession object to execute SQL statements to read the target business log data from the Hive table;

[0042] Perform data cleaning and transformation operations on the target business log data;

[0043] Based on the target business log data after data cleaning and transformation operations, calculate metrics according to requirements to obtain the metric data.

[0044] Preferably, the system further includes:

[0045] An analysis report generation unit, configured to generate an analysis report based on the log analysis report according to a template, and send the analysis report to the requester according to a preset file sending method.

[0046] The present invention provides a log analysis method and system under an object storage project, including: calling an object storage service to store business data in a Hadoop cluster, obtaining business log data generated by the call operation, and storing the business log data in a Tomcat container; checking the business log data in the Tomcat container to obtain target business log data, and uploading the target business log data to a distributed file system HDFS; loading HDFS data into a Hive data table, and using a Spark engine to calculate metrics for the target business log data in the Hive data table to obtain metric data, and storing it in a Hive warehouse; exporting the metric data in the Hive warehouse to other databases, and generating a log analysis report based on the metric data. The log analysis method based on the object storage project of the present invention can provide data support for marketing, sales strategies, etc. of a business platform, achieve secure and compliant auditing, protect business sensitive data, master each business status offline, and perform security monitoring on data. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] By referring to the following drawings, the exemplary embodiments of the present invention can be more completely understood:

[0048] Figure 1 FIG. 100 is a flowchart of a log analysis method under an object storage project according to an embodiment of the present invention;

[0049] Figure 2 FIG. 24 is a schematic diagram of a common open source technology architecture adopted according to an embodiment of the present invention;

[0050] Figure 3 FIG. 28 is a schematic diagram of a log analysis process according to an embodiment of the present invention;

[0051] Figure 4 FIG. 32 is a schematic diagram of the structure of a log analysis system 400 under an object storage project according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] Reference is now made to the accompanying drawings to describe exemplary embodiments of the present invention. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. These embodiments are provided to disclose the present invention in detail and completely, and to fully convey the scope of the present invention to those skilled in the art. The terms in the exemplary embodiments shown in the drawings are not intended to limit the present invention. In the drawings, the same units / components are denoted by the same reference numerals.

[0053] Unless otherwise specified, the terms used herein (including technical terms) have the ordinary meaning understood by those skilled in the art. In addition, it can be understood that the terms defined in a commonly used dictionary should be understood to have a meaning consistent with the context of their relevant fields, and should not be understood as idealized or overly formal meanings.

[0054] Figure 1 FIG. 100 is a flowchart of a log analysis method for an object storage project according to an embodiment of the present invention. As Figure 1 shown, the log analysis method for an object storage project provided by the embodiment of the present invention, based on the log analysis method for an object storage project, can provide data support for marketing, sales strategies, etc. of a business platform, implement security and compliance audits, protect business-sensitive data, offline master various business states, and perform security monitoring on data. The log analysis method 100 for an object storage project provided by the embodiment of the present invention starts at step 101. At step 101, an object storage service is called to store business data in a Hadoop cluster, business log data generated by the call operation is obtained, and the business log data is stored in a Tomcat container.

[0055] Preferably, the business log data is stored in the Tomcat container in text form.

[0056] At step 102, the business log data in the Tomcat container is checked to obtain target business log data, and the target business log data is uploaded to a distributed file system HDFS.

[0057] Preferably, the checking of the business log data in the Tomcat container to obtain target business log data includes:

[0058] Using a shell script program to send SSH commands regularly or at a preset time interval to check the Tomcat container based on the SSH commands, and filtering out the business log data corresponding to the call of the object storage service to obtain the target business log data.

[0059] In step 103, load the HDFS data into the Hive data table, and use the Spark engine to calculate metrics for the target business log data in the Hive data table, obtain the metric data, and store it in the Hive warehouse.

[0060] Preferably, the loading of the HDFS data into the Hive data table includes:

[0061] Create a corresponding table structure in Hive according to the data format and business requirements to receive the data loaded from HDFS;

[0062] Use the LOAD DATA statement to load the HDFS data into the Hive data table.

[0063] Preferably, the method further includes:

[0064] Before calculating the metrics for the target business log data in the Hive data table, perform an ELT operation on the data in the Hive data table.

[0065] Preferably, the using of the Spark engine to calculate metrics for the target business log data in the Hive data table and obtain the metric data includes:

[0066] Create a SparkSession object, and use the spark.sql() method through the SparkSession object to execute SQL statements to read the target business log data from the Hive table;

[0067] Perform data cleaning and transformation operations on the target business log data;

[0068] Based on the target business log data that has undergone data cleaning and transformation operations, calculate metrics according to requirements to obtain the metric data.

[0069] In step 104, export the metric data in the Hive warehouse to other databases, and generate a log analysis report based on the metric data.

[0070] Preferably, the method further includes:

[0071] Generate an analysis report based on the log analysis report according to a template, and send the analysis report to the requester according to a preset file sending method.

[0072] The method of the present invention is based on as Figure 2The implementation of the open-source technology architecture shown. The business system data consists of two parts: one is the business data of the business system, and the other is the business log data generated due to calls. The business data is stored on the Hadoop cluster by calling the object storage service. The business log data generated due to calls is stored in the Tomcat container. The data focused on analyzing in the present invention is based on the latter. The specific implementation process is as follows:

[0073] Step 1: Call the object storage service to generate business log data in the Tomcat container and store it in text form.

[0074] Step 2: The script periodically checks the Tomcat container and obtains the target log file. Among them, as shown in Figure 3 shown, use the shell script program to send SSH commands periodically or at preset time intervals to check the Tomcat container based on the SSH commands, and filter out the business log data corresponding to the call of the object storage service to obtain the target business log data.

[0075] Step 3: Upload the obtained target log file to HDFS.

[0076] Step 4: Load the HDFS data into the Hive data table and perform ETL-related operations. Specifically, according to the data format and business requirements, create the corresponding table structure in Hive to receive the data loaded from HDFS; use the LOAD DATA statement to load the HDFS data into the Hive data table, and perform ELT operations on the data in the Hive data table.

[0077] Step 5: Use the Spark engine to calculate the log data in Hive and store the calculated metrics. Specifically, create a SparkSession object, and use the spark.sql() method through the SparkSession object to execute SQL statements to read the target business log data from the Hive table; perform data cleaning and transformation operations on the target business log data; based on the target business log data after data cleaning and transformation operations, calculate metrics according to requirements to obtain the metric data.

[0078] Step 6: Export the metric data in the Hive warehouse to the MySQL database and generate a report, and send the report to the requester by email.

[0079] Figure 4 It is a schematic structural diagram of the log analysis system 400 under the object storage project according to the embodiment of the present invention. As shown in Figure 4As shown in the figure, the log analysis system 400 under the object storage project provided by the embodiment of the present invention includes: a business log data acquisition unit 401, a target business log data acquisition unit 402, a metric data calculation unit 403, and a log analysis report determination unit 404.

[0080] Preferably, the business log data acquisition unit 401 is used to call the object storage service to store business data in the Hadoop cluster, acquire the business log data generated by the call operation, and store the business log data in the Tomcat container.

[0081] Preferably, the target business log data acquisition unit 402 is used to check the business log data in the Tomcat container to acquire the target business log data, and upload the target business log data to the distributed file system HDFS.

[0082] Preferably, the metric data calculation unit 403 is used to load the HDFS data into the Hive data table, perform metric calculations on the target business log data in the Hive data table using the Spark engine, acquire the metric data, and store it in the Hive warehouse.

[0083] Preferably, the log analysis report determination unit 404 is used to export the metric data in the Hive warehouse to other databases and generate a log analysis report based on the metric data.

[0084] Preferably, the business log data is stored in the Tomcat container in text form.

[0085] Preferably, the target business log data acquisition unit checks the business log data in the Tomcat container to acquire the target business log data, including:

[0086] Using a shell script program to send SSH commands regularly or at preset time intervals to check the Tomcat container based on the SSH commands, and filtering out the business log data corresponding to the call of the object storage service to acquire the target business log data.

[0087] Preferably, the metric data calculation unit loads the HDFS data into the Hive data table, including:

[0088] According to the data format and business requirements, create the corresponding table structure in Hive to receive the data loaded from HDFS;

[0089] Use the LOAD DATA statement to load the HDFS data into the Hive data table.

[0090] Preferably, the index data calculation unit further includes:

[0091] Before calculating the index for the target business log data in the Hive data table, perform an ELT operation on the data in the Hive data table.

[0092] Preferably, the index data calculation unit uses the Spark engine to calculate the index for the target business log data in the Hive data table to obtain index data, including:

[0093] Create a SparkSession object and use the spark.sql() system through the SparkSession object to execute SQL statements to read the target business log data from the Hive table;

[0094] Perform data cleaning and transformation operations on the target business log data;

[0095] Based on the target business log data that has undergone data cleaning and transformation operations, perform index calculations according to requirements to obtain the index data.

[0096] Preferably, the system further includes:

[0097] An analysis report generation unit for generating an analysis report based on the log analysis report according to a template and sending the analysis report to the requester according to a preset file sending method.

[0098] The log analysis system 400 under the object storage project in the embodiment of the present invention corresponds to the log analysis method 100 under the object storage project in another embodiment of the present invention, and will not be elaborated here.

[0099] The present invention has been described by referring to a few embodiments. However, as is known to those skilled in the art, as defined by the appended patent claims, other embodiments equivalent to those disclosed above of the present invention equally fall within the scope of the present invention.

[0100] Generally, all terms used in the claims are interpreted according to their ordinary meanings in the technical field, unless otherwise clearly defined therein. All references to "a / the / this [device, component, etc.]" are open to interpretation as at least one instance of the device, component, etc., unless otherwise clearly stated. The steps of any method disclosed here need not be run in the exact order disclosed, unless clearly stated.

[0101] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0102] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0103] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0104] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: still can modify the specific implementation manners of the present invention or make equivalent replacements, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. A log analysis method under an object storage project, characterized in that: The method comprises: Call the object storage service to store the business data in the Hadoop cluster, obtain the business log data generated by the call operation, and store the business log data in the Tomcat container; Check the business log data in the Tomcat container to obtain target business log data, and upload the target business log data to the distributed file system HDFS; Load HDFS data into the Hive data table, and use the Spark engine to calculate the indicators of the target business log data in the Hive data table, obtain the indicator data, and store it in the Hive warehouse; The indicator data in the Hive warehouse is exported to other databases, and a log analysis report is generated based on the indicator data.

2. The method according to claim 1, characterized in that The business log data is stored in the Tomcat container in text form.

3. The method according to claim 1, characterized in that The checking of the business log data in the Tomcat container to obtain target business log data includes: The SSH command is sent regularly or at a preset time interval by using a shell script program to check the Tomcat container based on the SSH command, and the business log data corresponding to the call object storage service is filtered out to obtain the target business log data.

4. The method according to claim 1, characterized in that: The loading of HDFS data into the Hive data table includes: Create the corresponding table structure in Hive according to the data format and business requirements to receive the data loaded from HDFS; Use the LOAD DATA statement to load HDFS data into the Hive data table.

5. The method according to claim 1, characterized in that The method further comprises: Before performing indicator calculation on the target business log data in the Hive data table, an ELT operation is performed on the data in the Hive data table.

6. The method according to claim 1, characterized in that The using of the Spark engine to perform index calculation on the target business log data in the Hive data table to obtain the index data includes: Create a SparkSession object, and use the spark.sql() method through the SparkSession object to execute SQL statements to read target business log data from the Hive table; Performing data cleaning and conversion operations on the target business log data; Based on the target business log data that has undergone data cleaning and conversion operations, indicator calculation is performed according to demand to obtain the indicator data.

7. The method according to claim 1, characterized in that The method further comprises: An analysis report is generated according to a template based on the log analysis report, and the analysis report is sent to the demander according to a preset file sending method.

8. A log analysis system under an object storage project, characterized in that: The system comprises: A business log data acquisition unit is used to call the object storage service to store the business data in the Hadoop cluster, obtain the business log data generated by the call operation, and store the business log data in the Tomcat container; A target business log data acquisition unit is used to check the business log data in the Tomcat container to acquire the target business log data, and upload the target business log data to the distributed file system HDFS; An indicator data calculation unit is used to load HDFS data into a Hive data table, and use a Spark engine to perform indicator calculation on target business log data in the Hive data table, obtain indicator data, and store it in a Hive warehouse; The log analysis report determination unit is used to export the indicator data in the Hive warehouse to other databases and generate a log analysis report based on the indicator data.

9. The system according to claim 8, characterized in that The target business log data acquisition unit checks the business log data in the Tomcat container to acquire the target business log data, including: The SSH command is sent regularly or at a preset time interval by using a shell script program to check the Tomcat container based on the SSH command, and the business log data corresponding to the call object storage service is filtered out to obtain the target business log data.

10. The system according to claim 8, characterized in that The indicator data calculation unit loads the HDFS data into the Hive data table, including: Create the corresponding table structure in Hive according to the data format and business requirements to receive the data loaded from HDFS; Use the LOAD DATA statement to load HDFS data into the Hive data table.

11. The system according to claim 8, characterized in that The indicator data calculation unit further includes: Before performing indicator calculation on the target business log data in the Hive data table, an ELT operation is performed on the data in the Hive data table.

12. The system according to claim 8, characterized in that The indicator data calculation unit uses the Spark engine to perform indicator calculation on the target business log data in the Hive data table to obtain indicator data, including: Create a SparkSession object, and use the spark.sql() system to execute SQL statements through the SparkSession object to read target business log data from the Hive table; Performing data cleaning and conversion operations on the target business log data; Based on the target business log data that has undergone data cleaning and conversion operations, indicator calculation is performed according to demand to obtain the indicator data.

Citation Information

Patent Citations

  • Method for processing data offline in real time on the basis of Spark big data frame

    CN108874982A

  • Service index obtaining method and device, server and computer readable storage medium

    CN109800225A

  • Spark-based big data weblog acquisition, analysis and early warning method and system

    CN110690984A

  • Object storage flow auditing method and device

    CN112738221A

  • Securing data lakes via object store monitoring

    US20230142344A1