Log processing method, device, platform and system
By compressing log data packets and snapshot packets, the amount of log data is reduced and stored to remote devices, the problem of high log storage costs is solved and efficient and accurate log retrieval is achieved.
Patent Information
- Application Number
- CN202311498201.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-10
- Publication Date
- 2025-05-13
AI Technical Summary
The current log data storage volume is large, resulting in high hardware requirements and high storage costs.
By obtaining the compressed packets of log data packets and the compressed packets of snapshot packets, the amount of log data is reduced and the compressed data is stored in a remote storage device, reducing the demand for local hardware.
Reduces log storage costs and improves the efficiency and accuracy of log retrieval.
Smart Images

Figure CN119988341A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a log processing method, device, platform and system. Background Art
[0002] The process of auditing logs is to collect data information in the information system in a centralized manner, such as system security events, user access records, system operation logs, system operation status and other types of information, and then store and manage the data in a unified format of logs after standardized data processing, data filtering, data merging and data alarm analysis. Based on the correlation analysis of logs, a comprehensive audit of data information in the information system can be achieved.
[0003] At present, with the popularization of cloud computing technology and the rise of traffic-based applications, the amount of log data storage is relatively large, and the information security level protection clearly requires that logs need to be kept for a long time for later retrieval of data information. At present, logs are mainly stored and analyzed based on the search and analysis technology stack (composed of ElasticSearch, Logstash, and Kibana, referred to as the ELK technology stack), but when storing the full amount of logs based on ElasticSearch, in order to improve the performance of data retrieval, it is required to use a solid state drive (SSD) or a high input / output (I / O) disk. In addition, in order to ensure the reliability of logs, a copy mechanism is also required, that is, log backup storage, which increases the amount of log data storage. Due to the large amount of log data storage and high hardware requirements, the storage cost of logs is high. Summary of the invention
[0004] The present application provides a log processing method, device, platform and system, which can reduce the storage cost of logs.
[0005] In a first aspect, a log processing method is provided, which is applied to a log processing platform, and the method may include: obtaining a compressed package of a log data packet and a compressed package of a snapshot package of the log data packet, wherein the log data packet includes at least one log information, and one log information includes a parameter and a parameter value, and the parameters included in different log information are the same; then, sending the compressed package of the log data packet and the compressed package of the snapshot package of the log data packet to a remote storage device; wherein the package name of the compressed package of the log data packet includes a first key value, and the package name of the compressed package of the snapshot package includes a second key value, and the first key value and the second key value are key values corresponding to the dimensionality reduction key of the log data packet.
[0006] Through the above method, by obtaining a compressed package of the log data packet and a compressed package of the snapshot package of the log data packet, the amount of log data can be reduced. Subsequently, the log data packet and the snapshot package are stored through a remote storage device, that is, the compressed log data packet and the snapshot package are stored in other storage devices outside the log processing platform, thereby reducing the requirements for the hardware locally deployed on the log processing platform, thereby reducing the storage cost of the log.
[0007] In one possible design, the method also includes: obtaining log retrieval information, which is used to indicate the key to be retrieved and the key value to be retrieved; then, querying whether there is a target dimensionality reduction key corresponding to the dimensionality reduction index key of the log retrieval information; the dimensionality reduction index key is obtained based on the key to be retrieved and the key value to be retrieved; thereby, in the case where there is a target dimensionality reduction key, based on the target dimensionality reduction key, a search is performed to obtain log information corresponding to the log retrieval information from a compressed package of a log data packet stored in a remote storage device, and a compressed package of a snapshot package of the log data packet.
[0008] In this way, when the log processing platform needs to retrieve log information, it can obtain the dimensionality reduction index key based on the log retrieval information used to indicate the key to be retrieved and the key value to be retrieved, and then query the corresponding target dimensionality reduction key based on the dimensionality reduction index key, so that when the target dimensionality reduction key is retrieved, the required log information can be obtained from the remote storage device based on the target dimensionality reduction key, so that based on the log retrieval information, the corresponding log information can be retrieved step by step through the dimensionality reduction index key, the target dimensionality reduction key, and the compressed package of the log data packet, thereby improving the efficiency and accuracy of retrieving log information.
[0009] In another possible design, the target dimensionality reduction key corresponds to a first target key value and a second target key value; based on the target dimensionality reduction key retrieval, log information corresponding to the log retrieval information is obtained from a compressed package of a log data packet stored in a remote storage device, and a compressed package of a snapshot package of the log data packet, including: based on the first target key value corresponding to the target dimensionality reduction key, a target snapshot package corresponding to the target dimensionality reduction key is obtained; wherein the first key value included in the package name of the compressed package of the target snapshot package is the first target key value; then, when the target snapshot package matches the log retrieval information, based on the second target key value corresponding to the target dimensionality reduction key, a target log data packet is obtained; the second key value included in the package name of the compressed package of the target log data packet is the second target key value; and then, the log information corresponding to the log retrieval information is determined from the target log data packet.
[0010] In this way, the target snapshot package can be retrieved based on the first target key value corresponding to the target dimension reduction key, and when it is determined that the target snapshot package matches the log retrieval information, the target log data packet can be further obtained based on the second target key value corresponding to the target dimension reduction key, so that the log information corresponding to the log retrieval information can be directly determined from the target log data packet. The amount of data processed during the retrieval process can be reduced, and the corresponding log data packet can be determined by retrieving the determined log data packet to determine the log information corresponding to the log retrieval information. There is no need to retrieve other log data packets, which improves the efficiency of retrieving log information.
[0011] In another possible design, the reduced dimensionality keys are stored in a NoSQL database.
[0012] In another possible design, the association between the dimension reduction key of the log data packet and the address information of the remote storage device is saved.
[0013] In this way, by storing the dimension reduction key locally in the log processing platform, storing the compressed package of the log data packet and the compressed package of the snapshot package of the log data packet in the remote storage device, the storage cost of the log is reduced through the hierarchical storage method. In addition, based on the locally stored dimension reduction key and the association relationship between the dimension reduction key and the address information of the remote storage device, the corresponding log data packet can be accurately obtained, thereby improving the efficiency of retrieving log information.
[0014] In another possible design, the log information is packaged based on a preset data volume to obtain multiple log data packets, and one log data packet includes log information of the preset data volume.
[0015] In another possible design, based on the generation time of each log information, the log information is packaged to obtain multiple log data packets, and one log data packet includes log information generated within each period of a preset time interval.
[0016] In this way, the log processing platform can package the log information into multiple log data packets based on a preset data volume or a preset time period. Therefore, in the process of retrieving log information, the log data packet corresponding to the log information can be determined first, so that the corresponding log information can be retrieved by searching only a certain log data packet, which can improve the retrieval efficiency.
[0017] In a second aspect, a log processing device is provided. The log processing device is applied to a log processing platform. The log processing device includes: a communication module and a processing module.
[0018] The communication module is used to obtain a compressed package of a log data packet and a compressed package of a snapshot package of the log data packet, wherein the log data packet includes at least one log information, and one log information includes a parameter and a parameter value, and the parameters included in different log information are the same.
[0019] The above-mentioned communication module is also used to send a compressed package of the log data packet and a compressed package of the snapshot package of the log data packet to a remote storage device; wherein the package name of the compressed package of the log data packet includes a first key value, and the package name of the compressed package of the snapshot package includes a second key value, and the first key value and the second key value are key values corresponding to the dimensionality reduction key of the log data packet.
[0020] In a possible design, the communication module is further used to obtain log retrieval information, where the log retrieval information is used to indicate the key to be retrieved and the key value to be retrieved.
[0021] The above-mentioned processing module is used to query whether there is a target dimension reduction key corresponding to the dimension reduction index key of the log retrieval information; the dimension reduction index key is obtained based on the key to be retrieved and the key value to be retrieved.
[0022] The above-mentioned processing module is also used to obtain log information corresponding to the log retrieval information from the compressed package of the log data packet stored in the remote storage device and the compressed package of the snapshot package of the log data packet based on the target dimensionality reduction key retrieval when there is a target dimensionality reduction key.
[0023] In another possible design, the target dimensionality reduction key corresponds to a first target key value and a second target key value; the above-mentioned processing module is specifically used to obtain a target snapshot package corresponding to the target dimensionality reduction key based on the first target key value corresponding to the target dimensionality reduction key; wherein the first key value included in the package name of the compressed package of the target snapshot package is the first target key value.
[0024] The above-mentioned processing module is specifically used to obtain the target log data packet based on the second target key value corresponding to the target dimensionality reduction key when the target snapshot package matches the log retrieval information; the second key value included in the package name of the compressed package of the target log data packet is the second target key value.
[0025] The above-mentioned processing module is specifically used to determine the log information corresponding to the log retrieval information from the target log data packet.
[0026] In another possible design, the reduced dimensionality keys are stored in a NoSQL database.
[0027] In another possible design, the processing module is further used to save the association between the dimension reduction key of the log data packet and the address information of the remote storage device.
[0028] In another possible design, the processing module is further configured to package the log information based on a preset data volume to obtain a plurality of log data packets, wherein one log data packet includes log information of a preset data volume.
[0029] In another possible design, the processing module is further used to package the log information into multiple log data packets based on the generation time of each log information, and one log data packet includes the log information generated within each period of a preset time interval.
[0030] In a third aspect, a log processing platform is provided, which includes a memory and a processor, wherein the memory is coupled to the processor; the memory is used to store computer program code, and the computer program code includes computer instructions; when the computer instructions are executed by the processor, the log processing platform executes the method described in the first aspect and any possible design method thereof.
[0031] In a fourth aspect, a log processing system is provided, which includes a log processing platform and a remote storage device, and the log processing system executes the method described in the first aspect and any possible design thereof.
[0032] In a fifth aspect, a computer storage medium is provided, which includes computer instructions. When the computer instructions are executed on a log processing platform, the log processing platform executes the method described in the first aspect and any possible design thereof.
[0033] According to a sixth aspect, a computer program product is provided. When the computer program product is run on a computer, the computer is caused to execute the method as described in the first aspect and any possible design thereof.
[0034] It can be understood that the beneficial effects that can be achieved by the log processing device described in the second aspect and any possible design thereof, the log processing platform described in the third aspect and any possible design thereof, the log processing system described in the fourth aspect, the computer storage medium described in the fifth aspect, and the computer program product described in the sixth aspect can be referred to as the beneficial effects in the first aspect and any possible design thereof, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 A schematic diagram of a flow chart of storage and retrieval of a related log provided for this application;
[0036] Figure 2 A schematic diagram of the composition of a log processing system architecture provided for this application;
[0037] Figure 3A flow chart of a log processing method provided for this application;
[0038] Figure 4 A flow chart of another log processing method provided for this application;
[0039] Figure 5 A flow chart of another log processing method provided for this application;
[0040] Figure 6 A flow chart of another log processing method provided for this application;
[0041] Figure 7 A schematic diagram of an interface provided for this application;
[0042] Figure 8 A schematic diagram of another log processing device provided by the present application;
[0043] Fig. 9 A schematic diagram of the structural composition of a log processing platform provided for this application. DETAILED DESCRIPTION
[0044] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0045] In this application, the character " / " generally indicates that the objects before and after are in an "or" relationship. For example, A / B can be understood as A or B.
[0046] The terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this embodiment, unless otherwise specified, "plurality" means two or more.
[0047] In addition, the terms "including" and "having" and any variations thereof mentioned in the description of this application are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or modules is not limited to the listed steps or modules, but may optionally include other steps or modules that are not listed, or may optionally include other steps or modules that are inherent to these processes, methods, products or devices.
[0048] In addition, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present concepts in a specific way.
[0049] like Figure 1 As shown, after the cloud service in the cloud platform generates a log (i.e., an audit event), it needs to report the log to the log processing platform. The log processing platform can be replaced by an audit platform or a log processing system. After the log processing platform receives the log reported by the cloud service through the receiving component (Access), it needs to copy the log and deliver the copied log to the Kafka queue, so as to store the log in an open source distributed search and analysis engine (i.e., ElasticSearch) through the Kafka protocol, so as to store the full amount of logs based on ElasticSearch, so that users can retrieve the required logs from ElasticSearch through the browser.
[0050] Among them, ElasticSearch can be used to store and retrieve large amounts of structured and unstructured data, including log data, providing high performance, scalability, and full-text search capabilities.
[0051] However, when storing the full amount of logs based on ElasticSearch, it is generally required to use hardware storage devices such as SSD disks or high I / O disks to improve the query performance of the logs. In addition, to ensure reliability, it is also necessary to use a copy mechanism and deploy spare hardware storage devices to back up the logs through spare storage devices. Therefore, when storing the full amount of logs based on ElasticSearch, more hardware storage devices are required, the hardware cost is higher, and the storage cost of the logs is increased.
[0052] To this end, the present application provides a log processing method that can be applied to the process of storing and retrieving logs. In the method, the log processing platform obtains multiple log data packets by packaging the acquired logs; then, for each log data packet in the multiple log data packets, each log data packet is scanned to construct a snapshot package of each log data packet; thereby further, each log data packet and the snapshot package of each log data packet are compressed to obtain compressed data, and the compressed data is remotely stored.
[0053] Through the above method, by constructing a snapshot package of the log data packet and compressing the log data packet and the snapshot package, the amount of log data can be reduced. Subsequently, the log data is stored remotely, that is, the compressed log data packet and snapshot package are stored in other devices outside the log processing platform, which reduces the requirements for the hardware deployed locally on the log processing platform, thereby reducing the storage cost of the log. Taking full advantage of the highly structured characteristics of the log, a technical solution for two-level dimensionality reduction processing of the log is provided. Compared with the traditional method of storing the full amount of logs based on ElasticSearch, it can reduce storage costs while meeting basic log query requirements, helping customers meet the security requirements for logs at a lower storage cost.
[0054] Optionally, the log described in the embodiment of the present application has a highly structured feature. For example, the log described in the present application may include but is not limited to an audit log, and may also be other highly structured specific logs, etc., without limitation.
[0055] It should be noted that the highly structured feature of the log means that each log information included in the log to be audited has a unified data format, and each log information includes the same parameters, that is, the data form of the log information has a templated feature, and different log information has different parameter values. In addition, the logs in this application are not limited to the logs to be audited, and any other form of logs can be stored and retrieved based on the implementation method of this application.
[0056] The two-level dimensionality reduction processing can be understood as: based on the highly structured characteristics of logs, the logs to be audited are packaged to obtain multiple log data packets, and then a snapshot package of each log data packet is constructed for each log data packet, which is the first level of dimensionality reduction processing; and, a Bloom dimensionality reduction key for each log data packet is constructed, which is another level of dimensionality reduction processing. Based on this, a large amount of log information can be reduced in dimensionality (i.e., compressed), and logs can be stored and retrieved from the perspective of data packets, which can reduce storage costs.
[0057] Multi-level protection, or network security level protection, is a basic system in the field of network security and a general requirement for network security. It also adds extended requirements covering cloud computing security, mobile Internet security, the Internet of Things, industrial control security, big data security and other aspects.
[0058] refer to Figure 2 The log processing method provided in the embodiment of the present application can be applied to an implementation environment consisting of a cloud platform for cloud services, a log processing platform, a user terminal, and a remote storage device. Figure 2As shown, the implementation environment may include a cloud platform 201, a log processing platform 202, a user terminal 203, and a remote storage device 204. The cloud platform 201 includes cloud services, which may be applications used by users. In the process of users using cloud services, log information may be generated and reported to the log processing platform 202. The log processing platform 202 includes a log processing application, which is used to receive and store log information sent by the cloud service, and retrieve the required log information for the user terminal 203 when the user terminal 203 retrieves the required log information.
[0059] Exemplarily, the cloud platform in the embodiment of the present application can be a cloud server (or server group), which can provide a way to access the service logic for use by the application (system) of the user terminal. The cloud voucher can provide application services for users, such as the implementation of the Hyper Text Transfer Protocol (HTTP) and database connection management. The cloud platform has the function of implementing the method provided in the embodiment of the present invention. The function can be implemented by hardware, or it can be implemented by hardware executing the corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0060] In one possible design, the cloud platform includes a processor and a transmitter, and the processor is configured to support the cloud platform to perform the corresponding functions in the above method. The transmitter is used to support the communication between the cloud platform and the log processing platform, and send the log information involved in the above method to the log processing platform. The cloud platform may also include a memory, which is coupled to the processor and stores the necessary program instructions and data for the cloud platform. The cloud platform may also include a receiver, which is used to receive a request message sent by a user.
[0061] Exemplarily, the log processing platform in the embodiment of the present application can be a server (or server group), and the log processing platform has the function of implementing the above method. The function can be implemented by hardware, or the corresponding software can be implemented by hardware. The hardware or software includes one or more modules corresponding to the above functions, and the modules can be software and / or hardware.
[0062] In one possible design, the log processing platform includes a receiver and a processor, wherein the receiver is configured to support the log processing platform to receive messages or information from the cloud platform or the user terminal. The processor controls the log processing platform to execute a corresponding response according to the message or information received by the receiver.
[0063] In one possible design, the receiver may be a log collector, and the log processing platform further includes a log storage, an aggregation encoder, and a Bloom encoder. The log collector is used to collect log information generated by the cloud services included in the cloud platform, and then upload the collected log information to the local log storage. After a preset time or when the accumulated stored log information reaches a preset data volume, the collected log information is processed by the aggregation encoder and the Bloom encoder, and the processed data is remotely stored.
[0064] Exemplarily, the user terminal in the embodiments of the present application can be a tablet computer, a desktop, a laptop, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), an augmented reality (AR)\virtual reality (VR) device, etc. The embodiments of the present application do not impose any special restrictions on the specific form of the device.
[0065] The execution subject of the log processing method provided in the present application can be a central processing unit (CPU) in a log processing platform including a log processing application, or a storage control module for implementing logs in the log processing platform, or a log processing application in the log processing platform.
[0066] The technical solution provided in the embodiment of the present application can be applied to the above implementation environment. The implementation environment described in the embodiment of the present application is to more clearly illustrate the technical solution of the embodiment of the present application, and does not constitute a limitation on the technical solution provided in the embodiment of the present application. It is known to those skilled in the art that with the evolution of the implementation environment, the technical solution provided in the embodiment of the present application is also applicable to similar technical problems.
[0067] The methods in the following embodiments can all be implemented in a log processing platform having the above implementation environment. In the following embodiments, the execution subject of the log processing method provided by the present application is a log processing platform (or a server including an audit application) as an example to illustrate the methods of the embodiments of the present application.
[0068] The present application embodiment provides a log processing method, such as Figure 3 As shown, the log processing method may include S301-S303.
[0069] S301: The log processing platform obtains a compressed package of a log data packet and a compressed package of a snapshot package of the log data packet.
[0070] The log data packet includes at least one log information, and one log information includes parameters and parameter values. The parameters included in different log information are the same.
[0071] First, the log processing platform needs to obtain log information and package the log information into multiple log data packets.
[0072] Among them, the log information can be used to record the relevant execution process and / or execution result of a user performing a task. The log information may include multiple items and may also include the time when the log information is generated. It should be noted that this application uses the log information as an example of log information to be audited for exemplary description, so the log information in this application can be called an audit log, but the name of the "audit log" does not specifically limit the scope of protection of this application, and the log information in this application can also be any other form, log information in any scenario.
[0073] Optionally, each log message has a highly structured feature, for example, each log message has a unified data format, each log message includes the same parameters, that is, the data format of the log message has a templated feature, and the same parameter between different log messages can have different parameter values (which can be called Value), or can have the same parameter value, without restriction, but at least one parameter value is different between different log messages. Among them, the data format can include but is not limited to any of the following: table data format, JSON data format, XML data format, text, etc. Each log message can include at least one of the following six elements (i.e. parameters): executor, execution time, execution location, execution object, execution operation, and execution result.
[0074] It should be noted that the executor can be understood as the user who triggers the execution of the operation, the execution time can be understood as the time when the executor triggers the execution of the operation, the execution location can be understood as the IP address of the executor triggering the execution of the operation, the execution object can be understood as the device on which the operation is performed, the execution operation can be understood as what needs to be done, and the execution result can be understood as the result obtained after the execution of the operation. It should be understood that the logs in this application are not limited to the logs to be audited, and any other highly structured logs can be stored and retrieved based on the implementation method of this application. In addition, the execution location can also be replaced by the execution URL, etc.
[0075] Optionally, the log processing platform can collect the logs of each service included in the cloud platform through a log collector, and then upload the collected logs to a local log storage. Further, the log storage can package the stored logs based on a preset threshold to obtain multiple log data packets.
[0076] The preset threshold value can be determined based on the packaging rule. For example, the packaging rule can be the amount of data that can be contained in a predetermined log data packet. Under the packaging rule, the preset threshold value is the preset data amount, that is, the preset threshold value limits the maximum amount of data that can be contained in a log data packet. Alternatively, the packaging rule is based on the time sequence, and the logs generated within each interval of the preset duration are packaged into a log data packet. Under the packaging rule, the preset threshold value can be the preset duration.
[0077] Optionally, the log data packet obtained by packaging the logs may further include fields such as content that is not queried using a universally unique identifier (UUID) class, but these fields may be ignored during the process of processing the log data packet.
[0078] In a possible implementation, the log information can be packaged based on a preset data volume to obtain multiple log data packets, and one log data packet includes the log information of the preset data volume. In the case where the preset threshold includes the preset data volume, packaging the log to obtain multiple log data packets may include: packaging the log to obtain multiple log data packets based on the preset data volume, and each log data packet in the multiple log data packets includes the log information of the preset data volume.
[0079] Optionally, the log can be divided into multiple copies based on time sequence, each copy includes log information with a preset data volume, so that each log copy is packaged to obtain a log data package; or, based on the association relationship between the log information, the log information with an associated relationship can be divided into one copy, and the amount of data included in each log copy can be controlled, so as to package multiple log data packages.
[0080] It should be noted that log information with an associated relationship can be understood as: when the generation time of multiple log information is adjacent, the multiple log information can be called log information with an associated relationship; or, there is a causal relationship between the multiple log information, for example, multiple log information is log information corresponding to the same object, multiple log information is log information corresponding to the same audit event, etc.
[0081] In this way, the log processing platform can predetermine, based on preset packaging rules, that the log data packet can include a preset amount of data, and thus package the logs based on the preset data amount, so that the resulting multiple log data packets include logs of the same amount of data, which can improve the efficiency of log packaging.
[0082] A possible implementation method is to package the log information to obtain multiple log data packets based on the generation time of each log information, and one log data packet includes the log information generated within each time period of the preset time interval. In the case where the preset threshold includes the preset time interval, packaging the log to obtain multiple log data packets may include: based on the generation time of each log information in the log, each time interval of the preset time interval, packaging the log information generated within the preset time interval to obtain a log data packet. In this way, the log can be packaged to obtain multiple log data packets, so that each log data packet in the multiple log data packets includes the log information generated within each time period of the preset time interval, thereby improving the efficiency of packaging the log.
[0083] In a possible implementation manner, for each log data packet in the multiple log data packets, the log processing platform scans the log data packet and constructs a snapshot package of the log data packet.
[0084] The snapshot package of the log data packet may be a data packet of a Map structure. The Map structure is a collection that maps key objects and value objects. Each element of the Map structure contains a pair of key objects and value objects. As long as the key object is given, the value object corresponding to the key object will be returned.
[0085] In a possible implementation, the key object in the Map structure can be the parameters and parameter values included in the log information in the log data packet, specifically, a combination of processing fields and field values, for example, in the format of serveiceType_ecs, logLevel_warn, etc., where serveiceType and logLevel are processing fields, and ecs and warn are field values. The Map value object can be the number of log information with the same parameter value in the log data packet. If a key object is found for the first time, the value object is 1. Subsequently, each time the key object (i.e., parameter + parameter value) appears, the Map value object is incremented by 1.
[0086] Optionally, for each log data packet, the log processing platform can scan the log information included in the log data packet line by line, traverse each parameter, find the number of log information including the same parameter value, combine the number and the parameter value of the parameter to obtain a piece of information in the Map structure. Similarly, for other parameters, this method can also be used to obtain a piece of information in the Map structure corresponding to the other parameters. After the traversal is completed, a snapshot package of the log data packet is obtained.
[0087] It should be noted that in the process of constructing the snapshot package of each log data packet, fields such as the Universally Unique Identifier (UUID) class in the log data packet that are not queried can be ignored, and unimportant information fields (or invalid information) in the log data packet can be ignored, that is, parameters such as UUID that are not queried and parameter values will not be traversed to create an information in the Map structure corresponding to the UUID and other parameters, so as to reduce the number of invalid information included in the snapshot package of the Map structure and improve memory storage efficiency.
[0088] Exemplarily, as shown in Table 1, it is assumed that a log data packet includes 10 log messages, and each log message includes six elements: executor (Who), execution time (When), execution location (Where), execution object (What), execution operation (Action), and execution result (Result). For example, the six elements of the first log message are: executor: Zhang, execution time: 2023.11.07, execution location: 123.456.7.891, execution object: virtual machine (VM), execution operation: create, execution result: ok; the six elements of the second log message are: executor: Li, execution time: 2023.11.07, execution location: 123.456.7.892, execution object: VM, execution operation: create, execution result: ok, etc.
[0089] Table 1
[0090] serial number Executor Execution time Execution Location Execution Object Perform an action Execution Results 1 Zhang 2023.11.07 123.456.7.891 VM create OK 2 Li 2023.11.07 123.456.7.892 VM create OK 3 Wang 2023.11.07 123.456.7.893 VM create OK 4 Zhang 2023.11.07 123.456.7.891 VM delete OK 5 Li 2023.11.07 123.456.7.892 VM create fail 6 Wang 2023.11.07 123.456.7.893 VM delete fail 7 Zhang 2023.11.07 123.456.7.892 VM delete fail 8 Li 2023.11.07 123.456.7.893 VM delete fail 9 Wang 2023.11.07 123.456.7.891 VM create OK 10 Zhang 2023.11.07 123.456.7.892 VM create OK
[0091] Then scan the log data packet including 10 log information, for example, scan each parameter in the log data packet, such as executor (Who), execution time (When), execution location (Where), execution object (What), execution operation (Action), execution result (Result) and parameter value, and calculate the number of times the parameter with the same parameter value appears. For example, the executor can be scanned first to check the parameter value of the executor in each log information, and Zhang appears 4 times, then (Who_Zhang, 4) is obtained; Li appears 3 times, then (Who_Li, 3) is obtained; Wang appears 3 times, then (Who_Wang, 3) is obtained; similarly, (When_2023.11.07, 10), (Where 123.456.7.891, 3), (Where123.456.7.892, 4), (Where 123.456.7.893, 3), (What VM, 10), (Action_create, 6), (Action_delete, 4), (Result_ok, 6), (Result_fail, 4). Then, the snapshot package corresponding to the log data packet is constructed as shown in Table 2. The format of the snapshot package is Map format. The key object is the combination of the processing field and the field value, and the value object is the number of occurrences.
[0092] Table 2
[0093] serial number Key Object Value Objects 1 Who_Zhang 4 2 Who_Li 3 3 Who_Wang 3 4 When_2023.11.07 10 5 Where_123.456.7.891 3 6 Where_123.456.7.892 4 7 Where_123.456.7.893 3 8 What_VM 10 9 Action_create 6 10 Action_delete 4 11 Result_ok 6 12 Result_fail 4
[0094] It should be noted that the content and number of log information (10 log information) shown in Table 1 and Table 2 above are only exemplary illustrations. The technical solution disclosed in this application is not limited to the content shown in the above examples. This application does not limit the specific content of the log information and the number of log information.
[0095] S302: The log processing platform constructs a dimensionality reduction key for each log data packet.
[0096] Among them, the dimensionality reduction key is used to retrieve the snapshot package of each log data packet, and one dimensionality reduction key is used to retrieve the snapshot package of a log data packet. The dimensionality reduction key can specifically be a Bloom dimensionality reduction key, that is, a dimensionality reduction key obtained by a Bloom encoder. The present application does not limit the specific form of the dimensionality reduction key, and it can also be other forms of dimensionality reduction keys.
[0097] In one possible implementation method, the dimension reduction key is stored in a NoSQL database. The NoSQL database is a local database of the log processing platform, and specifically may be ElasticSearch.
[0098] It should be noted that dimensionality reduction processing of the log data packet to obtain a dimensionality reduction key can eliminate noise and adversarial data in the log data packet, and obtain the main low-dimensional features that describe the original data while maintaining the inherent structure of the original data of the log data packet as much as possible.
[0099] Optionally, for each log data packet in the multiple log data packets, when the log processing platform scans the information included in each log data packet line by line, it is also necessary to construct a Bloom dimensionality reduction key for each log data packet and store it in a NoSQL database, for example, the NoSQL database may be ElaticSearch.
[0100] It should be noted that the Bloom dimension reduction key for constructing each log data packet can be based on the Bloom encoder to encode the log data packet (which can be called the original log data packet), compress the original log data packet into a single field, and obtain the Bloom dimension reduction key. Among them, the number of digits of the Bloom dimension reduction key can be predetermined based on the actual situation, for example, it can be based on all the parameters included in the log data packet (i.e., the 60 parameters included in Table 1, determined by the product of the number of log lines and the number of fields), and calculated by the Bloom function applied in the Bloom encoder. Alternatively, it can be calculated by the Bloom function applied in the Bloom encoder through the non-repeated parameters included in the log data packet (i.e., the 12 different parameter values included in Table 1, Zhang, Li, Wang, 2023.11.07, 123.456.7.891, 123.456.7.892, 123.456.7.893, VM, create, delete, ok, _fail). The number of Bloom functions can also be determined based on actual conditions, and generally 3 to 5 functions are used.
[0101] Exemplarily, based on the 12 parameters included in the above Table 1 (i.e., Zhang, Li, Wang, 2023.11.07, 123.456.7.891, 123.456.7.892, 123.456.7.893, VM, create, delete, ok, _fail), the corresponding dimensionality reduction key can be calculated by combining multiple Bloom functions applied in the Bloom encoder with a preset confidence parameter. The dimensionality reduction key can be expressed in binary or hexadecimal, as shown in Table 3.
[0102] Table 3
[0103]
[0104] It should be noted that when the number and specific values of the parameters are different, the specific values of the dimensionality reduction keys calculated by the multiple Bloom functions applied in the Bloom encoder are different.
[0105] For further example, assuming that each log data packet includes 5,000 log information, each log information includes two parameters A and B and the parameter value corresponding to each parameter, the log data packet can be encoded by a Bloom encoder, and the original log data packet can be compressed into a single field, that is, the obtained Bloom dimensionality reduction key can be used to indicate the key value and the two parameters A and B.
[0106] Optionally, the data format of the Bloom dimensionality reduction key can be:<key,Value1,Value2> , key is the key value of the Bloom dimensionality reduction key, Value1 and Value2 are unique identification codes for each parameter information randomly generated according to the parameter information included in the log data packet.
[0107] Based on this, the log processing platform constructs a snapshot package and dimensionality reduction key for each log data packet. When the user retrieves log information, the corresponding snapshot package can be quickly retrieved through the dimensionality reduction key, so as to quickly determine whether the corresponding log data packet contains the retrieved log information based on the snapshot package. When the retrieved log information is contained, the log information is quickly obtained from the corresponding log data packet, thereby quickly retrieving the required log information and improving the efficiency of log retrieval.
[0108] S303: The log processing platform sends a compressed package of the log data packet and a compressed package of the snapshot package of the log data packet to a remote storage device.
[0109] The package name of the compressed package of the log data packet includes a first key value, and the package name of the compressed package of the snapshot package includes a second key value. The first key value and the second key value are key values corresponding to the dimensionality reduction key of the log data packet.
[0110] It can be understood that the log processing platform can compress each log data packet and the snapshot packet of each log data packet to obtain compressed data, and store the compressed data remotely.
[0111] It should be noted that remote storage of each log data packet and the compressed data corresponding to the snapshot packet of each log data packet may not occupy the local storage resources of the log processing platform, and remote storage also needs to meet the security of data information.
[0112] Exemplarily, data compression can be implemented using algorithms such as Zlib, Gzip, Bzip2, Deflater, Lz4, Lzo, and Snappy.
[0113] A possible implementation method also needs to save the association relationship between the dimension reduction key of the log data packet and the address information of the remote storage device. That is, it is also necessary to determine the remote storage address of the compressed data corresponding to each log data packet and the snapshot package of each log data packet, and construct the association relationship between the dimension reduction key of each log data packet and the remote storage address of the corresponding compressed data.
[0114] Optionally, the compressed data is stored remotely, and the compressed data can be uploaded to a remote bucket, which can be an object storage service (OBS) or a data storage device such as a third-party data storage server. The remote storage address can be the address of the OBS or the third-party data storage server.
[0115] It should be noted that Value1 in the Bloom dimensionality reduction key can be used as the prefix of the data packet corresponding to the compressed data after the snapshot packet of each log data packet is compressed, and Value2 in the Bloom dimensionality reduction key can be used as the prefix of the data packet corresponding to the compressed data after the log data packet is compressed.
[0116] When the log processing platform stores compressed data remotely, it can establish an association between the dimensionality reduction key of each log data packet and the remote storage address of the corresponding compressed data. Therefore, when the user retrieves log information, after determining the dimensionality reduction key corresponding to the retrieved log information, the remote storage address of the compressed data including the retrieved log information is first determined, and then the corresponding compressed data is obtained based on the remote storage address, so as to quickly obtain the required log information from the compressed data, thereby improving the efficiency of log retrieval.
[0117] An example, such as Figure 4 As shown, after the cloud service in the cloud platform reports the log information to the log processing platform, the log processing platform can perform pre-statistics on the log information through the aggregation encoder, and package the log information to obtain multiple log data packets, and then perform dimensionality reduction processing on each log data packet, and process the log data packet through the aggregation encoder to obtain the snapshot packet corresponding to each log data packet. In addition, the log data packet is processed through the Bloom encoder to obtain the dimensionality reduction key corresponding to each log data packet. The two encoders make full use of the highly structured and high compression ratio characteristics of the log to reduce the dimensionality and hierarchical storage of the original log information, thereby reducing the amount of data storage space occupied.
[0118] For example, the storage space occupied by the snapshot package corresponding to a log data packet can be reduced to 5%, the storage space occupied by the data packet can be reduced to 15%, the storage space occupied by the dimensionality reduction key can be reduced to 1%, and the overall storage space occupied is reduced to 21% of the log data packet itself.
[0119] For example, a log data packet contains 5,000 log messages, and the size of each log message is 1K. Then the size of a log data packet before compression is about 5M. By indexing the key value with a high compression ratio, a snapshot package of the log data packet is obtained with a compression ratio of 5%. Based on the engineering recommendation of 10 times the number of elements and an accuracy of about 90%, the size of the Bloom dimensionality reduction key is estimated to be: 5,000 * 10 index items * 0.1 removal rate * 10 times the number of elements / 8bit = 6.25K, that is, the compression ratio of the Bloom dimensionality reduction key is: 6.25K / 5M = 0.125%.
[0120] For example, assuming that the amount of new log data added by a user every day is 10T, and each log information needs to be stored for 180 days before it can be deleted, the log processing platform needs to store 1800T of data in total. According to 2 copies, 70% of the disk occupancy, and 16C128G10T node cost of 15,000 / month, the storage cost based on ElaticSearch is: 1.5*(10*180*2) / 10 / 0.7=7.71 million / month. Based on the above solution, according to the 7-day active data cache, the cost of OBS is estimated to be 92 yuan / TB per month, and the cost of cold data to be stored is: 92*1800T=165,600 / month; and the cost of storing hot data is: 1.5*(10*7*2) / 10 / 0.7=300,000 / month, a total of 465,600 / month, and it is expected to reduce 1-(15.56+30) / 771=95% of the storage cost.
[0121] After the log processing platform obtains the log, it can package the obtained log to obtain multiple log data packets. Then, each log data packet is scanned to construct a snapshot package of each log data packet; after obtaining the snapshot package of each log data packet, each log data packet and the snapshot package of each log data packet can be compressed, and the obtained compressed data can be stored remotely. Therefore, through the above method, by constructing a snapshot package of the log data packet, compressing the log data packet and the snapshot package, and then storing them in a remote storage manner, the amount of log data can be reduced, and the requirements for hardware can be reduced by remote storage, thereby reducing the storage cost of the log.
[0122] A possible implementation is Figure 5 As shown, the method also includes S501-S503.
[0123] S501: The log processing platform obtains log retrieval information.
[0124] The log retrieval information is used to indicate the key to be retrieved and the key value to be retrieved.
[0125] Optionally, when a user needs to query historical log information, the log processing platform can be triggered to retrieve the historical log information required by the user by sending log retrieval information to the log processing platform. The log processing platform can query the corresponding target dimensionality reduction key from all dimensionality reduction keys based on the index key of the constructed log retrieval information.
[0126] A possible implementation method also requires building a dimensionality reduction index key for log retrieval information based on the key to be retrieved and the key value to be retrieved, and the index key is used to retrieve the target dimensionality reduction key.
[0127] S502: The log processing platform queries whether there is a target dimension reduction key corresponding to the dimension reduction index key of the log retrieval information.
[0128] The dimension reduction index key is obtained based on the key to be retrieved and the key value to be retrieved.
[0129] In a possible implementation, when it is determined that a target dimensionality reduction key exists, required log information can be retrieved based on the target dimensionality reduction key.
[0130] S503: In the case where there is a target dimensionality reduction key, based on the target dimensionality reduction key search, obtain log information corresponding to the log search information from a compressed package of the log data packet and a compressed package of the snapshot package of the log data packet stored in a remote storage device.
[0131] For example, users can enter query conditions as log retrieval information. The query conditions can usually be<Key,Value> structure; when the log processing platform receives the log retrieval information, it processes the query conditions, calculates the dimensionality reduction index key of the query conditions, and generates a binary Bloom key (i.e., the target dimensionality reduction key). Further, based on the generated index key, the Bloom index library (including the dimensionality reduction key of each log data packet) is queried. If there is a dimensionality reduction key that meets the conditions, the corresponding<Value1,Value2> , based on the obtained<Value1,Value2> The required log information was retrieved.
[0132] When retrieving target log information, the log processing platform can construct an index key for the log retrieval information corresponding to the target log information, so as to quickly obtain the required target log information based on the dimensionality reduction key corresponding to the index key. The efficiency of log retrieval can be improved through the association between the index key and the dimensionality reduction key.
[0133] In a possible implementation method, the target dimensionality reduction key corresponds to a first target key value and a second target key value; the log processing platform can obtain a target snapshot package corresponding to the target dimensionality reduction key based on the first target key value corresponding to the target dimensionality reduction key; wherein the first key value included in the package name of the compressed package of the target snapshot package is the first target key value.
[0134] In one example, the log processing platform obtains the corresponding snapshot package from the remote device storing compressed data (i.e., OBS or a third-party data storage server) based on the target dimension reduction key, and reads the obtained target snapshot package to determine whether the target snapshot package contains the log information queried by the log retrieval information.
[0135] A possible implementation method is to obtain the target log data packet based on the second target key value corresponding to the target dimensionality reduction key when the target snapshot package matches the log retrieval information; the second key value included in the package name of the compressed package of the target log data packet is the second target key value.
[0136] It can be understood that when it is determined that the target snapshot package contains the log information queried by the log retrieval information, it can be determined that the target snapshot package matches the log retrieval information, and then the target log data packet corresponding to the target snapshot package can be further obtained.
[0137] In a possible implementation, the log processing platform may determine the log information corresponding to the log retrieval information from the target log data packet.
[0138] In one example, the log processing platform obtains<Value1,Value2> Afterwards, based on Value1, you can query the snapshot package with the prefix Value1 in the snapshot library to obtain the corresponding snapshot package, and then query the log data package with the prefix Value2 in the database based on Value2 to obtain the required log information from the log data package and return the retrieval result.
[0139] It should be noted that if multiple<Value1,Value2> , you need to determine the snapshot packet and log data packet corresponding to each subsidiary field respectively.
[0140] After determining the target dimension reduction key corresponding to the index key, the log processing platform can first obtain the target snapshot package corresponding to the target dimension reduction key, thereby obtaining the target log data packet corresponding to the target snapshot package, and then determining the target log information indicated by the log retrieval information from the target log data packet. Through the above method, based on the index key, dimension reduction key, snapshot package and corresponding log data packet, the target log information indicated by the log retrieval information is obtained step by step, which can improve the efficiency and accuracy of log retrieval.
[0141] For example, Figure 6 As shown, when the user queries log information on the log processing platform, the user enters the query conditions as the log retrieval information, triggering the log processing platform to retrieve the corresponding dimensionality reduction key, snapshot package and log data packet from the remote storage server, thereby retrieving the required log information in the log data packet and returning the retrieval results to the user.
[0142] It should be noted that this application is applied to the construction of a log processing platform, which uses a low-cost solution to comply with the security requirements and reduce the storage cost of logs. At the UI level, a function key with the words "data dimensionality reduction and compression" can be constructed, and a function key with the words "fast combination query" can also be included to trigger the log processing platform to quickly retrieve the log information required by the user.
[0143] For example, Figure 7 As shown, in the user interface (UI), when a user uses the cloud service in the cloud platform, the user can trigger the cloud platform to report the log information to the log processing platform based on the control with the words "data dimensionality reduction and compression" displayed in the interface. The log processing platform can then store the log information in a remote storage device after dimensionality reduction and compression. In addition, when a certain log information needs to be retrieved, the corresponding log information can be retrieved from the remote storage device based on the control with the words "fast combination query" displayed in the interface.
[0144] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of the log processing platform. It is understandable that the log processing platform includes hardware structures and / or software modules corresponding to the execution of each function in order to realize the above functions. Those skilled in the art should easily realize that, in conjunction with the steps of a log processing method of each example described in the embodiment disclosed in the present application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or electronic device software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present application.
[0145] The embodiment of the present application can divide the log processing platform into functional modules or functional units according to the above method example. For example, each functional module or functional unit can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of software functional modules or functional units. Among them, the division of modules or units in the embodiment of the present application is schematic, which is only a logical function division, and there may be other division methods in actual implementation.
[0146] Please refer to Figure 8 , which shows a schematic diagram of a log processing device provided in an embodiment of the present application, and the log processing device is applied to a log processing platform. Figure 8 As shown, the log processing device 800 may include: a communication module 801 and a processing module 802 .
[0147] Wherein, the communication module 801 is used for the log processing device 800 to execute: obtaining a compressed package of a log data packet and a compressed package of a snapshot package of the log data packet, wherein the log data packet includes at least one log information, a log information includes a parameter and a parameter value, and the parameters included in different log information are the same. For example, the communication module 801 is used to support the log processing device 800 to execute the steps in S301 in the above method embodiment, and / or other processes for the technology described herein.
[0148] The communication module 801 is also used for the log processing device 800 to execute: sending a compressed package of a log data packet and a compressed package of a snapshot package of the log data packet to a remote storage device; wherein the package name of the compressed package of the log data packet includes a first key value, and the package name of the compressed package of the snapshot package includes a second key value, and the first key value and the second key value are key values corresponding to the dimension reduction key of the log data packet. For example, the communication module 801 is used to support the log processing device 800 to execute the steps in S303 in the above method embodiment, and / or other processes for the technology described herein.
[0149] In a possible design, the communication module 801 is also used by the log processing device 800 to execute: obtaining log retrieval information, where the log retrieval information is used to indicate the key to be retrieved and the key value to be retrieved.
[0150] The processing module 802 is also used by the log processing device 800 to execute: query whether there is a target dimension reduction key corresponding to the dimension reduction index key of the log retrieval information; the dimension reduction index key is obtained based on the key to be retrieved and the key value to be retrieved.
[0151] The above-mentioned processing module 802 is also used for the log processing device 800 to execute: in the presence of a target dimensionality reduction key, based on the target dimensionality reduction key retrieval, log information corresponding to the log retrieval information is obtained from the compressed package of the log data packet stored in the remote storage device and the compressed package of the snapshot package of the log data packet.
[0152] In another possible design, the target dimensionality reduction key corresponds to a first target key value and a second target key value; the above-mentioned processing module 802 is specifically used for the log processing device 800 to execute: based on the first target key value corresponding to the target dimensionality reduction key, obtain the target snapshot package corresponding to the target dimensionality reduction key; wherein the first key value included in the package name of the compressed package of the target snapshot package is the first target key value.
[0153] The above-mentioned processing module 802 is specifically used for the log processing device 800 to execute: when the target snapshot package matches the log retrieval information, the target log data packet is obtained based on the second target key value corresponding to the target dimensionality reduction key; the second key value included in the package name of the compressed package of the target log data packet is the second target key value.
[0154] The processing module 802 is specifically used by the log processing device 800 to execute: determining the log information corresponding to the log retrieval information from the target log data packet.
[0155] In another possible design, the reduced dimensionality keys are stored in a NoSQL database.
[0156] In another possible design, the processing module 802 is also used by the log processing device 800 to execute: saving the association between the dimension reduction key of the log data packet and the address information of the remote storage device.
[0157] In another possible design, the processing module 802 is also used by the log processing device 800 to execute: based on a preset data volume, the log information is packaged to obtain multiple log data packets, and one log data packet includes log information of a preset data volume.
[0158] In another possible design, the processing module 802 is also used by the log processing device 800 to execute: based on the generation time of each log information, the log information is packaged to obtain multiple log data packets, and one log data packet includes the log information generated within each preset time period.
[0159] Some other embodiments of the present application provide a log processing platform. The log processing platform may include: a memory and one or more processors. The memory and the processor are coupled. The memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the log processing platform may perform each function or step performed in the above method embodiment.
[0160] The present application embodiment also provides a log processing platform 900, such as Fig. 9 As shown, the log processing platform 900 includes at least one processor 901 and at least one interface circuit 902. The processor 901 and the interface circuit 902 can be interconnected via lines. For example, the interface circuit 902 can be used to receive signals from other devices (such as a memory of an electronic device). For another example, the interface circuit 902 can be used to send signals to other devices (such as the processor 901). Exemplarily, the interface circuit 902 can read instructions stored in the memory and send the instructions to the processor 901. When the instructions are executed by the processor 901, the log processing platform can execute the various steps in the above embodiments. Of course, the log processing platform can also include other discrete devices, which are not specifically limited in the embodiments of the present application.
[0161] The embodiment of the present application also provides a computer storage medium, which includes computer instructions. When the computer instructions are executed on the above-mentioned log processing platform, the log processing platform executes each function or step executed in the above-mentioned method embodiment.
[0162] The embodiment of the present application also provides a computer program product. When the computer program product is run on a computer, the computer is enabled to execute each function or step executed in the above method embodiment.
[0163] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0164] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0165] The units described as separate components may or may not be physically separated, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple different places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0166] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0167] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.
[0168] The above contents are only specific implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application shall be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A log processing method, characterized in that: Applied to a log processing platform, the method includes: Obtaining a compressed package of a log data packet and a compressed package of a snapshot package of the log data packet, wherein the log data packet includes at least one log information, and one log information includes a parameter and a parameter value, and the parameters included in different log information are the same; Send a compressed package of the log data packet and a compressed package of a snapshot package of the log data packet to a remote storage device; wherein the package name of the compressed package of the log data packet includes a first key value, and the package name of the compressed package of the snapshot package includes a second key value, and the first key value and the second key value are key values corresponding to the dimensionality reduction key of the log data packet.
2. The method according to claim 1, characterized in that The method further comprises: Obtaining log retrieval information, where the log retrieval information is used to indicate a key to be retrieved and a key value to be retrieved; Query whether there is a target dimension reduction key corresponding to the dimension reduction index key of the log retrieval information; the dimension reduction index key is obtained based on the key to be retrieved and the key value to be retrieved; In the case where the target dimensionality reduction key exists, based on the target dimensionality reduction key retrieval, log information corresponding to the log retrieval information is obtained from the compressed package of the log data packet stored in the remote storage device and the compressed package of the snapshot package of the log data packet.
3. The method according to claim 2, characterized in that The target dimension reduction key corresponds to a first target key value and a second target key value; The retrieval based on the target dimension reduction key, obtaining the log information corresponding to the log retrieval information from the compressed package of the log data packet and the compressed package of the snapshot package of the log data packet stored in the remote storage device, includes: Based on a first target key value corresponding to the target dimensionality reduction key, a target snapshot package corresponding to the target dimensionality reduction key is acquired; wherein a first key value included in a package name of a compressed package of the target snapshot package is the first target key value; In the case where the target snapshot package matches the log retrieval information, a target log data packet is obtained based on a second target key value corresponding to the target dimension reduction key; the second key value included in the package name of the compressed package of the target log data packet is the second target key value; The log information corresponding to the log retrieval information is determined from the target log data packet.
4. The method according to any one of claims 1 to 3, characterized in that The dimensionality reduction keys are stored in a NoSQL database.
5. The method according to claim 4, characterized in that The method further comprises: The association relationship between the dimension reduction key of the log data packet and the address information of the remote storage device is saved.
6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: The log information is packaged based on a preset data volume to obtain a plurality of log data packets, and one log data packet includes log information of a preset data volume.
7. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Based on the generation time of each log information, the log information is packaged to obtain multiple log data packets, and one log data packet includes log information generated within each period of a preset time interval.
8. A log processing device, characterized in that: The log processing device is applied to a log processing platform, and the log processing device includes: A communication module, used for obtaining a compressed package of a log data packet and a compressed package of a snapshot package of the log data packet, wherein the log data packet includes at least one log information, and one log information includes a parameter and a parameter value, and the parameters included in different log information are the same; The communication module is also used to send a compressed package of the log data packet and a compressed package of a snapshot package of the log data packet to a remote storage device; wherein the package name of the compressed package of the log data packet includes a first key value, and the package name of the compressed package of the snapshot package includes a second key value, and the first key value and the second key value are key values corresponding to the dimensionality reduction key of the log data packet.
9. A log processing platform, characterized in that: The log processing platform includes a memory and a processor, the memory is coupled to the processor; the memory is used to store computer program code, and the computer program code includes computer instructions; when the computer instructions are executed by the processor, the log processing platform executes the method described in any one of claims 1-7.
10. A log processing system, characterized in that: The log processing system comprises a log processing platform and a remote storage device, and the log processing system is used to execute the method according to any one of claims 1 to 7.