Log processing method and device, electronic equipment and computer readable medium
By replacing Logstash with Flink, efficient cleaning and diversion of log data is achieved, and the problems of high computing resource costs and business iteration risks in the existing log system are solved, improving the flexibility and scalability of log processing.
Patent Information
- Application Number
- CN202410038977.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-10
- Publication Date
- 2025-07-11
AI Technical Summary
When the existing log system based on ElasticSearch+Logstash+Kibana processes massive log data, there is a risk of overall log processing services being unavailable due to high computing resource costs, low performance, and business changes and iterations. The log field type verification is not strict, which affects the writing of normal log data.
Logstash is replaced by Flink streaming data processing framework, and the original log is consumed and parsed through Flink, data cleaning and dynamic filtering are carried out according to component source requirements, dividing to multiple kafka topics, and writing to the corresponding log index dictionary through multiple Flink downstream jobs to achieve upstream and downstream data decoupling and flexible expansion.
It reduces the risk of overall log processing service unavailability caused by business changes and iterations, improves the flexibility and efficiency of log processing, avoids the problems of write blocking and abnormal exits, and supports efficient single topic or index extension iterative development.
Smart Images

Figure CN120295871A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of Internet technologies, and in particular, to a method for processing logs, an apparatus for processing logs, an electronic device, and a computer-readable medium. Background Art
[0002] Log data is an important manifestation of whether the system operation meets the expected goals. It is one of the important means to help discover system failures and verify execution results. The currently common log system is an exabyte-level massive log data collection, storage, and display platform built based on the ElasticSearch (distributed search and analysis engine) + Logstash (data processing pipeline) + Kibana (open-source analysis and visualization platform designed for Elasticsearch) system. It has characteristics such as low latency (about 10 - 30 seconds from the service log production end to Elasticsearch storage) and large data volume (20TB of data per day, 40TB during peak periods).
[0003] In such a large-scale massive data scenario, due to reasons such as excessive log quantity, too many dependent components, and high performance and cost requirements, the existing Logstash cannot well meet the requirements of business users in terms of performance and business expansion.
[0004] In view of this, there is an urgent need in the art for a more flexible and higher-performance log processing method.
[0005] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0006] The purpose of the present disclosure is to provide a method for processing logs, an apparatus for processing logs, an electronic device, and a computer-readable medium, so as to at least to some extent improve the efficiency and flexibility of log processing.
[0007] According to a first aspect of the present disclosure, there is provided a method for processing logs, including:
[0008] Obtain the collected original logs, and use flink to consume and parse the original logs to obtain log parsing information corresponding to each log data in the original logs;
[0009] According to the component source requirements corresponding to each log data, perform data cleaning on the log parsing information through flink processing operators, and perform data filtering according to dynamic data filtering configurations to obtain target information data corresponding to each log data;
[0010] The target information data corresponding to each piece of log data is respectively forwarded to multiple different Kafka topics in the Kafka cluster through a Flink write operator, obtaining shunted log data;
[0011] Multiple Flink downstream jobs respectively consume the shunted log data in each Kafka topic, and write the shunted log data in each Kafka topic into the corresponding log index dictionary.
[0012] According to a second aspect of the present disclosure, there is provided a log processing device, including:
[0013] An original log parsing module, configured to obtain the collected original log, consume and parse the original log using Flink, and obtain log parsing information corresponding to each piece of log data in the original log;
[0014] A data cleaning and filtering module, configured to perform data cleaning on the log parsing information through a Flink processing operator according to the component source requirements corresponding to each piece of log data, and perform data filtering according to dynamic data filtering configuration, obtaining target information data corresponding to each piece of log data;
[0015] A log data shunting module, configured to forward the target information data corresponding to each piece of log data to multiple different Kafka topics in the Kafka cluster through a Flink write operator, obtaining shunted log data;
[0016] A shunted data writing module, configured to respectively consume the shunted log data in each Kafka topic through multiple Flink downstream jobs, and write the shunted log data in each Kafka topic into the corresponding log index dictionary.
[0017] According to a third aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the log processing method according to any one of the above through executing the executable instructions.
[0018] According to a fourth aspect of the present disclosure, there is provided a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, the log processing method according to any one of the above is implemented.
[0019] The exemplary embodiments of the present disclosure may have the following beneficial effects:
[0020] In the method for processing logs according to the exemplary embodiments of the present disclosure, the raw logs are consumed and parsed by using Flink, and then according to the component source requirements corresponding to each log data, the log parsing information is data-cleaned by Flink processing operators, and data filtering is performed according to the dynamic data filtering configuration to obtain the target information data corresponding to each log data. Then, the target information data corresponding to each log data is forwarded to multiple different Kafka topics in the Kafka cluster by the Flink write operator, and finally, multiple Flink downstream jobs consume the shunted log data in each Kafka topic respectively and write the shunted log data. In the method for processing logs according to the exemplary embodiments of the present disclosure, on the one hand, the Flink job writes the log data into different Kafka topics, realizing the decoupling of upstream and downstream data, reducing the risk of the overall log processing service becoming unavailable due to business change iterations. By separating the upstream and downstream Kafka log data, the problem of the overall log processing function becoming unavailable caused by write blocking or abnormal exit for some reason can be avoided, effectively avoiding risks. On the other hand, by realizing the decoupling of data on different business sides and parsing and writing to different downstream Kafka topics, the extended iterative development for a single topic or a single index becomes more efficient, and the impact scope of the business's own processing logic is further controlled.
[0021] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0023] Figure 1 It shows a schematic flow chart of a method for processing logs according to a related embodiment of the present disclosure;
[0024] Figure 2 It shows a schematic flow chart of a method for processing logs according to the exemplary embodiments of the present disclosure;
[0025] Figure 3 It shows a schematic flow chart of a method for processing logs according to a specific embodiment of the present disclosure;
[0026] Figure 4 It shows a block diagram of a device for processing logs according to the exemplary embodiments of the present disclosure;
[0027] Figure 5 The figure shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present disclosure. Detailed implementation manners
[0028] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be used. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present disclosure.
[0029] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the figures denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the figures are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0030] In some related embodiments, based on the logstash-based logging system, the original logs collected by the self-developed log collection component can be uploaded to the corresponding topic of the message middleware kafka, and then logstash consumes the original log data from one or more topics, performs regular matching and parsing, and performs data cleaning and data integration according to special business logics. Finally, the processed data is batch-written into different indexes of elasticsearch according to business requirements. As Figure 1 shown, the above-mentioned method for processing logs may include the following steps:
[0031] Step S110. Obtain the original logs collected by k8s (Kubernetes, a container-based cluster management platform).
[0032] Step S120. Write to different kafka topics according to different k8s clusters.
[0033] Step S130. Logstash parses the original log information.
[0034] Logstash parses the log level, file name, IP (Internet Protocol) address, log data, reporting time, etc. in the original log information.
[0035] Step S140. Determine whether the parsing is successful.
[0036] If the parsing fails, go to step S150; if the parsing is successful, go to step S160.
[0037] Step S150. Write to the error log es (Elasticsearch) index.
[0038] If the parsing fails, add a failure label, supplement the stack error reason, and write to the dedicated error log es index.
[0039] Step S160. Perform data cleaning according to the requirements of the component.
[0040] Parse from the file name which component printed the log, and perform data cleaning according to the requirements of the component, such as format conversion, data information enhancement, etc.
[0041] Step S170. Write to the es index.
[0042] According to the parsed component information, determine which es index needs to be written to.
[0043] However, there are the following problems in the above ELK (ElasticSearch + Logstash + Kibana) log system:
[0044] 1. High computing resource cost. Currently, multiple physical machines are needed. Logstash is implemented based on Ruby (a dynamic programming language) with the design concept of flexible configuration and simple configuration. However, as an interpreted language, Ruby itself needs to rely on the underlying interpreter to perform additional interpretation processing on the execution object, and the execution speed is relatively slower than Java.
[0045] 2. The current log architecture, out of consideration for the performance of the collection end, does not do any data processing on the collected original logs, and directly reports them to Kafka. Therefore, in the log data processing process such as log data cleaning and data analysis, it undertakes all business processing work. Each business component has independent printing rules, independent field type requirements, and independent data writing logic. Therefore, in the above logstash configuration, there will be a large amount of rule judgment logic. Once there is a change in the business, the overall functional regression cost is high, and it will also have a great impact on the huge log processing architecture. This situation has a great risk point and should be avoided.
[0046] 3. In the writing logic of elasticsearch, elasticsearch will constrain the fields of log data under the same index, and will not allow different types of data to be written into the same field. However, in the above logstash log processing architecture, the field data of the log is not type-checked. For example, a log data has a foo field, which is registered as a string in elasticsearch. Then, when another log data is written in the integer form of foo=1, elasticsearch will reject this write. Then, if the container log prints the wrong field for some reason, this wrong log data will affect other normal logs.
[0047] In addition, the Flink streaming data processing framework has the following problems:
[0048] In the official implementation of Flink, the judgment logic of business rules is mostly in the form of static code or command line parameters. This will cause the Flink job to be restarted when the business changes, which is a very time-consuming operation. The Flink cluster needs to destroy the resources of the old job and reallocate the resources of the new job. During the restart of the Flink job, data cannot be continuously consumed from Kafka, which will cause a backlog of log data during the period, increase the latency of log processing, and is not conducive to the real-time requirements of some software for logs.
[0049] Based on the above problems, this example implementation uses flink to replace logstash as the log data processing component to reconstruct the log data processing business. This example implementation first provides a log processing method. Figure 2 As shown, the method for processing the above log may include the following steps:
[0050] Step S210: Get the collected original logs, use Flink to consume and parse the original logs, and obtain the log parsing information corresponding to each log data in the original logs.
[0051] Step S220. According to the component source requirements corresponding to each piece of log data, use Flink processing operators to clean the log parsing information and filter the data according to the dynamic data filtering configuration, so as to obtain the target information data corresponding to each piece of log data.
[0052] Step S230. Use Flink write operators to forward the target information data corresponding to each piece of log data to multiple different Kafka topics in the Kafka cluster, so as to obtain shunted log data.
[0053] Step S240. Consume the shunted log data in each Kafka topic through multiple Flink downstream jobs, and write the shunted log data in each Kafka topic into the corresponding log index dictionary.
[0054] In the log processing method of the exemplary embodiment of the present disclosure, by using Flink to consume and parse the original log, and then according to the component source requirements corresponding to each piece of log data, use Flink processing operators to clean the log parsing information and filter the data according to the dynamic data filtering configuration, so as to obtain the target information data corresponding to each piece of log data. Then use Flink write operators to forward the target information data corresponding to each piece of log data to multiple different Kafka topics in the Kafka cluster. Finally, consume the shunted log data in each Kafka topic through multiple Flink downstream jobs and write the shunted log data. In the log processing method of the exemplary embodiment of the present disclosure, on the one hand, the Flink job writes log data to different Kafka topics, realizing the decoupling of upstream and downstream data, reducing the risk of the overall log processing service being unavailable due to business change iterations. By separating upstream and downstream Kafka log data, it is possible to avoid problems such as write blocking or abnormal exit due to certain reasons leading to the unavailability of the overall log processing function, effectively avoiding risks. On the other hand, by realizing the decoupling of data on different business sides and parsing and writing to different downstream Kafka topics, the expansion and iterative development for a single topic or a single index become more efficient, and the impact scope of the business's own processing logic is further controlled.
[0055] Next, in combination with Figure 3 make a more detailed description of the above steps of this exemplary embodiment.
[0056] In step S210, obtain the collected original log, use Flink to consume and parse the original log, and obtain the log parsing information corresponding to each piece of log data in the original log.
[0057] In this exemplary embodiment, first, obtain the original logs collected by k8s, enable the Flink job, consume the original logs through the Flink consumption operator, and parse the log parsing information in the original log information through Logstash, including file name, IP address, log data, reporting time, and so on. If the parsing is successful, send it to the Flink processing operator for further processing.
[0058] In this exemplary embodiment, if the parsing of some log data in the original logs fails, write the failed part of the log data into the corresponding error log index dictionary. If the parsing fails, label the part of the log data as failed, supplement the stack error reason, and write it into the dedicated error log ES index.
[0059] In step S220, according to the component source requirements corresponding to each log data, perform data cleaning on the log parsing information through the Flink processing operator, and perform data filtering according to the dynamic data filtering configuration to obtain the target information data corresponding to each log data.
[0060] In this exemplary embodiment, it is possible to parse from the file name of each log data which component printed the log, that is, the component source of the log data. And according to the component source requirements corresponding to each log data, perform data cleaning on the log parsing information through the Flink processing operator, such as format conversion, data information enhancement, etc. Then perform data filtering according to the dynamic data filtering configuration to obtain the target information data corresponding to each log data.
[0061] In this exemplary embodiment, the dynamic data filtering configuration refers to the configuration information of the data filtering rules dynamically obtained from the dynamic configuration service. Since Flink itself cannot be flexibly configured, a scheduled task can be used to request the dynamic data filtering configuration from the dynamic configuration service at a preset time interval, and broadcast the dynamic data filtering configuration to the Flink processing operator through the broadcast mechanism of Flink. Among them, the dynamic configuration service is a web service that can perform dynamic management of data information and display information using a graphical interface. It can be used to maintain the data filtering rule database of Flink and can also provide a front-end web page to support the custom configuration of data filtering rules.
[0062] In this exemplary embodiment, a dynamic configuration service can be used to maintain the filtering rule database of Flink. For example, a web framework PythonApp (Python application service) based on Python 2 language can be used as the dynamic configuration service and exposed to the Flink job for query through an HTTP interface. The advantage of this is that, on the one hand, compared with other methods of directly accessing the database to query rules, managing Flink rules through an application service helps with subsequent business function iteration and can also be integrated into a series of middleware such as CICD (Continuous Integration, Continuous Delivery, and Continuous Deployment) processes and monitoring built around PythonApp. On the other hand, managing filtering rules through a service can avoid problems caused by SQL (Structured Query Language) adjusting rule data, improving the availability of the filtering service.
[0063] On this basis, the Flink job configures a scheduled task to periodically request filtering rules from the PythonApp service and broadcast the filtering rules to all Flink job processes. The Flink job processes can change the filtering strategy in real time according to the broadcast rule data. Therefore, the change of filtering rules is changed from reading from the original command line or the configuration file compiled by the Java jar (archive file) package to dynamically reading the data of the PythonApp service interface, and the cost of changing filtering rules becomes very low.
[0064] After the downstream log processing operator obtains the dynamic data filtering configuration, it can filter the log parsing information with matching fields by obtaining the filtering fields in the dynamic data filtering configuration when there are matching fields in the log parsing information.
[0065] The downstream log processing operator will match all incoming stream log dictionary data to check whether the log field data conforms to the filtering rules. If so, it will be filtered; otherwise, the data will be collected to the downstream Flink write operator.
[0066] In this exemplary embodiment, through the combination of the Flink job and the PythonApp service, the log data filtering rules and the filtering rules of log data fields are dynamically changed, so as to filter log data according to the exact match of a certain field in the log. In this way, on the one hand, it changes the old form of the Flink job that can only adjust the filtering strategy by changing the command line and the local configuration file, making the Flink respond to business needs more quickly. On the other hand, during the business peak period, some low-importance logs can be selectively filtered to ensure that key logs are effectively written to Elasticsearch.
[0067] In step S230, the Flink write operator forwards the target information data corresponding to each log data to multiple different Kafka topics in the Kafka cluster respectively, obtaining shunted log data.
[0068] In the embodiment of this example, when performing data cleaning on the log parsing information through the Flink processing operator, the write index corresponding to each log data can be obtained; the Flink write operator forwards the target information data corresponding to each log data to multiple different Kafka topics in the Kafka cluster according to the write index, obtaining shunted log data.
[0069] After receiving the incoming log, the Flink write operator writes it to N intermediate Kafka topics according to the configuration based on the different Elasticsearch indexes to be written. Among them, the performances of different Kafka topics, such as data throughput and data processing real-time performance, are different, and the priorities of the log data shunted to each Kafka topic are also different. The Flink job writes to different Kafka topics according to different Elasticsearch indexes to be written, realizing the decoupling of upstream and downstream data and reducing the risk of the overall log processing service being unavailable due to business change iterations. By separating the upstream and downstream Kafka log data, this log processing architecture can avoid the overall log processing function being unavailable caused by the blocking of a single Elasticsearch index write or abnormal exit for some reason, effectively avoiding risks.
[0070] In step S240, multiple Flink downstream jobs consume the shunted log data in each Kafka topic respectively, and write the shunted log data in each Kafka topic to the corresponding log index dictionary.
[0071] In the embodiment of this example, by starting N Flink downstream jobs to consume the N intermediate Kafka topics respectively, according to the parsed component information, it is judged which ES index needs to be written. The original log is consumed through multiple Kafka topics, and the corresponding information data of the original log is parsed regularly according to business needs and parsed into dictionary data in a specified format.
[0072] As Figure 3 shown is the complete flowchart of the log processing method in a specific embodiment of the present disclosure, which is an example of the above steps in the embodiment of this example. The specific steps of this flowchart are as follows:
[0073] Step S302. Obtain the original log collected by k8s.
[0074] Step S304. Consume the original log.
[0075] Step S306. Logstash parses the original log information.
[0076] Logstash parses the log level, file name, IP address, log data, reporting time, etc. in the original log information.
[0077] Step S308. Determine whether the parsing is successful.
[0078] If the parsing fails, go to step S316; if the parsing is successful, go to step S314.
[0079] Step S310. The scheduled task obtains the log filtering configuration data.
[0080] A scheduled task is formulated through the flink timerservice (time service) method to request the log filtering configuration data every minute, including the exact match based on a certain field to be filtered and the field whitelist for writing to a certain elasticsearch index.
[0081] Step S312. Broadcast the log filtering configuration data.
[0082] After the flink process obtains the log filtering configuration information, it broadcasts it to the downstream log processing operators through the flink's broadcast mechanism.
[0083] Step S314. Perform data cleaning according to the requirements of the components.
[0084] The flink processing operator parses which component printed the log from the file name and performs data cleaning according to the requirements of the component, such as format conversion, data information enhancement, etc.
[0085] Step S316. Write to the error log es index.
[0086] If the parsing fails, add a failure label, supplement the stack error reason, and write it to a dedicated error log es index.
[0087] Step S318. Log data filtering.
[0088] For the incoming log data, check if it matches the log filtering configuration. If a certain field matches the log filtering configuration, discard the log data.
[0089] Step S320. Forward to the downstream topic.
[0090] If the field does not match, forward the log data to the downstream topic.
[0091] Step S322. Kafka topic shunting.
[0092] The log data is shunted through multiple Kafka topics, and each Kafka topic is consumed by multiple Flink downstream jobs respectively. Among them, the performances of different Kafka topics are different, and the priorities of the log data shunted to each Kafka topic are also different. For example, if the performance of Kafka topic A is greater than that of Kafka topic B, and the performance of Kafka topic B is greater than that of Kafka topic C, then the priority of the log data shunted to Kafka topic A is greater than that of the log data shunted to Kafka topic B, and the priority of the log data shunted to Kafka topic B is greater than that of the log data shunted to Kafka topic C.
[0093] Step S324. Write to the ES index.
[0094] Based on the parsed component information, determine which ES index needs to be written to.
[0095] In the embodiment of this example, in terms of the processing architecture, the original entire rule configuration of Logstash is split into two, consuming all Kafka topics of the container clusters in a combined manner, and writing to different downstream Kafka topics as an intermediate state of data according to the ES write index, and then starting another Flink job to consume the downstream Kafka topic and write to the ES.
[0096] It should be noted that although the steps of the method in the present disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be executed in this specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0097] Furthermore, the present disclosure also provides a log processing device. Refer to Figure 4 As shown, the log processing device may include a raw log parsing module 410, a data cleaning and filtering module 420, a log data shunting module 430, and a shunted data writing module 440. Among them:
[0098] The raw log parsing module 410 can be used to obtain the collected raw logs, consume and parse the raw logs using Flink, and obtain the log parsing information corresponding to each log data in the raw logs;
[0099] The data cleaning and filtering module 420 can be used to perform data cleaning on the log parsing information through Flink processing operators according to the component source requirements corresponding to each log data, and perform data filtering according to the dynamic data filtering configuration to obtain the target information data corresponding to each log data;
[0100] The log data shunting module 430 can be used to forward the target information data corresponding to each log data to multiple different Kafka topics in the Kafka cluster through the Flink write operator to obtain shunted log data;
[0101] The shunted data writing module 440 can be used to consume the shunted log data in each Kafka topic through multiple Flink downstream jobs and write the shunted log data in each Kafka topic into the corresponding log index dictionary.
[0102] In some exemplary embodiments of the present disclosure, a log processing device provided by the present disclosure may further include a dynamic data filtering configuration acquisition module, which can be used to request and obtain the dynamic data filtering configuration from the dynamic configuration service at a preset time interval through a scheduled task, and broadcast the dynamic data filtering configuration to the Flink processing operator through the broadcast mechanism of Flink.
[0103] In some exemplary embodiments of the present disclosure, the dynamic configuration service is used to maintain the data filtering rule database of Flink.
[0104] In some exemplary embodiments of the present disclosure, the dynamic configuration service provides a front-end network page to support the custom configuration of data filtering rules.
[0105] In some exemplary embodiments of the present disclosure, the data cleaning and filtering module 420 may include a filtering field matching unit, which can be used to obtain the filtering fields in the dynamic data filtering configuration, and filter the log parsing information with matching fields when there are matching fields in the log parsing information.
[0106] In some exemplary embodiments of the present disclosure, the log data shunting module 430 may include a write index acquisition unit and a log data forwarding unit. Among them:
[0107] The write index acquisition unit can be used to obtain the write index corresponding to each log data when performing data cleaning on the log parsing information through the Flink processing operator;
[0108] The log data forwarding unit can be used to forward the target information data corresponding to each log data to multiple different Kafka topics in the Kafka cluster according to the write index through the Flink write operator to obtain shunted log data.
[0109] In some exemplary embodiments of the present disclosure, the original log parsing module 410 may include an error log parsing unit, which can be used to write the partially failed log data in the original log into the corresponding error log index dictionary if the parsing of some log data in the original log fails.
[0110] The specific details of each module / unit in the above log processing device have been described in detail in the corresponding method embodiment section, and will not be elaborated here.
[0111] Figure 5 The structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present disclosure is shown.
[0112] It should be noted that Figure 5 The computer system 500 of the electronic device shown is only an example, and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0113] As Figure 5 shown, the computer system 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage section 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for system operation are also stored. The CPU 501, ROM 502, and RAM 503 are connected to each other via a bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.
[0114] The following components are connected to the I / O interface 505: an input portion 506 including a keyboard, a mouse, etc.; an output portion 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 508 including a hard disk, etc.; and a communication portion 509 including a network interface card such as a LAN card, a modem, etc. The communication portion 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed, so that the computer program read from it can be installed into the storage portion 508 as needed.
[0115] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 509, and / or installed from the removable medium 511. When the computer program is executed by the central processing unit (CPU) 501, various functions defined in the system of the present disclosure are performed.
[0116] It should be noted that the computer-readable medium shown in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0117] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0118] As another aspect, the present disclosure also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by an electronic device, the electronic device is caused to implement the method as described in the above embodiments.
[0119] It should be noted that although several modules of devices for performing actions are mentioned in the above detailed description, such division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules described above can be embodied in one module. Conversely, the features and functions of one module described above can be further divided and embodied by multiple modules.
[0120] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure.
[0121] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for processing logs, characterized in that, including: Obtain the collected original logs, consume and parse the original logs using Flink, and obtain the log parsing information corresponding to each log data in the original logs; According to the component source requirements corresponding to each log data, perform data cleaning on the log parsing information through Flink processing operators, and perform data filtering according to the dynamic data filtering configuration, to obtain the target information data corresponding to each log data; Forward the target information data corresponding to each log data to multiple different Kafka topics in the Kafka cluster through Flink write operators, to obtain shunted log data; Consume the shunted log data in each Kafka topic through multiple Flink downstream jobs respectively, and write the shunted log data in each Kafka topic into the corresponding log index dictionary.
2. The method for processing a log according to claim 1, wherein The method further includes: Request and obtain the dynamic data filtering configuration from the dynamic configuration service at a preset time interval through a scheduled task, and broadcast the dynamic data filtering configuration to the Flink processing operator through the broadcast mechanism of Flink.
3. The method for processing a log according to claim 2, wherein The dynamic configuration service is used to maintain the data filtering rule database of Flink.
4. The method for processing a log according to claim 2, wherein, The dynamic configuration service provides a front-end web page to support the custom configuration of data filtering rules.
5. The method for processing a log according to claim 1, wherein The performing data filtering according to the dynamic data filtering configuration includes: Obtain the filtering fields in the dynamic data filtering configuration, and filter the log parsing information with the matching fields when there are matching fields in the log parsing information that match the filtering fields.
6. The method for processing a log according to claim 1, wherein The forwarding the target information data corresponding to each log data to multiple different Kafka topics in the Kafka cluster through Flink write operators, to obtain shunted log data, includes: When performing data cleaning on the log parsing information through Flink processing operators, obtain the write index corresponding to each log data; Forward the target information data corresponding to each log data to multiple different Kafka topics in the Kafka cluster through Flink write operators according to the write index, to obtain shunted log data.
7. The method for processing a log according to claim 1, wherein After consuming and parsing the original logs using Flink, the method further includes: If part of the log data in the original logs fails to be parsed, write the failed part of the log data into the corresponding error log index dictionary.
8. A processing device for logs, characterized in that, including: An original log parsing module, used to obtain the collected original logs, consume and parse the original logs using Flink, and obtain the log parsing information corresponding to each log data in the original logs; A data cleaning and filtering module, used to perform data cleaning on the log parsing information through Flink processing operators according to the component source requirements corresponding to each log data, and perform data filtering according to the dynamic data filtering configuration, to obtain the target information data corresponding to each log data; A log data shunting module, configured to forward the target information data corresponding to each piece of log data to multiple different Kafka topics in a Kafka cluster through a Flink write operator, so as to obtain shunted log data; A shunted data writing module, configured to consume the shunted log data in each Kafka topic through multiple Flink downstream jobs respectively, and write the shunted log data in each Kafka topic into a corresponding log index dictionary.
9. An electronic device, characterized in that, Comprising: A processor; And A memory, configured to store one or more programs, which when executed by the processor, cause the processor to implement the log processing method according to any one of claims 1 to 7.
10. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the log processing method according to any one of claims 1 to 7.
Citation Information
Cited By
Oracle data completion method and device based on Flink
CN121935326A