Data processing method and computing device

By adding labels to the to-process log data and assigning it to the corresponding processing process for analysis, cleaning and standardization, the problem of complexity of log data processing is solved, and efficient and accurate log data processing is achieved.

CN120179495APending Publication Date: 2025-06-20HENAN QINWEI DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510112897.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

With the increase in the amount of log data and the continuous changes in the log format, collecting, maintaining and updating log data and in-depth analysis has become more complicated. How to efficiently process log data has become an urgent problem for those skilled in the art.

Method used

Provide a data processing method, by obtaining the data to be processed, adding tags to it based on the data type, and assigning it to the corresponding processing process for processing, including parsing, cleaning and standardizing processing, and finally obtaining standardized business data.

Benefits of technology

It realizes efficient processing of different types of log data, improves the processing efficiency of data flow, ensures the accuracy and flexibility in the data processing process, and can adapt to the diversified needs of complex data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179495A_ABST
    Figure CN120179495A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method and computing equipment, and the method comprises the steps: obtaining at least one piece of to-be-processed data; based on the data type of the to-be-processed data, a label is added to the to-be-processed data, the label is used for marking a target processing flow used for processing the to-be-processed data, the target processing flow is one of a plurality of processing flows, the processing flow corresponds to the data type, and the target processing flow is used for processing the to-be-processed data. The different processing flows are used for processing the data of the corresponding data types through different processors; and based on the label, distributing the to-be-processed data to the target processing flow for processing to obtain a processing result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a data processing method and a computing device. Background Art

[0002] Currently, with the development of information technology, log data has become an important means for enterprises to monitor and analyze the operation status of systems. Since different device manufacturers and application programs may use different log formats, including but not limited to text files, XML (eXtensible Markup Language), JSON (JavaScript Object Notation), etc., log data has a high degree of diversity and complexity. Enterprises need to integrate log data from different sources for in-depth analysis to gain business insights or perform fault diagnosis.

[0003] However, with the increase in the amount of log data and the continuous change of log formats, it has become more and more complex to collect, maintain and update, and deeply analyze log data. Thus, how to efficiently process log data has become an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0004] Embodiments of this application provide a data processing method and a computing device, which can efficiently process different types of data to be processed such as log data.

[0005] Therefore, the following technical solutions are provided in the embodiments of this application:

[0006] In a first aspect, an embodiment of this application provides a data processing method, including: obtaining at least one data to be processed; adding a tag to the data to be processed based on the data type of the data to be processed, where the tag is used to mark a target processing flow for processing the data to be processed, and the target processing flow is one of multiple processing flows, the processing flow corresponds to the data type, and different processing flows are used to process data of the corresponding data type through different processors; and allocating the data to be processed to the target processing flow for processing based on the tag to obtain a processing result.

[0007] An embodiment of this application provides a dynamic, flexible, and easy-to-implement and maintain data processing method, aiming to process a variety of different data to be processed (such as log data) from different data sources, and through multiple steps such as marking, routing, and parsing processing, to achieve the purpose of efficiently processing log data.

[0008] As an implementable embodiment, the processing flow includes a first processor and a second processor; the step of allocating the data to be processed to the target processing flow for processing includes: using the first processor to parse the data to be processed to obtain an attribute value; using the second processor to process the attribute value to obtain a processing result.

[0009] As an implementable embodiment, the step of using the first processor to parse the data to be processed to obtain an attribute value includes: when the data type of the data to be processed is json, using an EvaluateJsonPath processor to obtain keywords in the data to be processed as attributes of the data to be processed, and encapsulating them into the attribute value of the data to be processed; and / or when the data type of the data to be processed is xml, using an EvaluateXPath processor to obtain element nodes in the data to be processed as attributes of the data to be processed, and encapsulating them into the attribute value of the data to be processed; and / or when the data type of the data to be processed is csv, using a ConvertRecord processor to obtain delimiters or header configurations in the data to be processed as attributes of the data to be processed, and encapsulating them into entity attributes; and / or when the data type of the data to be processed is a text file, using an ExtractGrok processor to obtain time fields in the data to be processed as attributes of the data to be processed, and encapsulating them into the attribute value of the data to be processed.

[0010] In this embodiment, by parsing and extracting the data to be processed in different data types, the information carried by the data to be processed is converted into standardized attributes, and the attribute values are encapsulated according to the attribute information of the corresponding data type. This not only improves the processing efficiency of the data stream, but also ensures the accuracy and flexibility in the data processing process. By combining dedicated processors and flexible routing rules, the technical solution of the present application can provide higher performance and scalability when processing complex data, meeting diverse data integration and processing requirements.

[0011] As an implementable embodiment, the step of using the second processor to process the attribute value includes: using a processing method such as cleaning, formatting, or standardization to process the attribute value of the data to be processed.

[0012] As an implementable embodiment, using the second processor to process the attribute value to obtain a processing result includes: using a ReplaceText processor to uniformly format the date and time fields in the attribute value, remove or replace irrelevant characters, error data, or duplicate data in the attribute value, or process missing data to standardize the data format of the attribute value, obtaining standard business data after cleaning and filtering the data to be processed, and storing the standard business data in the business library.

[0013] In this embodiment, by cleaning, filtering, and standardizing the data, data quality, format consistency, and analysis accuracy are ensured. For example, the data cleaning and filtering steps guarantee the quality of the data finally stored in the database, ensuring that the data meets the standards, is free of redundancy and errors, and optimizing the use of storage resources. By uniformly storing all processed data in the business library, it is convenient for subsequent data management, query, and analysis. Moreover, the standardized data can ensure the accuracy of subsequent data analysis. Specifically, it solves common quality problems in data processing and improves the usability of the data by removing irrelevant characters, fixing error data, removing duplicates, filling in missing data, and unifying the data format. Through the automated data cleaning and standardization process, not only is the processing efficiency improved, but also the reliability and consistency of subsequent analysis are guaranteed. At the same time, this application adds data cleaning and formatting to the process, reducing the storage of invalid or redundant data and optimizing the use of storage resources.

[0014] As an implementable embodiment, storing the standard business data in the business library includes: converting the standard business data into business library entity fields, classifying and storing them according to business requirements; and performing centralized analysis on the archived data.

[0015] In this embodiment, by converting the standardized business data into business library entity fields and storing them classified, a large amount of log data can be efficiently managed, and the data that meets the requirements can be used for report display. At the same time, this technical solution can ensure the efficient storage, classified management, and subsequent data analysis capabilities of the data.

[0016] As an implementable embodiment, the method further includes: storing the data to be processed with failed tag addition in the data lake; and obtaining the data type of the data to be processed stored in the data lake.

[0017] In this embodiment, the data to be processed for which the data type is not recognized can be stored in the data lake for archiving and used for centralized analysis of the data to be processed afterwards. In this way, by introducing the mechanism of the data lake and re-recognizing the data type, the integrity and accuracy of data processing are further improved, especially when facing unknown data that cannot be recognized, ensuring the flexibility and scalability of the processing process.

[0018] As an implementable embodiment, the method further includes: updating the tags stored in the template library according to the data types of the data to be processed stored in the acquired data lake.

[0019] In this embodiment, when there is no corresponding type tag in the template library for the data type of the data to be processed, after obtaining the data type, the type tags in the template library can be updated according to the data type, so as to facilitate the subsequent marking of the corresponding tags for the data to be processed of this type. In this way, dynamic parsing of log data is achieved. Even if the log format changes, it can adapt to the new format through configuration updates without stopping the service or making complex code modifications.

[0020] As an implementable embodiment, acquiring the data types of the data to be processed stored in the data lake includes: identifying the data types of the data to be processed according to the file suffixes of the data to be processed; and / or identifying the data types of the data to be processed according to the scripts in the ExecuteGroovyScript processor.

[0021] In this embodiment, the file suffix method can quickly determine the data type in a conventional scenario, while the Groovy scripts in the ExecuteGroovyScript processor can perform fine-grained analysis for complex or special data types. This combination not only improves the efficiency but also enhances the flexibility and scalability when processing files in different formats, and can adapt to the changing business requirements and data formats.

[0022] As an implementable embodiment, acquiring at least one piece of data to be processed includes: collecting source log data from different source devices according to the ListenTCP processor, the ListenUDP processor, and the ListenSyslog processor.

[0023] In this embodiment, when collecting logs, processors such as the ListenTCP processor, the ListenUDP processor, and the ListenSyslog processor can open different listening ports. These processors are respectively responsible for listening to the log data on specific communication protocols (TCP (transmission control protocol), UDP (user datagram protocol), Syslog (system log)), ensuring that log information can be efficiently collected from multiple data sources.

[0024] Embodiments of the present application provide a unified processing entry to collect complex data from different sources. By automatically identifying the data type and tagging it, the automation degree of data processing is improved, which enables the present application to efficiently process data streams from multiple sources with inconsistent formats. Subsequently, through the automatic routing and processing flow of the data stream, different types of data can be dynamically adapted, and different data will follow different processing flows, reducing manual intervention, thereby improving the efficiency, flexibility, and scalability of data processing. As an extensible streaming data processing platform, NiFi supports the connection of multiple data sources and targets and can handle various scenarios from simple text data to complex structured data. This solution is based on the data stream orchestration of NiFi, can flexibly respond to various data processing requirements, and can be extended in the future as the data type and processing requirements change. By configuring different processors and data processing flows in NiFi, dedicated receiving and processing logics can be designed for each data type. For example, a listener and parser for JSON data, a listener parser for CSV (Comma-Separated Values) data, etc. This targeted processing improves the accuracy of data parsing and the efficiency of processing.

[0025] In this way, the technical solution of the present application combines the advantages of data stream orchestration and automated processing, and can achieve efficient and flexible processing of data from different sources and formats. Through tag routing and targeted data parsing, standardized log data is finally generated, which can ensure data quality and provide highly scalable processing capabilities. Combining the streaming processing characteristics of NiFi, this solution is applicable to scenarios that require processing large-scale and diverse data, such as log data processing, IoT (Internet of Things) data stream processing, real-time business data integration, etc. That is to say, the present application can greatly improve the processing efficiency of data to be processed, such as log data.

[0026] As an implementable embodiment, the data type includes any one of xml, json, csv, and text files; the allocating the data to be processed to the target processing flow based on the tag includes: identifying the tags of the tag data of the xml data type by an EvaluateXPath processor and routing them to the processing flow corresponding to the tags; and / or identifying the tags of the tag data of the json data type by an EvaluateJsonPath processor and routing them to the processing flow corresponding to the tags; and / or identifying the tags of the tag data of the csv data type by a ConvertRecord processor and routing them to the processing flow corresponding to the tags; and / or identifying the tags of the tag data of the text file data type by an ExtractGrok processor and routing them to the processing flow corresponding to the tags.

[0027] In this embodiment, multiple dedicated processors are utilized, such as an EvaluateXPath processor, an EvaluateJsonPath processor, a ConvertRecord processor, and an ExtractGrok processor, to identify tags for different data formats and classify and route data streams, which can effectively support various data formats such as XML, JSON, CSV, and text files. The dynamic data routing based on tags can direct data to different processing flows according to the tag information of the data, thereby improving the flexibility, efficiency, and accuracy of system processing.

[0028] In a second aspect, an embodiment of the present application further provides a data processing device, including an acquisition module and a processing module. Among them, the acquisition module is used to acquire at least one data to be processed; the processing module is used to add a tag to the data to be processed based on the data type of the data to be processed, and the tag is used to mark the target processing flow for processing the data to be processed, where the target processing flow is one of multiple processing flows, and the processing flow corresponds to the data type, and different processing flows are used to process data of the corresponding data type through different processors; based on the tag, the data to be processed is allocated to the target processing flow for processing to obtain a processing result.

[0029] In a third aspect, an embodiment of the present application further provides a computing device, including at least one memory for storing a program; at least one processor for executing the program stored in the memory; wherein, the memory is coupled to the processor, and when the program stored in the memory is executed, the processor is used to execute the method described in the first aspect or any possible implementation manner of the first aspect.

[0030] Fourthly, an embodiment of the present application further provides a computing device cluster, including a management node and multiple computing nodes, where the management node is configured to execute the method described in the first aspect.

[0031] Fifthly, an embodiment of the present application further provides a computer-readable storage medium storing computer instructions, which, when executed by a computing device, cause the computing device to execute the method involved in the first aspect and its possible implementation manners.

[0032] Sixthly, an embodiment of the present application further provides a computer program product, which includes computer instructions, and when the computer instructions are executed by a computing device, the computing device is caused to execute the method involved in the first aspect and its possible implementation manners.

[0033] It can be understood that the beneficial effects of the above second aspect to the sixth aspect can refer to the relevant descriptions of the first aspect, and will not be elaborated here.

[0034] In summary, the present application has at least the following advantages:

[0035] 1. Log data can be collected through multiple data source processors (such as TCP, UDP, SYSLOG, etc.), and log data of multiple devices and systems can be collected. Relying on the horizontal scalability of the Nifi cluster, large-scale log data monitoring and collection can be achieved.

[0036] 2. Log data can be dynamically parsed. Even if the log format changes, it can adapt to the new format through configuration updates without stopping the service or making complex code modifications.

[0037] 3. Rich data conversion functions are provided, and the collected log data can be cleaned, formatted, and standardized.

[0038] 4. Based on the characteristics of Nifi dynamic routing, content parsing of multiple data protocols can be compatible. For example, according to the log content, data can be routed to the corresponding processing process in the form of tagging, and at the same time, identified / unidentified content can be distributed and saved for subsequent summary analysis and display.

[0039] 5. The process interface display is easy to operate and intuitive, facilitating the operation and maintenance of operation and maintenance personnel. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 The schematic diagram of the architecture of a data processing system provided by an embodiment of the present application is shown;

[0041] Figure 2Shown is a schematic structural diagram of a server cluster provided by an embodiment of the present application;

[0042] Figure 3 Shown is a schematic flowchart of a data processing method provided by an embodiment of the present application;

[0043] Figure 4 Shown is a schematic flowchart of another data processing method provided by an embodiment of the present application;

[0044] Figure 5 Shown is a schematic structural diagram of a data processing device provided by an embodiment of the present application;

[0045] Figure 6 Shown is a schematic structural diagram of a computing device provided by an embodiment of the present application;

[0046] Figure 7 Shown is a schematic architecture diagram of a computing device cluster provided in an embodiment of the present application;

[0047] Figure 8 Shown is a schematic architecture diagram of another computing device cluster provided in an embodiment of the present application. Detailed implementation manners

[0048] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application.

[0049] In the description of the embodiments of the present application, words such as "exemplary", "for example", or "for illustration" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary", "for example", or "for illustration" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example", or "for illustration" is intended to present relevant concepts in a specific manner.

[0050] In the description of the embodiments of the present application, the term "and / or" only describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, B exists alone, and A and B exist simultaneously. In addition, unless otherwise specified, the meaning of the term "plural" refers to two or more. For example, multiple systems refer to two or more systems, and multiple terminals refer to two or more terminals.

[0051] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the technical features indicated. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "comprise", "include", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0052] In the description of the embodiments of the present application, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0053] In the description of the embodiments of the present application, the terms "first / second / third, etc." or module A, module B, module C, etc. are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that, where permitted, the specific order or sequence can be interchanged so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0054] In the description of the embodiments of the present application, the reference numerals representing steps, such as S101, S102, etc., do not necessarily indicate that the steps will be executed in this order. Where permitted, the order of the front and back steps can be interchanged, or the steps can be executed simultaneously.

[0055] Related terms involved in the embodiments of the present application:

[0056] Cleaning and filtering: In the fields of data science and big data, cleaning and filtering refers to the process of preprocessing raw data, aiming to improve data quality and ensure data accuracy and consistency. This process includes removing incorrect, duplicate, incomplete or irrelevant data, as well as standardizing data types to make them more suitable for subsequent analysis and processing.

[0057] Cluster: It refers to a group of cooperating servers (or computing nodes) that are connected through a network to form a unified system to work together. Although these servers are independent physical machines, they cooperate through cluster management software or protocols to provide a combined service or resource to the outside world. Generally, to users and application programs, these servers appear to be a single system. Therefore, a cluster can be used as a single system to provide services, process tasks or store data. The purpose of a cluster is to improve the reliability, availability and performance of the system by increasing redundancy and scalability. Clusters are widely used in computing and storage tasks that require high performance and high reliability. Through the cluster architecture, the system can maintain stable and efficient operation in the face of failures, increased loads or a sharp increase in data volume.

[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. The terms used herein are for the purpose of describing embodiments of this application only and are not intended to limit this application.

[0059] Embodiments of this application provide a data processing method. When collecting log data, this method can use processors such as ListenTCP processor, ListenUDP processor, ListenSyslog processor, etc. to open different listening ports to collect log data of various source devices and application programs. After the log data enters through various entrances, it is uniformly judged for format through the script in the ExecuteGroovyScript processor, and a predefined tag value is added to the log data of each format. The tagged log data is identified by the RouteOnAttribute processor and routed to separate processing flows. Different processing flows can parse specific log protocols through the EvaluateJsonPath processor, EvaluateXPath processor, and ExtractGrok processor to obtain the required attributes. The ReplaceText processor is used to clean and filter the attributes of the log data and standardize the values to obtain standard business data. The processed standard business data is stored in the business table of the database as needed for subsequent retrieval and analysis. Log data for which the data type cannot be recognized is stored in the data lake for archiving and analyzed later and then re-integrated into the processing flow.

[0060] The data processing method provided by the embodiments of this application is not only dynamic and flexible, but also easy to implement and maintain, and can be used to efficiently process log data. Facing the diversification of current log devices, applications, and collection ends, as well as the challenges of complex formats in log parsing, this application provides a unified processing entry that can accept complex data from different sources and use a unified parsing mechanism to customize specific parsing strategies for various log formats. By this method, the complexity in the processes of log collection, analysis, and processing can be significantly reduced. Through automated log collection and parsing, manual intervention is reduced, and the efficiency of data processing is improved, thereby enhancing the efficiency and reliability of the entire system.

[0061] Next, an introduction is given to the architecture of a data processing system provided by the embodiments of this application.

[0062] Figure 1 Shown is a schematic diagram of the architecture of a data processing system provided by the embodiments of this application. As Figure 1As shown, the data processing system includes at least one terminal device 10, a storage device 11, and a server cluster. The terminal device 10, the storage device 11, and the server cluster can be directly or indirectly connected through a wired network or a wireless network for data transmission. The embodiments of the present application do not limit this.

[0063] Exemplarily, the server cluster may include multiple computing devices (or computing nodes) 12. For example, the multiple computing devices 12 may include computing device 12-1, computing device 12-2, computing device 12-3, etc. The number of server clusters may be more or less. The embodiments of the present application do not limit this. It should be noted that the computing device 12 may be a device such as a server, a computer, or a smart phone. The multiple computing devices 12 may be one type of device or a combination of multiple types of devices. The multiple computing devices 12 may be deployed with an application program for running the data processing system to execute the functions of the data processing system, such as data cleaning, data storage and management, data conversion and integration, data analysis and mining, real-time data processing, security and privacy protection, etc. For example, the server cluster in the embodiments of the present application may be a NiFi cluster deployed with the data processing system, and the NiFi cluster can be used for centralized processing of log data.

[0064] The computing device 12 may be provided with a storage unit. The storage unit may be a memory for temporarily storing data, such as a cache, a dynamic random access memory (DRAM), a static random access memory (SRAM), etc., and is used to store the data that needs to be temporarily stored during the operation of the data processing system for access or operation by a processor or other hardware. In the embodiments of the present application, the storage unit may store the data corresponding to the computing tasks to be executed by the data processing system, the data corresponding to the unexecuted computing tasks in the interrupted computing tasks, the intermediate data during the processing, the processed data, and other cached data.

[0065] The computing device 12 may be provided with a communication interface to enable data transmission with other computing devices, memories, and other devices. The communication interface may be an interface for wired transmission, such as a compute express link (CXL) interface, a peripheral component interconnect express (PCIe) interface, a universal serial bus (USB) interface, etc. The communication interface may be an interface for wireless transmission, such as a Bluetooth (BT) module, a wireless fidelity (WI-FI) module, a wireless communication module, etc.

[0066] In one embodiment, when the server cluster serves as a NiFi cluster, it can be used to collect and parse various data to be processed. Optionally, the data to be processed may include, but is not limited to, log data, IoT data, real-time business data, etc. In this embodiment, the server cluster is communicatively connected to a client that serves as a log production source. Exemplarily, the client can be implemented as different types of devices and system programs capable of generating various log data. For example, the client can be a server, a network device, a security device, and an application.

[0067] For example, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud business libraries, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, security services, and big data and artificial intelligence platforms.

[0068] Optionally, the network device can be a router, a switch, a sensor, or a load balancer.

[0069] Optionally, the security device can be a hardware device such as a hardware firewall, a secure boot device, or a hardware token. In some embodiments, the security device can also be a virtualized security system. For example, an intrusion detection system (IDS) / intrusion prevention system (IPS): generates logs about malicious activities, abnormal traffic, and attack signature matches. A security information and event management (SIEM) system: centrally collects security events and alert logs from various devices (such as firewalls, IDS / IPS, terminals, etc.). Antivirus software: records logs of virus detection, malware scanning results, cleaning operations, and system health status. An authentication system: records information such as user logins, authentication failures, and permission changes. A data loss prevention (DLP) system: records potential data leakage, file transfer, and interception events.

[0070] Optionally, when the client is an application, it can be an application program running on the user-side terminal device, or can be a web browser provided by the data processing system to the outside, etc. For example, the client can be a desktop application, a mobile application, a web application, or a web-based application, etc., which is not limited here. The user-side terminal device can be, but is not limited to, a smart phone, a tablet computer, a notebook computer, or a desktop computer, etc.

[0071] In some embodiments, on the client of the above-mentioned log production source, a log agent (logagent) for collecting log data on the client is deployed. In this embodiment, the log agent is a tool or software for collecting, gathering, and forwarding log data. Its main responsibility is to transmit log files, events, or data generated by the source side (the above-mentioned various devices and systems) to the NiFi cluster of the data processing system for subsequent processing, analysis, and storage. Correspondingly, a monitoring entry for receiving the log data sent by the log agent is set on the NiFi cluster.

[0072] Optionally, other log collection applications, such as Fluentd, Logstash, Filebeat, or Flume, can also be deployed on the client of the above-mentioned log production source, as long as they can collect and send the source-side log data to the NiFi cluster. This application does not make strict restrictions here.

[0073] The storage device 11 is an independent storage device. Exemplarily, the memory can be a hard disk drive (HDD), a solid state drive (SSD), a not and (NAND) flash memory, a disk, etc. Exemplarily, the storage device 11 can be divided into a business library and a data lake, which are used to store the overflow cache data when multiple computing devices 12 run the data processing system, as well as other data, such as the programs for running the data processing system and the data processed by the data processing system. Among them, the business library can be used to store the standard business data successfully processed by the data processing system; the data lake is used to store the to-be-processed data that the data processing system fails to recognize, and is used for centralized analysis of the data stored in the storage device after archiving. It should be noted that Figure 1 the memory 11 in is connected to the computing device 12, which not only restricts the memory 11 to establish a communication connection with the computing device 12, but also means that the memory 11 establishes a communication connection with the data processing system deployed on the computing device 12, so that the memory 11 can establish a communication connection with any computing device 12.

[0074] The storage device 11 may be provided with a communication interface to enable data transmission with multiple computing devices 12 and other devices. The communication interface may be an interface for wired transmission, such as a CXL interface, a PCIe interface, a USB interface, etc. The communication interface may be an interface for wireless transmission, such as a BT module, a WI-FI module, a wireless communication module, etc.

[0075] The terminal device 10 may be an entity on the user side for receiving or transmitting signals. The terminal device 10 may be referred to as a terminal, a user equipment (UE), a mobile station (MS), a mobile terminal (MT), an access terminal device, an industrial control terminal device, a UE unit, a UE station, a mobile station, a remote station, a remote terminal device, a mobile device, a wireless communication device, a UE agent, or a UE device, etc. The terminal device may be fixed or mobile. It should be noted that the terminal device may support at least one wireless communication technology, such as long time evolution (LTE), NR, 6th-generation (6G) mobile communication system, or next-generation wireless communication technology, etc.

[0076] For example, the terminal device may be a mobile phone, a pad, a desktop computer, a laptop computer, an all-in-one computer, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, a cellular phone, a cordless phone, a session initiation protocol (SIP) phone, a wireless local loop (WLL) station, a personal digital assistant (PDA), a handheld device with wireless communication capabilities, a computing device, or other processing devices connected to a wireless modem, a wearable device, a terminal device in a future mobile communication network, or a terminal device in a future evolved public land mobile network (PLMN), etc. Embodiments of the present application do not limit the specific technologies and specific device forms adopted by the terminal device. An application is arranged on the terminal device 10, and through this application, the user can obtain the log data processed by the NiFi cluster and stored in the storage device 11, as well as the log data not recognized by the NiFi cluster, so as to facilitate the retrieval and analysis of the above log data.

[0077] Exemplarily, the above-mentioned wired network or wireless network may use standard communication technologies and / or protocols, including but not limited to any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), mobile, wired networks, private networks, or virtual private networks.

[0078] Thus, the log data of various devices and systems at the source end is collected by the log proxy agent and sent to the listening entry of the Nifi cluster. Through the dynamic parsing and processing method of Nifi, unified log parsing and cleaning processing are performed, and the processed log data is stored in the storage device 11. The data that meets the requirements can be presented in the form of a report on the terminal device 10. After the unparsed log data is re-analyzed, it flows back to the Nifi cluster for data processing, forming a complete processing flow.

[0079] It should be understood that the data processing system described in the embodiments of this application is to more clearly illustrate the technical solutions of the embodiments of this application, and does not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.

[0080] To facilitate the understanding of the technical solutions of this application, next, a detailed introduction to the data processing principle of the embodiments of this application will be given.

[0081] Figure 2 Shown is a schematic structural diagram of a server cluster provided by an embodiment of this application. As Figure 2 shown, in some possible implementation manners, on the server cluster side of the data processing system, the server cluster is a NiFi cluster 20. The NiFi cluster 20 includes multiple NiFi nodes, and a NiFi instance is deployed on each NiFi node, and they cooperate with each other to process data flows (FlowFile). It is worth mentioning that among the multiple NiFi nodes, there is a management node 21 and multiple computing nodes. For example, the multiple computing nodes include computing node A, computing node B, and computing node C. In this application, the number of computing nodes is not limited. Among them, the multiple computing nodes coordinate their work through the management node 21, share states and data, and ensure the high availability and load balancing of the entire data processing system. In the NiFi cluster 20, the data flow (FlowFile) is shared by all nodes, and the management node 21 in the NiFi cluster will determine which nodes process which tasks. The management node 21 in the NiFi cluster 20 will automatically allocate tasks to each of the computing nodes A-C in the NiFi cluster 20 to improve processing efficiency and scalability. It should be noted that in the NiFi cluster 20, any one of the computing nodes can act as the management node 21. In other words, the management node 21 is elected on any one of the computing nodes in the NiFi cluster 20. The election process is automatically handled by the NiFi cluster 20 itself. Based on factors such as the health status and load of the computing nodes, the NiFi cluster 20 will select a node as the management node 21.

[0082] For example, when the NiFi cluster 20 starts up, all computing nodes will attempt to compete for the role of the management node 21. At a certain moment in the NiFi cluster 20, only one computing node will be elected as the management node 21 (usually through an election algorithm such as Zookeeper, Raft, etc.). If this computing node becomes unavailable due to certain reasons (such as a fault or a restart), other computing nodes will automatically initiate an election and elect a new management node 21. That is to say, the management node 21 is actually borne by any one of the computing nodes in the NiFi cluster 20. The role of this computing node is not only to serve as the "control center" of the NiFi cluster 20, but it is also a part of the NiFi cluster 20 and executes data flow tasks. Although only one management node 21 is responsible for coordination, all computing nodes will participate in actual data processing, task execution, and other work. In addition to executing data processing tasks, the management node 21 is also used to coordinate and manage the status of the NiFi cluster 20, task scheduling, resource allocation, data flow control, and other work.

[0083] Exemplarily, a plurality of processors (not shown in the figure) are deployed on multiple computing nodes in the NiFi cluster 20. Among them, the plurality of processors include, but are not limited to, a ListenTCP processor, a ListenUDP processor, a ListenSyslog processor, an ExecuteGroovyScript processor, a RouteOnAttribute processor, an EvaluateJsonPath processor, an EvaluateXPath processor, an ExtractGrok processor, and a ReplaceText processor, etc. It can be understood that according to requirements, the plurality of processors may also include other types of processors, which are not limited in this application. It should be noted that in the NiFi cluster 20, the above-mentioned processors are abstract. These processors do not physically correspond to independent hardware devices, but belong to the logical processing units of the NiFi cluster 20. The above-mentioned processors define the operation steps of the data flow. Each processor is used to execute a specific task, such as receiving data, parsing data, making conditional judgments, modifying data, etc. These processors are logically designed steps, but their actual execution is the responsibility of the corresponding computing nodes (or a single computing node) in the NiFi cluster 20. When the NiFi cluster 20 is running, the execution of the processors will be carried out in parallel on each computing node in the cluster, and the specific resource consumption (CPU, memory, disk I / O, etc.) is borne by the hardware of the computing node where it is located. Therefore, although they are called "processors", their execution does not depend on a single physical entity, but on the node resources in the distributed cluster.

[0084] Exemplarily, in the NiFi cluster 20, different processors can be distributed on different computing nodes. Each computing node in the NiFi cluster 20 may contain one or more processors. For example, Figure 2 if the NiFi cluster 20 in Figure 2 has multiple computing nodes (Computing Node A, Computing Node B, Computing Node C), and processors such as the ListenTCP processor, ExecuteGroovyScript processor, and ReplaceText processor are required to process log data, then the ListenTCP processor may run on Computing Node A, the ExecuteGroovyScript processor may run on Computing Node B, and the ReplaceText processor may run on Computing Node C. When the NiFi cluster 20 processes log data, all computing nodes will participate in the processing of the data stream. For example, the log data is first received by the ListenTCP processor and then flows downstream to processors such as the RouteOnAttribute processor, EvaluateJsonPath processor, ExtractGrok processor, etc. In addition, the computing nodes in the NiFi cluster 20 will coordinate to ensure the consistency of the data stream and the order of processing. For example, if a certain node of the RouteOnAttribute processor needs to rely on the processing result of the EvaluateJsonPath processor, it will wait for the upstream node to complete the work before continuing to process. In other words, each computing node in the NiFi cluster 20 can handle a specific task (such as parsing, cleaning, storing), and the NiFi cluster 20 will allocate tasks to different computing nodes in the NiFi cluster 20 according to the load balancing policy. Therefore, the NiFi cluster 20 will dynamically allocate tasks according to the load, resource status, task type, etc. of the computing nodes in the NiFi cluster 20.

[0085] Exemplarily, the ListenTCP processor is used to listen on a TCP port and receive TCP protocol data from a device or application. The ListenUDP processor is used to listen on a UDP port and receive log data of the UDP protocol. The ListenSyslog processor is used to listen on the Syslog protocol and is usually used to receive logs generated by network devices (such as routers, switches) or operating systems. That is to say, the ListenTCP processor / ListenUDP processor / ListenSyslog processor is used to receive data from the network (TCP, UDP, or Syslog). They receive the data stream through the network listening port and use the underlying network resources (such as ports and buffers) during the execution process, but their actual operations are executed by the NiFi node. The ports opened by these processors ensure that various formats of log data can be received from different data sources.

[0086] The ExecuteGroovyScript processor is used to execute Groovy scripts, process data, or perform custom operations. It is a script processing operation executed on NiFi nodes. Exemplarily, after the log data collected by the listening port enters the ExecuteGroovyScript processor, the Groovy script in the ExecuteGroovyScript processor judges the format of the log data. According to the format of the log, the script will identify which category the data belongs to and label each format of data with corresponding tags. For example: If the log is an access log from a web server, the tag can be web_access_log. If the log is from a database query, the tag can be db_query_log. These tags help the system classify and process according to the log type later.

[0087] The RouteOnAttribute processor: This processor performs conditional routing selection based on the attributes of the FlowFile. Its execution depends on the CPU and memory of the node, but essentially it is a logical judgment operation. Exemplarily, after label processing, the log data is routed through the RouteOnAttribute processor. This processor decides the data flow to different processing flows based on the previously assigned tag values. For example: Logs of the web_access_log type may be sent to the web log parsing process; Logs of the db_query_log type enter the database query log parsing process.

[0088] The EvaluateJsonPath processor / EvaluateXPath processor: These processors are used to parse JSON or XML data. They parse specific content through the configured paths according to the format of the input data. For example, the log protocol of the log data routed to different data processing flows according to different tags will be parsed. When the log data is in JSON format, the EvaluateJsonPath processor is used to extract specific attributes in the JSON data. When the log data is in XML format, the EvaluateXPath processor is used to extract fields in the XML.

[0089] The ExtractGrok processor: It is used to parse text data and extract information based on Grok patterns. It is a parser that performs data matching and extraction on NiFi nodes according to the given patterns during execution. For example, for traditional text logs (such as web access logs), Grok patterns are used to extract specific fields (such as request time, IP address, request path, etc.).

[0090] ReplaceText Processor: This processor is used for text replacement operations. When executed, it reads the content of the FlowFile and performs replacements to replace or clean invalid or non-standard values in the data. After the log data is parsed, it may contain irrelevant information or data with inconsistent formats, which need to be cleaned and standardized. For example, the ReplaceText processor can be used to uniformly format the date and time fields, remove or replace irrelevant characters, error data, or duplicate data in the attribute data, or handle missing data to obtain the processed log data. The processed log data can be stored as standardized business data in the storage device 11 for subsequent retrieval and analysis. Specific business data will be stored in different business tables as needed. For example, the access logs of a web service may be stored in a table named WebAccessLogs; the database query logs may be stored in the DBQueryLogs table. For data that cannot be recognized or has a mismatched format, the system will store it in the data lake for archiving. These data can be saved in the original format (such as JSON, CSV, text) for future further analysis and reprocessing. The data lake provides a flexible storage method that can accommodate undefined log data.

[0091] As described above, the data processing method in this application is executed by the NiFi cluster 20 in the data processing system and mainly involves the following steps:

[0092] Collect log data: Collect logs from different sources through different listeners (TCP / UDP / Syslog). Format judgment and tagging: Use Groovy scripts to judge the log format and tag different types of logs. Routing and protocol parsing: According to the tag values, the log data enters different processing flows and parses the log content through relevant processors (such as EvaluateJsonPath, ExtractGrok). Data cleaning and standardization: Clean the data through the ReplaceText processor to ensure that the data meets the business requirements. Storage and archiving: Store the processed data in the database for analysis, and store the unrecognized log data in the data lake for further processing. This method can flexibly process log data in multiple formats, ensuring the efficient storage, management, and analysis of the data.

[0093] The above is the introduction to the data processing system and process provided by the embodiments of this application. Based on the above content, the technical solutions of this application will be described in detail with specific embodiments below. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0094] Figure 3 The flowchart showing a data processing method provided by an embodiment of this application is as followsFigure 3 As shown, it can be understood that this method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities. Exemplarily, this method can be executed by a data processing device, where the device can be implemented by software and / or hardware, and can be, but is not limited to, configured in a computing device cluster including at least one computing device. Typically, it can be configured in a server. For ease of description, in this embodiment, any management node in the NiFi cluster will be used as the execution entity for description. As Figure 3 shown, this data processing method may include the following steps:

[0095] S101. Obtain at least one data to be processed.

[0096] In this embodiment, the data to be processed can be log data, IoT data, real-time business data, etc. In this embodiment, the data to be processed is log data for illustration. Exemplarily, log data is generated by different types of devices. Among them, log data from different data sources can be first collected by a log proxy agent deployed on the log source-side device / system, and then sent by the log proxy agent to the NiFi cluster listening entrance. The listening entrance is opened by various processors in the NiFi cluster. For example, logs are collected through listening processors for multiple data sources. It can be understood that there are various listeners and get processors provided in the NiFi cluster to collect logs and messages from different data sources. These processors allow NiFi to obtain data from different input channels and serve as the starting point for subsequent data processing and forwarding.

[0097] Specific examples: File system category: Such as ListenFTP processor and GetFTP processor, suitable for collecting logs from the File Transfer Protocol. Message queue category: Such as ConsumeKafka processor and ConsumeAMQP processor, used for real-time consumption of logs from the message queue. Database category: Such as QueryDatabaseTable processor, suitable for extracting log records from the database. Internet of Things device category: Such as ConsumeMQTT processor, to obtain log data from Internet of Things devices. Interface and protocol category: Such as ListenHCP processor, ListenUDP processor, ListenSyslog processor, and GetSplunk processor, used to receive and obtain log data published through protocols such as HTTP, UDP, Syslog, and Splunk.

[0098] As a specific example, in this embodiment, the changes of the log file are monitored by the ListenFTP processor and the GetFTP processor. Among them, the ListenFTP processor is used to monitor a specified directory on the FTP server and real-time monitor the changes of files, especially the upload of new files. It can obtain log files or other types of files through the FTP protocol and is applicable to scenarios where files need to be obtained in real-time or periodically. The GetFTP processor is similar to the ListenFTP processor, but it is used to pull files on the FTP server regularly instead of real-time monitoring file changes.

[0099] In one embodiment, for message-type logs, the log messages can be monitored through the ConsumeKafka processor and the ConsumeAMQP processor. For example, the ConsumeKafka processor is suitable for obtaining messages from the log data stream published by Kafka. Kafka is a distributed message stream platform widely used in real-time logs, event streams, and data transmission. The ConsumeAMQP processor is used to consume messages from an AMQP (advanced message queuing protocol) queue. AMQP is an open standard message middleware protocol that defines a message queue mechanism, enabling applications to send and receive messages through the message queue, thereby achieving asynchronous communication and decoupling. It supports multiple message queue systems such as RabbitMQ.

[0100] In one embodiment, for the recording of logs, the log records are queried through the QueryDatabaseTable processor. The QueryDatabaseTable processor is used to query records in a data table from a relational database (such as MySQL, PostgreSQL, Oracle, etc.) and is particularly suitable for querying log records.

[0101] Optionally, for iot devices, the logs are collected through the ConsumeMQTT processor. The ConsumeMQTT processor is used to obtain log or event data from an MQTT (message queuing telemetry transport) broker through the MQTT protocol. MQTT is a lightweight message transmission protocol commonly used for message transmission between IoT (Internet of Things) devices.

[0102] Optionally, for interface class logs, they are received through the ListenHCP processor. The ListenHCP processor is used to listen for data published via the HTTP protocol and is typically used to obtain log data from the interfaces of web services or applications. As a specific example, the ListenHCP processor is suitable for applications that publish log or event data through REST API interfaces. Through this processor, the NiFi cluster can receive and process interface class logs in real time.

[0103] Optionally, for broadcast class logs, they are received through the ListenUDP processor. The ListenUDP processor is used to listen for data streams on a UDP (User Datagram Protocol) port. UDP is a connectionless network protocol suitable for broadcast and multicast scenarios. As a specific example, the ListenUDP processor is used to receive log messages sent from various broadcast sources (such as system logs, network devices, applications, etc.) to a specific UDP port. For example, the log collection of network monitoring, routers, switches and other devices.

[0104] Optionally, for log service classes, such as syslog, they are received through the ListenSyslog processor. For example, the ListenSyslog processor is used to receive log data sent by the Syslog protocol. Syslog is a standard protocol widely used for logging devices and applications in IT infrastructure. For example, the ListenSyslog processor is used to receive Syslog - formatted log messages from network devices, servers, applications, etc. It is usually used to collect system - level event and error logs.

[0105] Exemplarily, Splunk receives log messages through the GetSplunk processor. Splunk is a powerful log search and analysis tool widely used for collecting and analyzing machine data. In this embodiment, the GetSplunk processor is used to obtain log data from the Splunk server. As a specific example, the GetSplunk processor is suitable for pulling log or event data from the Splunk server. Especially when Splunk is used as a log storage and analysis platform, the NiFi cluster can extract the log data in Splunk and perform subsequent processing.

[0106] S102. Based on the data type of the data to be processed, add a label to the data to be processed. The label is used to mark the target processing flow for processing the data to be processed, where the target processing flow is one of multiple processing flows, the processing flows correspond to the data types, and different processing flows are used to process data of the corresponding data types through different processors.

[0107] There are a wide variety of formats for log files, including common ones such as JSON, XML, CSV, text files, etc. Log data in different formats has different structures and parsing methods. When processing logs, it is necessary to accurately determine the log format in order to select the appropriate parsing and processing methods. To achieve this goal, the following steps and technical solutions can be adopted:

[0108] S1021. Judging by file suffix name.

[0109] First of all, the format of the log can be initially judged by the suffix name of the file (such as.json,.xml,.txt,.csv). Most file systems and log systems determine the format type according to the extension of the log file. The following is the correspondence between common file suffixes and format types:

[0110] .json → JSON format,.xml → XML format,.txt → plain text format,.csv → CSV format (comma-separated values).

[0111] The correspondence between file suffix and format type is a quick preliminary judgment method, but it is not always accurate because the file content may not match the extension (for example, the file suffix is changed, or the file is in a mixed format).

[0112] S1022. Parsing the message content.

[0113] If the file suffix cannot fully determine the log format or there is confusion in the file suffix, the format can be further parsed and judged according to the content of the file. For different types of log content, common parsing methods include the parsing of JSON, XML, and text files. The following is the specific judgment process:

[0114] JSON format judgment: Use JsonSlurper (a JSON parsing library in Groovy) to parse the log content. If the parsing is successful, it means that the file is in a valid JSON format. JSON data is usually wrapped by {} and the data structure consists of key-value pairs, usually in the form of {"key": "value"}.

[0115] XML format judgment: Use XmlSlurper (an XML parsing library in Groovy) to parse the file content. XML data is usually in the <tag>< / tag> form, which contains nested tags and attributes. If the file can be successfully parsed into XML, it means that the log file is in XML format.

[0116] Text format judgment: If the file content cannot be successfully parsed by a JSON or XML parser, the file may be in plain text format. In this case, it can be simply assumed that it is a plain text log, or the format of the data such as CSV or TSV can be inferred through specific delimiters in the content.

[0117] S1023. Obtain the root node and concatenate it into a key.

[0118] Once the log format (JSON, XML, etc.) is determined, the next step is to parse the data and extract relevant field information.

[0119] The specific steps are as follows:

[0120] JSON format: After parsing the JSON data, the root node and its attribute fields in the JSON object can be extracted. For example, {"user":{"name":"John","age":30}}, the root node is "user", and the fields are "name" and "age". Concatenate the root node and the attribute fields to form strings user.name, user.age, and these strings will become the keywords for matching in the subsequent template library.

[0121] XML format: Parse the XML data to obtain the root node and its attributes or child nodes. For example, <user> <name>John< / name> <age> 30< / age> < / user> , the root node is <user>, the field is <name>and <age>. Concatenate into strings such as user.name and user.age for subsequent comparison and tagging.

[0122] S1024. Compare with the template library and tag.

[0123] After obtaining the root node and attribute information of the data field, these need to be compared with the existing protocols in the template library. The template library contains specifications for different application systems, log formats, and protocols, which can help us identify the corresponding tags for the log content.

[0124] Template library maintenance: The template library needs to be accessed and maintained in advance by the application administrator of the log. The administrator pre-defines the structure information and key fields of different applications and log formats into the template library. For example, the JSON log of a specific application may have a fixed structure, where the root node is event and the fields include timestamp and message. These information will be stored in the template library as templates.

[0125] Compare templates: After parsing the root node and fields of the log, compare them with the entries in the template library through string matching or regular expressions. If the match is successful, tags can be assigned to the log data. Tags can be such as "log_style_1", "log_style_2", "error-log", "access-log", "transaction-log", etc.

[0126] S103. Based on the tags, allocate the data to be processed to the target processing flow for processing to obtain the processing result.

[0127] In this implementation step, the data to be processed, such as log data, is routed to different data processing flows according to the tags of the tag data. The processing flow includes a first processor and a second processor. For example, the data processing flow has preset processors based on the NiFi cluster, and the processors are used to parse and process the tag data of different data types.

[0128] In this embodiment, different processors (Processors) are used to parse the attributes or fields in the log data and perform appropriate routing according to different formats. Specifically, in this implementation, when the data type of the data to be processed is json, the EvaluateJsonPath processor is used to obtain the keywords in the data to be processed as the attributes of the data to be processed and encapsulate them into the attribute values of the data to be processed. JSON data is usually stored in the form of key-value pairs, and the EvaluateJsonPath processor can extract specific attributes in the log through JsonPath expressions.

[0129] Exemplarily, assume the log content is as follows:

[0130]

[0131] It is possible to use EvaluateJsonPath to extract attributes such as timestamp, user.name, user.age, etc.

[0132] In one implementation, when the data type of the data to be processed is xml, the EvaluateXPath processor is used to obtain the element nodes in the data to be processed as the attributes of the data to be processed and encapsulate them into the attribute values of the data to be processed. XML data is usually composed of nested tags, and the EvaluateXPath processor can parse and extract specific elements or attributes in the log through XPath expressions.

[0133] Exemplarily, assume the log content is as follows:

[0134]

[0135]

[0136] Using EvaluateXPath, fields such as timestamp, user / name, user / age, etc. can be extracted.

[0137] In one implementation, when the data type of the data to be processed is a text file, the ExtractGrok processor is used to obtain the time field in the data to be processed as the attribute of the data to be processed and encapsulate it into the attribute value of the data to be processed. The ExtractGrok processor is used to parse text-format logs (usually unstructured log data). Text logs usually do not have a fixed structure, so regular expressions or Grok patterns are needed to extract specific information from the logs. The ExtractGrok processor uses predefined Grok patterns or custom patterns to extract fields.

[0138] Exemplarily, assume the log content is as follows:

[0139] 2024-11-11 12:00:00User John logged in

[0140] It is possible to use a Grok pattern to extract fields such as timestamp, user, message, etc. For example, use the following Grok pattern:

[0141] %{TIMESTAMP_ISO8601:timestamp}%{WORD:user}%{GREEDYDATA:message}

[0142] This pattern decomposes the log into fields such as timestamp (2024-11-11 12:00:00), user (John), and message (logged in).

[0143] In one implementation, when the data type of the data to be processed is csv, the ConvertRecord processor is used to obtain the delimiter or header configuration in the data to be processed as an attribute of the data to be processed and encapsulate it into the entity attributes. The ConvertRecord processor is used to parse the CSV-formatted log. CSV-formatted log data usually consists of comma-separated fields, and the ConvertRecord processor will convert the CSV data into a structured record (e.g., JSON or Avro format) and can extract specific fields.

[0144] Exemplarily, assume the content of the CSV file is as follows:

[0145] timestamp,user,message

[0146] 2024-11-11T12:00:00Z,John,User logged in

[0147] Using the ConvertRecord processor, the CSV content can be converted into a structured record and fields such as timestamp, user, and message can be extracted.

[0148] It can be understood that after parsing logs in different formats through processors such as the EvaluateJsonPath processor, EvaluateXPath processor, ExtractGrok processor, and ConvertRecord processor, the data can be routed to different processing flows according to the extracted tag information (such as log type, key fields, etc.). For example: If the log format is JSON, after parsing with the EvaluateJsonPath processor, the result is routed to a process suitable for JSON format processing. If the log format is XML, after parsing with the EvaluateXPath processor, the result is routed to a process suitable for XML format processing. For plain text logs, after parsing with the ExtractGrok processor, the result is routed to a process suitable for text format processing. For CSV format logs, after converting and extracting attributes with the ConvertRecord processor, it is routed to the CSV processing flow.

[0149] In this step, different processors are used for field extraction and parsing according to the format of the log file (JSON, XML, text, CSV). Each format of log data has its unique parsing method. Through these processors, we can extract different formats of log data into standardized attribute fields and perform further log routing and processing based on this field information.

[0150] In one embodiment, specifically, the first processor is used to parse the data to be processed to obtain the attribute value. Exemplarily, the first processor is used to obtain the attributes of the data to be processed and encapsulate them into the corresponding attribute values, obtaining the attribute data encapsulated into the entity attributes.

[0151] In this step, the first processor extracts key information from different formats of data, encapsulates and converts it, facilitating subsequent operations, storage, or transmission. Exemplarily, for JSON format parsing and assignment, JSON is a lightweight data interchange format, usually representing data in the form of key-value pairs. The EvaluateJsonPath processor in the first processor uses a JSONPath expression (such as $.[pkName]) to extract the value of a field from JSON data and use it as an attribute of the FlowFile for subsequent operations. In other words, in JSON, if you need to extract a certain attribute and encapsulate it into an entity object, you can use a syntax similar to $.[pkName]. Here, the $.[pkName] syntax is part of the JSONPath expression, usually used to refer to a specific field or attribute of JSON data. For example, $ represents the root node of the JSON data, and.[pkName] represents extracting the value by the primary key name pkName.

[0152] Exemplarily, assume we have the following JSON data:

[0153]

[0154]

[0155] We want to extract the order.id field and encapsulate it into the orderId attribute in the entity object. The following JSONPath expression can be used in the EvaluateJsonPath processor: $.order.id. This will extract 12345 and encapsulate it into the orderId field of the object as a FlowFile attribute. In other words, when the data type of the tag data is json, obtain the keyword in the tag data as an attribute of the tag data and encapsulate it into the corresponding attribute value.

[0156] Exemplarily, XML format parsing and assignment: XML is a format for marking data, similar to HTML. The XML data structure uses tags to define elements. In this embodiment, the EvaluateXPath processor in the first processor uses XPath expressions to parse and extract XML data, and stores the extracted information as FlowFile attributes. As a specific example: "." selects the current node, and "@" selects the attribute node. Suppose we have the following XML data:

[0157]

[0158] These extracted values can be respectively encapsulated into the corresponding entity attributes.

[0159] In other words, when the data type of the tag data is xml, the element nodes in the tag data are obtained as the attributes of the tag data and encapsulated into the entity attributes.

[0160] Exemplarily, text format parsing and assignment (Grok regular expression): The ExtractGrok processor in the first processor is used to extract data from the FlowFile content using the Grok pattern, and store this data as FlowFile attributes for subsequent processors to use. Grok is a text parsing tool based on regular expressions, used to extract information from logs or other text data. By defining regular expressions, the required fields can be accurately extracted from the text. For example, the time field is extracted through the regular expression "\s+(?<request_time>\d+(?:\.\d+)?)\s+". This regular expression means to extract a floating number (possibly a decimal) and assign it to the request_time field. Regular expression explanation: "\s+" matches whitespace, "(?<request_time>...)" is a named capture group that extracts the number and stores it in the request_time field, and "\d+(?:\.\d+)?" matches an integer or a decimal.

[0161] Exemplarily, suppose we have the following text: "begin 123.456end". We want to extract 123.456 into the request_time field. The regular expression \s+(?<request_time>\d+(?:\.\d+)?)\s+ will successfully match and extract the value. In other words, when the data type of the tag data is a text file, the time field in the tag data is obtained as the data of the tag data and encapsulated into the corresponding attribute value.

[0162] Regarding CSV format parsing and assignment: CSV is a widely used data format, usually used for tabular data. In a CSV format file, data is separated by commas (other delimiters can also be used). When parsing CSV data, it is usually necessary to set the properties of the CSV parser, such as the delimiter, whether there is a header, etc., in order to correctly extract the data. During the CSV parsing process, the properties of the CSVReader processor in the first processor are configured to read the CSV file and extract the fields into the specified entity properties. Suppose we have the following CSV file:

[0163] id,product,quantity

[0164] 12345,Laptop,2

[0165] We want to extract these three columns of data into the corresponding entity properties. By configuring the properties of the CSVReader processor, we can configure options such as the delimiter (default is ",",) and whether to include the header (header: true), extract the required property values, and perform the encapsulation of entity properties.

[0166] In other words, in this embodiment, when the data type of the label data is csv, the delimiter or header configuration in the label data is obtained as the property of the label data and encapsulated into the entity property.

[0167] In this step, for JSON format: use a path expression similar to $.[pkName] to extract data. For XML format: use XPath syntax to extract data through label or attribute paths. For text format: extract specific text fields through regular expressions (such as Grok). For CSV format: configure the CSV parser to extract data according to the delimiter and header. In practical applications, appropriate parsing tools and methods can be selected according to the different data formats, and the extracted data can be assigned to the entity object to facilitate subsequent processing and use, and there is no strict limit in this embodiment.

[0168] In one embodiment, the second processor is used to process the property value to obtain a processing result. Exemplarily, the second processor is used to clean and filter the property value of the data to be processed to obtain standard business data.

[0169] In this embodiment, in the data processing flow, data cleaning can ensure the quality and consistency of the data. The goal of data cleaning is to make the subsequent data analysis or processing proceed more smoothly by removing incorrect data, filling in missing values, standardizing formats, etc. In a data flow processing framework (such as a NiFi cluster), multiple processors are used to clean different attribute data.

[0170] For example: Use the ReplaceText processor in the second processor to clean up the dirty data in the data. The ReplaceText processor is used to replace specific characters in the text, remove irrelevant characters, or incorrect formats according to regular expressions. The ReplaceText processor can remove irrelevant characters: for example, remove special characters in the log, remove spaces or line breaks in the data, etc.; replace incorrect data formats: for example, replace incorrect date formats with correct ones; clear invalid information: for example, delete some unnecessary information from the text.

[0171] Exemplarily, assume that there is a data stream containing date data, and the date formats of some data records are incorrect or contain unwanted characters, such as extra spaces. These incorrect date formats can be replaced by the ReplaceText processor. Before: "2023-11-11T10:15:00Z"; After: "2023-11-11". When configuring the ReplaceText processor, regular expressions can be used to remove the extra characters in the date string: Search pattern (regular expression): T.*Z; Replacement pattern: "" (replace it with an empty string). This operation will convert the ISO8601 format date (e.g., 2023-11-11T10:15:00Z) to the standard date format (e.g., 2023-11-11), removing the unwanted part.

[0172] Exemplarily, use the UpdateAttribute processor in the second processor to update the attributes in the file or FlowFile. In this embodiment, this processor is used to transform and format data, such as converting a timestamp to a readable date format, or formatting other data fields.

[0173] In the NiFi node, the UpdateAttribute processor can dynamically modify the attribute values through expression language. This is especially effective for date conversion and can convert timestamps into the specified date format. Exemplarily, assume that we have a timestamp attribute ExpectedDeliveryTime, whose value is a UNIX timestamp (e.g., 1697035600000), and it needs to be converted into a readable date format. The date function in the NiFi expression language can be used to achieve this. Original data: ExpectedDeliveryTime = 1697035600000 (Unix timestamp, representing October 14, 2023, 12:00:00). Target format: yyyy-MM-dd (standard date format). When configuring the UpdateAttribute processor, set a new attribute formattedDate and use the following expression for conversion:

[0174] ${ExpectedDeliveryTime:toNumber():toDate("yyyy-MM-dd"):format("yyyy-MM-dd")}

[0175] The explanation is as follows:

[0176] toNumber(): Converts ExpectedDeliveryTime to a number if it is a number in string form.

[0177] toDate("yyyy-MM-dd"): Converts the number (i.e., Unix timestamp) to a date type.

[0178] format("yyyy-MM-dd"): Formats the date into the specified date format, i.e., yyyy-MM-dd.

[0179] The converted data will be: formattedDate = 2023-10-14. In this way, the timestamp can be converted into a standard date format, making the data easier to understand and use.

[0180] It should be noted that in this embodiment, the attribute values also need to be standardized. Standardization is the process of uniformly converting data into a unified format. This is especially important among input data from different sources. The data may have different date formats, number formats, or timestamp formats, and standardization helps ensure data consistency. For example: Suppose there are date values in different formats: 2023-11-10, 10 / 11 / 2023, and 11 / 10 / 2023 15:30. To unify the date format, all date formats can be converted into a unified format (e.g., yyyy-MM-dd) during the processing through the ReplaceText or UpdateAttribute processor.

[0181] In this embodiment, the ReplaceText processor can be configured to identify these dates and convert them uniformly, or directly use an appropriate date conversion expression in the UpdateAttribute processor. The expression for unifying the date format: ${inputDate:toDate("MM / dd / yyyy"):format("yyyy-MM-dd")}, this expression will convert the date 10 / 11 / 2023 to 2023-10-11.

[0182] It is worth mentioning that if the data itself already conforms to the standard format, or has been cleaned and transformed in subsequent processing steps, certain steps can be skipped. In other words, in this embodiment, it is necessary to decide whether to execute a certain cleaning step according to the actual situation of the data. For example, in some data sources, the date is already in the standard format, so there is no need to perform date format conversion. By setting conditional judgments or control flows (for example, using the RouteOnAttribute processor or EvaluateXPath processor in the NiFi cluster to determine whether a certain field already conforms to the expected format), you can avoid unnecessary cleaning operations and improve processing efficiency.

[0183] In this step, the goal of data cleaning is to transform the data from its original, messy state into usable, standardized data. The NiFi cluster in the embodiments of this application provides a variety of tools and processors to achieve data cleaning, such as: the ReplaceText processor, which is used to replace unwanted characters or incorrect data formats. The UpdateAttribute processor, which is used for date conversion and other attribute updates according to expression language. Through these processors, it can be ensured that the data conforms to the expected format, making subsequent processing, analysis, and storage smoother.

[0184] In one embodiment, standard business data is stored in the business library. Exemplarily, the data can be converted into database entity fields and classified and stored in the database according to business requirements. At the same time, unrecognized data can also be stored in the data lake for archiving for subsequent centralized analysis of the data; for the archived data, the operation and maintenance personnel can analyze it through the big data platform, group and classify the keywords in the log, filter out different types of log data, summarize it, and feedback it to the business personnel for confirmation of the log data. If management is required, adding a processing flow branch for collecting log data can increase the data collection for this type of log protocol.

[0185] Generally, the original data appears in an unstructured or semi-structured format (such as log files, JSON, XML). For analysis, we need to convert it into a tabular form in a relational database. In this process, the data fields are mapped according to the defined business model. Exemplarily, assume we have a web service log, and the log content includes: request time, request path, response time, HTTP status code, etc. Map these fields to the database table:

[0186]

[0187] These fields are the entity fields of the database table. When storing, each log is converted into a row of data.

[0188] Classify and store according to business requirements: The data will be classified according to business requirements. For example, it can be classified and stored according to different log types (error logs, access logs, performance logs) or time periods. That is to say, different database tables can be created or identification fields can be added for different types of logs in the same table. Exemplarily, according to the category of the logs, different tables are created to store different types of log data:

[0189]

[0190] For unrecognized data or some data that does not meet the current business requirements, in this embodiment, these data are stored in the data lake. A data lake is a storage system that can store a large amount of raw and unprocessed data. It allows storing data in various formats, including structured data, semi-structured data, and unstructured data. The data lake can archive unrecognized or data that does not conform to business rules for more in-depth analysis or processing in the future. The main advantage of the data lake is its ability to flexibly store different types of data without the need to define a strict structure in advance. The data in the data lake can be stored in its original format and then processed and analyzed through a big data platform to find new data patterns or potential business values.

[0191] In one embodiment, the data to be processed with failed tag addition is stored in the data lake. For unrecognized data or some data that does not meet the current business requirements, in this embodiment, these data are stored in the data lake, and the data type of the data to be processed stored in the data lake is obtained. Exemplarily, assume that there is a log file containing some data with non-standard or unresolvable data formats. These data will not be immediately classified into the database but are archived in the data lake. The log data may be stored in JSON format, and these logs will not be lost but are saved in the big data platform for analysis when the subsequent requirements change.

[0192] In one embodiment, it is also necessary to obtain the data type of the data to be processed stored in the data lake, and update the tags stored in the template library based on the obtained data type of the data to be processed stored in the data lake. For example, identify the data type of the data to be processed based on the file suffix of the data to be processed; and / or identify the data type of the data to be processed according to the script in the ExecuteGroovyScript processor. When there is no corresponding type tag in the template library for the data type of the data to be processed, after obtaining the data type, the type tags in the template library can be updated according to the data type to facilitate subsequent marking of corresponding tags for the data to be processed of this type.

[0193] Exemplarily, if the corresponding item of the log format is not found in the template library, the log format can be marked as unrecognized or a new format, and the unmatched format can be recorded in the template library, and the type tag in the template library can be updated to facilitate adding the processing flow for this type of data type. For example, record the key obtained by getting the root node and concatenating it in the above steps in the template library. Then, identify the data type of the data to be processed according to the file suffix of the data to be processed; or identify the data type of the data to be processed according to the script in the ExecuteGroovyScript processor; or, the template library can be accessed and maintained in advance by the application administrator of the log. The system can display the unrecognized data. After manual recognition, the application administrator inputs the recognition result, and records the recognition result input after the manual recognition as the type tag for parsing the data type of the data to be processed. In this way, the system can adapt to the changing log format without facing frequent code modifications and system upgrades.

[0194] In this way, different formats of log data can be effectively judged and parsed. By combining the preliminary judgment based on the file suffix and the parsing of the message content, meaningful field information can be finally extracted and compared with the definitions in the template library. This can not only automatically process various types of log formats, but also provide feedback when the log format is undefined, facilitating the timely update of the template library and improving the flexibility and accuracy of the log processing system. Through the tagged log data, the NiFi cluster can perform corresponding processing according to the tags, such as storage, forwarding, alerting, etc. Thus, dynamic parsing of log data is achieved. Even if the log format changes, it can adapt to the new format through configuration updates without stopping the service or making complex code modifications. In addition, after the data is archived in the data lake, the operation and maintenance personnel or data analysts can analyze this data through big data platforms (such as Apache Hadoop, Spark, etc.). The big data platforms provide powerful computing capabilities, which can help users efficiently process massive log data. The operation and maintenance personnel can perform the following operations on the data: Grouping of log key fields: Analyze the key fields in the log (such as error codes, request paths, response times, etc.) and group them to more clearly understand the distribution of different categories of logs. Exemplarily, if log analysis is required, the operation and maintenance personnel may be concerned about the occurrence frequency of each HTTP status code:

[0195]

[0196] This can help them identify which HTTP status codes occur frequently, such as the 500 error (server error), so as to respond and process in a timely manner.

[0197] Log filtering and summarization: By filtering and summarizing log data of different categories, operation and maintenance personnel can obtain more valuable information, such as the number of accesses, request error rate, performance bottlenecks, etc. within a certain period of time. Exemplarily, access log data within a certain time period can be filtered to calculate the average response time:

[0198]

[0199] Feedback to business personnel: After analyzing the log data, operation and maintenance personnel can feedback the key analysis results to business personnel to help them make decisions. For example, if it is found that certain error logs occur frequently, it may be necessary to adjust the system configuration or optimize the service.

[0200] In this step, data collection of log protocols can also be increased. If it is found that the current data collection process is insufficient to meet business requirements, new data collection logic can be added to the data pipeline. By designing new processing flow branches to collect specific types of log or protocol data. Exemplarily, if only HTTP access logs are currently collected and it is found that business requirements need to collect database query logs, operation and maintenance personnel can add a new processing flow for collecting database query logs in the data stream: a new data source processing branch can be added at the data collection layer, specifically responsible for collecting database query logs. Store this type of log data in a new data table or data lake.

[0201]

[0202] In this way, the scope of data collection can be flexibly extended to ensure the comprehensiveness and coverage of data.

[0203] In this step, the cleaned and processed data is stored in the database according to business requirements, while the unrecognized data is stored in the data lake for archiving. Through big data platforms to analyze the archived data, operation and maintenance personnel can group, filter, and summarize the log data and feedback the results to business personnel. If the current data collection process is insufficient to meet the requirements, the scope of data collection can be extended by adding processing flow branches to ensure that all key data can be collected and processed. Through the above process, enterprises can effectively perform data analysis while maintaining data integrity, help business make decisions, and improve the flexibility and efficiency of data management.

[0204] Next, a data processing method provided by a specific embodiment of the present application will be described in conjunction with the accompanying drawings.

[0205] Embodiment 1

[0206] Exemplarily, Figure 4 shows a flowchart of another data processing method provided by an embodiment of the present application. As Figure 4 As shown, this method is applied to a NiFi cluster and includes:

[0207] S201. Log data collection

[0208] First, the system collects log data from various source devices and applications through different listeners:

[0209] The ListenTCP processor: listens on a TCP port and is used to receive TCP protocol data from devices or applications.

[0210] The ListenUDP processor: listens on a UDP port and receives log data in the UDP protocol.

[0211] The ListenSyslog processor: listens on the Syslog protocol and is usually used to receive logs generated by network devices (such as routers, switches) or operating systems.

[0212] The ports opened by these processors ensure that various formats of log data can be received from different data sources.

[0213] S202. Log format judgment and tagging

[0214] The collected log data enters the ExecuteGroovyScript processor. In this step, the log data is judged in format through a Groovy script. According to the format of the log, the script will identify which category the data belongs to and tag the corresponding tags for each format of data. For example:

[0215] If the log is an access log from a web server, the tag may be web_access_log.

[0216] If the log comes from a database query, the tag may be db_query_log.

[0217] These tags help the system classify and process according to the log type later.

[0218] S203. Data routing and specific protocol parsing

[0219] After tag processing, the log data is routed through the RouteOnAttribute processor. This processor decides the data flow to different processing processes based on the previously tagged tag values. For example:

[0220] Logs of the web_access_log type may be sent to the web log parsing process;

[0221] Logs of the db_query_log type enter the database query log parsing process.

[0222] Next, the specific log protocol will be parsed:

[0223] EvaluateJsonPath Processor: When the log data is in JSON format, this processor is used to extract specific attributes from the JSON data.

[0224] EvaluateXPath Processor: When the log data is in XML format, XPath is used to extract fields from the XML data.

[0225] ExtractGrok Processor: For traditional text logs (such as Apache access logs), Grok patterns are used to extract specific fields (such as request time, IP address, request path, etc.).

[0226] S204. Data Cleaning and Standardization

[0227] After the log data is parsed, it may contain irrelevant information or data with inconsistent formats, which need to be cleaned and standardized:

[0228] ReplaceText Processor: This processor is used to replace or clean invalid or non-standard values in the data. For example, unify the formatting of datetime fields, remove redundant information from IP addresses, or desensitize certain sensitive fields.

[0229] S205. Data Storage

[0230] The processed standardized data is stored in a database for subsequent retrieval and analysis. Specific business data will be stored in different business tables as needed. For example:

[0231] Access logs of web services may be stored in a table named WebAccessLogs;

[0232] Database query logs may be stored in the DBQueryLogs table.

[0233] For data that cannot be recognized or data with mismatched formats, the system will store it in a data lake for archiving. These data can be saved in their original formats (such as JSON, CSV, text) for future further analysis and reprocessing. The data lake provides a flexible storage method that can accommodate undefined log data.

[0234] In summary, the data processing method provided by the embodiments of the present application involves the following steps: Collecting log data: Collect logs from different sources through different listeners (TCP / UDP / Syslog). Format judgment and tagging: Use Groovy scripts to judge the log format and tag different types of logs. Routing and protocol parsing: According to the tag values, the log data enters different processing flows and the log content is parsed through relevant processors (such as EvaluateJsonPath processor, ExtractGrok processor). Data cleaning and standardization: Clean the data through the ReplaceText processor to ensure that the data meets the business requirements. Storage and archiving: Store the processed data in a database for analysis, and store the unrecognized log data in a data lake for further processing. This method can flexibly process log data in various formats, ensuring the efficient storage, management, and analysis of data.

[0235] It can be understood that the magnitudes of the sequence numbers of the steps in the above various embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. In addition, in some possible implementation manners, the steps in the above embodiments can be selectively executed according to actual situations, can be partially executed, or can be fully executed, which is not limited herein. Additionally, all or part of any feature of any of the above embodiments can be freely combined in any way without conflict; the combined technical solutions are also within the scope of the present application.

[0236] As Figure 5 shown, the embodiments of the present application also provide a data processing device 500, including: an acquisition module 501 and a processing module 502. Among them, the acquisition module 501 is used to acquire at least one data to be processed; the processing module 502 is used to add a tag to the data to be processed based on the data type of the data to be processed, where the tag is used to mark the target processing flow for processing the data to be processed, and the target processing flow is one of multiple processing flows, and the processing flow corresponds to the data type, and different processing flows are used to process data of the corresponding data type through different processors; based on the tag, the data to be processed is allocated to the target processing flow for processing to obtain a processing result.

[0237] In a possible implementation manner, the processing flow includes a first processor and a second processor; the processing module 502 allocates the data to be processed to the target processing flow for processing, specifically: using the first processor to parse the data to be processed to obtain an attribute value; using the second processor to process the attribute value to obtain a processing result.

[0238] In a possible implementation, the processing module 502 uses a first processor to parse the data to be processed to obtain an attribute value, specifically: when the data type of the data to be processed is json, the EvaluateJsonPath processor is used to obtain keywords in the data to be processed as attributes of the data to be processed, and encapsulate them into the attribute value of the data to be processed; and / or when the data type of the data to be processed is xml, the EvaluateXPath processor is used to obtain element nodes in the data to be processed as attributes of the data to be processed, and encapsulate them into the attribute value of the data to be processed; and / or when the data type of the data to be processed is csv, the ConvertRecord processor is used to obtain delimiters or header configurations in the data to be processed as attributes of the data to be processed, and encapsulate them into entity attributes; and / or when the data type of the data to be processed is a text file, the ExtractGrok processor is used to obtain time fields in the data to be processed as attributes of the data to be processed, and encapsulate them into the attribute value of the data to be processed.

[0239] In a possible implementation, the processing module 502 uses a second processor to process the attribute value, specifically: using a processing method of cleaning, formatting, or standardizing to process the attribute value of the data to be processed.

[0240] In a possible implementation, the processing module 502 uses a second processor to process the attribute value to obtain a processing result, specifically: using the ReplaceText processor to uniformly format date and time fields in the attribute value, remove or replace irrelevant characters, error data, or duplicate data in the attribute value, or process missing data to standardize the data format of the attribute value, obtain the standard business data after cleaning and filtering the data to be processed, and store the standard business data in the business library.

[0241] In a possible implementation, the processing module 502 is further configured to store the data to be processed for which adding tags fails in the data lake; and obtain the data type of the data to be processed stored in the data lake.

[0242] In a possible implementation, the processing module 502 is further configured to update the tags stored in the template library according to the data type of the data to be processed stored in the data lake.

[0243] In a possible implementation, the acquisition module 501 is configured to acquire at least one piece of data to be processed, specifically: collecting source log data of different source end devices according to the ListenTCP processor, ListenUDP processor, and ListenSyslog processor.

[0244] In a possible implementation, the data type includes any one of xml, json, csv, and text files. The processing module 502 allocates the data to be processed to the target processing flow based on the tags, specifically: identifying the tags of the tag data of the xml data type by the EvaluateXPath processor and routing them to the processing flow corresponding to the tags; and / or identifying the tags of the tag data of the json data type by the EvaluateJsonPath processor and routing them to the processing flow corresponding to the tags; and / or identifying the tags of the tag data of the csv data type by the ConvertRecord processor and routing them to the processing flow corresponding to the tags; and / or identifying the tags of the tag data of the text file data type by the ExtractGrok processor and routing them to the processing flow corresponding to the tags.

[0245] In a possible implementation, the processing module 502 stores the standard business data in the business library, specifically: converting the standard business data into the entity fields of the business library and classifying and storing them according to business requirements; and centrally analyzing the archived data.

[0246] It should be understood that both the communication module 501 and the processing module 502 can be implemented by software or can be implemented by hardware. Exemplarily, next, taking the communication module 501 as an example, the implementation manner of the communication module 501 will be introduced. Similarly, the implementation manner of the processing module 502 can refer to the implementation manner of the communication module 501.

[0247] As an example of a software functional unit, the communication module 501 can include code running on a computing instance. Among them, the computing instance can include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance can be one or more. For example, the communication module 501 can include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the code can be distributed in the same region or can be distributed in different regions. Further, the multiple hosts / virtual machines / containers for running the code can be distributed in the same availability zone (AZ) or can be distributed in different AZs, and each AZ includes one data center or multiple geographically close data centers. Among them, generally, one region can include multiple AZs.

[0248] Similarly, multiple hosts / virtual machines / containers for running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Usually, one VPC is set up within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set up within each VPC, and the interconnection between VPCs is achieved through the communication gateway.

[0249] As an example of a hardware functional unit, the communication module 501 may include at least one computing device, such as a server. Alternatively, the communication module 501 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0250] The multiple computing devices included in the communication module 501 may be distributed in the same region or in different regions. The multiple computing devices included in the communication module 501 may be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the communication module 501 may be distributed within the same VPC or across multiple VPCs. Among them, the multiple computing devices may be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0251] Based on the same inventive concept, the principle and beneficial effects of the server provided in the embodiments of the present application for solving problems can be referred to the principle and beneficial effects of the method implementation. For the sake of brevity, it will not be elaborated here.

[0252] As Figure 6 shown, an embodiment of the present application also provides a computing device 600. Exemplarily, the computing device 600 may be a server or a terminal. When the computing device 600 runs, the computing device 600 may execute the method in the above embodiments.

[0253] The computing device includes: a bus 601, a processor 602, a memory 603, and a communication interface 604. The processor 602, the memory 603, and the communication interface 604 communicate with each other through the bus 601. The computing device 600 may be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 600.

[0254] The bus 601 can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of illustration, only one line is used in the figure to represent it, but it does not mean that there is only one bus or one type of bus. The bus 601 can include a path for transmitting information between various components of the computing device 600 (for example, the processor 602, the memory 603, the communication interface 604).

[0255] The processor 602 can include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0256] The memory 603 can include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0257] The memory 603 stores executable program instructions, and the processor 102 executes the executable program instructions to respectively implement the test methods involved in the above embodiments.

[0258] The communication interface 604 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement the communication between the computing device 600 and other devices or a communication network.

[0259] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0260] As Figure 7 shown, the computing device cluster includes at least one computing device 600. The same instructions for executing the data processing method can be stored in the memory 603 of one or more computing devices 600 in the computing device cluster.

[0261] In some possible implementation manners, parts of instructions for executing the data processing method may also be separately stored in the memories 603 of one or more computing devices 600 in the computing device cluster. In other words, the combination of one or more computing devices 600 may jointly execute the instructions for executing the data processing method.

[0262] It should be noted that the memories 603 in different computing devices 600 in the computing device cluster may store different instructions, which are respectively used to execute partial functions of the above-mentioned multiple modules. That is to say, the instructions stored in the memories 603 of different computing devices 600 may implement the functions of one or more of the above-mentioned multiple modules.

[0263] In some possible implementation manners, one or more computing devices in the computing device cluster may be connected through a network. Among them, the network may be a wide area network, a local area network, or the like. Figure 8 A possible implementation manner is shown. As Figure 8 shown, two computing devices, namely computing device 600A and computing device 600B, are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation manner, the memory 603 in computing device 600A stores instructions for executing partial functions of some of the above-mentioned multiple modules. At the same time, the memory 603 in computing device 600B stores instructions for executing partial functions of some other modules of the above-mentioned multiple modules.

[0264] Figure 8 The connection manner between the computing device clusters shown may be considered as the data processing method provided in this application requires a large amount of data storage. Therefore, it is considered to hand over the functions implemented by some other modules of the above-mentioned multiple modules to computing device 600B for execution.

[0265] It should be understood that Figure 8 the functions of computing device 600A shown may also be completed by multiple computing devices 600. Similarly, the functions of computing device 600B may also be completed by multiple computing devices 600.

[0266] The embodiments of this application also provide another computing device cluster. The connection relationship between the computing devices in this computing device cluster may be similarly referred to Figure 7 and Figure 8 the connection manner of the computing device cluster described above. The difference is that the memories 603 in one or more computing devices 600 in this computing device cluster may store the same instructions for executing the data processing method.

[0267] In some possible implementations, parts of the instructions for executing the data processing method may also be separately stored in the memories 603 of one or more computing devices 600 in the computing device cluster. In other words, the combination of one or more computing devices 600 can jointly execute the instructions for executing the data processing method.

[0268] It should be noted that the memories 603 in different computing devices 600 in the computing device cluster may store different instructions for executing partial functions of the computing device 600. That is, the instructions stored in the memories 603 of different computing devices 600 can implement the functions of one or more of the above-mentioned multiple modules.

[0269] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium is used to store computer program instructions. When the computer program instructions run on a computing device, the computing device is enabled to execute the data processing method involved in the above embodiment. The computer-readable storage medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive), etc.

[0270] The embodiment of the present application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, at least one computing device is enabled to execute the data processing method.

[0271] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application. Those of ordinary skill in the art should understand that although the present application has been described in detail with reference to the foregoing embodiments, they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions in the embodiments of the present application.< / age> < / name> < / user>

Claims

1. A data processing method, characterized in that: include: Obtain at least one piece of data to be processed; Based on the data type of the data to be processed, adding a label to the data to be processed, the label is used to mark a target processing flow for processing the data to be processed, wherein the target processing flow is one of a plurality of processing flows, the processing flow corresponds to the data type, and different processing flows are used to process data of corresponding data types through different processors; Based on the label, the data to be processed is assigned to the target processing flow for processing to obtain a processing result.

2. The method according to claim 1, characterized in that The processing flow includes a first processor and a second processor; the allocating the data to be processed to the target processing flow for processing includes: Using a first processor to parse the data to be processed to obtain an attribute value; The attribute value is processed by a second processor to obtain a processing result.

3. The method according to claim 2, characterized in that The adopting the first processor to parse the data to be processed to obtain the attribute value includes: In the case where the data type of the data to be processed is json, using the EvaluateJsonPath processor to obtain keywords in the data to be processed as attributes of the data to be processed, and encapsulating them into attribute values ​​of the data to be processed; and / or In the case where the data type of the data to be processed is xml, using the EvaluateXPath processor to obtain element nodes in the data to be processed as attributes of the data to be processed, and encapsulating them into attribute values ​​of the data to be processed; and / or In the case where the data type of the data to be processed is csv, a ConvertRecord processor is used to obtain a separator or a header configuration in the data to be processed as an attribute of the data to be processed, and encapsulates it into an entity attribute; and / or In the case that the data type of the data to be processed is a text file, an ExtractGrok processor is used to obtain the time field in the data to be processed as the attribute of the data to be processed, and encapsulate it into the attribute value of the data to be processed.

4. The method according to claim 2 or 3, characterized in that: The adopting a second processor to process the attribute value includes: The attribute values ​​of the data to be processed are processed by a cleaning, formatting or standardization processing method.

5. The method according to claim 4, characterized in that The using a second processor to process the attribute value to obtain a processing result includes: The ReplaceText processor is used to uniformly format the date and time fields in the attribute values, remove or replace irrelevant characters, erroneous data or duplicate data in the attribute values, or process missing data to standardize the data format of the attribute values, obtain standard business data after the data to be processed is cleaned and filtered, and store the standard business data in the business library.

6. The method according to claims 1-5, characterized in that: Also includes: Store the pending data that failed to add labels into the data lake; Get the data type of the data to be processed stored in the data lake.

7. The method according to claim 6, characterized in that Also includes: The labels stored in the template library are updated according to the data types of the data to be processed stored in the acquired data lake.

8. The method according to claim 6, characterized in that The data type of the data to be processed stored in the data lake is obtained, including: Identifying the data type of the data to be processed according to the file suffix of the data to be processed; and / or The data type of the data to be processed is identified according to the script in the ExecuteGroovyScript processor.

9. The method according to any one of claims 1 to 8, characterized in that: The obtaining of at least one piece of data to be processed comprises: Collect source log data from different source devices based on the ListenTCP processor, ListenUDP processor, and ListenSyslog processor.

10. A computing device, characterized in that include: at least one memory for storing a program; at least one processor, configured to execute the program stored in the memory; The memory is coupled to the processor, and when the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1 to 9.