Data processing method and server

Through the data flow template configuration, the inefficiency problem caused by the complex visual configuration in the prior art is solved, and the simplification and efficiency improvement of data flow configuration are achieved.

CN120295703APending Publication Date: 2025-07-11HENAN QINWEI DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510212937.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing visual configuration methods are complex, resulting in low data configuration efficiency.

Method used

By templated the data stream, users need to configure the identification of the data stream template and the data acquisition strategy, create a processing unit and a workflow for data acquisition, and connect the processing unit to complete the configuration of the data stream.

Benefits of technology

Improves the efficiency of data flow configuration, simplifies the configuration process, and improves the efficiency of data flow configuration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295703A_ABST
    Figure CN120295703A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and a server. In the embodiment, the method comprises the following steps: acquiring a data stream configuration request; wherein the data stream configuration request comprises an identifier of a data processing template and first configuration information, the data processing template is used for indicating a first data stream, the first data stream is a plurality of first processing units connected in sequence, and the first data stream is used for indicating a strategy for processing target data; the first configuration information is used for indicating the second processing unit to acquire a strategy of target data from the data provider; acquiring the data processing template according to the identifier of the data processing template; according to the data processing template and the first configuration information, a second data stream is determined, and the second data stream comprises a second processing unit and a first data stream connected with the second processing unit; and performing data processing according to the second data stream. The data stream is templated, and the user configures the identifier of the data stream template and the data acquisition strategy to complete the configuration of the data stream, so that the configuration efficiency of the data stream is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of servers, and in particular, to a data processing method and a server. Background Art

[0002] Under the background of enterprise digitalization and the rapid increase in business demand, data-driven business and data visualization systems have grown rapidly. The development method of visual configuration based on modules and templates abstracts heavy business requirements and configures module data.

[0003] However, in the existing methods, the visual configuration method is complex, resulting in low data configuration efficiency. Summary of the Invention

[0004] Embodiments of this application provide a data processing method and a server, which improve the data stream configuration efficiency.

[0005] In a first aspect, an embodiment of this application provides a data processing method, which is applied to a server. The method includes:

[0006] Obtain a data stream configuration request; wherein, the data stream configuration request includes an identifier of a data processing template and first configuration information. The data processing template is used to indicate a first data stream, and the first data stream is a plurality of first processing units connected in sequence. The first data stream is used to indicate a strategy for processing target data, and the first configuration information is used to indicate a strategy for the second processing unit to obtain target data from a data provider; obtain the data processing template according to the identifier of the data processing template; determine a second data stream according to the data processing template and the first configuration information, where the second data stream includes a second processing unit and the first data stream connected to the second processing unit; perform data processing according to the second data stream.

[0007] In this solution, by templatizing the data stream, the user only needs to configure the identifier of the data stream template and the strategy for data acquisition, then the processing unit and workflow for data acquisition can be created, and connecting the processing unit and the workflow can complete the configuration of the data stream, improving the data stream configuration efficiency.

[0008] In a possible implementation manner, determining the second data stream according to the data processing template and the first configuration information includes: creating a first data stream according to the data processing template; creating a second processing unit according to the first configuration information; connecting the created second processing unit to the first processing unit at the head of the created first data stream to obtain the second data stream.

[0009] In this solution, the processing unit is created based on the configuration information, the workflow is created based on the template, and by connecting the processing unit and the data stream, the configuration efficiency of the data stream is improved.

[0010] In a possible implementation, the data stream configuration request includes second configuration information, and the second configuration information is used to instruct the third processing unit to store the data after the first data stream is processed; determining a second data stream according to the data processing template and the first configuration information includes: determining the second data stream according to the data processing template, the first configuration information, and the second configuration information.

[0011] In a possible implementation, the first data stream includes a storage processing unit;

[0012] Determining a second data stream according to the data processing template, the first configuration information, and the second configuration information includes: creating a third processing unit according to the second configuration information; connecting the created third processing unit to the storage processing unit.

[0013] In this solution, a processing unit is created based on the configuration information, a workflow is created based on the template, and by connecting the processing unit and the data stream, the configuration efficiency of the data stream is improved.

[0014] In a possible implementation, the strategy for processing target data is used to convert multiple fields in the target data to multiple preset fields.

[0015] In a possible implementation, performing data processing according to the second data stream includes: receiving target data through the second processing unit; receiving the target data sent by the second processing unit through the first data stream, parsing the target data to obtain the first field values of multiple fields, and based on the mapping relationship between the multiple fields and the multiple preset fields, mapping the first field values to the multiple preset fields

[0016] In a possible implementation, mapping the first field values to the multiple preset fields based on the mapping relationship between the multiple fields and the multiple preset fields includes: performing arrangement processing on the first field values to obtain the second field values of the multiple fields; mapping the second field values to the multiple preset fields based on the mapping relationship between the multiple fields and the multiple preset fields.

[0017] In a possible implementation, the data stream configuration request further includes an identifier of the server.

[0018] In a possible implementation, the data processing template includes a description file and a script of the first data stream, and the script is used to describe the configuration information corresponding to the first data stream stored in the database.

[0019] In a second aspect, an embodiment of the present application provides a data processing method, which is applied to a terminal, and the method includes:

[0020] Obtain a data stream configuration request; wherein, the data stream configuration request includes first configuration content of a user for a first configuration item in a first interface, the first configuration content includes an identifier of a data processing template and first configuration information, the data processing template is used to indicate a first data stream, the first data stream is a plurality of first processing units connected in sequence, the first data stream is used to indicate a policy for processing target data, and the first configuration information is used to indicate a policy for a second processing unit to obtain target data from a data provider;

[0021] Send the data stream configuration request to a server, so that the server obtains a data processing template according to the identifier of the data processing template; determine a second data stream according to the data processing template and the first configuration information, the second data stream includes a second processing unit and the first data stream connected to the second processing unit; perform data processing according to the second data stream.

[0022] In a possible implementation manner, the method further includes:

[0023] Obtain second configuration content of the user for a second configuration item in a second interface; wherein, the second configuration content includes target data and a policy for processing the target data; generate a data processing template based on the second configuration content of the second configuration item.

[0024] In a third aspect, an embodiment of the present application provides a data processing device, the data processing device includes several modules, and each module is used to execute each step in the data processing method provided in the first aspect of the embodiment of the present application. The division of the modules is not limited here. For the specific functions executed by each module of the data processing device and the beneficial effects achieved, please refer to the functions of each step in the data processing method provided in the first aspect of the embodiment of the present application, which will not be elaborated here.

[0025] Exemplarily, the data processing device, applied to a server, includes:

[0026] A request acquisition module, configured to obtain a data stream configuration request; wherein, the data stream configuration request includes an identifier of a data processing template and first configuration information, the data processing template is used to indicate a first data stream, the first data stream is a plurality of first processing units connected in sequence, the first data stream is used to indicate a policy for processing target data, and the first configuration information is used to indicate a policy for a second processing unit to obtain target data from a data provider;

[0027] A template acquisition module, configured to obtain a data processing template according to the identifier of the data processing template;

[0028] A data stream determination module, configured to determine a second data stream according to the data processing template and the first configuration information, the second data stream includes a second processing unit and the first data stream connected to the second processing unit;

[0029] A data processing module for processing data according to the second data stream.

[0030] In a possible implementation, a data stream determination module is configured to create a first data stream according to a data processing template, create a second processing unit according to first configuration information, and connect the created second processing unit to a first processing unit at the head of the created first data stream to obtain a second data stream.

[0031] In a possible implementation, the data stream configuration request includes second configuration information for instructing a third processing unit to store data after the first data stream is processed. The data stream determination module is configured to determine the second data stream according to the data processing template, the first configuration information, and the second configuration information.

[0032] In a possible implementation, the first data stream includes a storage processing unit;

[0033] The data stream determination module is configured to create a third processing unit according to the second configuration information and connect the created third processing unit to the storage processing unit.

[0034] In this solution, by connecting the processing unit to the existing data stream, the configuration efficiency of the data stream is improved.

[0035] In a possible implementation, the strategy for processing target data is used to convert multiple fields in the target data to multiple preset fields.

[0036] In a possible implementation, the data processing module is configured to receive target data through the second processing unit, receive the target data sent by the second processing unit through the first data stream, parse the target data to obtain first field values of multiple fields, and map the first field values to multiple preset fields based on the mapping relationship between the multiple fields and the multiple preset fields.

[0037] In a possible implementation, the data processing module is configured to perform arrangement processing on the first field values to obtain second field values of multiple fields, and map the second field values to multiple preset fields based on the mapping relationship between the multiple fields and the multiple preset fields.

[0038] In a possible implementation, the data stream configuration request further includes an identifier of the server.

[0039] In a possible implementation, the data processing template includes a description file and a script of the first data stream, and the script is used to describe the configuration information corresponding to the first data stream stored in the database.

[0040] In a third aspect, an embodiment of the present application provides a data processing device. The data processing device includes a number of modules, and each module is used to execute each step in the data processing method provided in the first aspect of the embodiment of the present application. The division of the modules is not limited herein. For the specific functions executed by each module of the data processing device and the beneficial effects achieved, please refer to the functions of each step in the data processing method provided in the first aspect of the embodiment of the present application, which will not be elaborated herein.

[0041] Exemplarily, the data processing device, which is applied to a terminal, includes:

[0042] A request acquisition module, configured to acquire a data stream configuration request; wherein, the data stream configuration request includes first configuration content of a user for a first configuration item in a first interface, and the first configuration content includes an identifier of a data processing template and first configuration information. The data processing template is used to indicate a first data stream, and the first data stream is a plurality of first processing units connected in sequence. The first data stream is used to indicate a strategy for processing target data, and the first configuration information is used to indicate a strategy for a second processing unit to acquire target data from a data provider;

[0043] A request sending module, configured to send the data stream configuration request to a server, so that the server acquires a data processing template according to the identifier of the data processing template; determine a second data stream according to the data processing template and the first configuration information, where the second data stream includes a second processing unit and the first data stream connected to the second processing unit; and perform data processing according to the second data stream.

[0044] In a possible implementation manner, the device further includes:

[0045] A template generation module, configured to acquire second configuration content of a user for a second configuration item in a second interface; wherein, the second configuration content includes target data and a strategy for processing the target data; and generate a data processing template based on the second configuration content of the second configuration item.

[0046] In a fifth aspect, an embodiment of the present application provides a data processing device, including: at least one memory for storing a program; at least one processor for executing the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the methods provided in the first aspect and the methods provided in the second aspect.

[0047] In a sixth aspect, an embodiment of the present application provides a data processing device. The device runs computer program instructions to execute the methods provided in the second aspect. Exemplarily, the device may be a chip or a processor.

[0048] In one example, the device may include a processor, which may be coupled to a memory, read instructions from the memory, and execute the method provided in the second aspect or the method provided in the third aspect according to the instructions. Wherein, the memory may be integrated in the chip or the processor, or may be independent of the chip or the processor.

[0049] In a seventh aspect, an embodiment of the present application provides a server, including: at least one memory for storing a program; at least one processor for executing the program stored in the memory, and when the program stored in the memory is executed, the processor is used to execute the method provided in the first aspect.

[0050] In a seventh aspect, an embodiment of the present application provides a terminal, including: at least one memory for storing a program; at least one processor for executing the program stored in the memory, and when the program stored in the memory is executed, the processor is used to execute the method provided in the second aspect.

[0051] In an eighth aspect, an embodiment of the present application provides a computer storage medium, in which instructions are stored, and when the instructions run on a computer, the computer is caused to execute the method provided in the second aspect or the method provided in the third aspect.

[0052] In a ninth aspect, an embodiment of the present application provides a computer program product containing instructions, and when the instructions run on a computer, the computer is caused to execute the method provided in the second aspect or the method provided in the third aspect. Description of the Drawings

[0053] Figure 1a is a system architecture diagram of a data flow system provided by an embodiment of the present application;

[0054] Figure 1b is a system architecture diagram of another data flow system provided by an embodiment of the present application;

[0055] Figure 2 is a schematic flowchart of a data processing method provided by an embodiment of the present application;

[0056] Figure 3 is a schematic diagram of a data flow provided by an embodiment of the present application;

[0057] Figure 4 is a schematic diagram of determining a data processing template provided by an embodiment of the present application;

[0058] Figure 5a is a schematic diagram of a configuration interface of a data sample provided by an embodiment of the present application;

[0059] Figure 5bIt is a schematic diagram of a configuration interface and background processing logic for data recognition provided by an embodiment of the present application;

[0060] Figure 5c It is a schematic diagram of a configuration interface and background processing logic for an analysis rule provided by an embodiment of the present application;

[0061] Figure 5d It is a schematic diagram of a configuration interface and background processing logic for field layout processing provided by an embodiment of the present application;

[0062] Figure 5e It is a schematic diagram of a configuration interface and background processing logic for field mapping provided by an embodiment of the present application;

[0063] Figure 6 It is a schematic flowchart of a data processing template configuration provided by an embodiment of the present application;

[0064] Figure 7 It is a schematic diagram of an interface for data docking provided by an embodiment of the present application;

[0065] Figure 8 It is a schematic diagram of a second data stream provided by an embodiment of the present application;

[0066] Figure 9 It is a schematic flowchart of another data processing method provided by an embodiment of the present application;

[0067] Figure 10 It is a schematic diagram of a third processor provided by an embodiment of the present application;

[0068] Figure 11 It is a schematic diagram of another second data stream provided by an embodiment of the present application;

[0069] Figure 12a It is a schematic diagram of a data stream configuration scenario provided by an embodiment of the present application;

[0070] Figure 12b It is a schematic diagram of another data stream configuration scenario provided by an embodiment of the present application;

[0071] Figure 13 It is a schematic structural diagram of a data stream configuration device provided by an embodiment of the present application;

[0072] Figure 14 It is a schematic structural diagram of a computing device provided by an embodiment of the present application;

[0073] Figure 15 It is a schematic structural diagram of another data stream configuration device provided by an embodiment of the present application. Detailed implementation manners

[0074] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be described below with reference to the accompanying drawings.

[0075] In the description of the embodiments of this application, words such as "exemplary", "for example", or "for illustration purposes" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary", "for example", or "for illustration purposes" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example", or "for illustration purposes" is intended to present relevant concepts in a specific manner.

[0076] In the description of the embodiments of this application, the term "and / or" merely describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, B exists alone, and both A and B exist simultaneously. Additionally, unless otherwise specified, the meaning of the term "plural" refers to two or more. For example, multiple systems refer to two or more systems, and multiple terminals refer to two or more terminals.

[0077] Furthermore, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0078] The following explains some terms in this embodiment. It should be noted that these explanations are for the convenience of those skilled in the art and do not constitute a limitation on the scope of protection required by this application.

[0079] Data model: Composed of many fields, used to store the normalized data of each manufacturer.

[0080] NiFi: A software for data processing, mainly used for the processing and distribution of data streams. It supports highly configurable directed graphs for data routing, transformation, and system transfer logic. NiFi provides a web-based graphical interface through which users can complete process-based programming by dragging, connecting, and configuring, thus realizing functions such as data collection.

[0081] Data stream: In NiFi, a data stream is divided into multiple processing units, and each processing unit will perform certain operations on the data, such as extraction, transformation, merging, etc. These processing units are connected together in a certain order to form a complete data processing flow.

[0082] Structured Query Language (SQL): A programming language used for managing and operating relational databases. It allows users to communicate with the database by defining, querying, modifying, and controlling data. The SQL language has various functions such as data manipulation and data definition, providing great convenience for users.

[0083] SQL script: A file containing SQL statements used to interact with the database and automate complex database operations. It usually consists of a series of SQL commands to complete specific tasks, such as creating database objects, inserting data, performing batch updates, etc.

[0084] Kafka: An open-source distributed streaming data platform, also known as a distributed message queue. Used for high-throughput and low-latency data publishing and subscribing.

[0085] Pulsar: A distributed messaging system developed and maintained by the Apache Software Foundation, featuring high throughput, low latency, and horizontal scalability. Pulsar is designed with a layered architecture that separates storage and computing, allowing seamless expansion. It natively supports multi-tenancy, can isolate data for multiple tenants on the same cluster, and supports flexible message retention policies and dynamic expansion.

[0086] Transmission Control Protocol (TCP): A connection-oriented (connection-oriented), reliable, IP-based transport layer protocol.

[0087] Internet Protocol (IP): The network layer protocol in the TCP / IP architecture. The purpose of designing IP is to improve the scalability of the network: one is to solve Internet problems and achieve interconnection of large-scale and heterogeneous networks; the other is to separate the coupling relationship between top-level network applications and underlying network technologies to facilitate their independent development. According to the end-to-end design principle, IP only provides a connectionless, unreliable, best-effort packet transmission service for hosts.

[0088] Internet Protocol Address (IP address): Also translated as Internet Protocol Address. An IP address is a unified address format provided by the IP protocol. It assigns a logical address to each network and each host on the Internet to mask the differences in physical addresses.

[0089] User Datagram Protocol (UDP): Provides connectionless and unreliable delivery services, suitable for real-time services such as live broadcasts and unreliable delivery services. UDP is very simple and only adds port delivery and error detection functions to the IP datagram service.

[0090] Extensible Markup Language (XML), a subset of the Standard Generalized Markup Language, can be used to mark data and define data types. It is a source language that allows users to define their own markup languages.

[0091] XML Path Language (XPath): It is a language used to determine the location of a certain part in an XML document.

[0092] JSONPath: It is a query language used to locate and extract specific data in JSON data. It is inspired by XPath but is specifically designed to handle JSON-formatted data. Through JSONPath, users can write expressions to navigate and extract specific parts of a JSON structure.

[0093] Elasticsearch: An open-source, distributed, and real-time search and analysis engine designed to provide fast, scalable, and high-performance search solutions. It supports multiple data formats, including text, numbers, geographical locations, etc., and provides a flexible query language to meet various search needs. Elasticsearch is mainly used in scenarios such as large-scale text search, log analysis, and real-time data analysis. Elasticsearch has a database where data can be stored, and Elasticsearch can query data from its database.

[0094] Data structure: It is a format for organizing, managing, and storing data. It is a collection of data elements that have one or more specific relationships with each other. Usually, a carefully selected data structure can bring higher running or storage efficiency.

[0095] Illegal values in data: Usually refer to those data values that do not conform to the data model or constraints.

[0096] Next, an introduction will be given to the data flow system to which the data processing method provided in the embodiments of the present application may be applied. Figure 1a Shows an architecture example diagram of a data flow system provided in the embodiments of the present application. The data processing method provided in the embodiments of the present application can be applied to the Figure 1a system architecture diagram shown. As shown inFigure 1a As shown in the figure, the data flow system may further include a terminal 110 and a server cluster 120. Among them, the terminal 110 communicates with the server cluster 120 through a network. The network can be a wired network and / or a wireless network. It can be understood that the network can use any known network communication protocol to achieve different communications, and the above network communication protocol can be various wired and / or wireless communication protocols. The network can use any known network communication protocol to achieve different communications, and the above network communication protocol can be various wired communication protocols.

[0097] Among them, the terminal 110 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. Exemplary embodiments of the terminal devices involved in this solution include, but are not limited to, electronic devices equipped with iOS, android, Windows, Harmony OS, or other operating systems. The embodiments of the present application do not specifically limit the type of electronic devices.

[0098] Among them, the server cluster 120 can be implemented by an independent server or a device cluster composed of multiple servers. The servers involved in this solution can be hardware servers or can be implanted into a virtualized environment. For example, the servers involved in this solution can be virtual machines running on a hardware server including one or more other virtual machines. In addition, the servers involved in this solution can be used to provide cloud services, and it can be a server or a super terminal that can establish a communication connection with other devices and provide computing functions and / or storage functions for other devices.

[0099] Optionally, the server cluster 120 may include a data flow control platform 121 and a node cluster 122; the data flow control platform 121 is used to manage the node cluster 122. The node cluster 122 is composed of one or more nodes, and the nodes can be NiFi nodes, and the nodes are used for data processing. The main purpose of the node cluster 122 is to achieve high availability and load balancing of the data flow. By connecting multiple nodes together, the cluster mode can distribute the processing of the data flow, improve the overall processing capacity and stability of the system. Each node performs the same task but uses different data sets. One of the nodes will be automatically selected as the cluster coordinator to coordinate and synchronize the working status and data flow of each node. All nodes will send heartbeat information to the coordinator, and the coordinator is responsible for disconnecting the nodes that have not had a heartbeat for a long time and providing the latest flow information when new nodes are added.

[0100] The above system architecture is only an example and does not constitute a specific limitation. In some other possible implementation manners, such as Figure 1bAs shown in the figure, the server cluster 120 may further include a management node 123. The management node 120 is used to manage the node cluster 122. For example, it configures the data stream to be executed for the node cluster 122. Exemplarily, the management node 120 may be an object server or a virtual machine. The management node 123 can share the management pressure of the data stream management platform 121. There is no need for the data stream management platform 121 to manage a large number of node clusters 122. A management node 120 is configured for each node cluster 122, thereby improving the management efficiency.

[0101] Next, in combination with the data stream system provided above, a data processing method provided in an embodiment of the present application will be introduced in detail.

[0102] Figure 2 It is a schematic flowchart of the data processing method provided in an embodiment of the present application. This embodiment can be applied to an electronic device, specifically, it can be applied to the server cluster 120. As Figure 2 shown, the data processing method provided in an embodiment of the present application at least includes the following steps:

[0103] Step 201, the terminal 110 obtains a data stream configuration request; wherein, the data stream configuration request includes an identifier of a data processing template and first configuration information. The data processing template is used to indicate a first data stream. The first data stream is a plurality of first processing units connected in sequence. The first data stream is used to indicate a strategy for processing target data. The first configuration information is used to indicate a strategy for the second processing unit to obtain target data from a data provider.

[0104] Among them, the data processing template is used to indicate the first data stream. Optionally, in one example, the data processing template may include a description file of the first data stream. For example, when the first data stream is processed by NIFI, the description file may be a NIFI file. The description file may include the type of each first processor and the configuration content of each first processor among a plurality of first processors; Optionally, in another example, the data processing template may include a description file of the first data stream and a script. The script is used to indicate the storage structure of the configuration information (strategy for processing target data) of the first data stream. For example, the script may be used to describe the configuration information corresponding to the first data stream stored in a database. The database may be SQL, and the script may be an SQL script. It should be noted that the script can facilitate users to view the configuration information of the first data stream and can modify the first data stream, thereby obtaining other data processing templates. The identifier of the data processing template may be assigned by the data stream management platform 121 or determined by the user.

[0105] The first data stream is a plurality of first processing units connected in sequence, where each first processing unit is used to illustrate a processing method; for example, as Figure 3As shown, the first data stream includes 22 processing units, and the 22 processing units are connected in a certain order. The types of the first processing unit can include ExtracText, judge-classify, ExtractGrok, EvaluateJsonPath, EvaluateXpath, UpdateAttribute, ExecuteScript, ReplaceText, original-json, PulishKafka, package-save-json, MergeContent. Among them, ExtracText is used to extract specific text data, supports regular expressions and fixed text patterns, and can extract different formats of data according to different requirements; judge-classify-52 is used to classify and judge data, can judge categories according to attributes or content, and perform different operations accordingly. For example, it can perform classification processing according to conditions such as file type, file size, content keywords, etc.; ExtractGrok is mainly used for log analysis and data processing to extract key information in the log; EvaluateJsonPath is used to extract specified fields in JSON data according to JSONPath expressions; EvaluateXpath is based on the XPath expression language and can select and extract nodes and attributes in XML or HTML documents by specifying a path; UpdateAttribute is used to update the attributes of the flow file through the attribute expression language and can delete attributes that meet the conditions according to regular expressions; ExecuteScript is used to allow users to perform data processing and conversion through custom scripts; ReplaceText is used to find and replace specified text in the data stream. It supports using regular expressions for matching and replacement and can handle various complex matching requirements; original-json is used to split JSON strings and retain the original JSON data; PulishKafka is used to send data to the Kafka cluster; package-save-json is used to package and save the first data stream (including configurations such as processing units and connections) as a JSON file, and this JSON file can be used for backup, migration, or import; MergeContent is used to merge multiple files into one file, can merge all necessary files into one snapshot, and can select the merge method according to requirements, such as merging by file name, file size, etc.

[0106] Note that there are multiple data processing templates stored in the server cluster 120. Optionally, the multiple data processing templates may adopt the same data processing procedure for different data. In other words, the processing units in the data flow of the multiple data processing templates and the connection relationships between the processing units are the same, but the processed data is different. For example, there may be three data processing templates, one for processing log data, one for processing network data, and another for processing data collected by sensors; optionally, the multiple data processing templates may adopt different data processing procedures for the same data. In other words, the processing units in the data flow of the multiple data processing templates and the connection relationships between the processing units are different, but the processed data is the same. For example, there may be three data processing templates, all three processing log data, with different numbers, types of processing units, and connection relationships between the processing units among the three data processing templates; optionally, the multiple data processing templates may include multiple data processing templates that adopt the same data processing procedure for different data, and multiple data processing templates that adopt different data processing procedures for the same data; the data processing template in the data flow configuration request may be one or more. For any data processing template in the data flow configuration request, the data processing template may be any one of the multiple data processing templates stored in the server cluster 120.

[0107] The first data flow in the data processing template is used to indicate the strategy for processing the target data. Optionally, in some scenarios, the strategy for processing the target data may be to map multiple fields in the target data to multiple preset fields. Among them, for each preset field in the multiple preset fields, the name of the preset field may be the same as the name of the field in the target data, or may be a synonymous field of the field in the target data. A synonymous field can be understood as a field with the same meaning but different expression as the field.

[0108] In one example, the strategy for processing the target data may include parsing the target data to obtain the first field values of multiple fields. The first field values are the field values of each field in the multiple fields, and the first field values will vary according to the actual situation. Based on the mapping relationship between the multiple fields and the multiple preset fields, the first field values of the multiple fields are mapped to the multiple preset fields to obtain the field values of the multiple preset fields. Correspondingly, the configuration information of the first data flow may include the target data, the parsing rule, and the mapping relationship between the multiple fields and the multiple preset fields. Among them, the parsing rule may be xml, Json, or plain text, etc. The mapping relationship between the multiple fields and the multiple preset fields may be recorded in the form of key-value. The key is the key value, which can be a preset field, and the value can be the field in the target data corresponding to the preset field.

[0109] Optionally, in one example, mapping the first field values of multiple fields to multiple preset fields based on the mapping relationship between the multiple fields and the multiple preset fields may include: performing an arrangement process on the first field values, for example, performing an arrangement process on at least some of the field values in the first field values respectively to obtain the second field values of the multiple fields; and mapping the second field values to the multiple preset fields based on the mapping relationship between the multiple fields and the multiple preset fields. Correspondingly, the configuration information of the first data stream may further include the arrangement process of at least some of the multiple fields respectively. Among them, the arrangement process of the field values may include an arrangement method and a processing method. The arrangement method may be custom, field splicing, or regular matching, which can be specifically determined according to actual requirements; among them, custom means that the user sets the field values by himself; among them, field splicing is composed of several fields spliced together, which can express more content. For example, the multiple fields may be the attack method and the attack result, and the spliced field may be the attack method + the attack result; among them, regular matching extracts the field values in the way of regular expressions. Among them, the processing method may be desensitization processing, such as encryption, case conversion, timestamp conversion, which can be specifically determined according to actual requirements.

[0110] Optionally, in one example, the policy for processing target data may further include parsing the target data when it is determined that the target data belongs to the target data identifier. Among them, the target data identifier is any field in the multiple fields of the target data, or a combination of some fields in the target data. Correspondingly, the configuration information of the first data stream may include the target data identifier of the target data.

[0111] The data processing template may be provided by the data stream management platform 121 or customized by the user. In some possible scenarios, optionally, the user may log in to the data stream management platform 121 through the terminal 110, for example, by logging in with an account, and view multiple data processing templates pre-configured on the management side of the data stream management platform 121 (that is, the developer of the data stream management platform 121). When the user uses the multiple data processing templates, the user may import the description file of the data stream in the data processing template in the interface provided by the management side of the node cluster 122 (different from the data stream management platform 121, which is a management platform specifically for the node cluster 122), and import the script of the data stream into the database corresponding to the node cluster 122, so that the configuration content of the data stream is stored in the database. Subsequently, the data processing template can be used normally; the user can log in to the data stream management platform 121 and import the data processing template; optionally, as Figure 4As shown, the user can log in to the data stream management platform 121 through the terminal 110, for example, by logging in with an account, and can customize the data stream using the data stream configuration interface (which can also be referred to as the second interface) provided by the data stream management platform 121 to obtain a data processing template. It should be noted that the data processing template includes the script of the data stream, so that the user can directly view the configuration information of the data stream and can modify the configuration information of the data stream, so as to quickly obtain other data processing templates.

[0112] In an optional implementation, customizing the data stream using the data stream management platform 121 may include obtaining the second configuration content of the user for the second configuration item in the second interface; wherein, the second configuration content includes the target data and the strategy for processing the target data; based on the second configuration content of the second configuration item, generate a data processing template.

[0113] In an example, the second configuration item is used to indicate the parameters required to configure the strategy for data processing. For example, the second configuration item may include data samples, parsing rules, field arrangement processing, field mapping, and further may include data identifiers. The second configuration content may include the target data, parsing rules, the mapping relationship between multiple fields and multiple preset fields, the arrangement processing of fields, and the target data identifier.

[0114] Optionally, the second interface may include an interface for data samples, such as Figure 5a As shown, the interface for data samples may include an input box, and the user can enter multiple fields in the input box, or enter multiple fields and the field values of each field.

[0115] Optionally, the second interface may include an interface for data identification, such as Figure 5b As shown, the parameters that can be configured in the interface for data identification are the identification method: automatic identification or custom identification. Among them, for automatic identification, the system assigns an identification field; for custom identification, the user customizes the identification field. For example, it is defined that the data content includes field 1 in a certain log, and subsequently, in the manner of OR or AND, it is continued to define that the data content includes field 2 in the log, and field 2 can also be deleted. Continue to refer to Figure 5b, the processing logic corresponding to the data recognition interface is implemented by processing units 1 to 6. Among them, processing units 1 and 1 are configured by the user side, that is, the processing units are configured with the content configured by the user in the data recognition interface. Processing units 2 to 4 are configured by the management side without the user's awareness. The management side configuration means that it is configured by the provider of the data stream management platform 121. Among them, processing unit 1 can be used to extract the field value of a field according to a regular expression. Processing unit 2 is used to determine whether there is a field value. If there is, processing units 3 and 4 replace the illegal values in the data, such as the backslash \. Processing unit 5 constructs a data structure. Processing unit 6 is used to store the data into the Elasticsearch database.

[0116] Optionally, the second interface may include an interface for parsing rules, such as Figure 5c As shown, the parameters that can be configured in the interface for parsing rules are the parsing type and the rule syntax. The rule syntax can be a regular expression. For example, the parsing type is JOSN parsing and the rule syntax is %{data=json}. Continue to refer to Figure 5c , the interface for parsing rules is implemented by processing units 7 to 12. Among them, processing units 7 to 12 are all configured by the user side, that is, the processing units are configured with the content configured by the user in the interface for parsing rules. Among them, processing unit 7 can be used for preprocessing of extracting json data, such as converting plain text data into a json object. Processing unit 8 is used to extract the field value from data, such as a log file. Processing unit 9 is used to extract the field value according to a regular expression. Processing unit 10 can process the extracted field value. For example, for an array matched by a regular expression, it can be configured to use the nth element in the array. Processing unit 10 is used to extract XML data. Processing unit 12 is used to process complex logic, such as processing arrays in the data.

[0117] Optionally, the second interface may include an interface for field value arrangement and processing, such as Figure 5dAs shown, the parameters configured in the interface for field value arrangement and processing include field name, field type, arrangement method, and processing method; the field type can be text; the field assignment is used to describe the arrangement method of the field. For example, the arrangement method is used to describe how to obtain the field value, and 3 methods can be provided: custom, field concatenation, and regular expression matching (the regular expression for extracting the field value). In the input box below the arrangement method, custom content, the concatenated field, and the regular expression used for regular expression matching can be entered; the processing method is used to describe how the field value is processed, and 2 options can be provided: no processing and processing. In the case of selecting no processing, only the field value needs to be obtained; after selecting the processing option, the specific processing method of the field value needs to be configured in the input box below. Subsequently, after obtaining the field value, the field value needs to be processed according to the processing method. Multiple specific processing methods can be configured, such as encryption, decryption, desensitization. The desensitization processing is used to describe the encryption method such as Base64_Decode, or deleting sensitive data such as personal names. In addition, the parameters configured during the field arrangement and processing can also include whether the field value is enumerated, and whether it is a multi-value. Continue to refer to Figure 5d , the processing logic corresponding to the interface for field value arrangement and processing is implemented by processing units 13 to 21. Among them, processing units 13 to 18 are configured by the user side, that is, the processing units are configured using the content configured by the user in the interface for field value arrangement and processing. Processing units 19 to 21 are configured by the management side without the user's awareness. The management side configuration means that it is configured by the provider of the data stream management platform 121; among them, processing unit 13 can be used to perform field concatenation according to: combining the extracted field values, and processing units 14 to 18 are used to process the field values: such as encryption, decryption, desensitization, case conversion, etc.; processing unit 19 is used for text replacement, processing unit 20 is used for data merging, and processing unit 21 is used for sending data.

[0118] Optionally, the second interface can include the interface for field value arrangement and processing, such as Figure 5e As shown, the parameters that can be configured for field arrangement and processing can include the fields in the data sample, the preset fields for mapping, and the unmapped and mapped fields in the data sample can also be displayed for the user to view. Continue to refer to Figure 5e , the interface for field value arrangement and processing is implemented by processing unit 21. Among them, processing unit 21 is configured by the user side, that is, processing unit 21 is configured using the content configured by the user in the interface for field value arrangement and processing; among them, processing unit 21 is used to perform normalized data, and based on the mapping relationship between the fields in the data and the standard fields, the field value is associated with the preset field.

[0119] It should be noted that the data processing template includes the script of the data stream, and this script can restore Figures 5a to 5eThe configuration of the interface shown facilitates the user to view the configuration of the data stream interface and allows the information in the interface to be modified.

[0120] In some optional examples, such as Figure 6 shown, the target data is a log file, and multiple fields in the target data can be called log fields. If the configuration process of the data processing template is applied to the data stream management platform 121, the configuration process of the data processing template may include the following steps:

[0121] 1. Upload data samples. As Figure 5a shown, the data samples include multiple log fields. For example, log_time_str, log_time, device_sn, src_mac, dst_mac, src_location, src_location, app_protocol, app_name, DNS, session_dir, src_ip, ip_protocol, src_port, dst_port, interface, trans_id, flags, opcode, req_type, req_class, req_name, rsp_addr, answer_cname, rsp_ttl, rsp_code, rsp_content, rsp_soa, dns_type, quest_count, answer_count.

[0122] 2. Define the basis for data recognition. As Figure 5b shown, the data recognition interface may include recognition methods, and the recognition methods include automatic recognition and custom recognition; among them, for automatic recognition, the system assigns identification fields; for custom recognition, the user defines the identification fields. For example, it is defined that the data content contains a certain log field such as log_time_str, and subsequently, in the way of OR or AND, it continues to be defined that the data content contains other log fields such as DNS, src_location, and req_name.

[0123] 3. Define the data parsing method. As Figure 5c shown, the parsing type and rule syntax can be configured in the parsing rule interface. For example, the parsing type is JSON parsing, and the rule syntax is %{data=json}.

[0124] 4. Perform field value arrangement processing on the log fields in the data samples. As Figure 5dAs shown, the parameters that can be configured in the interface for field value arrangement processing can include field name, field type, arrangement method, and processing method. The field type can be text. Field assignment is used to describe the arrangement method of the field. For example, the arrangement method is used to describe how to obtain the field value, and three methods can be provided: custom, field splicing, and regular expression matching (the regular expression for extracting the field value). For custom, field splicing, and regular expression matching (the regular expression for extracting the field value), custom content, the spliced field, and the regular expression used for regular expression matching can be entered in the input box below the arrangement method. The processing method is used to describe how the field value is processed, and two options can be provided: do not process and process. In the case of selecting not to process, the field value can be obtained directly. After selecting the option of processing, the processing method of the field value needs to be configured in the input box below. Subsequently, after obtaining the field value, the field value needs to be processed according to the processing method. One or more processing methods can be configured, such as encryption, decryption, desensitization, case conversion, etc. Desensitization processing is used to describe the encryption method such as Base64_Decode, or to delete sensitive data such as personal names. In addition, the parameters configured during the field arrangement processing can also include whether the field value is enumerated, and whether it is a multi-value field.

[0125] 5. Select a suitable data model, which can be understood as a collection of multiple preset fields. A suitable data model includes the preset fields corresponding to each field in the data sample. The data model can be pre-configured by the management side in the data stream management platform 121, or can be customized and configured by the user when logging in to the data stream management platform 121. The user can view the data model to select a suitable data model for the data sample.

[0126] 6. Perform mapping processing on the preset fields in the data model and the fields in the data sample. The mapping processing is used to determine the preset fields in the data model to which the fields in the data sample need to be converted. In specific implementation, as Figure 5e shown, the fields in the data sample and the mapped preset fields (the standard fields in the data model) can be entered in the input box in the field mapping interface. When entering the fields in the data sample, the unmapped fields and the mapped fields in the data sample can be referred to.

[0127] 7. Generate a data processing template. For example, after the user completes the configuration of the above steps 1 to 6, the user can click OK. The data stream management platform 121 sends the configuration content of the interfaces in steps 1 to 6 to the node cluster 122. The node cluster 122 generates a description file of the data stream based on the configuration content of the interfaces in steps 1 to 6. The data stream management platform 121 sends the configuration content of the interfaces in steps 1 to 6 to the database corresponding to the node cluster 122, and an SQL script of the data stream is generated in the database. Here, the description file and the SQL script can be used as data processing templates. It should be noted that the above steps 1 and 6 are pre-corresponded to the processing units in the data stream. The data stream management platform 121 can send the data stream pre-corresponding to steps 1 and 6 to the node cluster 122, and the node cluster 122 configures the configuration content of the interfaces in steps 1 to 6 into the processing units in the corresponding data stream to generate a description file of the data stream.

[0128] Among them, the first configuration information is used for the second processor to obtain target data from a data provider (the party that provides the target data, which can be a device enterprise, software, etc.). Among them, the data provider can be a device or software. Exemplarily, the device can be a switch or a router, and the software can be a log collection software, and the target data is a log file. Optionally, in one example, the first configuration information may include the type of the second processing unit, the data collection method, and the data collection configuration. The type of the second processing unit can be ConsumePulsar, and ConsumePulsar is used to process and consume messages in Apache Pulsar. The data collection method can be passive collection (that is, passively receive data), or active collection (that is, actively obtain data). The data collection configuration is used to describe the configuration parameters required in the data collection method, and is used to describe the data provider and the target data. Exemplarily, the data collection method can be Kafka, Pulsar, etc., and the configuration parameters can include the receiving address, receiving port, authentication method, security protocol, account, password, Topic, etc. Exemplarily, the data receiving method can be Tcp, Udp, and the configuration parameters can be the source IP address, source port, destination IP address, destination port, and transport layer protocol. The above specific data collection methods are only examples and do not constitute specific limitations, and can be designed specifically in combination with actual requirements.

[0129] In some possible implementation manners of this embodiment, the user can access the data stream management platform 121 through the terminal 110. The data stream management platform 121 can display a first interface, and the user can operate the first configuration item in the first interface to obtain the first configuration content, thereby generating a data stream configuration request.

[0130] In some possible examples, in the data stream system, Figure 1bIn the scenario of the architecture shown, the data stream configuration information may further include the address of the management node 120, such as an IP address. Exemplarily, when the terminal 110 can access the data stream control platform 121, such as by logging in with an account, it displays as Figure 7 the data docking interface shown. The data docking interface includes a data docking name, a collection node, a collection method, a receiving method, a parsing plugin, and configuration information. Among them, the collection node is used to describe the IP address of the management node 120. The collection methods include passive reception and active collection. The receiving methods can be Kafka, Pulsar, Tcp, Udp. Exemplarily, when the receiving method is Pulsar, the configuration parameters can be a receiving address, a receiving port, an authentication method, a security protocol, an account, a password, and a Topic. The data processing template is used as a plugin, and the parsing plugin is used to select the identifier of the data processing template, and one or more identifiers of the data processing template can be selected.

[0131] Step 202: The terminal 110 sends a data stream configuration request to the server cluster 120.

[0132] Step 203: The server cluster 120 obtains the data stream configuration request.

[0133] Step 204: The server cluster 120 obtains the data processing template according to the identifier of the data processing template.

[0134] The server cluster stores the data processing template and the identifier of the data processing template, and the server cluster can obtain the data processing template based on the identifier of the data processing template.

[0135] Step 205: The server cluster 120 determines a second data stream according to the data processing template and the first configuration information. The second data stream includes a second processing unit and a first data stream connected to the second processing unit.

[0136] In some optional implementation manners of this embodiment, the server cluster 120 creates a first data stream according to the data processing template: a plurality of first processing units connected in sequence; creates a second processing unit according to the first configuration information. Considering that the second processing unit is used as a data entry, and the first processing unit at the head of the first data stream is used as the starting point for data entry processing. In order for the first data stream to process the target data obtained by the second processor, the created second processing unit is connected to the first processing unit at the head of the created first data stream to obtain a second data stream.

[0137] Exemplarily, the second processing unit is processing unit 23: ConsumePulsar, the processing unit at the head of the first data stream is processing unit 1, and processing unit 23 is connected to processing unit 1 to obtain Figure 8 the second data stream shown.

[0138] Among them, the server cluster 120 creating multiple first processing units connected in sequence according to the data processing template may include: The server cluster 120 creates multiple first processing units connected in sequence to be configured based on the data processing template, and configures the configuration information of the first data stream indicated by the data processing template (including the configuration of each first processing unit among the multiple first processing units) into the multiple first processing units, thereby successfully creating multiple first processing units.

[0139] Among them, creating a second processing unit according to the first configuration information may include: creating a second processing unit to be configured based on the type of the second processing unit in the first configuration information, and configuring the configuration information of the second processing unit in the first configuration information into the second processing unit, successfully creating the second processing unit.

[0140] It should be noted that the second data stream is deployed in one or more nodes in the server cluster 120, and a second processing unit and / or one or more first processing units in the first data stream can be created in one node. In this embodiment, after the server cluster 120 creates the second processing unit based on the first configuration information and creates the first data stream based on the data processing template (which can also be understood as instantiating the data processing template to obtain a real data stream of the data processing template), the second processing unit and the first data stream are connected, thereby improving the configuration efficiency of the data stream.

[0141] Step 206, the server cluster 120 processes data according to the second data stream.

[0142] In some optional implementation manners of this embodiment, the server cluster 120 receives target data through the second processing unit; receives the target data sent by the second processing unit through the first data stream, and converts multiple fields in the target data into multiple preset fields.

[0143] In an optional example, converting multiple fields in the target data into multiple preset fields may include parsing the target data to obtain the first field values of the multiple fields, and based on the mapping relationship between the multiple fields and the multiple preset fields, mapping the first field values to the multiple preset fields. For example, assume that the multiple fields are field 1, field 2, and field 3, field 1 is mapped to preset field 1, field 2 is mapped to preset field 2, field 3 is mapped to preset field 3, and the field values of field 1, field 2, and field 3 are field value 1, field value 2, and field value 3 respectively. Then, associate field value 1 with preset field 1, field value 2 with preset field 2, and field value 3 with preset field 3, thereby realizing data normalization.

[0144] Exemplarily, based on the mapping relationship from multiple fields to multiple preset fields, mapping the first field value to multiple preset fields may include: performing an arrangement process on the first field value to obtain the second field values of multiple fields; based on the mapping relationship from multiple fields to multiple preset fields, mapping the second field values to multiple preset fields. For example, assume that the multiple fields are field 1, field 2, and field 3, field 1 is mapped to preset field 1, field 2 is mapped to preset field 2, field 3 is mapped to preset field 3, the respective field values of field 1, field 2, and field 3 are field value 1, field value 2, and field value 3, and field value 1 needs to be encrypted to obtain the encrypted field value 1. Then, associate the encrypted field value 1 with preset field 1, field value 2 with preset field 2, and field value 3 with preset field 3, thereby achieving data normalization.

[0145] In some optional implementation manners of this embodiment, there may be multiple data processing templates in the data flow configuration request. Multiple data processing templates respectively generate corresponding second data flows through the above step 205. If any one of the second data flows fails to execute, then execute other second data flows. Optionally, the server cluster 120 may serially execute all the second data flows. If one fails to execute, then execute the next data flow. It should be noted that this embodiment does not intend to limit the execution order of the second data flows. For the convenience of description and understanding, take 2 data processing templates as an example for description, which may be referred to as data processing template 1 and data processing template 2. Data processing template 1 generates data flow 1 in the manner of the above step 205, and data processing template 2 generates data flow 2 in the manner of the above step 205. After data flow 1 fails to execute, execute data flow 2.

[0146] In this solution, by templatizing the data flow, the user only needs to configure the identifier of the data flow template and the data acquisition strategy, then the processing unit and workflow for data acquisition can be created, and connecting the processing unit and the workflow can complete the configuration of the data flow, improving the data flow configuration efficiency.

[0147] Figure 2 The shown is only the basic embodiment of the method of the embodiments of the present application. Based on it, with certain optimizations and expansions, other preferred embodiments of the method can also be obtained.

[0148] In one embodiment, on the basis of the foregoing embodiment, a more specific description and a certain degree of optimization are made for the data processing process. As Figure 9 shown, the method in this embodiment includes:

[0149] Step 901: The terminal 110 obtains a data stream configuration request. The data stream configuration request includes an identifier of a data processing template and first configuration information. The data processing template is used to indicate a first data stream. The first data stream is a plurality of first processing units connected in sequence, and the first data stream is used to indicate a policy for processing target data. The first configuration information is used to indicate a policy for the second processing unit to obtain target data from a data provider, and the second configuration information is used to indicate a policy for the third processing unit to store the data after the first data stream is processed.

[0150] In some optional implementation manners of this embodiment, the second configuration information may include the type of the third processing unit and the configuration of the third processing unit. Optionally, there may be one or more third processing units. The type of the third processing unit may include PutElasticsearchHttp, ReplaceText, and MergeContent. PutElasticsearchHttp is used to write data into an Elasticsearch database, supports batch writing operations, and can write the data in multiple files into the Elasticsearch database at one time, improving the efficiency of data processing.

[0151] Exemplarily, as Figure 10 shown, the third processing units may be processing unit 24 to processing unit 29. Processing unit 24 is used to store data (the data after the first data stream is processed) into an Elasticsearch database. Processing unit 25 is used to extract the original data (i.e., the target data) when an error occurs during the process of storing data into the Elasticsearch database. Processing units 25 - 27 are used to process illegal values in the original data. Processing unit 28 is used to construct a new data structure for the processed original data. Processing unit 29 is used to store data into the Elasticsearch database.

[0152] Step 902: The terminal 110 sends the data stream configuration request to the server cluster 120.

[0153] Step 903: The server cluster 120 obtains the data stream configuration request.

[0154] Step 904: The server cluster 120 obtains the data processing template according to the identifier of the data processing template.

[0155] Step 905: The server cluster 120 determines a second data stream according to the data processing template, the first configuration information, and the second configuration information. The second data stream includes a second processing unit, the first data stream connected to the second processing unit, and the third processing unit connected to the first data stream.

[0156] In some optional implementation manners of this embodiment, the server cluster 120 creates a first data stream according to a data processing template: a plurality of first processing units connected in sequence; creates a second processing unit according to first configuration information; considering that the second processing unit is used as a data entry, and the first processing unit at the head of the first data stream is used as the starting point for data entry processing, in order for the first data stream to process the target data obtained by the second processor, the created second processing unit is connected to the first processing unit at the head of the created first data stream; creates a third processing unit according to second configuration information. In some optional implementation manners of this embodiment, the server cluster 120 creates a first data stream according to a data processing template: a plurality of first processing units connected in sequence; creates a second processing unit according to first configuration information; considering that the third processing unit is used to store the data processed by the first data stream, and the first data stream includes a storage processing unit as a data outlet, in order for the third processing unit to store the data processed by the first data stream, the created third processing unit is connected to the storage processing unit in the created first data stream to obtain a second data stream.

[0157] Exemplarily, the second processing unit is processing unit 23: ConsumePulsar, the processing unit at the head of the first data stream is processing unit 1, processing unit 23 is connected to processing unit 1, the storage processing unit is processing unit 21 (package-save-json), the third processing unit is processing units 22 - 29, and processing unit 21 and processing unit 24 are connected to obtain the second data stream as shown in Figure 11 the following figure.

[0158] It should be noted that the second data stream is deployed in one or more nodes of the server cluster 120. One or more of the second processing unit, the third processing unit, and / or the first processing units in the first data stream can be created in one node. In this embodiment, after the server cluster 120 creates the second processing unit based on the first configuration information, creates the third processing unit based on the second configuration information, and creates the first data stream based on the data processing template (it can also be understood as instantiating the data processing template to obtain a real data stream of the data processing template), the second processing unit, the third processing unit, and the first data stream are connected, thereby improving the configuration efficiency of the data stream.

[0159] Step 906: The server cluster 120 processes data according to the second data stream.

[0160] For the detailed content, please refer to the description of step 205 above and will not be elaborated here.

[0161] In this solution, by templatizing the data stream, the user needs to configure the identifier of the data stream template, the strategy for data acquisition, and the strategy for data storage, and then can create a processing unit for data acquisition, a processing unit for data storage, and a workflow. Connecting the processing units and the workflow can complete the configuration of the data stream, improving the efficiency of data stream configuration.

[0162] Based on the above-provided data processing method, a specific application of the data processing method will be described.

[0163] Figure 12a It is a schematic flowchart of a specific application of a data processing method provided for the implementation of this application. As Figure 12a shown, the specific content includes:

[0164] The terminal 110 configures the identifier of the data processing template, the first configuration information for obtaining target data (for the detailed description of the first configuration information, refer to the description in step 201), and the second configuration information for storing the data after the first data stream is processed (for the detailed description of the second configuration information, refer to the description in step 901), obtains a data stream configuration request, and sends the data stream configuration request to the data stream control platform 121.

[0165] The data stream control platform 121 creates a first data stream in the node cluster 122 based on the description file of the first data stream in the data processing template, creates a first processing unit in the node cluster 122 based on the first configuration information, creates a third processing unit in the node cluster 122 based on the second configuration information. The first processing unit is connected to the first processing unit at the beginning of the first data stream, and the processing unit at the end of the first data stream, such as the storage processing unit, is connected to the third processing unit, thereby obtaining a second data stream.

[0166] Figure 12b It is a schematic flowchart of a specific application of a data processing method provided for the implementation of this application. As Figure 12b shown, the difference from Figure 12a is that the data stream configuration request includes the identifier of the management node 123, such as the IP address. The data stream management platform 121 sends the data stream configuration request to the management node 123, and the management node 123 configures the second data stream in the node cluster 122.

[0167] Based on the same concept as the method embodiments of the present application, the embodiments of the present application further provide a data processing device, which is applied to the server cluster 120. The data processing device includes a number of modules, and each module is used to execute each step in the data processing method provided by the embodiments of the present application. The division of the modules is not limited herein. Those skilled in the art can clearly understand that in practical applications, each step in the data processing method provided by the embodiments of the present application can be allocated to different modules according to needs, that is, the internal structure of the device is divided into different modules to complete all or part of the functions described above. Each module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more modules can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the modules are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the modules in the above device can refer to the corresponding process in the foregoing method embodiments and will not be repeated here.

[0168] Exemplarily, the data processing device is used to execute the data processing method provided by the embodiments of the present application. Figure 13 It is a schematic structural diagram of the data processing device provided by the embodiments of the present application. As Figure 13 shown, the data processing device provided by the embodiments of the present application includes:

[0169] A request acquisition module 1301, configured to acquire a data stream configuration request; wherein, the data stream configuration request includes an identifier of a data processing template and first configuration information, the data processing template is used to indicate a first data stream, the first data stream is a plurality of first processing units connected in sequence, the first data stream is used to indicate a strategy for processing target data, and the first configuration information is used to indicate a strategy for the second processing unit to acquire the target data from a data provider;

[0170] A template acquisition module 1302, configured to acquire the data processing template according to the identifier of the data processing template;

[0171] A data stream determination module 1303, configured to determine a second data stream according to the data processing template and the first configuration information, the second data stream includes the second processing unit and the first data stream connected to the second processing unit;

[0172] A data processing module 1304, configured to perform data processing according to the second data stream.

[0173] Based on the same concept as the method embodiments of the present application, the embodiments of the present application further provide a computing device. The computing device can be a server. Figure 14It is a schematic structural diagram of a computing device provided by an embodiment of the present application.

[0174] As Figure 14 shown, the computing device 1400 includes a processor 1401, a memory 1402, and a network interface 1403.

[0175] The processor 1401 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0176] The memory 1402 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0177] Exemplarily, a computer program may be stored on the memory 1402. When the processor 1401 executes the computer program, the steps in the above-described embodiments of the data processing method are implemented. For example Figure 2 the steps 203 to 205 shown, or Figure 9 the steps 903 to 905 shown. Alternatively, when the processor 1401 executes the computer program, the functions of the various modules in the above-described device embodiments are implemented. Exemplarily, the computer program may be divided into one or more modules / units. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions. The one or more modules / units are stored in the memory 1402 and executed by the processor 1401 to complete the present application. For example, the computer program may be divided into a request acquisition module 1301, a template acquisition module 1302, a data stream determination module 1303, and a data processing module 1304. For the specific functions of each module, refer to the above description.

[0178] The network interface 1403 is used for sending and receiving data. For example, the data processed by the processor 1401 is sent to other computing devices, or data sent by other computing devices is received, etc.

[0179] Of course, for simplicity Figure 14 only some of the components related to the present application in the computing device 1400 are shown in [the figure], and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the computing device 1400 may further include any other appropriate components. Additionally, the computing device may be a desktop computer, a notebook, a palm computer, a cloud server, or other computing devices. Those skilled in the art can understand that Figure 14 merely an example of the computing device 1400, which does not constitute a limitation on the computing device. It may include more or fewer components than shown in the figure, or combine certain components, or have different components. For example, the computing device may further include an input device, an output device, a network access device, a bus, etc. Exemplarily, the input device may be a microphone array and may also include, for example, a keyboard, a mouse, etc. Exemplarily, the output device may output various information to the outside and may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0180] Based on the same concept as the method embodiment of the present application, an embodiment of the present application further provides a data processing device, which is applied to the terminal 110. The data processing device includes a number of modules, and each module is used to execute each step in the data processing method provided by the embodiment of the present application. The division of the modules is not limited herein. Those skilled in the art can clearly understand that in practical applications, each step in the data processing method provided by the embodiment of the present application can be allocated to different modules as needed, that is, the internal structure of the device is divided into different modules to complete all or part of the functions described above. Each module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more modules can be integrated in a unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the modules are only for the convenience of distinguishing from each other and do not limit the protection scope of the present application. The specific working process of the modules in the above device can refer to the corresponding process in the foregoing method embodiment and will not be elaborated herein.

[0181] Exemplarily, the data processing device is used to execute the data processing method provided by the embodiment of the present application. Figure 15 is a schematic structural diagram of the data processing device provided by the embodiment of the present application. As Figure 15 shown, the data processing device provided by the embodiment of the present application includes:

[0182] A request acquisition module 1501, configured to acquire a data stream configuration request; wherein, the data stream configuration request includes first configuration content of the user for a first configuration item in a first interface, and the first configuration content includes an identifier of the data processing template and first configuration information. The data processing template is used to indicate a first data stream, the first data stream is a plurality of first processing units connected in sequence, the first data stream is used to indicate a policy for processing target data, and the first configuration information is used to indicate a policy for the second processing unit to acquire the target data from a data provider;

[0183] A request sending module 1502, configured to send the data stream configuration request to a server, so that the server acquires the data processing template according to the identifier of the data processing template; determine a second data stream according to the data processing template and the first configuration information, the second data stream includes the second processing unit and the first data stream connected to the second processing unit; and perform data processing according to the second data stream.

[0184] Based on the same concept as the method embodiment of the present application, an embodiment of the present application further provides a terminal. A schematic structural diagram of the terminal can be seen in Figure 14, the difference is that the terminal may further include components such as a display screen. The terminal may include a processor, a memory, and a network card, and the network card is used to send a data stream configuration request to a computing device such as a server.

[0185] Exemplarily, a computer program may be stored on the memory. When the processor executes the computer program, the steps in the data processing method embodiment described above are implemented. For example Figure 2 the steps 201 and 202 shown, or, Figure 9 the steps 901 and 902 shown. Alternatively, when the processor executes the computer program, the functions of each module in the above device embodiment are implemented. Exemplarily, the computer program may be divided into one or more modules / units. The one or more modules / units may be a series of computer program instruction segments capable of completing specific functions. The one or more modules / units are stored in the memory and executed by the processor to complete the present application. For example, the computer program may be divided into a request acquisition module 1501 and a request sending module 1502. For the specific functions of each module, refer to the above description.

[0186] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0187] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution is prior or posterior. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0188] The basic principles of the present application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, advantages, effects, etc. mentioned in the present application are only examples and not limitations. It cannot be considered that these advantages, advantages, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-disclosed specific details are only for the purposes of illustration and easy understanding, and are not limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0189] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present disclosure are only exemplary examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended words, meaning "including but not limited to", and can be used interchangeably with each other. The word "or" and "and" used herein refer to the word "and / or", and can be used interchangeably with each other, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with each other.

[0190] It should also be noted that in the devices, equipment and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations shall be regarded as equivalent solutions of the present disclosure.

[0191] The above description has been given for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

[0192] It can be understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application.

Claims

1. A data processing method, characterized in that, Applied to a server, the method includes: Obtain a data stream configuration request; wherein, the data stream configuration request includes an identifier of a data processing template and first configuration information, the data processing template is used to indicate a first data stream, the first data stream is a plurality of first processing units connected in sequence, the first data stream is used to indicate a strategy for processing target data, and the first configuration information is used to indicate a strategy for a second processing unit to obtain the target data from a data provider; Obtain the data processing template according to the identifier of the data processing template; Determine a second data stream according to the data processing template and the first configuration information, the second data stream includes the second processing unit and the first data stream connected to the second processing unit; Perform data processing according to the second data stream.

2. The method according to claim 1, wherein The determining the second data stream according to the data processing template and the first configuration information includes: Create the first data stream according to the data processing template; Create the second processing unit according to the first configuration information; Connect the created second processing unit to the first processing unit at the head of the created first data stream to obtain a second data stream.

3. The method according to claim 1 or 2, characterized in that, The data stream configuration request includes second configuration information, and the second configuration information is used to indicate that a third processing unit stores the data after the first data stream is processed; The determining the second data stream according to the data processing template and the first configuration information includes: Determine a second data stream according to the data processing template, the first configuration information, and the second configuration information.

4. The method according to claim 3, characterized in that, The first data stream includes a storage processing unit; The determining the second data stream according to the data processing template, the first configuration information, and the second configuration information includes: Create the third processing unit according to the second configuration information; Connect the created third processing unit to the storage processing unit.

5. The method according to any one of claims 1 to 4, characterized in that, The strategy for processing the target data is used to convert multiple fields in the target data into multiple preset fields.

6. The method according to any one of claims 1 to 5, characterized in that The performing data processing according to the second data stream includes: Receive the target data through the second processing unit; Receive the target data sent by the second processing unit through the first data stream, parse the target data to obtain first field values of multiple fields, and map the first field values to multiple preset fields based on the mapping relationship between the multiple fields and the multiple preset fields.

7. The method according to claim 6, characterized in that, The mapping the first field values to multiple preset fields based on the mapping relationship between the multiple fields and the multiple preset fields includes: Perform arrangement processing on the first field values to obtain second field values of the multiple fields; Map the second field values to multiple preset fields based on the mapping relationship between the multiple fields and the multiple preset fields.

8. A data processing method, characterized in that, Applied to a terminal, the method includes: Obtain a data stream configuration request; wherein, the data stream configuration request includes first configuration content of the user for a first configuration item in a first interface, and the first configuration content includes an identifier of the data processing template and first configuration information. The data processing template is used to indicate a first data stream, the first data stream is a plurality of first processing units connected in sequence, the first data stream is used to indicate a strategy for processing target data, and the first configuration information is used to indicate a strategy for the second processing unit to obtain the target data from a data provider; Send the data stream configuration request to a server, so that the server obtains the data processing template according to the identifier of the data processing template; determine a second data stream according to the data processing template and the first configuration information, the second data stream includes the second processing unit and the first data stream connected to the second processing unit; perform data processing according to the second data stream.

9. The method according to claim 8, characterized in that The method further includes: Obtain second configuration content of the user for a second configuration item in a second interface; wherein, the second configuration content includes the target data and a strategy for processing the target data; Generate the data processing template based on the second configuration content of the second configuration item.

10. A server, characterized in that, Comprising a processor and a memory; wherein, The memory is used to store a program; The processor is used to execute the program stored in the memory, and when the program stored in the memory is executed, the method according to any one of claims 1 to 7 is executed.