Data processing system, method, device, medium and program product
By creating multiple controller objects in the container orchestration system to distribute data processing tasks, the problem of observable nodes being prone to downtime when processing high volumes of log data is solved, and more efficient data processing and system stability are achieved.
Patent Information
- Application Number
- CN202411890736.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-12-19
AI Technical Summary
When observable nodes process up to 6-10 million log data, a stand-alone system is prone to memory overflow (OOM), causing downtime.
By creating a target number of controller objects in the container orchestration system, and assigning configuration files of the to-process data to these controller objects, distributed data processing is implemented and data processing pressure for a single node is reduced.
It improves the throughput of system data processing, achieves the purpose of processing massive data, and reduces the probability of observable nodes crashing, and improves system stability.
Smart Images

Figure CN119336587B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a data processing system, method, device, medium and program product. Background Art
[0002] Observability services can help users comprehensively monitor and diagnose their applications and infrastructure, thereby improving system reliability and performance. Observability services can build a comprehensive and three-dimensional monitoring and analysis system by collecting and analyzing various observable data of the system, which can help the operation and maintenance team understand the health status, performance and fault causes of the system in real time in a complex and dynamic Internet environment, and make accurate decisions based on this information to achieve rapid problem location, preventive maintenance and continuous optimization.
[0003] Observable nodes provide observable services. In actual use, observable nodes need to process 6 to 10 million logs per minute at a customer site, which causes the single-machine system corresponding to a single observable node to frequently experience out of memory (OOM) and crash. Summary of the invention
[0004] Multiple aspects of the present application provide a data processing system, method, device, medium and program product to reduce the probability of downtime of observable nodes.
[0005] An embodiment of the present application provides a data processing system, including: a first observable node and a first controller running in a container orchestration system; the first observable node is used to provide an observable service;
[0006] The first observable node is used to obtain a configuration file of data to be processed; send a first creation request to the first controller, the first creation request is used to request the first controller to create a target number of first controller objects;
[0007] The first controller is used to create the target number of first controller objects in response to the first creation request;
[0008] The first observable node is used to assign the configuration file of the data to be processed to the target number of first controller objects;
[0009] The target number of first controller objects are used to process the data to be processed according to the configuration file of the data to be processed.
[0010] The embodiment of the present application also provides a data processing method, which is applicable to an observable node running in a container orchestration system, wherein the observable node provides an observable service; the method comprises:
[0011] Get the configuration file of the data to be processed;
[0012] Sending a first creation request to a first controller in the container orchestration system; the first creation request is used to request the first controller to create a target number of first controller objects;
[0013] The configuration file of the data to be processed is distributed to the target number of first controller objects, so that the target number of first controller objects perform data processing on the data to be processed according to the configuration file of the data to be processed.
[0014] The embodiment of the present application further provides a data processing method, which is applicable to a first controller in a container orchestration system, wherein the container orchestration system further runs an observable node, and the observable node is used to provide an observable service, and the method includes:
[0015] Obtaining a first creation request for a first controller object sent by the observable node; the first creation request is used to request the first controller to create a target number of first controller objects;
[0016] In response to the first creation request, the target number of first controller objects are created so that the observable node can assign the configuration files of the data to be processed to the target number of first controller objects, so that the target number of first controller objects can process the data to be processed according to the configuration files of the data to be processed.
[0017] The embodiment of the present application further provides an electronic device, comprising: a memory and a processor; wherein the memory is used to store a computer program;
[0018] The processor is coupled to the memory and is configured to execute the computer program to perform the steps in the aforementioned data processing methods.
[0019] An embodiment of the present application also provides a computer-readable storage medium storing computer instructions. When the computer instructions are executed by one or more processors, the one or more processors are caused to execute the steps in the aforementioned data processing methods.
[0020] An embodiment of the present application also provides a computer program product, including a computer program. When the computer program is executed by one or more processors, the one or more processors are caused to execute the steps in the aforementioned data processing methods.
[0021] In an embodiment of the present application, a single observable node is used, and the capabilities of the container orchestration system itself are utilized to create a target number of controller objects; thereafter, the configuration files of the data to be processed are assigned to the target number of controller objects, and the target number of controller objects process the data to be processed according to the configuration files of the data to be processed, thereby realizing distributed data processing with the aid of the capabilities of the container orchestration system. Compared with the data processing method of a single observable node, the throughput of system data processing can be improved, which helps to achieve the purpose of processing massive amounts of data. On the other hand, the data processing load of the observable node is unloaded to the target number of controller objects, which reduces the data processing pressure of the observable node, helps to reduce the probability of downtime of the observable node, and thus helps to improve the stability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0023] Figure 1 It is a schematic diagram of the structure of a traditional distributed processing system;
[0024] Figure 2 and Figure 3 A schematic diagram of the structure of a data processing system provided in an embodiment of the present application;
[0025] Figure 4 and Figure 5 A flowchart of a data processing method provided in an embodiment of the present application;
[0026] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0027] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.
[0028] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0029] The concepts or terms involved in the embodiments of the present application are first explained below.
[0030] Container orchestration system: A container orchestration system is a container orchestration platform used to automatically deploy, scale, and manage containerized applications. In the embodiments of the present application, the container orchestration system can be a self-developed container orchestration system or an open source container orchestration system, such as Kubernetes (K8s).
[0031] Controller: A controller is a component in the control plane of a container orchestration system that monitors the state of the cluster and makes necessary adjustments to ensure that the actual state of the cluster is consistent with the desired state defined by the user. The controller ensures that the actual state of the cluster is consistent with the desired state defined by the user by continuously monitoring the state of the cluster and creating, updating, or deleting resource objects as needed.
[0032] Controller Object: A controller object is a resource object in a container orchestration system that is used to define and manage a group of related resource objects. A controller object usually contains detailed information about the desired state, and the controller is responsible for achieving these desired states. A controller object can be a job, a daemon set, a replica set, a deployment, or a stateful set.
[0033] Job: A Job is a resource object in a container orchestration system. It is a controller object or controller resource that manages tasks. It can create one or more container groups (such as Pods) and ensure that a specified number of container groups are successfully completed. It is suitable for one-time tasks and batch task processing. Once these container groups successfully complete the task, the Job is also completed. A Job can be created and completed by a Job Controller.
[0034] DaemonSet: DaemonSet is a resource object in the container orchestration system. It is a controller object that ensures that a copy of a container group runs on each node (or certain specific stages) in the cluster. When a new node joins the cluster, DaemonSet automatically creates a container group on the new node; when a node is removed from the cluster, the container group is also deleted accordingly.
[0035] ReplicaSet: ReplicaSet is another controller resource in the container orchestration system that ensures the normal operation of a specified number of container groups.
[0036] Deployment: Deployment is a high-level resource object used to define and manage container group replica sets in a container orchestration system. It is a controller resource. Deployment provides a declarative way to create, update, and delete container group replicas, ensure that a specified number of container group replicas run in the cluster, and automatically handle the expansion and reduction of container groups. It is mainly used to manage the replica sets, deploy and manage stateless applications.
[0037] StatefulSet: StatefulSet is another controller resource in the container orchestration system, which is used to manage container groups of stateful applications. It ensures that the container group has stable identity, network identity, and storage during deployment and expansion.
[0038] Stateless Applications: Stateless applications are applications that do not save any session or client state. Each request is independent of the previous and subsequent requests and does not depend on any context information.
[0039] Stateful applications: Stateful applications are applications that need to save and manage session or client state. These applications rely on previous state information to process current requests.
[0040] Observability: A cloud service platform that helps users comprehensively monitor and diagnose their applications and infrastructure. It can build a comprehensive and three-dimensional monitoring and analysis system by collecting and analyzing various observable data of the system. It can help the operation and maintenance team understand the health status, performance and causes of failures within the system in real time in a complex and dynamic Internet environment, and make accurate decisions based on this information to achieve rapid problem location, preventive maintenance and continuous optimization.
[0041] OneAgent: A lightweight service for collecting host hardware metrics and log data.
[0042] Log data source: A service or platform that collects or stores log data and can provide log data to the observable service. For example, the log data source can be: the Application Programming Interface (API) provided by the observable service, the Simple Log Service (SLS) or a distributed log system (such as Kafka, etc.). Users can provide log data to the observable service through the API provided by the observable service and wait for the data to be processed.
[0043] Kafka: A fast, highly scalable, high-throughput distributed logging system that provides distributed logging services.
[0044] The following is an explanation of the traditional solution for distributed data processing of log data and the like.
[0045] In the actual use of observable services, the observable node at a certain customer site needs to process 5 to 7 million logs per minute, which causes the single-machine system to frequently experience OOM and downtime.
[0046] In some traditional solutions, multiple observable nodes are deployed for log processing. Figure 1 As shown in the figure, in order to maintain the consistency of multiple observable nodes, multiple observable nodes can select a master node through the master election operation, and the master node will allocate log processing tasks. Figure 1 In the Observable Node, the observable node is mainly used for log data processing to obtain the required indicator data. Figure 1 The method for obtaining other indicators to be monitored is also shown in FIG. Figure 1 As shown, the data collection node can collect the status data and log data of the monitored object. Among them, the status data of the monitored object may include: host hardware indicators, such as CPU resource information, memory resource information, etc. Further, as Figure 1 As shown in step 1 "status reporting" and step 2 "status sending", the data collection node reports the status data of the monitored object to the message server, and the message server sends the status data to the observable node. Figure 1 As shown in step 3 "Pull indicators", the data collection node can also push the indicator data in these status data to the time series database.
[0047] like Figure 1 As shown in steps 3 and 4, the time series database can obtain indicator data from these status data from the observable nodes. Figure 1 As shown in step 3 "Pull indicators", the time series database can actively pull indicator data from observable nodes. Figure 1As shown in step 4 "Push Indicators", the observable node can also actively push the indicator data to the time series database. The time series database can collect and store time series data, and can perform performance monitoring and alarms on the monitored objects based on the time series data. The acquisition path of the indicator data is irrelevant to the data processing solution provided in the embodiment of the present application, so it will not be described in detail. Of course, if Figure 1 As shown in step 5 “Log storage”, the observable node can also write log data to the search engine to facilitate subsequent user searches.
[0048] exist Figure 1 In the scheme shown, the observable node mainly obtains the indicators that cannot be directly obtained in the above steps 1-4 by processing the log data. Specifically, by designing the database table of the master node, multiple observable nodes can compete for the master node. Each observable node can use the optimistic lock preemption method to elect the master. If multiple observable nodes preempt the master node at the same time, the node with the successful optimistic lock update will be the master node. The master election program periodically detects the online status of the master node. If the master node is found to be offline, the master election logic will be retriggered there.
[0049] In this traditional distributed data processing solution, the master node schedules the log data task according to the resource information of multiple observable nodes, and schedules the log processing task to multiple observable nodes. At the same time, the master node needs to regularly detect the survival status of all observable nodes. Specifically, a heartbeat data table can be designed, which is used to record the heartbeat reporting status of each observable node. Multiple observable nodes write heartbeat information to their corresponding heartbeat data tables according to the set heartbeat cycle (such as 5 seconds). The master node determines whether the observable node is alive by reading the heartbeat information from the heartbeat data table. If the heartbeat information of the observable node is not updated for more than a set time (such as 15 seconds, etc.), it is determined that the observable node has been offline. If the number of observable nodes changes, the task is rescheduled.
[0050] After the heartbeat report is interrupted, the node offline situation needs to be considered, and the master node and slave node situations can be distinguished. For the master node: if the master node fails to write the heartbeat information three times in a row, the master election logic will be re-triggered to select a new master node to continue the task scheduling operation. For the slave node, if the heartbeat information fails to be written three times in a row, it is considered that the observable node network is abnormal, and the master node will reschedule the log processing task on the observable node. Each observable node can also periodically check the task assigned to it. If it is different from the currently running task, the task hot loading logic is triggered.
[0051] The above distributed data processing solution involves master-slave election, and task scheduling depends on the master node. If the master node fails, the entire distributed processing system will re-elect the master and re-schedule tasks, which will make the abnormal recovery time of the distributed processing system longer and the system stability lower.
[0052] In some embodiments of the present application, a new distributed data processing solution is proposed, the main principle of which is: using a single observable node, using the capabilities of the container orchestration system itself, to create a target number of controller objects; then, assigning the configuration file of the data to be processed to the target number of controller objects, and the target number of controller objects process the data to be processed according to the configuration file of the data to be processed, thereby realizing distributed data processing with the help of the capabilities of the container orchestration system. Compared with the data processing method of a single observable node, the throughput of system data processing can be improved, which helps to achieve the purpose of processing massive amounts of data. On the other hand, the data processing load of the observable node is unloaded to the target number of controller objects, which reduces the data processing pressure of the observable node, helps to reduce the probability of downtime of the observable node, and thus helps to improve the stability of the system.
[0053] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.
[0054] It should be noted that the same reference numerals represent the same object or the same step in the following drawings and embodiments, and therefore, once an object or a step is defined in one drawing or embodiment, there is no need to further discuss it in the subsequent drawings and embodiments.
[0055] Figure 2 and Figure 3 This is a schematic diagram of the structure of the data processing system provided in the embodiment of the present application. The data processing system runs on the container orchestration system. Figure 2 and Figure 3As shown, the data processing system mainly includes: an observable node 10 and a controller 20. The controller 20 refers to a control component used to manage controller resources in a container orchestration system. For example, assuming that the controller resource object is a task (Job), the controller 20 is a task controller (JobController). For another example, if the controller resource object is a daemon set (DaemonSet), the controller 20 is a daemon set controller (DaemonSetController). For another example, if the controller resource object is a replica set (ReplicaSet), the controller 20 is a replica set controller (ReplicaSetController). For another example, if the controller resource object is a deployment (Deployment), the controller 20 is a deployment controller (DaemonSetController). For another example, if the controller resource object is a deployment (Deployment), the controller 20 is a deployment controller (DeploymentController), and so on.
[0056] In an embodiment of the present application, the observable node 10 can provide observable services. Among them, observable services refer to a series of services such as data monitoring, problem analysis, system diagnosis, etc. derived from system metrics, logs and link tracking. Metrics are used to record quantitative information of various dimensions over a period of time to observe certain states and trends of the system. Logs are used to record some discrete events generated during the running of the program. Link tracking is used to record the call link of a request from the receipt to the completion of the entire life cycle.
[0057] In this embodiment, the observable node 10 is a single node. The observable node 10 can be implemented as a controller resource (i.e., a controller object) of a container orchestration system, such as a deployment, a daemon set, a replica set, or a job. The observable node 10 is implemented as a controller object of a container orchestration system, and the corresponding controller performs health status monitoring and lifecycle management.
[0058] In the embodiment of the present application, in order to realize distributed data processing, the distributed processing capability of the container orchestration system and the resource isolation capability between the controller objects can be used to distribute the data processing tasks received by the observable node 10 to the controller objects in the container orchestration system to realize distributed data processing. Specifically, Figure 2As shown, the observable node 10 can obtain a configuration file of the data to be processed in response to a data processing task. The data to be processed may be log data or data of other contents. The configuration file of the data to be processed is used to indicate the way to obtain the data to be processed, and may include: where to obtain the data to be processed and how to obtain the data to be processed, etc. Correspondingly, the configuration file of the data to be processed may include: the access address of the data source instance to which the data to be processed belongs, and / or, the authentication information for accessing the data source instance, etc. The access address of the data source instance may be the network link address of the data source instance, or the Internet Protocol (IP) address of the data source instance, etc. The authentication information for accessing the data source instance refers to the identity credentials for verifying and authorizing access to the data source. These credentials ensure that only authorized users or applications can access the data in the data source. The authentication information may include a user name, password, and certificate, etc.
[0059] Since the configuration file of the data to be processed can reflect which data comes from the same data source instance, the data storage format of the same data source instance is generally the same. In order to reduce the data processing complexity of the controller object and reduce the input / output (IO) access frequency of the data source, the data of the same data source instance is generally handed over to the same controller object for processing. Based on this, Figure 2 and Figure 3 As shown, the observable node 10 can determine the target number of controller objects to be created based on the configuration file of the data to be processed, which is recorded as M. M ≥ 1 and is an integer. In the embodiment of the present application, for the convenience of description and distinction, the controller object used to process the data processing task of the observable node 10 is defined as the first controller object; and the controller object implemented by the aforementioned observable node is defined as the second controller object. Correspondingly, the controller 20 corresponding to the first controller object is defined as the first controller 20; and the controller corresponding to the second controller object is defined as the second controller. Among them, the first controller object and the second controller object can be controller objects of the same type or different types. For example, the first controller object can be a task (Job), the second controller object can be a deployment (Deployment), etc.
[0060] In the embodiment of the present application, the specific implementation method of the observable node 10 determining the target number M of the first controller objects to be created is not limited. In some embodiments, the observable node 10 is pre-configured with the target number M of the first controller objects to be created, and accordingly, the observable node 10 can obtain the pre-configured target number M corresponding to the first controller to be created.
[0061] In other embodiments, the data source instance to which the data to be processed belongs is determined according to the configuration file of the data to be processed. Figure 3 As shown in "File Splitting" and "Task Grouping", the data source instance to which the data to be processed belongs can be determined based on the access address of the data source instance to which the data to be processed belongs. Furthermore, the configuration files of the data to be processed can be allocated to at least one task group based on the data source instance to which the data to be processed belongs. Specifically, it can be determined that the data to be processed with the same access address belongs to the same data source instance; and the configuration sub-files of the data to be processed belonging to the same data source instance are allocated to the same task group to obtain at least one task group.
[0062] In some embodiments, the configuration subfile of the to-be-processed data belonging to the same data source instance can be directly assigned to the same task group to obtain at least one task group. In other embodiments, such as in the log processing scenario, different data sources provide log data of different magnitudes, some data source instances provide a small amount of log data, and some data source instances provide a large amount of log data. For example, the observable node 10 can provide an API, and the user can call the API to provide the observable node 10 with the log data to be processed, and the amount of log data provided by the API is small. The amount of log data provided by the log service instance is relatively large. In order to efficiently utilize resources and reduce the idleness and waste of resources, the configuration subfile of the first to-be-processed data belonging to the API in the to-be-processed data can be assigned to the same task group, and the first to-be-processed data of the task group is subsequently assigned to the same first controller object for processing, that is, all the to-be-processed data provided by the API are assigned to a single controller object for processing, which can reduce the idleness and waste of resources of the first controller object. Since the amount of log data provided by the log service instance is relatively large, the configuration subfile of the second to-be-processed data belonging to the same log service instance in the to-be-processed data can be assigned to the same task group. Subsequently, the to-be-processed data belonging to the same log service instance can be assigned to an independent first controller object for processing, rather than assigning the to-be-processed data of all log service instances to the same first controller object for processing. This can prevent the first controller object from crashing due to overload and ensure the stable operation of the log processing task.
[0063] After the configuration file of the data to be processed is allocated to at least one task group, the number of task groups can be used as the target number M of first controller objects to be created. Accordingly, the number of task groups is also the target number, that is, M task groups.
[0064] Further, if Figure 2 and Figure 3As shown in step 1, the observable node 10 may send a creation request to the first controller 20. The creation request is used to request the first controller 20 to create a target number of first controller objects. The creation request may carry the target number of first controller objects to be created. The first controller 20 may respond to the creation request and create a target number of first controller objects 30 in the container orchestration system, that is, create M first controller objects 30.
[0065] In some embodiments, the observable node 10 may specify a resource amount of the first controller object. Figure 1 The distributed processing system shown can only assign data processing tasks to specified observable nodes, but cannot control the amount of resources that can be used by data processing tasks. The amount of resources may include: computing resources, memory resources, and storage resources. Computing resources can be represented by the number of cores of the processor. Storage resources specifically refer to the number and storage capacity of persistent storage media such as disks.
[0066] In some embodiments of the present application, the configuration file of the data to be processed is assigned to a target number of task groups, and the configuration sub-file assigned to each task group is assigned to a first controller object. In this embodiment, the observable node 10 can determine the resource amount corresponding to each of the target number of task groups (i.e., M task groups). Specifically, the amount of resources required for the historical data processing tasks of the data source instance can be counted; and the amount of resources required for the data processing tasks of the data source instance can be determined based on the amount of resources required for the historical data processing tasks of the data source instance. For example, the average amount of resources required for the historical data processing tasks of the data source instance can be calculated as the amount of resources required for the data processing tasks of the data source instance. Alternatively, the maximum amount of resources required for the historical data processing tasks of the data source instance can be taken as the amount of resources required for the data processing tasks of the data source instance, etc. Further, for any task group, the amount of resources required for the data processing tasks of the data source instance to which the task group belongs can be determined from the amount of resources required for the data processing tasks of each data source instance as the amount of resources corresponding to the task group.
[0067] Alternatively, for any task group, the amount of logs contained in the task group can be counted; based on the amount of logs contained in the task group, the amount of resources required by the task group is determined as the amount of resources corresponding to the task group. Among them, the task group with a larger amount of logs has a larger amount of resources corresponding to it. For example, the amount of logs can be represented by the number of log entries. In some embodiments, the amount of resources required for a single log entry can be pre-set, and then the amount of resources required for the task group can be determined based on the amount of logs contained in the task group and the amount of resources required for a single log entry, as the amount of resources corresponding to the task group.
[0068] Further, the observable node 10 may use the amount of resources corresponding to each of the target number of task groups as the amount of resources corresponding to the target number of first controller objects. Further, the observable node 10 may generate a creation request based on the amount of resources corresponding to each of the target number of task groups and the target number. The creation request includes the target number of first controller objects to be created, and the amount of resources corresponding to each of the target number of first controller objects. Further, the observable node 10 may send the creation request to the first controller 20 to request the first controller 20 to create the target number of first controller objects, thereby enabling the observable node to specify the amount of resources allocated to the data processing task of the first controller object, and realizing refined control of the resources required for the data processing task.
[0069] Accordingly, if Figure 2 and Figure 3 As shown in step 2, the first controller 20 may, in response to the creation request, create a target number of first controller objects 30 with corresponding resource quantities according to the resource quantities corresponding to each of the target number of task groups. Specifically, for any task group A, the first controller 20 may create a first controller object with the resource quantities corresponding to the task group A according to the resource quantities corresponding to the task group A. Specifically, the first controller 20 may determine, from the working cluster in the container orchestration system, a target working node with an idle resource quantity greater than or equal to the resource quantity corresponding to the task group A according to the resource quantity corresponding to the task group A; thereafter, create a first controller object 30 with the resource quantity corresponding to the task group A on the target working node.
[0070] Specifically, the observable node 10 may send a creation request for the first controller object to the API server in the container orchestration system, and the creation request includes configuration information of the first controller object to be created, and the configuration information may include the resource quantity of task group A and the resource definition of the first controller object. Among them, the resource definition of the first controller object may include: metadata and specifications (Spec) of the first controller object, etc. Among them, the metadata may include: the name and namespace (Namespace) of the first controller object, etc. The specifications may include: the parallelism (Parallelism) and completions (Completions) and template (Template) of the first controller object, etc. Among them, parallelism refers to the number of container groups (such as Pods) running simultaneously. The number of completions refers to the number of container groups required for the first controller object to complete. The template is used to define the specifications of the container group (such as Pod) created by the first controller object.
[0071] The API server receives a creation request for the first controller object. The first controller 20 obtains the resource definition of the first controller object from the API server, and creates a container group with the resource quantity corresponding to the task group A according to the resource definition of the first controller object, so as to obtain the first controller object corresponding to the task group A. Using the same method, the target number of first controller objects can be obtained.
[0072] Further, if Figure 2 and Figure 3 As shown in step 3, the observable node 10 can assign the configuration file of the data to be processed to the target number of first controller objects 30. Correspondingly, the target number of first controller objects 30 can process the data to be processed according to the configuration file of the data to be processed. In this way, the observable node can distribute the data processing tasks to the target number of first controller objects for processing, and realize distributed data processing with the help of the capabilities of the container orchestration system. Compared with the data processing method of a single observable node, the throughput of system data processing can be improved, which helps to achieve the purpose of processing massive data. On the other hand, the data processing load of the observable node is unloaded to the target number of first controller objects, which reduces the data processing pressure of the observable node, helps to reduce the probability of downtime of the observable node, and thus helps to improve the stability of the system.
[0073] On the other hand, this embodiment uses a single observable node and uses the ability of the container orchestration cluster's own controller object to achieve distributed data processing, avoiding the above Figure 1 The development and integration of distributed consensus algorithms in the case of multiple observable nodes in the scheme shown can reduce the workload of system development.
[0074] In addition, if Figure 3 As shown in "Dynamic Expansion", every time a new data processing task arrives, the observable node and the controller cooperate with each other and can create a first controller object to process the corresponding data processing task in the above manner, so as to realize the dynamic expansion of the first controller object according to the data processing task. Therefore, by using the capabilities of the controller object of the container orchestration system to run data processing tasks, the data processing volume can be continuously increased within the scope allowed by the hardware capabilities of the working cluster of the container orchestration system, realizing massive data processing.
[0075] In some embodiments of the present application, when the observable node 10 assigns the configuration file of the data to be processed to the target number of first controller objects 30, it can determine the configuration sub-files of the data to be processed to which the target number of task groups are respectively assigned from the configuration file of the data to be processed. Furthermore, the configuration sub-files corresponding to the target number of task groups can be respectively assigned to the target number of first controller objects; each first controller object is assigned to a configuration sub-file. In this way, the observable node 10 can distribute the data processing tasks to the target number of first controller objects for processing, which reduces the data processing pressure of the observable node, helps to reduce the probability of downtime of the observable node, and thus helps to improve the stability of the system. Each first controller object undertakes the data processing tasks of a task group, which helps to achieve load balancing between the target number of first controller objects.
[0076] Specifically, since the amount of resources corresponding to each task group can be the same or different, and the first controller object also has the amount of resources, for any task group A, the configuration subfile corresponding to task group A can be allocated to the first controller object (defined as the target controller object) having the same amount of resources as the amount of resources corresponding to task group A. That is, the target controller object refers to: among the target number of first controller objects, the first controller object to which the amount of resources allocated is the amount of resources corresponding to task group A. In the same way, the configuration subfiles of the target number of task groups A can be allocated to the target number of first controller objects. This task allocation method can make the amount of resources required by task group A match the amount of resources provided by the first controller object. On the one hand, it can prevent the amount of resources of the first controller object from being greater than the amount of resources required by task group A, resulting in a waste of resources; on the other hand, it can also prevent the amount of resources of the first controller object from being less than the amount of resources required by task group A, causing the first controller object to be overloaded and downtime.
[0077] Specifically, the configuration sub-file corresponding to task group A can be assigned to the target controller object with the help of the configuration table (ConfigMap) in the container orchestration system. Among them, the configuration table (ConfigMap) is an API object for storing configuration data in the container orchestration system. ConfigMap can be used to save non-sensitive configuration information, such as configuration files, command line parameters, environment variables, etc. The data of ConfigMap can be used by the container group in the controller object in the form of files or environment variables. Based on this, the observable node 10 can write the configuration sub-file corresponding to task group A into the target configuration table corresponding to the target controller object in the container orchestration system. Correspondingly, the target controller object can obtain the configuration sub-file corresponding to task group A from the target configuration table, thereby realizing the configuration sub-file corresponding to task group A and assigning it to the target controller object. In the same way, the configuration sub-files of the target number of task groups A can be assigned to the target number of first controller objects. This task allocation method reuses the data transmission path of the container orchestration system itself, which can reduce the development difficulty and development cost of the distributed data processing system.
[0078] Further, for the first controller object, the data processing task can be executed according to the assigned target configuration subfile. Specifically, the first controller object 30 can obtain the access address of the target data source instance to which the target data to be processed by the first controller object 30 belongs and the authentication information for accessing the target data source instance from the assigned target configuration subfile. The data to be processed includes the target data. Further, the first controller object 30 can obtain the target data from the target data source instance according to the access address and authentication information of the target data source instance.
[0079] In some embodiments, all data stored in the target data source instance are target data. Accordingly, the first controller object 30 can access the target data source instance according to the access address of the target data source instance; and obtain the permission to read data from the target data source instance according to the authentication information; and then, all data stored in the target data source instance can be read from the target data source instance as target data.
[0080] In other embodiments, the data processing task processes part of the data in the target data source instance. Accordingly, the configuration file of the data to be processed may also include: the identification of the data to be processed (such as an index), etc. Accordingly, the target configuration subfile may also include: the identification of the target data (such as an index), etc. Therefore, accordingly, the first controller object 30 may access the target data source instance according to the access address of the target data source instance; and obtain the permission to read data from the target data source instance according to the authentication information; and then, the target data may be read from the target data source instance according to the identification of the target data.
[0081] In an embodiment of the present application, the configuration file of the data to be processed may also include: a data processing method for the data to be processed. The data processing method may be determined by the observable node, or may be sent to the observable node by other nodes. In some embodiments, the observable node may determine the data processing method based on the indicator to be obtained by the data processing task. Among them, the data processing method is different for different indicators to be obtained for the data processing task. In some embodiments, for example, if the indicator to be obtained for the data processing task is the number of requests processed per second (Queries-per-second, QPS) by the monitored object (such as a server, etc.), then the data processing method is: counting the number of requests processed per unit time (such as per second). For another example, if the indicator to be obtained for the data processing task is the average response time of the monitored object, then the data processing method is: taking the average of the response time of the monitored object. For another example, if the indicator to be obtained for the data processing task is the peak value of the processor utilization rate of the monitored object, then the data processing method is: taking the maximum value of the processor utilization rate of the monitored object, and so on.
[0082] After determining the data processing method of the data to be processed, the observable node 10 can write the data processing method of the data to be processed into the configuration file of the data to be processed. Accordingly, the configuration subfile assigned to the first controller object also includes: the data processing method of the target data to be processed by the first controller object.
[0083] In other embodiments, the data processing method of the data to be processed is sent by other nodes to the observable node 10. Other nodes carry the data processing method of the data to be processed in the configuration file of the data to be processed and send it to the observable node 10. For example, some open source performance indicator collection, processing and export tools (such as OpenTelemetry) can send the configuration file of the data to be processed to the observable node 10. The configuration file includes: the data processing method of the data to be processed. Among them, OpenTelemetry is an open source project for collecting, processing and exporting application performance indicators and distributed tracing information, aiming to provide a unified tool chain to facilitate application developers and operation and maintenance personnel to easily obtain system monitoring data. OpenTelemetry can send the configuration file of the data to be processed to the observable node 10 through its internal pipeline (Pipeline). Among them, Pipeline is a modular process in OpenTelemetry for collecting, processing and exporting monitoring data.
[0084] When the observable node 10 allocates the configuration file of the data to be processed to multiple task groups, the configuration sub-file allocated to each task group also includes: the data processing method of the data corresponding to the configuration sub-file. Based on this, the first controller object 30 can also process the target data according to the data processing method carried by the target configuration sub-file allocated to it, so as to obtain the data processing result of the target data.
[0085] If the indicator to be obtained by the data processing task is the number of requests processed per second (QPS) by the monitored object (such as a server, etc.), the data processing method is: count the number of requests processed within a unit time (such as per second), then the first controller object 30 can obtain the target data belonging to the same monitored object according to the identifier of the monitored object corresponding to the target data; then, the number of requests processed by the monitored object within a unit time can be determined from the target data belonging to the same monitored object. For another example, if the indicator to be obtained by the data processing task is the peak value of the processor usage of the monitored object, then the data processing method is: take the maximum value of the processor usage of the monitored object. Accordingly, the first controller object 30 can obtain the target data belonging to the same monitored object according to the identifier of the monitored object corresponding to the target data; then, the processor usage of the monitored object can be determined from the target data belonging to the same monitored object, and the maximum processor usage can be determined from the processor usage of the monitored object.
[0086] In some embodiments of the present application, Figure 3 As shown, after the first controller object 30 processes the target data and obtains the target indicator required for the data processing task, the target indicator can also be written into the time series database. In some embodiments, the first controller object 30 can write the target indicator into the time series database by means of a remote procedure call. The time series database can perform performance monitoring and alarm on the monitored object based on the time series data.
[0087] like Figure 3 As shown, the first controller object 30 can also store the acquired target data (such as log data) in the search engine. Since the search engine can provide powerful search capabilities, the target data is stored in the search engine for subsequent retrieval by the user. For example, if the user wants to review the processing results of the target data, the target data can be retrieved from the search engine, and further, the processing results can be reviewed based on the target data.
[0088] The data processing methods shown in the above embodiments are only exemplary and not limiting. Of course, the first controller object 30 can also use other data processing methods to process the target data, which is determined by the requirements of the data processing task.
[0089] Compared to Figure 1 In the distributed processing system shown, since the observable node in the embodiment of the present application does not undertake data processing tasks and task scheduling, if the distributed processing system provided by the present embodiment has an abnormality in the observable node, there is no need to re-elect the master, which can improve the speed of abnormal recovery of the observable node. The abnormal recovery process of the observable node is exemplarily described below.
[0090] Specifically, the observable node 10 can be implemented as a second controller object in the container orchestration system. Accordingly, the container orchestration system also includes: a controller of the second controller object (defined as a second controller). For the specific implementation form of the second controller object and the controller of the second controller object, please refer to the relevant content of the aforementioned embodiment, which will not be repeated here.
[0091] The second controller may monitor the health status of the observable node 10. In some embodiments, the second controller may use a live probe to perform a health check on the observable node 10, that is, monitor the health status of the observable node 10. The second controller may periodically send a heartbeat message to the observable node 10 according to a set check cycle; if a response message to the heartbeat message from the observable node 10 is received within a set time period, it is determined that the observable node 10 is operating normally; if a response message to the heartbeat message from the observable node 10 is not received within a set time period, it is determined that the health status of the observable node 10 is abnormal.
[0092] Further, when an abnormal health status of the observable node 10 is detected, the second controller may create a new observable node in the container orchestration system according to the resource definition of the observable node 10. In the embodiment of the present application, for the convenience of description and distinction, the original observable node is defined as the first observable node, and the newly created observable node is defined as the second observable node. Regarding the specific implementation method of the second controller creating a new observable node in the container orchestration system according to the resource definition of the observable node 10, please refer to the creation process of the first controller object mentioned above, which will not be repeated here.
[0093] Furthermore, the load of the first observable node can be migrated to the newly created second observable node to achieve abnormal recovery of the observable node. The load of the first observable node may include the currently received data processing tasks and the configuration files of the data to be processed.
[0094] In this embodiment, since the observable node does not undertake data processing tasks, it is only responsible for task grouping and the creation trigger of controller objects, and there is no need to perform master selection operations and task scheduling. Therefore, the abnormal recovery of the observable node is faster, which can improve the abnormal recovery speed of the observable node and improve the system stability.
[0095] In the embodiment of the present application, in addition to the abnormality of the observable node, the first controller object may also be abnormal. In order to be able to timely discover the abnormal situation of the first controller object, and timely perform abnormal recovery on the first controller object to ensure the normal execution of the data processing task, in some embodiments of the present application, assuming that the observable node in the current data processing system is still the first observable node, the first controller object can periodically send heartbeat data to the first observable node 10 according to the set heartbeat cycle. The first observable node can update the heartbeat time of the first controller object stored in the database according to the received heartbeat data of the first controller object 30; and check whether the heartbeat time of the first controller object is expired. For example, if the state of mind time of the first controller object exceeds the set time and is not updated, it means that the heartbeat time of the first controller object has expired. Further, if the heartbeat time of the first controller object expires, the first controller 20 is requested to recreate the first controller object. Specifically, the first observable node can send another creation request to the first controller 20, which is used to request the re-creation of the aforementioned abnormal first controller object. For ease of description and distinction, the aforementioned creation request for requesting to create a target number of first controller objects is defined as a first creation request; and the creation request for requesting to recreate the aforementioned abnormal first controller object is defined as a second creation request.
[0096] Accordingly, the first controller 20 can create a new first controller object in response to the second creation request, and migrate the data processing tasks of the original first controller object to the new first controller object, so as to realize abnormal recovery of the first controller object, ensure the continuity of the data processing tasks, and help meet the high availability requirements of the system.
[0097] In some embodiments of the present application, in order to realize the migration of the data processing task of the abnormal first controller object to the new first controller object, for the original first controller object, after being assigned to the target configuration subfile, a resource entry of the first controller object can also be added to the database. The resource entry includes: the identifier of the configuration table corresponding to the first controller object in the container orchestration system and the resource information of the first controller object.
[0098] Accordingly, the first controller 20 reads the resource information of the first controller object from the database in response to the second creation request; and creates a new first controller object according to the resource information of the first controller object. The resource amount of the new first controller object is the same as the resource amount of the original first controller object. Further, the first controller 20 may provide the identifier of the configuration table to the new first controller object.
[0099] The new first controller object can obtain the target configuration sub-file assigned to the original first controller object from the configuration table according to the identifier of the configuration table; and execute the data processing task of the original first controller object according to the target configuration sub-file, thereby realizing the migration of data processing tasks and realizing abnormal recovery of the first controller object, which can ensure the continuity of data processing tasks and help meet the high availability requirements of the system.
[0100] In some embodiments of the present application, in addition to realizing dynamic expansion of the first controller object, such as Figure 3 As shown in “Dynamic Scaling”, the first controller object can also be dynamically scaled down. Specifically, Figure 3 As shown in "Dynamic Scaling", the observable node 10 can send a destruction request of the first controller object to the first controller 20 when detecting a destruction event for the first controller object. The first controller 20 can respond to the destruction request by deleting the first controller object and releasing the resources occupied by the first controller object to achieve dynamic scaling of the first controller object. In this way, the resources occupied by the destroyed first controller object can be allocated to other controller objects for use, which helps to improve resource utilization.
[0101] In the embodiments of the present application, the specific implementation form of the destruction event is not limited. In some embodiments, the observable node 10 may determine that a destruction event of the first controller object has occurred when it is monitored that the task execution of the first controller object is completed. In other embodiments, the general data source provides a continuous stream of data to be processed. Only when a certain data source instance is deleted from the data processing system provided in the embodiments of the present application, the observable node 10 will consider that the task execution corresponding to the data source instance is completed. Accordingly, the observable node 10 may determine that a destruction event of the first controller object (defined as the first controller object X) that executes the processing task of the data corresponding to the target data source instance has occurred when it is monitored that the target data source instance is deleted, and the observable node 10 may send a destruction request of the first controller object X to the first controller 20. The first controller 20 may delete the first controller object X in response to the destruction request and release the resources occupied by the first controller object X.
[0102] In some embodiments of the present application, the first controller 20 may also delete the configuration table corresponding to the first controller object X, and delete the resource entry of the first controller object X stored in the database, which helps to save storage resources.
[0103] It is worth noting that if Figure 3 As shown, for an embodiment in which the data to be processed is log data, in addition to obtaining indicator data by using the data processing method provided in the aforementioned embodiment, it is also possible to Figure 3 Steps 6-8 shown in the figure obtain indicator data. Figure 3 The description of steps 6-8 can be found in the previous Figure 1 The relevant contents of steps 1-4 in the above will not be repeated here.
[0104] In addition to the above data processing system, the embodiment of the present application also provides a data processing method. The following is an exemplary description of the data processing method provided in the embodiment of the present application from the perspectives of the observable node and the first controller.
[0105] Figure 4 The flowchart of the data processing method provided in the embodiment of the present application is shown in FIG. The data processing method is applicable to observable nodes. The observable nodes provide observable services. Figure 4 As shown, the method mainly includes:
[0106] 401. Get the configuration file of the data to be processed.
[0107] 402. Send a first creation request to a first controller in the container orchestration system, where the first creation request is used to request the first controller to create a target number of first controller objects.
[0108] 403. Allocate the configuration file of the data to be processed to a target number of first controller objects, so that the target number of first controller objects process the data to be processed according to the configuration file of the data to be processed.
[0109] Figure 5 The flowchart of the data processing method provided in the embodiment of the present application is shown in FIG. The data processing method is applicable to the first controller in the container orchestration system. The container orchestration system also runs an observable node, which is used to provide observable services. Figure 5 As shown, the method mainly includes:
[0110] 501. Obtain a first creation request for a first controller object sent by an observable node; the first creation request is used to request the controller to create a target number of first controller objects.
[0111] 502. In response to the first creation request, a target number of first controller objects are created, so that the observable node can assign configuration files of the data to be processed to the target number of first controller objects, so that the target number of first controller objects can process the data to be processed according to the configuration files of the data to be processed.
[0112] Regarding the implementation forms of the observable node, the first controller and the first controller object, please refer to the relevant content of the aforementioned system embodiment, which will not be repeated here.
[0113] In the embodiment of the present application, in order to realize distributed data processing, the distributed processing capability of the container orchestration system and the resource isolation capability between the controller objects can be used to distribute the data processing tasks received by the observable node to the controller objects in the container orchestration system to realize distributed data processing. Specifically, Figure 4 As shown in step 401, a configuration file of the data to be processed may be obtained. The data to be processed may be log data or data of other contents. For a description of the configuration file of the data to be processed, please refer to the relevant contents of the above-mentioned embodiment.
[0114] Since the configuration file of the data to be processed can reflect which data comes from the same data source instance, the data storage format of the same data source instance is generally the same. In order to reduce the data processing complexity of the controller object and reduce the IO access frequency to the data source, the data of the same data source instance is generally handed over to the same controller object for processing. Based on this, the target number of controller objects to be created can be determined according to the configuration file of the data to be processed, which is recorded as M. M ≥ 1 and is an integer.
[0115] In some embodiments, the data source instance to which the data to be processed belongs can be determined based on the configuration file of the data to be processed. Specifically, the data source instance to which the data to be processed belongs can be determined based on the access address of the data source instance to which the data to be processed belongs. Further, the configuration file of the data to be processed can be assigned to at least one task group based on the data source instance to which the data to be processed belongs. Specifically, it can be determined that the data to be processed with the same access address belongs to the same data source instance; and the configuration subfiles of the data to be processed belonging to the same data source instance are assigned to the same task group to obtain at least one task group.
[0116] In some embodiments, the configuration subfile of the to-be-processed data belonging to the same data source instance can be directly assigned to the same task group to obtain at least one task group. In other embodiments, for example, in the log processing scenario, different data sources provide log data of different magnitudes, and some data source instances provide a small amount of log data, while some data source instances provide a large amount of log data. In order to efficiently utilize resources and reduce the idleness and waste of resources, the configuration subfile of the first to-be-processed data belonging to the API in the to-be-processed data can be assigned to the same task group, and the first to-be-processed data of the task group can be subsequently assigned to the same first controller object for processing, that is, all the to-be-processed data provided by the API can be assigned to a single controller object for processing, which can reduce the idleness and waste of resources of the first controller object. Since the amount of log data provided by the log service instance is relatively large, therefore, the configuration subfile of the second to-be-processed data belonging to the same log service instance in the to-be-processed data can be assigned to the same task group, and the to-be-processed data belonging to the same log service instance can be subsequently assigned to an independent first controller object for processing, rather than assigning all the to-be-processed data of the log service instance to the same first controller object for processing, which can prevent the first controller object from being down due to overload and ensure the stable operation of the log processing task.
[0117] After the configuration file of the data to be processed is allocated to at least one task group, the number of task groups can be used as the target number M of first controller objects to be created. Accordingly, the number of task groups is also the target number, that is, M task groups.
[0118] Further, in step 402, a first creation request for requesting to create a target number of first controller objects may be sent to the first controller. The first creation request may carry the target number of first controller objects to be created.
[0119] Accordingly, for the first controller, in step 501, a first creation request may be received, and in step 502, in response to the first creation request, a target number of first controller objects may be created in the container orchestration system, that is, M first controller objects may be created.
[0120] In some embodiments, the observable node may specify the amount of resources of the first controller object. In some embodiments of the present application, the configuration file of the data to be processed is assigned to a target number of task groups, and the configuration sub-file assigned to each task group is assigned to a first controller object. The observable node determines the amount of resources required for each of the target number of task groups as the amount of resources for each of the first controller objects. Further, the observable node may generate a first creation request based on the amount of resources corresponding to each of the target number of first controller objects and the target number. The first creation request includes the target number of first controller objects to be created, and the amount of resources corresponding to each of the target number of first controller objects. Further, the first creation request may be sent to the first controller to request the first controller to create a target number of first controller objects, thereby enabling the observable node to specify the amount of resources for the data processing tasks assigned to the first controller objects, and achieving refined control over the resources required for the data processing tasks.
[0121] Accordingly, the aforementioned step 502 can be implemented as follows: in response to the first creation request, according to the resource amounts corresponding to the target number of task groups, a target number of first controller objects with corresponding resource amounts are created. Specifically, for any task group A, according to the resource amounts corresponding to task group A, a first controller object with the resource amounts corresponding to task group A can be created. Specifically, according to the resource amounts corresponding to task group A, a target working node with an idle resource amount greater than or equal to the resource amount corresponding to task group A can be determined from the working cluster in the container orchestration system; thereafter, a first controller object with the resource amount corresponding to task group A is created on the target working node.
[0122] Further, in step 403, the configuration file of the data to be processed can be assigned to the target number of first controller objects. Accordingly, the target number of first controller objects can process the data to be processed according to the configuration file of the data to be processed. In this way, the observable node can distribute the data processing tasks to the target number of first controller objects for processing, and realize distributed data processing with the help of the capabilities of the container orchestration system. Compared with the data processing method of a single observable node, the throughput of system data processing can be improved, which helps to achieve the purpose of processing massive data. On the other hand, the data processing load of the observable node is unloaded to the target number of first controller objects, which reduces the data processing pressure of the observable node, helps to reduce the probability of downtime of the observable node, and thus helps to improve the stability of the system.
[0123] On the other hand, this embodiment uses a single observable node and uses the ability of the container orchestration cluster's own controller object to achieve distributed data processing, avoiding the above Figure 1The development and integration of distributed consensus algorithms in the case of multiple observable nodes in the scheme shown can reduce the workload of system development.
[0124] In addition, every time a new data processing task arrives, the observable node and the controller cooperate with each other and can create a first controller object to process the corresponding data processing task in the above manner, so as to realize the dynamic expansion of the first controller object according to the data processing task. Therefore, by using the capabilities of the controller object of the container orchestration system to run data processing tasks, the data processing volume can be continuously increased within the scope allowed by the hardware capabilities of the working cluster of the container orchestration system, realizing massive data processing.
[0125] In some embodiments of the present application, step 403 can be implemented as follows: from the configuration file of the data to be processed, determine the configuration sub-file of the data to be processed to which each of the target number of task groups is assigned. Further, the configuration sub-files corresponding to each of the target number of task groups can be respectively assigned to the target number of first controller objects; each first controller object is assigned to a configuration sub-file. In this way, the observable node can distribute the data processing tasks to the target number of first controller objects for processing, which reduces the data processing pressure of the observable node, helps to reduce the probability of downtime of the observable node, and thus helps to improve the stability of the system. Each first controller object undertakes the data processing tasks of a task group, which helps to achieve load balancing between the target number of first controller objects.
[0126] Specifically, since the amount of resources corresponding to each task group can be the same or different, and the first controller object also has the amount of resources, for any task group A, the configuration subfile corresponding to task group A can be allocated to the first controller object (defined as the target controller object) having the same amount of resources as the amount of resources corresponding to task group A. That is, the target controller object refers to: among the target number of first controller objects, the first controller object to which the amount of resources allocated is the amount of resources corresponding to task group A. In the same way, the configuration subfiles of the target number of task groups A can be allocated to the target number of first controller objects. This task allocation method can make the amount of resources required by task group A match the amount of resources provided by the first controller object. On the one hand, it can prevent the amount of resources of the first controller object from being greater than the amount of resources required by task group A, resulting in a waste of resources; on the other hand, it can also prevent the amount of resources of the first controller object from being less than the amount of resources required by task group A, causing the first controller object to be overloaded and downtime.
[0127] Specifically, the configuration sub-file corresponding to task group A can be assigned to the target controller object with the help of the configuration table (ConfigMap) in the container orchestration system. Among them, the configuration table is an object used to store configuration data in the container orchestration system, which can contain data in the form of key-value pairs. Based on this, the observable node can write the configuration sub-file corresponding to task group A into the target configuration table corresponding to the target controller object in the container orchestration system. Correspondingly, the target controller object can obtain the configuration sub-file corresponding to task group A from the target configuration table, thereby allocating the configuration sub-file corresponding to task group A to the target controller object. In the same way, the configuration sub-files of the target number of task groups A can be assigned to the target number of first controller objects. This task allocation method reuses the data transmission path of the container orchestration system itself, which can reduce the development difficulty and development cost of the distributed data processing system.
[0128] Further, for the first controller object, a data processing task can be executed according to the assigned target configuration subfile. Specifically, the first controller object can obtain the access address of the target data source instance to which the target data to be processed by the first controller object belongs and the authentication information for accessing the target data source instance from the assigned target configuration subfile. The data to be processed includes the target data. Further, the first controller object can obtain the target data from the target data source instance according to the access address and authentication information of the target data source instance.
[0129] In some embodiments, all data stored in the target data source instance are target data. Accordingly, the first controller object can access the target data source instance according to the access address of the target data source instance; and obtain the permission to read data from the target data source instance according to the authentication information; and then, all data stored in the target data source instance can be read from the target data source instance as target data.
[0130] In other embodiments, the data processing task processes part of the data in the target data source instance. Accordingly, the configuration file of the data to be processed may also include: the identification of the data to be processed (such as an index), etc. Accordingly, the target configuration subfile may also include: the identification of the target data (such as an index), etc. Therefore, accordingly, the first controller object may access the target data source instance according to the access address of the target data source instance; and obtain the permission to read data from the target data source instance according to the authentication information; and then, the target data may be read from the target data source instance according to the identification of the target data.
[0131] In an embodiment of the present application, the configuration file of the data to be processed may also include: a data processing method for the data to be processed. The data processing method may be determined by the observable node, or may be sent to the observable node by other nodes. In some embodiments, the observable node may determine the data processing method based on the indicators to be obtained in the data processing task. After determining the data processing method for the data to be processed, the observable node may write the data processing method for the data to be processed into the configuration file for the data to be processed. Accordingly, the configuration sub-file assigned to the first controller object also includes: a data processing method for the target data to be processed by the first controller object.
[0132] In other embodiments, the data processing method of the data to be processed is sent to the observable node by other nodes. Other nodes carry the data processing method of the data to be processed in the configuration file of the data to be processed and send it to the observable node. When the observable node distributes the configuration file of the data to be processed to multiple task groups, the configuration sub-file distributed to each task group also includes: the data processing method of the data corresponding to the configuration sub-file. Based on this, the first controller object can also process the target data according to the data processing method carried by the target configuration sub-file distributed to it, so as to obtain the data processing result of the target data.
[0133] The data processing methods shown in the above embodiments are only exemplary and do not constitute limitations. Figure 1 In the distributed processing system shown, since the observable node in the embodiment of the present application does not undertake data processing tasks and task scheduling, if the distributed processing system provided by the present embodiment has an abnormality in the observable node, there is no need to re-elect the master, which can improve the speed of abnormal recovery of the observable node. The abnormal recovery process of the observable node is exemplarily described below.
[0134] Specifically, the observable node can be implemented as a second controller object in the container orchestration system. Accordingly, the container orchestration system also includes: a controller of the second controller object (defined as a second controller). For the specific implementation form of the second controller object and the controller of the second controller object, please refer to the relevant content of the aforementioned embodiment, which will not be repeated here.
[0135] The second controller may monitor the health status of the observable node. In some embodiments, the second controller may use a live probe to perform a health check on the observable node, that is, monitor the health status of the observable node. The second controller may periodically send a heartbeat message to the observable node according to a set inspection cycle; if a response message to the heartbeat message from the observable node is received within a set time period, it is determined that the observable node is operating normally; if a response message to the heartbeat message from the observable node is not received within a set time period, it is determined that the health status of the observable node is abnormal.
[0136] Furthermore, when an abnormal health status of an observable node is detected, the second controller may create a new observable node in the container orchestration system according to the resource definition of the observable node. In the embodiment of the present application, for ease of description and distinction, the original observable node is defined as the first observable node, and the newly created observable node is defined as the second observable node. Regarding the specific implementation method of the second controller creating a new observable node in the container orchestration system according to the resource definition of the observable node, please refer to the creation process of the first controller object mentioned above, which will not be repeated here.
[0137] Furthermore, the load of the first observable node can be migrated to the newly created second observable node to achieve abnormal recovery of the observable node. The load of the first observable node may include the currently received data processing tasks and the configuration files of the data to be processed.
[0138] In this embodiment, since the observable node does not undertake data processing tasks, it is only responsible for task grouping and the creation trigger of controller objects, and there is no need to perform master selection operations and task scheduling. Therefore, the abnormal recovery of the observable node is faster, which can improve the abnormal recovery speed of the observable node and improve the system stability.
[0139] In an embodiment of the present application, in addition to the observable node having an abnormality, the first controller object may also have an abnormality. In order to be able to promptly discover the abnormal situation of the first controller object, and promptly perform abnormal recovery on the first controller object to ensure the normal execution of the data processing task, in some embodiments of the present application, assuming that the observable node in the current data processing system is still the first observable node, the first controller object can periodically send heartbeat data to the first observable node according to the set heartbeat cycle. The first observable node can update the heartbeat time of the first controller object stored in the database based on the received heartbeat data of the first controller object; and check whether the heartbeat time of the first controller object has expired. Further, if the heartbeat time of the first controller object expires, the controller is requested to recreate the first controller object. Specifically, the first observable node can send a second creation request to the first controller requesting the re-creation of the aforementioned abnormal first controller object.
[0140] Accordingly, the first controller can create a new first controller object in response to the second creation request, and migrate the data processing tasks of the original first controller object to the new first controller object, thereby realizing abnormal recovery of the first controller object, ensuring the continuity of the data processing tasks, and helping to meet the high availability requirements of the system.
[0141] In some embodiments of the present application, in order to realize the migration of the data processing task of the abnormal first controller object to the new first controller object, for the original first controller object, after being assigned to the target configuration subfile, a resource entry of the first controller object can also be added to the database. The resource entry includes: the identifier of the configuration table corresponding to the first controller object in the container orchestration system and the resource information of the first controller object.
[0142] Correspondingly, the first controller responds to the second creation request by reading the resource information of the first controller object from the database; and creates a new first controller object based on the resource information of the first controller object. The resource amount of the new first controller object is the same as the resource amount of the original first controller object. Further, the first controller may provide the identifier of the configuration table to the new first controller object. The new first controller object may obtain the target configuration subfile assigned to the original first controller object from the configuration table based on the identifier of the configuration table; and execute the data processing task of the original first controller object based on the target configuration subfile, thereby realizing the migration of data processing tasks and realizing abnormal recovery of the first controller object, which can ensure the continuity of data processing tasks and help meet the requirements of high availability of the system.
[0143] In some embodiments of the present application, in addition to realizing dynamic expansion of the first controller object, dynamic shrinkage of the first controller object can also be realized. Specifically, the observable node can send a destruction request of the first controller object to the first controller when a destruction event for the first controller object is detected. The first controller can respond to the destruction request by deleting the first controller object and releasing the resources occupied by the first controller object to realize dynamic shrinkage of the first controller object. In this way, the resources occupied by the destroyed first controller object can be allocated to other controller objects for use, which helps to improve resource utilization. For the specific implementation methods for the destruction event, please refer to the relevant content of the aforementioned embodiment, which will not be repeated here.
[0144] In some embodiments of the present application, the controller may also delete the configuration table corresponding to the destroyed first controller object, and delete the resource entry of the destroyed first controller object stored in the database, which helps to save storage resources.
[0145] It should be noted that the execution subject of each step of the method provided in the above embodiment can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 401 and 402 can be device A; for another example, the execution subject of step 401 can be device A, and the execution subject of step 402 can be device B; and so on.
[0146] In addition, in some of the processes described in the above embodiments and the accompanying drawings, multiple operations appearing in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel, and the sequence numbers of the operations, such as 401, 402, etc., are only used to distinguish between different operations, and the sequence numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel.
[0147] Accordingly, an embodiment of the present application also provides a computer-readable storage medium storing computer instructions. When the computer instructions are executed by one or more processors, the one or more processors are caused to execute the steps in the aforementioned data processing methods.
[0148] The embodiment of the present application also provides a computer program product, including a computer program, which, when executed by one or more processors, causes the one or more processors to execute the steps in each data processing method. In the embodiment of the present application, the specific implementation form of the computer program product is not limited. In some embodiments, the computer program product can be implemented as an application (Application, APP), a mini-program, a computer-side client, a program module, a plug-in, an installation package, a software development kit (Software Development Kit, SDK), an image file of a CD (such as an ISO file), a plug-in or software as a service (Software as a Service, SaaS) software, etc., but is not limited to this.
[0149] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 6 As shown, the electronic device includes: a memory 60a and a processor 60b. The memory 60a is used to store computer programs.
[0150] The processor 60b is coupled to the memory 60a and is used to execute the computer program to execute the steps in the data processing methods provided in the above embodiments. The specific implementation of each step can be found in the relevant description of the above embodiments, which will not be repeated here.
[0151] In some optional embodiments, such as Figure 6 As shown, the electronic device may also include optional components such as a communication component 60c, a power component 60d, a display component 60e and an audio component 60f. Figure 6 The components are shown schematically only, and do not mean that the electronic device must include Figure 6 The components shown do not necessarily mean that the electronic device can only include Figure 6 Components shown.
[0152] in addition, Figure 6 The components in the dashed box are optional components, not mandatory components, and may depend on the product form of the electronic device. The electronic device of this embodiment may be implemented as a terminal device such as a desktop computer, a laptop computer, a mobile phone, or an IoT device; or it may be various server devices such as a traditional server, a cloud server, or a server cluster.
[0153] In an embodiment of the present application, the memory is used to store a computer program and can be configured to store various other data to support operations on the device where it is located. Among them, the processor can execute the computer program stored in the memory to implement the corresponding control logic. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random-Access Memory, SRAM), electrically erasable programmable read only memory (Electrically Erasable Programmable Read Only Memory, EEPROM), erasable programmable read only memory (Electrical Programmable Read Only Memory, EPROM), programmable read only memory (Programmable Read Only Memory, PROM), read only memory (Read Only Memory, ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0154] In the embodiment of the present application, the processor may be any hardware processing device that can execute the logic of the above method. Optionally, the processor may be a central processing unit (CPU), a graphics processing unit (GPU) or a microcontroller unit (MCU); it may also be a field programmable gate array (FPGA), a programmable array logic device (PAL), a general array logic device (GAL), a complex programmable logic device (CPLD) or other programmable devices; or it may be an advanced reduced instruction set (RISC) processor (Advanced RISC Machines, ARM) or a system on chip (System on Chip, SoC), etc., but not limited thereto.
[0155] In an embodiment of the present application, the communication component is configured to facilitate wired or wireless communication between the device in which it is located and other devices. The device in which the communication component is located can access a wireless network based on a communication standard, such as Wireless Fidelity (WiFi), 2G or 3G, 4G, 5G or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component can also be implemented based on Near Field Communication (NFC) technology, Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology or other technologies.
[0156] In an embodiment of the present application, the display component may include a liquid crystal display (LCD) and a touch panel (TP). If the display component includes a touch panel, the display component may be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.
[0157] In an embodiment of the present application, a power supply component is configured to provide power to various components of the device in which it is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply component is located.
[0158] In an embodiment of the present application, the audio component may be configured to output and / or input audio signals. For example, the audio component includes a microphone (Microphone, MIC), and when the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal may be further stored in a memory or sent via a communication component. In some embodiments, the audio component also includes a speaker for outputting an audio signal. For example, for a device with a language interaction function, voice interaction with a user may be achieved through an audio component.
[0159] It should be noted that the descriptions such as “first” and “second” in this article are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence, nor do they limit “first” and “second” to different types.
[0160] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, compact disc read-only memory (CD-ROM), optical storage, etc.) containing computer-usable program codes.
[0161] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0162] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0163] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0164] In a typical configuration, a computing device includes one or more processors (CPU, etc.), input / output interfaces, network interfaces, and memory.
[0165] Memory may include non-permanent storage in a computer-readable medium, in the form of random-access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0166] Computer storage media is readable storage media, also known as readable media. Readable storage media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic cassettes, disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0167] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of further restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the above elements.
[0168] The above contents are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A data processing system, characterized in that: include: A first observable node and a first controller running in a container orchestration system; The first observable node is used to provide observable services; The first observable node is a single node; The first observable node is used to obtain a configuration file of the data to be processed; Sending a first creation request to the first controller, wherein the first creation request is used to request the first controller to create a target number of first controller objects; the configuration file is used to indicate a way to obtain the data to be processed and includes a data processing method for the data to be processed; The first controller is used to create the target number of first controller objects in response to the first creation request; The first observable node is used to assign the configuration file to the target number of first controller objects; The target number of first controller objects are used to process the data to be processed according to the configuration file.
2. The system according to claim 1, characterized in that The first observable node is further used for: According to the configuration file, determining the data source instance to which the data to be processed belongs; Allocating the configuration file to at least one task group according to the data source instance to which the data to be processed belongs; Taking the number of the at least one task group as the target number; The first observable node, when allocating the configuration file to the target number of first controller objects, is specifically used to: Determine, from the configuration file, configuration subfiles of the to-be-processed data to which the target number of task groups are respectively allocated; The configuration sub-files respectively allocated to the target number of task groups are respectively allocated to the target number of first controller objects; each first controller object is allocated to one configuration sub-file.
3. The system according to claim 1, characterized in that Any first controller object among the target number of first controller objects is specifically used for: Obtaining the access address of the target data source instance and the authentication information for accessing the target data source instance from the allocated target configuration subfile; the configuration file includes the target configuration subfile; The target data source instance is a data source instance to which the target data to be processed by any first controller object belongs; the data to be processed includes the target data; Acquire the target data from the target data source instance according to the access address and the authentication information; The target data is processed according to the data processing method carried by the target configuration sub-file to obtain a data processing result of the target data.
4. The system according to any one of claims 1 to 3, characterized in that: The first observable node is implemented as a second controller object in the container orchestration system; the container orchestration system further includes: a second controller corresponding to the second controller object; The second controller is used to monitor the health status of the first observable node; when it is detected that the health status of the first observable node is abnormal, create a second observable node in the container orchestration system according to the resource definition of the first observable node; and migrate the load of the first observable node to the second observable node.
5. A data processing method, applicable to an observable node running in a container orchestration system, characterized in that: The observable node is a single node, providing observable services; the method comprises: Acquire a configuration file of the data to be processed; the configuration file is used to indicate a way to acquire the data to be processed and includes a data processing method for the data to be processed; Sending a first creation request to a first controller in the container orchestration system; the first creation request is used to request the first controller to create a target number of first controller objects; The configuration file of the data to be processed is distributed to the target number of first controller objects, so that the target number of first controller objects perform data processing on the data to be processed according to the configuration file of the data to be processed.
6. The method according to claim 5, characterized in that Before sending the first creation request to the first controller in the container orchestration system, the method further includes: According to the configuration file, a target number of first controller objects to be created is determined.
7. The method according to claim 6, characterized in that The step of determining a target number of first controller objects to be created according to the configuration file includes: According to the configuration file, determining the data source instance to which the data to be processed belongs; Allocating the configuration file to at least one task group according to the data source instance to which the data to be processed belongs; Taking the number of the at least one task group as the target number; The step of allocating the configuration file to the target number of first controller objects includes: Determine, from the configuration file, configuration subfiles of the to-be-processed data to which the target number of task groups are respectively allocated; The configuration sub-files assigned to the target number of task groups are respectively assigned to the target number of first controller objects; each first controller object is assigned to a configuration sub-file, so that the target number of first controller objects can process the target data corresponding to the target configuration sub-file according to the assigned target configuration sub-file.
8. The method according to claim 7, characterized in that The data source instance includes an application programming interface and / or a log service instance provided by the observable node; and the configuration file is assigned to at least one task group according to the data source instance to which the data to be processed belongs, including: Allocating the configuration subfile corresponding to the first data to be processed belonging to the application programming interface among the data to be processed to the same task group; and / or, The configuration subfile corresponding to the second data to be processed belonging to the same log service instance among the data to be processed is allocated to the same task group.
9. The method according to claim 7, characterized in that: The sending a first creation request to the first controller includes: Determine the resource amount corresponding to each of the target number of task groups as the resource amount corresponding to each of the target number of first controller objects; Generate the first creation request according to the resource amounts corresponding to the target number of task groups and the target number; the first creation request includes the target number and the resource amounts corresponding to the target number of first controller objects; Sending the first creation request to the first controller; The step of respectively allocating the configuration subfiles to which the target number of task groups are respectively allocated to the target number of first controller objects comprises: For any task group, the configuration subfile corresponding to the task group is allocated to a second controller object; the second controller object refers to the first controller object among the target number of first controller objects, the amount of resources allocated to which is the amount of resources corresponding to the task group.
10. The method according to claim 9, characterized in that The step of allocating the configuration subfile corresponding to any one of the task groups to the target controller object includes: The configuration sub-file corresponding to any one of the task groups is written into the target configuration table corresponding to the target controller object in the container orchestration system, so that the target controller object can obtain the configuration sub-file corresponding to any one of the task groups from the target configuration table, and assign the configuration sub-file corresponding to any one of the task groups to the target controller object.
11. The method according to any one of claims 5 to 10, characterized in that: Also includes: According to the received heartbeat data of any first controller object, updating the heartbeat time of any first controller object stored in the database; The heartbeat data is periodically sent by any first controller object to the observable node according to a set heartbeat period; Check whether the heartbeat time of any of the first controller objects has expired; If the heartbeat time of any of the first controller objects expires, a second creation request is sent to the first controller, and the second creation request is used to request the re-creation of any of the first controller objects, so that the first controller can respond to the second creation request, create a new first controller object, and migrate the data processing tasks of any of the first controller objects to the new first controller object.
12. A data processing method, applicable to a first controller in a container orchestration system, characterized in that: The container orchestration system also runs an observable node, which is a single node and is used to provide observable services. The method includes: Obtaining a first creation request for a first controller object sent by the observable node; the first creation request is used to request the first controller to create a target number of first controller objects; In response to the first creation request, the target number of first controller objects are created so that the observable node can assign the configuration files of the data to be processed to the target number of first controller objects, so that the target number of first controller objects can process the data to be processed according to the configuration files of the data to be processed; the configuration files are used to indicate the way to obtain the data to be processed and include the data processing method of the data to be processed.
13. The method according to claim 12, characterized in that Also includes: Obtain the observable node to send a second creation request; The second creation request is used to request to recreate any first controller object, and the second creation request is sent when the observable node detects that the heartbeat time of any first controller object has expired; In response to the second creation request, creating a new first controller object; Migrate the data processing task of any one of the first controller objects to the new first controller object.
14. The method according to claim 13, characterized in that Creating a new first controller object in response to the second creation request includes: In response to the second creation request, read a resource entry of any first controller object from a database; the resource entry is added by any first controller object in the database; the resource entry includes: an identifier of a configuration table corresponding to any first controller object in the container orchestration system and resource information of any first controller object; Creating the new first controller object according to the resource information of any first controller object; The step of migrating the data processing task of any one of the first controller objects to the new first controller object comprises: The identifier of the configuration table is provided to the new first controller object so that the new first controller object can obtain the target configuration sub-file assigned to any of the first controller objects from the configuration table according to the identifier of the configuration table, so as to migrate the data processing task of any of the first controller objects to the new first controller object.
15. An electronic device, characterized in that: include: A memory and a processor; wherein the memory is used to store a computer program; The processor is coupled to the memory and configured to execute the computer program to perform the steps of the method according to any one of claims 5 to 14.
16. A computer-readable storage medium storing computer instructions, characterized in that: When the computer instructions are executed by one or more processors, the one or more processors are caused to execute the steps in the method according to any one of claims 5 to 14.
17. A computer program product, characterized in that The method comprises a computer program which, when executed by one or more processors, causes the one or more processors to execute the steps of the method according to any one of claims 5 to 14.
Citation Information
Patent Citations
Application program configuration method and computing device
CN118885168A