Data detection method and system
By performing multiple data filtering operations at the data acquisition and detection nodes, the problems of data detection node crashes and memory overflows in the data processing cluster were solved, enabling effective detection of the operational status of a large number of data processing units.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ALIBABA CLOUD COMPUTING CO LTD
- Filing Date
- 2024-10-18
- Publication Date
- 2026-04-21
AI Technical Summary
The excessive number of data processing units and the large amount of data in the data processing cluster caused the data detection nodes to crash and experience memory overflow, making it impossible to effectively detect their operational status.
By performing the first data filtering at the data acquisition node to reduce the amount of data in the target unit, and performing the second data filtering at the data detection node to determine the data of the unit to be detected, the operation status detection of a small number of data processing units can be achieved.
This avoids the problems of data detection node downtime and memory overflow, and enables effective detection of the operating status of a large number of data processing units.
Smart Images

Figure CN121901063A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of data processing technology, and in particular to a data detection method; one or more embodiments of this specification also relate to a data detection system, a computing device, a computer-readable storage medium, and a computer program product. Background Technology
[0002] With the continuous development of computer technology, data processing clusters containing multiple data processing units can be used to process data using computers.
[0003] In the current data processing cluster, to ensure the smooth operation of data processing units, data detection nodes are needed to monitor the operational status of these units. However, due to the large number of data processing units and the large amount of data per unit, problems such as data detection nodes crashing or memory overflow may occur, making it impossible to monitor the operational status of the data processing units. Therefore, how to monitor the operational status of data processing units with a large number of units and a large amount of data per unit has become an urgent problem to be solved. Summary of the Invention
[0004] In view of this, embodiments of this specification provide a data detection method. One or more embodiments of this specification also relate to a data detection system, another data detection method, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
[0005] According to a first aspect of the embodiments of this specification, a data detection method is provided, applied to a data detection system, the data detection system including a data detection node and a data acquisition node, wherein... The data acquisition node acquires unit data from multiple data processing units within the data processing cluster, filters the unit data of each data processing unit based on the unit attribute information of each data processing unit, obtains the target unit data of the target data processing unit, and sends the target unit data to the data detection node. The data detection node filters the target unit data based on the variable attribute information in the target unit data to obtain the unit data to be detected within the target unit data. Based on the data of the unit to be detected, a data processing unit to be detected is determined from the plurality of data processing units, and based on the unit data of the data processing unit to be detected, the operating status of the data processing unit to be detected is detected to obtain the operating status detection result.
[0006] According to a second aspect of the embodiments of this specification, a data detection system is provided, including a data detection node and a data acquisition node, wherein... The data acquisition node is used to acquire unit data from multiple data processing units contained in the data processing cluster, filter the unit data of each data processing unit according to the unit attribute information of each data processing unit, obtain the target unit data of the target data processing unit, and send the target unit data to the data detection node. The data detection node is used to filter the target unit data based on the variable attribute information in the target unit data to obtain the unit data to be detected from the target unit data. Based on the data of the unit to be detected, a data processing unit to be detected is determined from the plurality of data processing units, and based on the unit data of the data processing unit to be detected, the operating status of the data processing unit to be detected is detected to obtain the operating status detection result.
[0007] According to a third aspect of the embodiments of this specification, another data detection method is provided, applied to a data detection system, the data detection system including a data detection node and a data acquisition node, wherein... The data acquisition node acquires unit data of multiple Pods contained in the Kubernetes cluster, filters the unit data of each Pod according to the unit attribute information of each Pod, obtains the target unit data of the target Pod, and sends the target unit data to the data detection node. The data detection node filters the target unit data based on the variable attribute information in the target unit data to obtain the unit data to be detected within the target unit data. Based on the unit data to be detected, the Pod to be detected is determined from the plurality of Pods, and the running status of the Pod to be detected is detected based on the unit data of the Pod to be detected, so as to obtain the running status detection result.
[0008] According to a fourth aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the above-described data detection method.
[0009] According to a fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the data detection method described above.
[0010] According to a sixth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the data detection method described above.
[0011] The data processing method for a data detection system provided in one or more embodiments of this specification, during the operation status detection process, utilizes a data acquisition node to perform a first data filtering on the acquired unit data and sends the target unit data obtained from the data filtering to a data detection node. During the operation status detection process, the data detection node performs a second data filtering on the target unit data, thereby obtaining a smaller number of unit data to be detected, thus avoiding the problem of too much detection data. Then, based on the unit data to be detected, a smaller number of data processing units to be detected are determined from multiple data processing units, and the operation status of these data processing units is detected, thus avoiding the problem of a large number of data processing units to be detected. Based on this, this data processing method enables operation status detection of a large number of data processing units with a large amount of unit data, avoiding problems such as data detection node crashes and memory overflows caused by excessive number of data processing units and excessive unit data volume. Attached Figure Description
[0012] Figure 1 This is a schematic diagram illustrating the application of a data detection method provided in one embodiment of this specification; Figure 2 This is a flowchart of a data detection method provided in one embodiment of this specification; Figure 3 This is a flowchart illustrating the processing procedure of a data detection method provided in one embodiment of this specification; Figure 4 This is a schematic diagram of the structure of a data detection system provided in one embodiment of this specification; Figure 5 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0013] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0014] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0015] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0016] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0017] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0018] Prometheus: A cluster monitoring system.
[0019] K8s: Also known as Kubernetes cluster, K8s can be used to manage containerized applications on multiple hosts in a cloud platform. K8s can be a cluster scheduling system.
[0020] Service discovery (discover targets): A feature in Prometheus used to discover which Kubernetes pods need to be monitored.
[0021] relebal: Prometheus uses this feature to filter whether a pod needs to be monitored.
[0022] list / watch: Prometheus's ability to retrieve Kubernetes pods.
[0023] pod: The smallest unit of container group running in Kubernetes.
[0024] namespace: refers to a namespace. Pods in Kubernetes are grouped by namespace; list can specify a namespace parameter to retrieve pods.
[0025] Kruise: The core of the OpenKruise project, it is a set of controllers that extend and complement the Kubernetes core controllers in application workload management.
[0026] list / watchkuberenetesapi: List / watch the Kubernetes API.
[0027] keep / drop: Keep / discard.
[0028] annotation: comment.
[0029] pprof: Performance analysis.
[0030] oteltracing: OpenTelemetry tracing or "OpenTelemetry tracking".
[0031] Short-lived jobs refer to short-lived or batched tasks that do not have enough time to be deleted and cannot be retrieved by pull. They need to be retrieved by push and pushgeteway.
[0032] Push Gateway: A component in the Prometheus ecosystem, primarily used to address the issue that Prometheus's default pull mode may fail to retrieve data in certain situations.
[0033] OOM (Out Of Memory): This refers to a memory overflow, which means that there is unrecoverable memory or excessive memory usage in the application system, ultimately causing the program to require more memory than the maximum available memory.
[0034] HDD (Hard Disk Drive) refers to a hard disk, which is the most important storage device in a computer.
[0035] SSD (Solid State Drive): refers to a solid-state drive.
[0036] With the continuous development of computer technology, data processing clusters containing multiple data processing units can be used to process data. In order to ensure the smooth operation of the data processing units, data detection nodes are needed to monitor the running status of the data processing units. However, due to the large number of data processing units and the large amount of data in each unit, problems such as data detection nodes crashing and memory overflow may occur, making it impossible to monitor the running status of the data processing units.
[0037] For example, some internet service platforms are extremely large, and the Kubernetes clusters running these services might have over 300,000 pods. This can cause Prometheus to consume excessive memory, making it impossible to monitor all pods. Especially in the event of a Kubernetes cluster failure, the fault-handling component, Kruise, needs to monitor the online / offline status of all pods in the cluster and perform dynamic service discovery to promptly update the IP address of the monitored Kruise. However, due to the large number of pods and the massive amount of data in the cluster, the data volume quickly reaches 40GB after Prometheus deployment, followed by continuous OutOfMemoryError (OOM) events, making it impossible to monitor core critical components like Kruise.
[0038] Based on this, a data detection method is provided in this specification. One or more embodiments of this specification also relate to a data detection system, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0039] See Figure 1 , Figure 1 The diagram illustrates an application of a data detection method according to an embodiment of this specification. This data detection method can be applied to a data detection system, which includes a Prometheus client and a service discovery module. The service discovery module can obtain the filtered metric data from multiple pods in a Kubernetes cluster based on preset pod IPs, and then send the filtered metric data to the Prometheus client. This Prometheus client can trim some variables in the filtered metric data to obtain the variable information needed by the Prometheus client. Then, based on this variable information, it can dynamically discover the IPs of key pods, and then start collecting detection data of key pods based on these IPs. By detecting this detection data, the detection results can be obtained.
[0040] See Figure 2 , Figure 2A flowchart of a data detection method according to an embodiment of this specification is shown. The data detection method is applied to a data detection system, which includes a data detection node and a data acquisition node, and specifically includes the following steps.
[0041] Step 202: The data acquisition node acquires the unit data of multiple data processing units contained in the data processing cluster, filters the unit data of each data processing unit according to the unit attribute information of each data processing unit, obtains the target unit data of the target data processing unit, and sends the target unit data to the data detection node. The data processing cluster can be understood as a computer cluster used for data processing. For example, the data processing cluster can be a Kubernetes cluster, a distributed server cluster, a database cluster, etc.; the data processing can be understood as data computation, data storage, data backup, etc.
[0042] A data processing unit can be understood as a unit used to perform data processing operations. For example, the data processing unit can be a pod, container, server, cloud server, virtual machine, database, etc., without specific restrictions.
[0043] Unit data can be understood as performance indicator data, log data, configuration parameters, etc., determined by each data processing unit during the data processing process, or the unit data can be the source code data of the data processing unit.
[0044] Unit attribute information can be understood as information representing specific attributes of a data processing unit, such as the unit identifier, IP address, unit type, and interface information corresponding to the data processing unit.
[0045] A data detection node can be understood as a node used to detect data processing units. This data detection node can be a client, service system, server, or server; for example, this data detection node can be Prometheus.
[0046] The data acquisition node can be understood as a node or functional module used to acquire unit data from the data processing cluster and forward the unit data to the data detection node; for example, the data acquisition node can be a module that implements the kubernetesdiscovertargets function, which can be set in a server, client or cloud server.
[0047] In one or more embodiments provided in this specification, the step of filtering the unit data of each data processing unit based on the unit attribute information of each data processing unit to obtain the target unit data of the target data processing unit includes: The data acquisition node determines the unit identifier of each data processing unit from the unit attribute information of each data processing unit, and determines a preset unit identifier for the target data processing unit. If the unit identifier is consistent with the preset unit identifier, the data processing unit corresponding to the unit identifier is determined as the target data processing unit, and the target unit data of the target data processing unit is obtained.
[0048] The unit identifier can be understood as information that uniquely identifies a data processing unit, such as the unit name and unit number of the data processing unit.
[0049] The preset unit identifier can be understood as a pre-set unit identifier, which is used to identify the key target data processing unit from multiple data processing units.
[0050] Taking the application of the data detection method provided in this manual in a large-scale Kubernetes full pod service discovery scenario as an example, the data detection method is explained. In this case, the data processing cluster is a Kubernetes cluster, the data processing unit is a pod unit, the data detection node is Prometheus, and the data acquisition node is the Kubernetes discovertargets module.
[0051] Based on this, in the Prometheus architecture, the Kubernetes discover targets module is also called the Prometheus service discovery module. Prometheus and Kubernetes use the list / watch Kubernetes API to perform discover target operations.
[0052] The `list` function within the `list / watch` feature. When no `namespace` parameter is passed during data retrieval, the first time the Prometheus client synchronizes the metric data of all pods in the Kubernetes cluster into the client's memory. After successful execution, the Prometheus memory caches the metric data of all pods and uses this data as input to the Prometheus retrieval module. However, when Kubernetes metrics consume a significant amount of memory, such as data from 300,000 pods, a network request can cause excessive memory usage on the client side, leading to an OutOfMemoryError (OOM) and process termination. This problem is known as the listAll problem.
[0053] To address the aforementioned issues, since excessive memory consumption was observed after the initial deployment of the Prometheus architecture, the data processing method provided in this manual configures the data filtering function into the Prometheus service discovery module, thereby performing data filtering on the unit data.
[0054] Specifically, this method moves the relabeling module down to the service discovery module, thereby pre-filtering 730,000 pods. This relabeling module is a component of the scrape function within retrieval. The scrape function configured within this retrieval can use this input data to determine which critical pods to collect for detection. The specific decision on which critical pods to collect is achieved through the keep / drop functionality of the scrape's relabeling module, combined with annotations added to the pods. This allows for the dynamic discovery of critical pod IPs; then, the collection of detection data for these critical pods begins. For example, out of 300,000 pods, only 3 are kruisepods (critical pods). In other words, the relabeling function can filter out over 290,000 pods.
[0055] The relabel module sends the IP of a critical pod in the following way: it determines the pod IP from the pod's metric data. If the pod IP matches the preset critical pod IP, then the pod is a critical pod.
[0056] In the above embodiments, by configuring the data filtering function to the data acquisition node, the unit data is filtered for the first time, thus avoiding the occurrence of a large amount of unit data to the data detection node, which could cause problems such as memory overflow and crashes in the data detection node.
[0057] In one or more embodiments provided in this specification, obtaining unit data of multiple data processing units included in the data processing cluster includes: The data detection node sends an initial data acquisition request to the data acquisition node, wherein the initial data acquisition request carries multiple unit group identifiers; The data acquisition node generates multiple target data acquisition requests based on the multiple unit group identifiers carried in the initial data acquisition request, and acquires the unit data of the multiple data processing units contained in the data processing cluster according to each target data acquisition request.
[0058] In this context, a unit group can be understood as a group consisting of multiple data processing units. In a data processing cluster, multiple data processing units can be managed through unit groups, and a unit group can include one or more data processing units. For example, a unit group can be a namespace.
[0059] A unit group identifier can be understood as information used to uniquely identify a unit group, such as the unit group name or unit group number.
[0060] Specifically, to avoid excessively large amounts of unit data, the data detection node in this method, after sending an initial data acquisition request to the data acquisition node, can determine that the initial data acquisition request carries multiple unit group identifiers, and generate multiple target data acquisition requests based on the multiple unit group identifiers. This decomposes an initial data acquisition request into multiple target data acquisition requests, and then sequentially acquires the unit data of the multiple data processing units contained in the data processing cluster according to each target data acquisition request, thereby avoiding the problem of excessively large amounts of unit data acquired in a single request.
[0061] In one or more embodiments provided in the specification, the data acquisition node generates multiple target data acquisition requests based on the multiple unit group identifiers carried in the initial data acquisition request, and acquires the unit data of the multiple data processing units included in the data processing cluster according to each target data acquisition request, including: The data acquisition node determines the multiple unit group identifiers carried in the initial data acquisition request, and generates multiple target data acquisition requests based on the multiple unit group identifiers, wherein the unit group identifier is the identifier of the unit group in the data processing cluster, and each target data acquisition request corresponds to each unit group in the data processing cluster; Based on the multiple target data acquisition requests, determine the unit group corresponding to each unit group identifier from the data processing cluster; A target unit group is determined from multiple unit groups, and unit data of the data processing units contained in the target unit group is obtained, wherein the target unit group is any one of the multiple unit groups.
[0062] Continuing with the previous example, after shifting the relevant filtering to the data acquisition node, the amount of data in memory can be reduced, for example, to below 10GB. However, occasionally there are cases where the memory capacity is exceeded (e.g., 10GB); through analysis, it was determined that this was because the data acquisition node was loading too much pod's metric data at once.
[0063] Therefore, this method references the pagination implementation of databases, decomposes a single request into multiple requests, and limits the amount of pod metrics data retrieved each time by executing the request one by one; alternatively, it enables multi-threaded execution of the request to avoid excessive time consumption.
[0064] In the process of breaking down a single request into multiple requests, this method uses namspace as the pagination basis. Each request in the decomposed multiple requests corresponds to a namspace, thereby enabling the acquisition of pod metric data in the namspace.
[0065] In one or more embodiments provided in this specification, the data detection system includes a data transmission node, and the method further includes: The data sending node acquires short-term task data from the multiple data processing units, wherein the short-term task data is data generated by the data processing units executing short-term tasks, and the short-term task is a task whose execution time is less than a preset time threshold. The short-term task data is sent as unit data to the data detection node.
[0066] A data sending node can be understood as a unit used to acquire short-term task data and push the short-term task data to the data detection node. For example, the data sending node is pushgatway.
[0067] Short-lived jobs can be understood as tasks that exist for a short period of time.
[0068] Continuing with the previous example, Push Gateway is primarily used for short-lived jobs (i.e., short-term tasks). Because these jobs have a short lifespan, they may disappear before Prometheus performs a pull operation. Therefore, to obtain metric data for these tasks, Push Gateway can retrieve the metric data for short-lived jobs and push this metric data (i.e., short-term task data) directly to the Prometheus client, thereby avoiding data omissions.
[0069] Step 204: The data detection node filters the target unit data according to the variable attribute information in the target unit data to obtain the unit data to be detected in the target unit data. Based on the unit data to be detected, it determines the data processing unit to be detected from the plurality of data processing units, and performs operation status detection on the data processing unit to be detected based on the unit data of the data processing unit to be detected to obtain the operation status detection result.
[0070] Among them, variable attribute information can be understood as information that represents specific attributes of variable data. For example, variable attribute information includes variable name, variable ID, etc.
[0071] The data to be detected can be understood as the data of the units that need to be detected and processed. For example, by using the keep / drop function in the relabel module of the scrape function module in the data detection node, together with the annotations in the pod, the IP of the key pod can be dynamically discovered. Then, all the indicator data of the key pod are collected for detection to obtain the detection results.
[0072] In one or more embodiments provided in this specification, the variable attribute information is a variable identifier; The step of filtering the target unit data based on the variable attribute information in the target unit data to obtain the target unit data to be detected includes: The data detection node determines variable data and variable identifiers of the variable data from the target unit data, and determines a preset variable identifier for the variable data; If the variable identifier is consistent with the preset variable identifier, the variable data corresponding to the variable identifier is clipped to obtain the retained variable data; The retained variable data is determined as the data of the unit to be detected in the target unit data.
[0073] Continuing with the previous example, when the Prometheus client was consuming excessive memory, analysis revealed that a cluster with 300,000 pods had an array in memory that had ballooned to 730,000. This array was caused by retrieving all Kubernetes pod metric data without any filtering. Based on this, this method reduces the number of pod variables, retaining only the pod variables required by Prometheus; see the following data for details: f len(c.Ports) == 0 { / / We don't have a port so we just set the address label to the podIP. / / The user has to add a port manually. tg.Targets = append(tg.Targets, model.LabelSet{ model.AddressLabel: lv(pod.Status.pod ip), podContainerNameLabel: lv(c.Name), podContainerIDLabel: lv(cID), podContainerImageLabel: lv(c.Image), podContainerIsInit: lv(strconv.FormatBool(isInit)), }) continue }".
[0074] As can be seen from the above data, it contains various variables such as pod IP, Name, and Image; and these variables occupy a large amount of memory space. Therefore, this method prunes some of the pod variables, keeping only the pod variables required by Prometheus (such as Name), thereby reducing the memory usage of variable data.
[0075] In one or more embodiments provided in this specification, after performing operational status detection on the data processing unit to be detected based on the unit data of the data processing unit to be detected and obtaining the operational status detection result, the method further includes: The data detection node receives a detection data acquisition request sent by the client, wherein the detection data acquisition request is sent by the client when the user performs a detection data acquisition operation on the data acquisition page; Determine the detection data identifier carried in the detection data acquisition request, and obtain the detection data corresponding to the detection data identifier from the time series database, wherein the detection data includes the unit data of the data processing unit to be detected and / or the operation status detection result; The detection data is sent to the client, so that the client can display the detection data using the data acquisition page. For example, the detection data can be displayed to the user.
[0076] Following the previous example, this Prometheus architecture can receive query statements sent by clients via an HTTP service. Prometheus can then retrieve corresponding indicator data or detection results from the time series database based on the query statement and return them to the client.
[0077] This client can display indicator data or test results to users through a visual interactive interface or user interface.
[0078] In one or more embodiments provided in this specification, after performing operational status detection on the data processing unit to be detected based on the unit data of the data processing unit to be detected and obtaining the operational status detection result, the method further includes: When the operation status detection result is consistent with the preset risk detection result, the data detection node generates a risk warning message based on the operation status detection result and sends the risk warning message to the risk processing node.
[0079] In this context, the risk processing node can be understood as a device that processes risks. For example, the risk processing node can be an alarm manager or an operation and maintenance device.
[0080] Continuing with the previous example, during the detection process, this Prometheus architecture can push the risk alert to the alert manager (i.e., the risk handling node) or to the operations and maintenance equipment (i.e., the risk handling node) when it detects a risk in a pod; specifically, it can push the alert information to the operations and maintenance equipment through the alert manager.
[0081] The data processing method for a data detection system provided in one or more embodiments of this specification, during the operation status detection process, utilizes a data acquisition node to perform a first data filtering on the acquired unit data and sends the target unit data obtained from the data filtering to a data detection node. During the operation status detection process, the data detection node performs a second data filtering on the target unit data, thereby obtaining a smaller number of unit data to be detected, thus avoiding the problem of too much detection data. Then, based on the unit data to be detected, a smaller number of data processing units to be detected are determined from multiple data processing units, and the operation status of these data processing units is detected, thus avoiding the problem of a large number of data processing units to be detected. Based on this, this data processing method enables operation status detection of a large number of data processing units with a large amount of unit data, avoiding problems such as data detection node crashes and memory overflows caused by excessive number of data processing units and excessive unit data volume.
[0082] The following is in conjunction with the appendix Figure 3 Taking the data detection method provided in this specification as an example of its application in the discovery of a large-scale Kubernetes full pod service, the data detection method will be further explained. Among other things, Figure 3The present specification illustrates a process flowchart of a data detection method according to an embodiment. This data detection method can be applied to a Prometheus architecture, which includes a Prometheus client and a service discovery module (discovertargets module).
[0083] In the Prometheus architecture, the Prometheus client and the Kubernetes cluster use the list / watch Kubernetes API to perform the discover targets operation.
[0084] In the list / watch functionality, when no namespace parameter is passed during data retrieval, the Prometheus client needs to synchronize the metric data of all pods in the Kubernetes cluster on the first attempt. After the operation is successful, Prometheus memory will cache the metric data of all pods and input all the metric data into the retrieval module in Prometheus. However, when the Kubernetes cluster's metrics consume a significant amount of memory, such as data from 300,000 pods, a network request can cause excessive memory usage on the client side, leading to an OutOfMemoryError (OOM) and process termination. This problem is known as the listAll problem.
[0085] To address the above issues, this solution provides three measures: relabeling filtering pushdown, minimizing pod object information, and batching lists.
[0086] The specific implementation of the relabel filter pushdown is as follows: Because the memory consumption was too high after the initial deployment of the Prometheus architecture, this solution configures the data filtering function into the Prometheus service discovery module to perform data filtering on the unit data.
[0087] Specifically, this solution moves the relabeling module down to the service discovery module, thereby pre-processing the filtering of 730,000 pods. The relabeling module is a component of the scrape function within retrieval; the scrape function configured within retrieval can then use this input data to determine which critical pods to collect for inspection.
[0088] Specifically, deciding which critical pods to collect data from can be achieved using the keep / drop functionality of scrape's relabel module, combined with annotations added to the pods. This allows for the dynamic discovery of critical pod IPs. Then, based on these IPs, the collection of detection data for the critical pods begins. For example, out of 300,000 pods, there might only be 3 kruisepods (critical pods). In other words, the relabel function can filter out over 290,000 pods.
[0089] The relabel module sends the IP address of the critical pod in the following way: The pod IP is determined from the pod's metric data. If the pod IP matches the preset critical pod IP, then the pod is a critical pod. For details, please refer to the comparison of information between ordinary pods and critical pods below.
[0090] A regular pod: "apiVersion:v1" kind:Pod metadata: name:my-pod1” The key pod: "apiVersion:v1" kind:Pod metadata: name: kruise annotations: prometheus.io / scrape: "true" # Enable Prometheus scraping prometheus.io / port: "9102" # Specifies the port that Prometheus crawls on, the default is 9102. Based on the data comparison above, it can be seen that the critical pod contains specific key information (such as prometheus.io / scrape: "true"). Based on this key information, the critical pod and the critical pod IP can be identified from multiple pods.
[0091] The specific implementation of batching lists is as follows: After shifting the relevant filtering to the data acquisition node, the amount of data in memory can be reduced, for example, to below 10GB. However, occasionally there are cases where the memory capacity is exceeded (e.g., 10GB); through analysis, it was determined that this was because the data acquisition node was loading too much pod's metric data at once.
[0092] Therefore, this solution references the pagination implementation in databases, breaking down a single request into multiple requests and executing each request sequentially to limit the amount of pod metrics data retrieved each time. Furthermore, multi-threading can be enabled to execute the request, thus preventing excessive processing time.
[0093] In the process of breaking down a single request into multiple requests, this solution uses namspace as the basis for pagination. Each request in the decomposed multiple requests corresponds to a namspace, thereby enabling the acquisition of pod metric data within the namspace.
[0094] The specific implementation method for minimizing pod object information is as follows: Analysis revealed that in a cluster with 300,000 pods, an array in memory had ballooned to 730,000 elements. This was caused by retrieving all Kubernetes pod metric data without any filtering. Based on this, this method reduces the number of pod variables, retaining only the pod variables required by Prometheus; see the following data for details: f len(c.Ports) == 0 { / / We don't have a port so we just set the address label to the podIP. / / The user has to add a port manually. tg.Targets = append(tg.Targets, model.LabelSet{ model.AddressLabel: lv(pod.Status.pod ip), podContainerNameLabel: lv(c.Name), podContainerIDLabel: lv(cID), podContainerImageLabel: lv(c.Image), podContainerIsInit: lv(strconv.FormatBool(isInit)), }) continue }".
[0095] As can be seen from the above data, it contains various variables such as pod IP, Name, and Image; and these variables occupy a large amount of memory space. Therefore, this method prunes some of the pod variables, keeping only the pod variables required by Prometheus (such as Name), thereby reducing the memory usage of variable data.
[0096] Furthermore, based on Figure 3 As can be seen, the Prometheus architecture of this solution also includes push gataway, alarm manager, and data visualization and export.
[0097] Push gataway is primarily used for short-lived jobs; these jobs have a short lifespan and may disappear before Prometheus performs a pull operation. Therefore, to obtain metric data for these jobs, they can directly push their metric data to the Prometheus client.
[0098] The alarm manager is used for alarm notifications. When the Prometheus client detects a risk in a pod, it can push the risk alarm to the alarm manager, which in turn pushes the alarm information to the maintenance equipment.
[0099] The data visualization and export functions provide data query capabilities. This Prometheus architecture can receive query statements sent by clients via HTTP service. Prometheus can then retrieve corresponding indicator data or detection results from the time series database based on the query statement and return them to the client. The client can then display the indicator data or detection results to the user through a visual interactive interface or user interface.
[0100] Based on the above steps, the data detection method provided in this manual offers a low-memory-consumption method for large-scale Kubernetes full pod service discovery using Prometheus. This method continuously tracks the consumption and changes of various variable objects in memory for Prometheus Kubernetes service discovery through detailed memory analysis and tracing. It designs three effective improvement techniques: relabel filtering pushdown, minimizing pod object information, and batch listing. This frees up memory space and reduces memory consumption, for example, reducing memory consumption from 40GB to 4GB.
[0101] The above method addresses the issue of Prometheus monitoring for Kubernetes clusters with 300,000 pods. The problem is that when the number of pods in a Kubernetes cluster is large, the Prometheus service discovery module requires over 40GB of memory, while the memory of a typical Node.js ECS instance typically doesn't exceed 40GB. This prevents Kubernetes from monitoring all pods. This solution, however, allows monitoring of the original 40GB cluster to run with only 4GB of memory, demonstrating significant application value.
[0102] This specification provides another data detection method according to one or more embodiments. This data detection method is applied to a data detection system, which includes a data detection node and a data acquisition node. The data acquisition node acquires unit data of multiple Pods contained in the Kubernetes cluster, filters the unit data of each Pod according to the unit attribute information of each Pod, obtains the target unit data of the target Pod, and sends the target unit data to the data detection node. The data detection node filters the target unit data based on the variable attribute information in the target unit data to obtain the unit data to be detected within the target unit data. Based on the unit data to be detected, the Pod to be detected is determined from the plurality of Pods, and the running status of the Pod to be detected is detected based on the unit data of the Pod to be detected, so as to obtain the running status detection result.
[0103] The data detection method provided in one or more embodiments of this specification, during the runtime status detection process, utilizes a data acquisition node to obtain unit data of Pods in a Kubernetes cluster, performs a first data filtering on the unit data, and sends the target unit data obtained from the data filtering to the data detection node. During the runtime status detection process, the data detection node performs a second data filtering on the target unit data, thereby obtaining a smaller number of unit data to be detected, thus avoiding the problem of too much detection data. Then, based on the unit data to be detected, a smaller number of Pods to be detected are determined from multiple Pods, and runtime status detection is performed on these Pods, thus avoiding the problem of too many Pods needing to be detected. Based on this, this data detection system enables runtime status detection of Pods with an excessive number of Pods and a large amount of unit data, avoiding problems such as data detection node crashes and memory overflows caused by excessive number of Pods and excessive amount of unit data.
[0104] The above is an illustrative scheme of another data detection method in this embodiment. It should be noted that the technical solution of this other data detection method belongs to the same concept as the technical solution of the data detection method described above. For details not described in detail in the technical solution of the other data detection method, please refer to the description of the technical solution of the data detection method described above.
[0105] Corresponding to the above method embodiments, this specification also provides embodiments of a data detection system. Figure 4 A schematic diagram of the structure of a data detection system provided in one embodiment of this specification is shown. Figure 4 As shown, the data detection system includes a data detection node 404 and a data acquisition node 402, wherein, The data acquisition node 402 is used to acquire unit data of multiple data processing units contained in the data processing cluster, filter the unit data of each data processing unit according to the unit attribute information of each data processing unit, obtain the target unit data of the target data processing unit, and send the target unit data to the data detection node 404. The data detection node 404 is used to filter the target unit data according to the variable attribute information in the target unit data to obtain the unit data to be detected in the target unit data, determine the data processing unit to be detected from the plurality of data processing units according to the data to be detected, and perform operation status detection on the data processing unit to be detected based on the unit data of the data processing unit to be detected to obtain the operation status detection result.
[0106] Optionally, the data acquisition node 402 is further configured to determine the unit identifier of each data processing unit from the unit attribute information of each data processing unit, and determine a preset unit identifier for the target data processing unit. If the unit identifier is consistent with the preset unit identifier, the data processing unit corresponding to the unit identifier is determined as the target data processing unit, and the target unit data of the target data processing unit is obtained.
[0107] Optionally, the variable attribute information is a variable identifier; The data detection node 404 is also used to determine variable data and the variable identifier of the variable data from the target unit data, and to determine a preset variable identifier for the variable data; If the variable identifier is consistent with the preset variable identifier, the variable data corresponding to the variable identifier is clipped to obtain the retained variable data; The retained variable data is determined as the data of the unit to be detected in the target unit data.
[0108] Optionally, the data detection node 404 is further configured to send an initial data acquisition request to the data acquisition node 402, wherein the initial data acquisition request carries multiple unit group identifiers; The data acquisition node 402 is further configured to generate multiple target data acquisition requests based on the multiple unit group identifiers carried in the initial data acquisition request, and acquire the unit data of the multiple data processing units contained in the data processing cluster according to each target data acquisition request.
[0109] Optionally, the data acquisition node 402 is further configured to determine the plurality of unit group identifiers carried in the initial data acquisition request, and generate a plurality of target data acquisition requests based on the plurality of unit group identifiers, wherein the unit group identifier is the identifier of a unit group in the data processing cluster, and each target data acquisition request corresponds to a unit group in the data processing cluster; Based on the multiple target data acquisition requests, determine the unit group corresponding to each unit group identifier from the data processing cluster; A target unit group is determined from multiple unit groups, and unit data of the data processing units contained in the target unit group is obtained, wherein the target unit group is any one of the multiple unit groups.
[0110] Optionally, the data detection node 404 is further configured to receive a detection data acquisition request sent by the client, wherein the detection data acquisition request is sent by the client when the user performs a detection data acquisition operation on the data acquisition page; Determine the detection data identifier carried in the detection data acquisition request, and obtain the detection data corresponding to the detection data identifier from the time series database, wherein the detection data includes the unit data of the data processing unit to be detected and / or the operation status detection result; The detection data is sent to the client so that the client can display the detection data using the data acquisition page.
[0111] Optionally, the data detection node 404 is further configured to generate risk warning information based on the operation status detection result when the operation status detection result is consistent with the preset risk detection result, and send the risk warning information to the risk processing node.
[0112] Optionally, the data detection system further includes a data transmission node; The data sending node is used to acquire short-term task data from the plurality of data processing units, wherein the short-term task data is data generated by the data processing unit executing short-term tasks, and the short-term task is a task whose execution time is less than a preset time threshold. The short-term task data is sent as unit data to the data detection node.
[0113] The data detection system provided in one or more embodiments of this specification, during the operation status detection process, utilizes a data acquisition node to perform a first data filtering on the acquired unit data and sends the target unit data obtained from the data filtering to the data detection node. During the operation status detection process, the data detection node performs a second data filtering on the target unit data, thereby obtaining a smaller number of unit data to be detected, thus avoiding the problem of too much detection data. Then, based on the unit data to be detected, a smaller number of data processing units to be detected are determined from multiple data processing units, and the operation status of these data processing units is detected, thus avoiding the problem of a large number of data processing units to be detected. Based on this, this data detection system enables operation status detection of a large number of data processing units with a large amount of unit data, avoiding problems such as data detection node crashes and memory overflows caused by excessive number of data processing units and excessive unit data volume.
[0114] The above is an illustrative scheme of a data detection system according to this embodiment. It should be noted that the technical solution of this data detection system and the technical solution of the data detection method described above belong to the same concept. For details not described in detail in the technical solution of the data detection system, please refer to the description of the technical solution of the data detection method described above.
[0115] Figure 5 A structural block diagram of a computing device 500 according to one embodiment of this specification is shown. The components of the computing device 500 include, but are not limited to, a memory 510 and a processor 520. The processor 520 is connected to the memory 510 via a bus 530, and a database 550 is used to store data.
[0116] The computing device 500 also includes an access device 540, which enables the computing device 500 to communicate via one or more networks 560. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 540 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.
[0117] In one embodiment of this specification, the above-described components of the computing device 500 and Figure 5 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 5 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0118] The computing device 500 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 500 can also be a mobile or stationary server.
[0119] The processor 520 is used to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the above-described data detection method.
[0120] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computing device embodiments are basically similar to the data detection method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the data detection method embodiments.
[0121] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the data detection method described above.
[0122] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. In particular, the computer-readable storage medium embodiments are relatively simple in description because they are fundamentally similar to the data detection method embodiments; relevant parts can be referred to in the description of the data detection method embodiments.
[0123] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described data detection method.
[0124] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the data detection method described above belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the data detection method described above.
[0125] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0126] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0127] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0128] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0129] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A data detection method, applied to a data detection system, the data detection system comprising a data detection node and a data acquisition node, wherein, The data acquisition node acquires unit data from multiple data processing units within the data processing cluster, filters the unit data of each data processing unit based on the unit attribute information of each data processing unit, obtains the target unit data of the target data processing unit, and sends the target unit data to the data detection node. The data detection node filters the target unit data based on the variable attribute information in the target unit data to obtain the unit data to be detected within the target unit data. Based on the data of the unit to be detected, a data processing unit to be detected is determined from the plurality of data processing units, and based on the unit data of the data processing unit to be detected, the operating status of the data processing unit to be detected is detected to obtain the operating status detection result.
2. The data detection method according to claim 1, wherein the step of filtering the unit data of each data processing unit based on the unit attribute information of each data processing unit to obtain the target unit data of the target data processing unit includes: The data acquisition node determines the unit identifier of each data processing unit from the unit attribute information of each data processing unit, and determines a preset unit identifier for the target data processing unit. If the unit identifier is consistent with the preset unit identifier, the data processing unit corresponding to the unit identifier is determined as the target data processing unit, and the target unit data of the target data processing unit is obtained.
3. The data detection method according to claim 1 or 2, wherein the variable attribute information is a variable identifier; The step of filtering the target unit data based on the variable attribute information in the target unit data to obtain the target unit data to be detected includes: The data detection node determines variable data and variable identifiers of the variable data from the target unit data, and determines a preset variable identifier for the variable data; If the variable identifier is consistent with the preset variable identifier, the variable data corresponding to the variable identifier is clipped to obtain the retained variable data; The retained variable data is determined as the data of the unit to be detected in the target unit data.
4. The data detection method according to claim 1, wherein acquiring the unit data of multiple data processing units included in the data processing cluster comprises: The data detection node sends an initial data acquisition request to the data acquisition node, wherein the initial data acquisition request carries multiple unit group identifiers; The data acquisition node generates multiple target data acquisition requests based on the multiple unit group identifiers carried in the initial data acquisition request, and acquires the unit data of the multiple data processing units contained in the data processing cluster according to each target data acquisition request.
5. The data detection method according to claim 4, wherein the data acquisition node generates multiple target data acquisition requests based on the multiple unit group identifiers carried in the initial data acquisition request, and acquires the unit data of the multiple data processing units included in the data processing cluster according to each target data acquisition request, including: The data acquisition node determines the multiple unit group identifiers carried in the initial data acquisition request, and generates multiple target data acquisition requests based on the multiple unit group identifiers, wherein the unit group identifier is the identifier of the unit group in the data processing cluster, and each target data acquisition request corresponds to each unit group in the data processing cluster; Based on the multiple target data acquisition requests, determine the unit group corresponding to each unit group identifier from the data processing cluster; A target unit group is determined from multiple unit groups, and unit data of the data processing units contained in the target unit group is obtained, wherein the target unit group is any one of the multiple unit groups.
6. The data detection method according to claim 1, after performing operational status detection on the data processing unit to be detected based on the unit data of the data processing unit to be detected and obtaining the operational status detection result, further includes: The data detection node receives a detection data acquisition request sent by the client, wherein the detection data acquisition request is sent by the client when the user performs a detection data acquisition operation on the data acquisition page; Determine the detection data identifier carried in the detection data acquisition request, and obtain the detection data corresponding to the detection data identifier from the time series database, wherein the detection data includes the unit data of the data processing unit to be detected and / or the operation status detection result; The detection data is sent to the client so that the client can display the detection data using the data acquisition page.
7. The data detection method according to claim 1, after performing operational status detection on the data processing unit to be detected based on the unit data of the data processing unit to be detected and obtaining the operational status detection result, further comprising: When the operation status detection result is consistent with the preset risk detection result, the data detection node generates a risk warning message based on the operation status detection result and sends the risk warning message to the risk processing node.
8. The data detection method according to claim 1, wherein the data detection system includes a data sending node, and the method further includes: The data sending node acquires short-term task data from the multiple data processing units, wherein the short-term task data is data generated by the data processing units executing short-term tasks, and the short-term task is a task whose execution time is less than a preset time threshold. The short-term task data is sent as unit data to the data detection node.
9. A data detection method, applied to a data detection system, the data detection system comprising a data detection node and a data acquisition node, wherein, The data acquisition node acquires unit data of multiple Pods contained in the Kubernetes cluster, filters the unit data of each Pod according to the unit attribute information of each Pod, obtains the target unit data of the target Pod, and sends the target unit data to the data detection node. The data detection node filters the target unit data based on the variable attribute information in the target unit data to obtain the unit data to be detected within the target unit data. Based on the unit data to be detected, the Pod to be detected is determined from the plurality of Pods, and the running status of the Pod to be detected is detected based on the unit data of the Pod to be detected, so as to obtain the running status detection result.
10. A data detection system, comprising a data detection node and a data acquisition node, wherein, The data acquisition node is used to acquire unit data from multiple data processing units contained in the data processing cluster, filter the unit data of each data processing unit according to the unit attribute information of each data processing unit, obtain the target unit data of the target data processing unit, and send the target unit data to the data detection node. The data detection node is used to filter the target unit data based on the variable attribute information in the target unit data to obtain the unit data to be detected from the target unit data. Based on the data of the unit to be detected, a data processing unit to be detected is determined from the plurality of data processing units, and based on the unit data of the data processing unit to be detected, the operating status of the data processing unit to be detected is detected to obtain the operating status detection result.
11. A computing device, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 9.
12. A computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 9.
13. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 9.