Cluster information processing method and device, storage medium and program product
By identifying and displaying resource usage and topological distribution information of target hardware units in the cluster, the problem of inefficient cluster monitoring in the existing technology is solved, resource observation and analysis from a cluster perspective is realized, and management efficiency is improved.
Patent Information
- Application Number
- CN202510887959.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing cluster monitoring methods mainly focus on a single node, lacking resource observation and performance analysis capabilities in the cluster dimension, resulting in inefficient cluster management and requiring manual intervention to investigate problems.
By identifying the target hardware unit of the object to be observed in the cluster, obtaining its resource usage information and topological distribution information, and using topological distribution information to display resource usage, realizing observation and analysis from a clustered perspective.
It reduces human intervention, improves cluster information observation efficiency, provides intuitive display and analysis of resource use within the cluster, and improves the accuracy and efficiency of cluster management.
Smart Images

Figure CN120386691A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information processing technologies, and in particular, to a method, device, storage medium, and program product for processing cluster information. Background Art
[0002] Since the artificial intelligence (AI) generative model has extremely high requirements for computing power and the computing power of a single machine cannot meet the needs, the model is generally split into multiple parts and run in a distributed manner. Therefore, it has become a necessity to provide cluster computing power resources with an AI infrastructure cluster composed of multiple machines and multiple cards.
[0003] In order to ensure the stable and maximum supply of computing power in the cluster, in the daily cluster management process, it is more necessary to pay attention to the specific situation of AI applications at the entire cluster level, such as the cluster usage and running conditions of AI jobs. Therefore, the ability to uniformly observe cluster resources and track performance when problems occur is particularly important. However, the existing observation and performance analysis methods mainly focus on individual nodes. The cluster monitoring information mainly provides the hardware information of the nodes with relatively high usage in the cluster (such as the nodes ranked in the top several in terms of usage rate), or only rough resource utilization statistical information. Therefore, for further performance troubleshooting, cluster management and development personnel still need to manually identify the target nodes to be troubleshot through the cluster monitoring information, and then manually switch to the target nodes to view resources, perform performance analysis on specific nodes, or perform other troubleshooting operations such as offline replacement. It can be seen that the current cluster observation method is very inefficient. Summary of the Invention
[0004] The main purpose of the embodiments of this application is to provide a method, device, storage medium, and program product for processing cluster information, which realizes observing and analyzing the hardware units used by the object to be observed in the cluster and the resource usage information of the object to be observed on each hardware unit from a cluster perspective, intuitively presenting the resource usage situation of the object to be observed in the entire cluster, reducing manual intervention, and improving the efficiency of cluster information observation.
[0005] In a first aspect, the embodiments of this application provide a method for processing cluster information, including: determining an object to be observed running in a cluster; identifying at least one target hardware unit used by the object to be observed in the cluster; determining the resource usage information of the object to be observed on the at least one target hardware unit and the topological distribution information of the at least one target hardware unit in the cluster; and displaying the resource usage information of the object to be observed on the target hardware unit according to the topological distribution information.
[0006] In one embodiment, the object to be observed is a specified data processing task; identifying at least one target hardware unit used by the object to be observed in the cluster includes: obtaining a task identifier of the data processing task; and identifying at least one first target node and / or at least one first target processor in the cluster for executing the data processing task according to the task identifier and resource allocation information of the cluster, where the target hardware unit includes the first target node and / or the first target processor.
[0007] In one embodiment, determining resource usage information of the object to be observed on the at least one target hardware unit includes: obtaining resource occupancy information of the data processing task on each of the first target processors; and / or, if the data processing task is processed across hardware units in the cluster, obtaining communication traffic information during the cross-hardware unit processing of the data processing task in the cluster; aggregating the resource occupancy information and / or the communication traffic information according to a preset dimension to obtain the resource usage information corresponding to the data processing task, where the preset dimension includes one or more of a container dimension, a node dimension, and a process dimension.
[0008] In one embodiment, the topology distribution information includes: the topological relationship between a target container corresponding to the data processing task in the cluster, the first target node, a target process, and the first target processor, where the target process is at least one process occupied by the target container on the first target node in the cluster.
[0009] In one embodiment, the object to be observed is a specified session request; the method further includes: injecting a tracking identifier for the session request at an ingress gateway of the session request; setting a data point in an inference code corresponding to the session request according to the tracking identifier, and tracking link information of the session request in the cluster according to the data point; and identifying at least one target hardware unit used by the object to be observed in the cluster includes: determining a second target node and / or a second target processor allocated by the cluster for the session request according to the link information, where the target hardware unit includes the second target node and / or the second target processor.
[0010] In one embodiment, tracking the link information of the session request in the cluster further includes: if the session request has cross-hardware unit propagation in the cluster, tracking the link information of the context cross-hardware unit propagation of the session request according to the tracking identifier.
[0011] In one embodiment, determining the resource usage information of the object to be observed on the at least one target hardware unit includes: determining the latency information during the propagation of the session request between different target hardware units according to the link information, and the resource usage information includes the latency information.
[0012] In one embodiment, it further includes: obtaining the communication connection relationship and / or physical topology relationship between different hardware units in the cluster, and displaying a topology graph of the cluster on the interaction interface, where the topology graph includes the communication connection relationship and / or the physical topology relationship; in response to a query instruction for the topology graph of the cluster on the interaction interface, displaying information of an object indicated by the query instruction in the cluster on the interaction interface.
[0013] In a second aspect, an embodiment of the present application provides a cluster information processing device, including:
[0014] A first determination module, configured to determine an object to be observed running in the cluster;
[0015] An identification module, configured to identify at least one target hardware unit used by the object to be observed in the cluster;
[0016] A second determination module, configured to determine the resource usage information of the object to be observed on the at least one target hardware unit and the topological distribution information of the at least one target hardware unit in the cluster;
[0017] A first display module, configured to display the resource usage information of the object to be observed on the target hardware unit according to the topological distribution information.
[0018] In one embodiment, the object to be observed is a specified data processing task; the identification module is configured to obtain a task identifier of the data processing task; according to the task identifier and the resource allocation information of the cluster, identify at least one first target node and / or at least one first target processor in the cluster for executing the data processing task, and the target hardware unit includes the first target node and / or the first target processor.
[0019] In one embodiment, the second determination module is configured to obtain the resource occupancy information of the data processing task on each first target processor; and / or, if the data processing task is processed across hardware units in the cluster, obtain the communication traffic information during the cross-hardware-unit processing of the data processing task in the cluster; aggregate the resource occupancy information and / or the communication traffic information according to a preset dimension to obtain the resource usage information corresponding to the data processing task, where the preset dimension includes one or more of a container dimension, a node dimension, and a process dimension.
[0020] In one embodiment, the topology distribution information includes: the topology relationship among the target container corresponding to the data processing task in the cluster, the first target node, the target process, and the first target processor, where the target process is at least one process occupied by the target container on the first target node in the cluster.
[0021] In one embodiment, the object to be observed is a specified session request; the apparatus further includes:
[0022] A tracing module, configured to inject a tracing identifier for the session request at the ingress gateway of the session request; set tracing points in the inference code corresponding to the session request according to the tracing identifier, and trace the link information of the session request in the cluster according to the tracing points; and
[0023] An identification module, further configured to determine a second target node and / or a second target processor allocated by the cluster for the session request according to the link information, where the target hardware unit includes the second target node and / or the second target processor.
[0024] In one embodiment, the tracing module is further configured to, if the session request has cross-hardware unit propagation in the cluster, trace the link information of the context cross-hardware unit propagation of the session request according to the tracing identifier.
[0025] In one embodiment, a second determination module is configured to determine the delay information during the propagation of the session request between different target hardware units according to the link information, where the resource usage information includes the delay information.
[0026] In one embodiment, the apparatus further includes:
[0027] A second display module, configured to obtain the communication connection relationship and / or the physical topology relationship between different hardware units in the cluster, and display a topology map of the cluster on an interaction interface, where the topology map includes the communication connection relationship and / or the physical topology relationship;
[0028] A query module, configured to, in response to a query instruction for the topology map of the cluster on the interaction interface, display the information of the object indicated by the query instruction in the cluster on the interaction interface.
[0029] In a third aspect, an embodiment of the present application provides an electronic device, including:
[0030] At least one processor; and
[0031] A memory communicatively connected to the at least one processor;
[0032] Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to cause the electronic device to execute the method described in any of the above aspects.
[0033] In a fourth aspect, an embodiment of the present application provides a cloud device, including:
[0034] At least one processor; and
[0035] A memory communicatively connected to the at least one processor;
[0036] Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to cause the cloud device to execute the method described in any of the above aspects.
[0037] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when a processor executes the computer-executable instructions, the method described in any of the above aspects is implemented.
[0038] In a sixth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in any of the above aspects is implemented.
[0039] In a seventh aspect, an embodiment of the present application provides a cluster information processing system, including a server and a terminal, and the terminal and the server perform data interaction to implement the method described in any of the above aspects.
[0040] The cluster information processing method, device, storage medium, and program product provided by the embodiments of the present application identify at least one target hardware unit used by an object to be observed in a cluster, and accurately locate the specific hardware unit used by the object to be observed. By obtaining the resource usage information of the object to be observed on these target hardware units and the topological distribution information of the target hardware units in the cluster, finally, by means of the topological distribution information of the target hardware units, the resource usage information of the object to be observed on the target hardware units is displayed, realizing the observation and analysis of the target hardware units used by the object to be observed in the cluster and the resource usage information of the object to be observed on each target hardware unit from a cluster perspective, and visualizing and displaying the resource usage information based on the topological distribution information of the hardware units, intuitively presenting the resource usage situation of the object to be observed in the entire cluster, reducing manual intervention, and improving the efficiency of cluster information observation. Description of the Drawings
[0041] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1 Schematic diagram of the structure of an electronic device provided by an embodiment of this application;
[0043] Figure 2 Schematic diagram of the application scenario of a cluster information processing system provided by an embodiment of this application;
[0044] Figure 3 Schematic diagram of the application system framework of a cluster information processing system provided by an embodiment of this application;
[0045] Figure 4 Schematic flow chart of a cluster information processing method provided by an embodiment of this application;
[0046] Figure 5 Schematic diagram of the application scenario of a method for locating hardware units using job correlation analysis provided by an embodiment of this application;
[0047] Figure 6 Schematic diagram of the structure of a cluster information processing device provided by an embodiment of this application;
[0048] Figure 7 Schematic diagram of the structure of a cloud device provided by an embodiment of this application.
[0049] Through the above accompanying drawings, specific embodiments of this application have been shown, and there will be more detailed descriptions later. These drawings and textual descriptions are not intended to limit the scope of the concept of this application in any way, but to illustrate the concept of this application to those skilled in the art by referring to specific embodiments. Detailed implementation manners
[0050] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with this application.
[0051] The term "and / or" in this article is used to describe the association relationship of associated objects, and specifically represents three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0052] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0053] To clearly describe the technical solutions of the embodiments of this application, the following is an interpretation of the terms involved in this application:
[0054] AI: Artificial Intelligence, artificial intelligence.
[0055] AI infrastructure: Refers to the infrastructure and resources that support the development, deployment, and operation of artificial intelligence (AI) technology. It covers multiple aspects such as hardware, software, data, and networks, providing necessary support for the research and development and commercialization of AI applications.
[0056] Cluster: Refers to a computer cluster, which is a computer system that is connected by a group of loosely integrated computer software or hardware and collaborates highly closely to complete computing tasks. In a cluster system, a single computer can be called a node.
[0057] K8s cluster: A Kubernetes cluster consists of a master node and worker nodes, and each component collaborates to achieve the automated management of containerized applications.
[0058] CPU: Central Processing Unit, central processing unit.
[0059] GPU: Graphics Processing Unit, graphics processing unit.
[0060] Profiling tool: A technical means to locate resource bottlenecks and optimize execution efficiency by dynamically collecting and analyzing the performance data of distributed jobs during cluster operation.
[0061] OpenTelemetry: An open-source observability framework used to collect, process, and export telemetry data (such as traces, metrics, and logs) in distributed systems. It provides a unified API (Application Programming Interface) and SDK (Software Development Kit), supports multiple programming languages, supports multiple data formats and export targets, and the main applicable scenario is the full-link tracing of distributed systems.
[0062] NVLink: A high-speed interconnect technology used to achieve efficient data transfer between GPUs, between a GPU and a CPU, or with other peripherals (such as network interface cards).
[0063] RDMA: remote direct memory access, Remote Direct Memory Access.
[0064] GPUDirect RDMA: Enables a direct data path between the GPU's video memory and other devices (such as RDMA network cards) through hardware-level optimization. In a computer cluster, the NVLink technology can be used for communication in a single machine with multiple GPUs scenario, and the GPUDirect RDMA technology can be used to achieve GPU communication between nodes in a distributed scenario with multiple machines.
[0065] Jaeger: An open-source distributed tracing system used to monitor and diagnose request flows in a microservices architecture. It has the functions of tracing the call chain of requests between multiple services and providing a visual interface to help developers analyze performance bottlenecks and errors. Its main application scenarios are performance monitoring in microservices architectures and fault troubleshooting in distributed systems.
[0066] preffto: An open-source data visualization tool for tracing events, mainly used in the scenario of single-node AI job processes for performance analysis of call relationships based on time series.
[0067] SM: Streaming Multiprocessor, Streaming Multiprocessor. The cluster information processing method in the embodiments of this application can be applied to any field that requires observing the information of a distributed cluster.
[0068] For AI generative models, due to their particularly large demand for computing power, the computing power of a single machine cannot meet the requirements. Therefore, the model is generally divided into multiple parts and run in a distributed manner. Thus, an AI infrastructure cluster composed of multiple machines and multiple GPUs to provide clustered computing power resources has become a necessity.
[0069] To ensure the stable and maximized supply of computing power in the cluster, during the daily cluster management process, it is more necessary to pay attention to the specific situation of AI applications at the entire cluster level, such as the cluster usage and running status of AI jobs. Therefore, the ability to uniformly observe cluster resources and trace performance when problems occur is particularly important.
[0070] However, the existing observation and performance analysis methods mainly focus on a single node. For example, monitoring tools for collecting resources on a single node only provide hardware information of nodes with relatively high usage in the cluster (such as the top several nodes in terms of usage rate), or only rough resource utilization statistics. They do not have the ability to centrally analyze AI tasks in the cluster and the relationships between nodes from the cluster dimension. Therefore, for further performance troubleshooting, cluster management and developers still need to manually identify the target nodes to be troubleshot through cluster monitoring information, and then manually switch to the target nodes to view resources, perform performance analysis on specific nodes, or carry out other troubleshooting operations such as offline replacement. It can be seen that the current cluster observation method is very inefficient.
[0071] To solve at least one of the above problems, an embodiment of the present application provides a cluster information processing solution. By identifying at least one target hardware unit used by the object to be observed in the cluster, the specific hardware unit where the resources are used can be accurately located. By obtaining the resource usage information of the object to be observed on these target hardware units and the topological distribution information of the target hardware units in the cluster, finally, the resource usage information of the object to be observed on the target hardware units is displayed through the topological distribution information of the target hardware units, realizing the observation and analysis of the target hardware units used by the object to be observed in the cluster and the resource usage information of the object to be observed on each target hardware unit from a cluster perspective, and visualizing and displaying the resource usage information based on the topological distribution information of the hardware units, intuitively presenting the resource allocation and usage situation of the object to be observed in the entire cluster, reducing human intervention, and improving the efficiency of cluster information observation.
[0072] The following will describe some embodiments of the present application in detail with reference to the accompanying drawings. Without conflict between the embodiments, the embodiments and the features in the embodiments can be combined with each other. In addition, the step timing in the following method embodiments is only an example and is not strictly limited.
[0073] As Figure 1 shown, this embodiment provides an electronic device 1, including: at least one processor 11 and a memory 12. Figure 1 Taking one processor as an example. The processor 11 and the memory 12 are connected through a bus 10. The memory 12 stores instructions executable by the processor 11. When the instructions are executed by the processor 11, the electronic device 1 can execute all or part of the processes of the methods in the following embodiments, so as to realize the observation and analysis of the hardware units used by the object to be observed in the cluster and the resource usage information of each hardware unit from a cluster perspective, intuitively presenting the resource allocation and usage situation of the object to be observed in the entire cluster, reducing human intervention, and improving the efficiency of cluster information observation.
[0074] In one embodiment, the electronic device 1 may be a mobile phone, a tablet computer, a laptop computer, a desktop computer, or a large computing system composed of multiple computers.
[0075] Figure 2 FIG. 200 is a schematic diagram of an application scenario of a cluster information processing system provided by an embodiment of the present application. As Figure 2 shown, the system includes: a server 210 and a terminal 220, where:
[0076] The server 210 may be a data center that provides cluster information processing services, such as a data center of an AI infrastructure cluster. In an actual scenario, a data center may have multiple servers 210, Figure 2 and one server 210 is taken as an example herein.
[0077] The terminal 220 may be an electronic device that interacts with the data center, such as a computer, a mobile phone, a tablet, etc. used when accessing the data center. There may also be multiple terminals 220, Figure 2 and two terminals 220 are taken as an example for illustration herein.
[0078] Information can be transmitted between the terminal 220 and the server 210 through the Internet so that the terminal 220 can access the data on the server 210. The above terminal 220 and / or server 210 can be implemented by the electronic device 1.
[0079] The cluster information processing solution of the embodiment of the present application can be deployed on the server 210, or can be deployed on the terminal 220, or partially deployed on the server 210 and partially deployed on the terminal 220. In an actual scenario, it can be selected based on actual needs, and this embodiment does not make any limitations.
[0080] When the cluster information processing solution is fully or partially deployed on the server 210, a call interface can be opened to the terminal 220 to provide algorithm support for the terminal 220.
[0081] The method provided by the embodiment of the present application can be implemented by the electronic device 1 executing corresponding software code and by interacting with the server for data. Among them, the electronic device 1 can be a local terminal device. When the method runs on the server, the method can be implemented and executed based on a cloud interaction system, where the cloud interaction system includes a server and a client device.
[0082] In a possible implementation manner, the method provided by the embodiment of the present application provides a graphical user interface through a terminal device, where the terminal device can be the local terminal device mentioned above or the client device in the cloud interaction system mentioned above.
[0083] As Figure 3As shown in the figure, it is a framework schematic diagram of an application system 300 for a cluster information processing method provided by an embodiment of the present application. Taking the distributed cluster of the AI generative model as an example, the system framework mainly includes an AI infrastructure, a data collection module, a data storage module, a clustered data analysis module, and a visualization display module, where:
[0084] The AI infrastructure may include a computer cluster that provides computing power resources for AI applications, such as a GPU cluster, a K8s cluster.
[0085] The data collection module can be implemented by a data collection agent. For example, the data collection agent can be deployed on each node in the GPU cluster. The data collection agent can be integrated with GPU monitoring and management tools, CPU / GPU profiling tools, and trace tools. Among them, the GPU monitoring and management tools are used to collect GPU hardware performance metric information (such as GPU temperature, power, fan, health status, configuration policy, etc. information), the CPU / GPU profiling tools are used to collect AI job metadata and job profiling data, and the trace tools are used to collect tracing data (referring to data that constructs the complete execution trace of a request in a distributed system by recording events, call relationships, and context states during program execution).
[0086] The data storage module is used to store the data collected by the data collection module. For example, cloud storage can be used to store the collected time series metrics, profiling, and tracing data.
[0087] The clustered data analysis module is used to obtain relevant data from the data storage module for analysis. Specifically, first, determine the object to be observed in the cluster. Identify at least one target hardware unit used by the object to be observed in the cluster. Determine the resource usage information of the object to be observed on at least one target hardware unit and the topological distribution information of at least one target hardware unit in the cluster. Realize observing and analyzing the hardware units used by the object to be observed in the cluster and the resource usage information of each hardware unit from a clustered perspective.
[0088] The visualization display module: is used to display the resource usage information of the object to be observed in the cluster according to the topological distribution information, intuitively presenting the resource allocation and usage situation of the object to be observed in the entire cluster, reducing human intervention, and improving the cluster information observation efficiency. For example, the topological presentation can be realized by using a custom Web UI (Website User Interface), or the GPU hardware topology visualization method can be used for presentation.
[0089] Please refer to Figure 4, which is a method for processing cluster information according to an embodiment of the present application. This method can be executed by the Figure 1 shown electronic device 1 and can be applied to the Figures 2 - 3 shown cluster information processing application scenario. In particular, by combining the cluster data analysis module and the visualization display module, it is possible to observe and analyze the hardware units used by the object to be observed in the cluster and the resource usage information of the object to be observed on each hardware unit from a cluster perspective, intuitively presenting the resource allocation and usage situation of the object to be observed in the entire cluster, reducing manual intervention, and improving the efficiency of cluster information observation. Taking the terminal 220 as the execution end in this embodiment, the method includes the following steps:
[0090] Step 401: Determine the object to be observed running in the cluster.
[0091] In this step, the object to be observed refers to a business object running in a computer cluster that needs to be observed and analyzed, including but not limited to data processing tasks and session tasks. For example, the object to be observed can be a certain AI job or a certain session request. The object to be observed can be given automatically by the system. For example, a certain type of AI job task or session request is preset as the object to be observed. When the cluster processes these objects to be observed, performance analysis and observation are performed on the hardware units involved in these objects to be observed. The object to be observed can also be specified by the user. For example, the system administrator can specify to observe a certain AI job task or session request by inputting a job task identifier or a session request identifier.
[0092] Step 402: Identify at least one target hardware unit used by the object to be observed in the cluster.
[0093] In this step, the hardware unit refers to a computer hardware unit deployed in the cluster for data processing. The types of hardware units include but are not limited to processors, memories, GPU cards, chips, nodes, etc. After determining the object to be observed, identify the target hardware unit used by the cluster when executing the object to be observed. Here, the target hardware unit may be one. For example, a certain AI job task only uses one node or GPU for execution. The target hardware unit may also be multiple. For example, some complex AI job tasks require multiple GPUs or multiple nodes to cooperate to complete. After determining the target hardware unit, further identify the topological distribution information of these target hardware units in the cluster. This topological distribution information can intuitively represent the connection relationship between the target hardware unit and other hardware in the cluster, facilitating the intuitive presentation of the resource usage information of the target hardware unit.
[0094] In one embodiment, the object to be observed is a specified data processing task. Identifying at least one target hardware unit used by the object to be observed in the cluster in step 402 includes: obtaining a task identifier of the data processing task. According to the task identifier and the resource allocation information of the cluster, identifying at least one first target node and / or at least one first target processor in the cluster for executing the data processing task, and the target hardware unit includes the first target node and / or the first target processor.
[0095] In this embodiment, if the object to be observed is a certain specified data processing task, such as an AI job task specified by the user through the task identifier, at this time, by obtaining the task identifier of the data processing task and based on the association between the task identifier and the cluster resource allocation information, the specific running location of the data processing task in the cluster can be quickly located, and at least one first target node and / or at least one first target processor in the cluster for executing the data processing task can be identified. The target hardware unit includes but is not limited to the first target node and the first target processor, making resource monitoring and management more flexible and accurate. The resource allocation information of the cluster includes but is not limited to the allocation information of resource types such as computing resources, storage resources, and network resources, and may also include resource scheduling policies and management and monitoring information. For example, it may include the hierarchical relationship of resource pools in the cluster, node grouping (such as availability zone division), replica distribution policies, etc. Therefore, by combining the task identifier of the data processing task with the resource allocation information, it is possible to accurately identify which hardware units in the cluster execute the data processing task.
[0096] Optionally, the job association analysis method can be used to determine the target containers in the cluster for executing the data processing task based on the task identifier and the resource allocation information, and at least one target process occupied by each target container on the first target node in the cluster and the first target processor used by each target process. Through this multi-level identification process, a comprehensive understanding of the task execution environment can be ensured. In order to obtain detailed resource usage information about task execution, so as to provide data support for optimizing the resource allocation strategy and reducing resource conflicts.
[0097] As Figure 5 shown, it is a schematic diagram of an application scenario for locating hardware units using the job association analysis method provided by the embodiment of the present application. Taking the K8s cluster as an example, assuming that the object to be observed is an AI job task, and the task identifier is represented in the form of JobID (Job Identifier, a string used to uniquely identify a job or task), first obtain the JobID of the AI job task, and then determine the target hardware unit through the following process:
[0098] JobID-->Container((K8s container))
[0099] Container --> PID1(Process 10)
[0100] Container --> PID2(Process 20)
[0101] PID1 --> GPU1[[GPU0]]
[0102] PID1 --> GPU2[[GPU1]]
[0103] PID2 --> GPU3[[GPU2]]
[0104] As described above, the ID (Identity document) of the target container Container used by the AI job in the cluster is located through the JobID. Through the target container, the process IDs (such as Process 10 and Process 20) on the specific first target node used by the AI job in the cluster are located. Further, the target processors used by Process 10 and Process 20 can be located. For example, the GPU cards used by Process 10 (PID1) are GPU0 and GPU1, and the GPU card used by Process 20 (PID2) is GPU2. Among them, GPU1[[GPU0]] means that the configuration of GPU1 inherits the parameters of GPU0, GPU2[[GPU1]] means that the configuration of GPU2 inherits the parameters of GPU1, and GPU3[[GPU2]] means that the configuration of GPU3 inherits the parameters of GPU2.
[0105] Step 403: Determine the resource usage information of the object to be observed on at least one target hardware unit and the topological distribution information of at least one target hardware unit in the cluster.
[0106] In this step, by obtaining the resource usage information of the object to be observed on these target hardware units and the topological distribution information of the target hardware units in the cluster, the observation and analysis of the hardware units used by the object to be observed in the cluster and the resource usage information of each hardware unit are realized from a cluster perspective, and the resource allocation and usage situation of the object to be observed in the entire cluster are intuitively presented based on the topological distribution of the hardware units.
[0107] In one embodiment, determining the resource usage information of the object to be observed on at least one target hardware unit in step 403 includes: obtaining the resource occupancy information of the data processing task on each first target processor. And / or, if the data processing task is processed across hardware units in the cluster, obtaining the communication traffic information during the process of the data processing task being processed across hardware units in the cluster. Aggregating the resource occupancy information and / or the communication traffic information according to a preset dimension to obtain the resource usage information corresponding to the data processing task, where the preset dimension includes one or more of the container dimension, the node dimension, and the process dimension.
[0108] In this embodiment, if the object to be observed is a specified data processing task, the resource occupancy information includes but is not limited to GPU utilization / video memory utilization. By comprehensively determining the resource usage information of the data processing task on the target hardware unit, the refinement and resource optimization capabilities of cluster management are significantly improved. First, by obtaining the resource occupancy information of the data processing task on each first target processor, the resource consumption of the data processing task at the processor level can be detailedly recorded. If the data processing task is processed across hardware units in the cluster, such as the case of cross-node processing, the communication traffic information during its cross-hardware unit processing is further obtained. Here, the communication traffic information can characterize the communication resource usage between different hardware units during the cross-hardware unit processing of the data processing task, thereby providing a complete view of the interaction of the data processing task between different hardware units. By aggregating the resource occupancy information and / or the communication traffic information according to a preset dimension, a multi-dimensional resource usage information report is generated, where the preset dimension includes one or more of the container dimension, the node dimension, and the process dimension. This aggregation method not only makes the resource usage situation more intuitive and easy to analyze, but also provides a basis for system administrators to optimize resource configuration and improve task execution efficiency. It realizes detailed resource monitoring and analysis from a cluster perspective, helps improve the overall performance of the cluster, reduces resource waste, and enhances the stability and reliability of the system.
[0109] Taking the above K8s cluster scenario as an example, after locating the ID of the target container Container used by the AI job in the cluster through the JobID, and then locating the process IDs (such as process 10 and process 20) of the AI job on the specific first target node in the cluster through the target container, and further locating that process 10 and process 20 use GPU cards, according to the configuration relationship between the container, process and node, key metrics are aggregated to obtain the resource usage information of the AI job task. For example, the resource usage information of the AI job task can include the GPU utilization / VRAM utilization at the card level and the GPU utilization / VRAM utilization at the process dimension of the AI job task. The single physical GPU card can be used as the basic monitoring unit to separately count the resource usage metrics of each card to obtain the GPU utilization / VRAM utilization at the card level of the AI job task. Or the single application process can be used as the monitoring unit to separately count the GPU resources consumed by each process to obtain the GPU utilization / VRAM utilization at the process dimension of the AI job task. In the case of cross-nodes, the resource usage information of the AI job task can include the inter-node communication traffic matrix (such as the communication traffic and data throughput statistics through the AllReduce (global reduction) operation), and the inter-card communication throughput (such as the communication throughput of NVLink). In one embodiment, the topology distribution information includes: the topology relationship between the target container, the first target node, the target process and the first target processor corresponding to the data processing task in the cluster.
[0110] In this embodiment, the target process is at least one process occupied by the target container on the first target node in the cluster. If the object to be observed is a certain specified data processing task, the topology distribution information details the topology relationship between the target container, the first target node, the target process and the first target processor used by the data processing task in the cluster. It not only reveals the association and dependency relationship between the data processing task and each computing unit in the cluster (here the computing unit includes but is not limited to the hardware unit and the virtualized computing unit, and the virtualized computing unit can be, for example, a container, a process, etc. computing unit), but also provides a clear picture of the task execution path. It is convenient for the system administrator to identify potential performance bottlenecks and resource conflict points, so as to formulate more accurate resource allocation and scheduling strategies. In addition, this cluster topology perspective also provides an important basis for fault diagnosis and performance tuning, helping to quickly locate the root cause of the problem and take effective measures. By introducing the topology distribution information of the target hardware unit in the cluster, a panoramic insight into the running environment of the data processing task in the cluster is achieved, greatly enhancing the visualization management and optimization capabilities of the system.
[0111] In one embodiment, the object to be observed is a specified session request. The method further includes: injecting a tracking identifier for the session request at the ingress gateway of the session request. Setting breakpoints in the inference code corresponding to the session request according to the tracking identifier, and tracking the link information of the session request in the cluster according to the breakpoints. And the at least one target hardware unit used by the object to be observed in the cluster identified in step 402 may include: determining, according to the link information, a second target node and / or a second target processor allocated by the cluster for the session request, and the target hardware unit includes the second target node and / or the second target processor.
[0112] In this embodiment, the object to be observed may be a specified session request. For example, if a user wants to track and observe a certain session request from the perspective of the cluster, the specified session request can be used as the object to be observed. The system can inject a tracking identifier at the ingress gateway of the session request, so that the transfer path of each session request in the cluster can be uniquely identified and tracked. By setting breakpoints in the inference code corresponding to the session request, the system can capture and record the link information of the request in the cluster in real time. Ensure comprehensive monitoring of the execution path of the session request. Furthermore, based on the obtained link information, the second target node and / or the second target processor allocated by the cluster for the session request can be accurately identified, so as to clarify the target hardware unit used by the session request in the cluster. This refined identification and tracking mechanism not only improves the visualization of session request processing, provides important support for fault diagnosis and performance optimization, and helps quickly locate and solve potential problems. It also provides a basis for system administrators to optimize resource allocation and improve processing efficiency.
[0113] In one embodiment, tracking the link information of the session request in the cluster further includes: if the session request has cross-hardware-unit propagation in the cluster, tracking the link information of the context cross-hardware-unit propagation of the session request according to the tracking identifier.
[0114] In this embodiment, when the session request propagates across multiple hardware units in the cluster, such as cross-node propagation, the system uses the tracking identifier to track its context propagation path in detail, ensuring that each propagation step in the propagation path is accurately recorded. It can not only reveal the interaction relationship between the session request among different hardware units, but also provide a clear view of the overall execution of the session request from the perspective of the cluster. Furthermore, it helps the system administrator to better understand the transfer process of the session request across hardware units in the cluster, identify potential performance bottlenecks and resource conflicts, and improve the cluster observation efficiency.
[0115] In one embodiment, determining the resource usage information of the object to be observed on at least one target hardware unit in step 403 includes: determining the latency information during the propagation of the session request between different target hardware units according to the link information, and the resource usage information includes the latency information.
[0116] In this embodiment, based on the link information, the propagation path of the session request among various hardware units can be accurately identified, and the delay information of each propagation step is recorded in detail. As an important part of the resource usage information, these delay information provide a quantitative evaluation of the request processing efficiency. This delay information not only provides key data from a clustered perspective for fault diagnosis and performance tuning to help quickly locate and solve potential problems. Moreover, by carefully monitoring the delay information of the propagation steps, it can also help enhance the performance analysis ability of the cluster system, improving the overall resource utilization efficiency and service quality.
[0117] Optionally, the trace analysis process of the session request can be as follows:
[0118] 1. Inject a TraceID (i.e., trace identifier) at the ingress gateway of the session request.
[0119] 2. Instrument the inference code of the session request through the OpenTelemetry SDK to trace the link information of the session request.
[0120] 3. If the session request propagates across nodes, trace its context.
[0121] 4. Centrally process the trace data and finally parse it into a renderable data format such as jaeger or preffto for subsequent display. Taking an HTTP request as an example, the resource usage information finally traced for this HTTP request can be as follows:
[0122]
[0123] Among them, the link information of [HTTP request] indicates that [HTTP request] propagates from GPU0 of Node1 to GPU1 of Node2 through RDMA, and the delay information is 2ms. In addition, [HTTP request] is also propagated to GPU0 of Node2, and the delay information is 3ms.
[0124] Step 404: Display the resource usage information of the object to be observed on the target hardware unit according to the topology distribution information.
[0125] In this step, the topology distribution information characterizes the overall distribution status of the target hardware units used by the object to be observed in the cluster. The resource usage information of the object to be observed in the cluster is displayed through the topology distribution information of the target hardware units. The topology distribution information can be displayed in the form of an image. For example, a topology distribution map of the target hardware units in the cluster is drawn according to the topology distribution information, and the topology distribution map of the target hardware units in the cluster is displayed on the interaction interface. Based on the topology distribution map, the resource usage information of each target hardware unit is displayed. It realizes observing and analyzing the hardware units used by the object to be observed in the cluster and the resource usage information of each hardware unit from a cluster perspective, and visually presenting the resource allocation and usage situation of the object to be observed in the entire cluster by visualizing the topology distribution of the hardware units, reducing human intervention and improving the efficiency of cluster information observation.
[0126] In one embodiment, the method further includes: obtaining the communication connection relationship and / or physical topology relationship between different hardware units in the cluster, and displaying the topology map of the cluster on the interaction interface, where the topology map includes the communication connection relationship and / or physical topology relationship. In response to a query instruction for the topology map of the cluster on the interaction interface, the information of the object indicated by the query instruction in the cluster is displayed on the interaction interface.
[0127] In this embodiment, taking the K8s cluster as an example, the communication connection relationship and physical topology relationship between different nodes in the cluster can be obtained through the K8s API, or the NCCL (NVIDIA Collective Communications Library) communication library can be used to detect the communication connection relationship and / or physical topology relationship between GPUs through NVLink / InfiniBand (Infinite Bandwidth). For example, through the encapsulated function of NCCL communication monitoring, when the ncclSend function of NCCL (the core function for cross-GPU or cross-node data transmission in NCCL) is called to send data, custom communication monitoring logic (such as recording the communication start time, data volume, target GPU, etc.) is inserted, and then the physical topology relationship between GPUs is obtained. Then, the topology relationship of each unit in the cluster is displayed in the form of a topology map on the interaction interface, realizing the observability of multi-GPU communication from a cluster perspective.
[0128] When the user issues a query instruction for the topology map of the cluster on the interaction interface, the object indicated by the query instruction is determined as the object to be observed. Furthermore, the relevant information of the object to be observed in the cluster can be displayed on the interaction interface, such as the topology distribution information and resource usage information of the target hardware units used, improving the interaction performance of the system.
[0129] Optionally, in the interactive interface, the hardware units in the cluster can be displayed differently according to their resource usage. For example, a preset coloring strategy can be used for differential display. The coloring strategy can be as follows:
[0130] For a GPU card with a GPU utilization rate greater than 90% or a video memory usage rate greater than 95%, the icon of the GPU card will be marked red in the cluster topology displayed in the interactive interface. For a GPU card with a GPU utilization rate greater than 70%, the icon of the GPU card will be marked orange in the cluster topology displayed in the interactive interface. For other GPU cards, the GPU card will be marked green in the cluster topology displayed in the interactive interface. By adopting the differential display method, the resource usage of different hardware units in the cluster can be presented more intuitively.
[0131] Optionally, the topology diagrams of the hardware units in the cluster displayed in the interactive interface can be hierarchically displayed according to their physical topology relationships. For example, the topology diagram LOD (Level of Detail) grading can be as follows:
[0132] Cluster level: The node is a cube.
[0133] Node level: Display the GPU board layout.
[0134] Chip level: Display the SM unit.
[0135] Optionally, filtering and display of the object to be observed can be implemented in the interactive interface. For example, when the user selects a certain AI job task, the target nodes used by the selected AI job task in the cluster topology will be highlighted, and the connections between the target nodes can also be highlighted, which is convenient for differential display with other content and improves the intuitiveness of information display.
[0136] Optionally, the topology distribution information and / or resource usage information of the hardware units in the cluster can be displayed by floating a detailed information card. For example, information such as the video memory at the card level, the GPU utilization rate at the card level, and the details of the AI jobs running on the card can be displayed by floating a detailed information card.
[0137] Optionally, the user can query in the interactive interface in terms of session request dimensions. For example, when the user enters a session identifier in the interactive interface to query a specified session request, a visual link flame graph of the session request can be obtained through jaeger or preffto. The flame graph is a tool for visualizing performance bottleneck analysis, which displays the resource consumption (such as CPU time, memory allocation, etc.) of the function call stack through an intuitive hierarchical graph, helping developers quickly locate the performance hotspots of the system.
[0138] The solution of the embodiment of the present application realizes presenting a cluster in the form of a GPU topology diagram. In the topology diagram, information such as the communication throughput details between cards, latency information, card-level utilization, video memory size, and details of running AI jobs can be intuitively displayed. It supports cluster-based tracking and analysis of the distribution details and resource usage details of a specified AI application in the cluster, and supports analyzing the running status of a session request in the link of each node in the cluster and the overall latency. This can help developers and operation and maintenance personnel observe and analyze the running status of AI jobs from a cluster perspective, timely discover and solve performance bottlenecks, thereby improving the overall efficiency and reliability of cluster computing power utilization.
[0139] Please refer to Figure 6 , which is a cluster information processing device 600 according to an embodiment of the present application. This device can be applied to an electronic device 1 and can be applied to Figures 2 - 3 the cluster information processing application scenario shown in, and in particular, by combining a cluster data analysis module and a visualization display module, it can realize observing and analyzing the hardware units used by an object to be observed in the cluster and the resource usage information of each hardware unit from a cluster perspective, intuitively presenting the resource allocation and usage situation of the object to be observed in the entire cluster, reducing human intervention, and improving the efficiency of cluster information observation. The device includes: a first determination module 601, an identification module 602, a second determination module 603, and a first display module 604. The functional principles of each module are as follows:
[0140] The first determination module 601 is used to determine an object to be observed running in the cluster.
[0141] The identification module 602 is used to identify at least one target hardware unit used by the object to be observed in the cluster.
[0142] The second determination module 603 is used to determine the resource usage information of the object to be observed on at least one target hardware unit and the topological distribution information of at least one target hardware unit in the cluster.
[0143] The first display module 604 is used to display the resource usage information of the object to be observed on the target hardware unit according to the topological distribution information.
[0144] In one embodiment, the object to be observed is a specified data processing task. The identification module 602 is used to obtain the task identifier of the data processing task. According to the task identifier and the resource allocation information of the cluster, at least one first target node and / or at least one first target processor used to execute the data processing task in the cluster are identified, and the target hardware unit includes the first target node and / or the first target processor.
[0145] In one embodiment, the second determination module 603 is configured to obtain the resource occupancy information of the data processing task on each first target processor. And / or, if the data processing task is processed across hardware units in the cluster, obtain the communication traffic information during the process of the data processing task being processed across hardware units in the cluster. Aggregate the resource occupancy information and / or the communication traffic information according to a preset dimension to obtain the resource usage information corresponding to the data processing task, where the preset dimension includes one or more of a container dimension, a node dimension, and a process dimension.
[0146] In one embodiment, the topology distribution information includes: the topology relationship between the target container, the first target node, the target process, and the target processor corresponding to the data processing task in the cluster, where the target process is at least one process occupied by the target container on the first target node in the cluster.
[0147] In one embodiment, the object to be observed is a specified session request. The apparatus further includes:
[0148] A tracing module, configured to inject a tracing identifier for the session request at the ingress gateway of the session request. Set breakpoints in the inference code corresponding to the session request according to the tracing identifier, and trace the link information of the session request in the cluster according to the breakpoints. And
[0149] The identification module 602 is further configured to determine the second target node and / or the second target processor allocated by the cluster for the session request according to the link information, and the target hardware unit includes the second target node and / or the second target processor.
[0150] In one embodiment, the tracing module is further configured to, if the session request propagates across hardware units in the cluster, trace the link information of the context of the session request propagating across hardware units according to the tracing identifier.
[0151] In one embodiment, the second determination module 603 is configured to determine the latency information during the propagation of the session request between different target hardware units according to the link information, and the resource usage information includes the latency information.
[0152] In one embodiment, the apparatus further includes:
[0153] A second display module, configured to obtain the communication connection relationship and / or the physical topology relationship between different hardware units in the cluster, and display a topology diagram of the cluster on the interaction interface, where the topology diagram includes the communication connection relationship and / or the physical topology relationship.
[0154] A query module, configured to, in response to a query instruction for the topology diagram of the cluster on the interaction interface, display the information of the object indicated by the query instruction in the cluster on the interaction interface.
[0155] For a detailed description of the above cluster information processing device 600, please refer to the description of the relevant method steps in the above embodiments. Their implementation principles and technical effects are similar, and will not be elaborated here in this embodiment.
[0156] Figure 7 FIG. 4 is a schematic structural diagram of a cloud device 70 provided by an exemplary embodiment of the present application. The cloud device 70 can be used to run the method provided in any of the above embodiments. As Figure 7 shown, the cloud device 70 may include: a memory 704 and at least one processor 705, Figure 7 Taking one processor as an example.
[0157] The memory 704 is used to store computer programs and can be configured to store various other data to support operations on the cloud device 70. The memory 704 may be an Object Storage Service (OSS).
[0158] The memory 704 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0159] The processor 705 is coupled to the memory 704 and is used to execute the computer program in the memory 704 to implement the solution provided in any of the above method embodiments. The specific functions and achievable technical effects will not be elaborated here.
[0160] Further, as Figure 7 shown, the cloud device further includes: a firewall 701, a load balancer 702, a communication component 706, a power supply component 703, and other components. Figure 7 Only some components are schematically shown in FIG. 4, which does not mean that the cloud device only includes Figure 7 the components shown in FIG. 4.
[0161] In one embodiment, the above Figure 7The communication component 706 therein is configured to facilitate communication, either wired or wirelessly, between the device where the communication component 706 is located and other devices. The device where the communication component 706 is located can access a communication standard-based wireless network, such as WiFi, 2G, 3G, 4G, LTE (Long Term Evolution), 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component 706 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 706 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology and other technologies.
[0162] In one embodiment, the above-mentioned Figure 7 power supply component 703 supplies power to various components of the device where the power supply component 703 is located. The power supply component 703 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located.
[0163] The embodiments of the present application also provide a computer-readable storage medium storing computer-executable instructions, and when the processor executes the computer-executable instructions, the methods of any of the foregoing embodiments are implemented.
[0164] The embodiments of the present application also provide a computer program product including a computer program, and when the computer program is executed by the processor, the methods of any of the foregoing embodiments are implemented.
[0165] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0166] The integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above-mentioned software functional modules are stored in a storage medium and include several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods of various embodiments of the present application.
[0167] It should be understood that the above-mentioned processor may be a central processing unit (Central Processing Unit, abbreviated as CPU), and may also be other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as DSP), application specific integrated circuits (Application Specific Integrated Circuit, abbreviated as ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor. The memory may include high-speed RAM (Random Access Memory) memory, and may also include non-volatile storage NVM (Nonvolatile memory, abbreviated as NVM), such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.
[0168] The above-mentioned storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random-Access Memory, abbreviated as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable read only memory, abbreviated as EEPROM), erasable programmable read-only memory (Erasable Programmable Read-Only Memory, abbreviated as EPROM), programmable read-only memory (Programmable read-only memory, abbreviated as PROM), read-only memory (Read-OnlyMemory, abbreviated as ROM), magnetic memory, flash memory, magnetic disk or optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0169] An exemplary storage medium is coupled to a processor, enabling the processor to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an Application Specific Integrated Circuit (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a master control device.
[0170] It should be noted that in this document, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article of manufacture or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article of manufacture or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the presence of additional identical elements in the process, method, article of manufacture or device including such element.
[0171] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.
[0172] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that makes a contribution to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods of the various embodiments of the present application.
[0173] In the technical solution of the present application, the processing of collection, storage, use, processing, transmission, provision, and disclosure of user data and other information complies with the provisions of relevant laws and regulations and does not violate public order and good customs.
[0174] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A method for processing cluster information, characterized in that, Including: Determine the object to be observed running in the cluster; Identify at least one target hardware unit used by the object to be observed in the cluster; Determine the resource usage information of the object to be observed on the at least one target hardware unit and the topological distribution information of the at least one target hardware unit in the cluster; Display the resource usage information of the object to be observed on the target hardware unit according to the topological distribution information.
2. The method according to claim 1, wherein The object to be observed is a specified data processing task; identifying at least one target hardware unit used by the object to be observed in the cluster includes: Obtain the task identifier of the data processing task; According to the task identifier and the resource allocation information of the cluster, identify at least one first target node and / or at least one first target processor in the cluster for executing the data processing task, and the target hardware unit includes the first target node and / or the first target processor.
3. The method according to claim 2, characterized in that, Determining the resource usage information of the object to be observed on the at least one target hardware unit includes: Obtain the resource occupancy information of the data processing task on each of the first target processors; and / or, if the data processing task is processed across hardware units in the cluster, obtain the communication traffic information during the cross-hardware unit processing of the data processing task in the cluster; Aggregate the resource occupancy information and / or the communication traffic information according to a preset dimension to obtain the resource usage information corresponding to the data processing task, where the preset dimension includes one or more of the container dimension, the node dimension, and the process dimension.
4. The method according to claim 2, wherein The topological distribution information includes: the topological relationship between the target container, the first target node, the target process, and the first target processor corresponding to the data processing task in the cluster, where the target process is at least one process occupied by the target container on the first target node in the cluster.
5. The method according to claim 1, characterized in that, The object to be observed is a specified session request; the method further includes: Inject a trace identifier for the session request at the ingress gateway of the session request; Set breakpoints in the inference code corresponding to the session request according to the trace identifier, and trace the link information of the session request in the cluster according to the breakpoints; and Identifying at least one target hardware unit used by the object to be observed in the cluster includes: Determine the second target node and / or the second target processor allocated by the cluster for the session request according to the link information, and the target hardware unit includes the second target node and / or the second target processor.
6. The method according to claim 5, wherein Tracking the link information of the session request in the cluster further includes: If the session request has cross-hardware unit propagation in the cluster, track the link information of the context cross-hardware unit propagation of the session request according to the trace identifier.
7. The method according to claim 5, characterized in that Determining the resource usage information of the object to be observed on the at least one target hardware unit includes: Determine the latency information during the propagation of the session request among different target hardware units according to the link information, where the resource usage information includes the latency information.
8. The method according to any one of claims 1-7, characterized in that Further comprising: Obtain the communication connection relationship and / or physical topology relationship among different hardware units in the cluster, and display the topology diagram of the cluster on the interaction interface, where the topology diagram includes the communication connection relationship and / or the physical topology relationship; In response to a query instruction for the topology diagram of the cluster on the interaction interface, display the information of the object indicated by the query instruction in the cluster on the interaction interface.
9. An electronic device, characterized in that, Comprising: At least one processor; And A memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to cause the electronic device to execute the method according to any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the processor executes the computer-executable instructions, the method according to any one of claims 1-8 is implemented.
11. A computer program product, characterized in that, Including a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1-8 is implemented.
12. A cluster information processing system, characterized in that, Comprising: A server and a terminal, and the terminal and the server perform data interaction to implement the method according to any one of claims 1-8.
Citation Information
Patent Citations
Micro-service performance monitoring and anomaly diagnosis method
CN112181759A
Resource use data display method and device
CN118550649A
Topological graph generation method and computing device
CN118890282A