Method and device for monitoring computing power routing state by computing power use high precision of intelligent computing cloud platform

By collecting and binding multi-dimensional indicator data with tags in real time in the intelligent computing cloud platform, the problem of insufficient accuracy in computing power routing status monitoring has been solved, and high-precision computing power routing status monitoring has been achieved.

CN122633370APending Publication Date: 2026-08-25DATACANVAS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610517560.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-17
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

In existing technologies, the monitoring accuracy of computing power routing status is insufficient, and it is impossible to accurately capture the real-time operating status of computing power routing, resulting in the monitoring accuracy remaining at a low level for a long time.

Method used

Through the intelligent computing cloud platform, the built-in indicator exposure service of the computing power gateway collects multi-dimensional indicator data in real time and binds corresponding tag combinations to the indicator data of different dimensions. The external data collection service pulls and stores the data in time sequence according to a preset frequency, supporting single-dimensional and multi-dimensional cross-queries.

Benefits of technology

It enables comprehensive and fine-grained monitoring of the entire computing power routing and forwarding process, improves the flexibility of monitoring data usage, and enhances the monitoring accuracy of computing power routing status.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122633370A_ABST
    Figure CN122633370A_ABST
Patent Text Reader

Abstract

The application provides a method and device for monitoring the state of computing power routing by intelligent computing cloud platform through high-precision computing power use, and relates to the technical fields of intelligent computing center, intelligent computing center, computing power infrastructure and intelligent computing cloud, which comprises the following steps: S1, collecting multi-dimensional index data generated by computing power gateway in the process of computing power routing forwarding computing power running task through the index exposure service built in the computing power gateway, including computing power routing class, upstream service class and gateway state class index data; S2, exposing the multi-dimensional index data and binding the corresponding dimension label combination for different dimensions of index data when exposing; S3, pulling the index data according to the preset frequency through the external data collection service, and storing the index data in time sequence according to the label combination; S4, processing the time-series stored index data and showing the visual result when receiving the query statement based on the label combination, which realizes the accurate monitoring of the state of computing power routing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of intelligent computing centers, smart computing centers, computing infrastructure, and smart cloud technologies, specifically to a method and apparatus for high-precision monitoring of computing power routing status through computing power usage in an intelligent computing cloud platform. Background Technology

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "smart computing centers" have emerged.

[0003] An "intelligent computing center" refers to a facility that provides the necessary computing power, data, and algorithms for artificial intelligence applications (such as the development, training, and inference of deep learning models) by utilizing large-scale heterogeneous computing resources, including general-purpose and intelligent computing power. Intelligent computing centers encompass facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enablement.

[0004] "Intelligent computing center" includes, but is not limited to, "intelligent computing center".

[0005] "Intelligent computing center" or artificial intelligence computing center is a type of computing infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications, based on artificial intelligence theory and adopting artificial intelligence computing architecture.

[0006] "Computing power" is the core of "intelligent computing center" and "smart computing center". It is the ability of computer equipment or computing / data center to process parameters. It is the ability of computer hardware and software to work together to execute a certain computing requirement. It is the computing power to achieve the target result output by processing parameter data. It is a new type of productivity that integrates parameter computing power, network carrying capacity and data storage capacity. It mainly provides services to society through computing power infrastructure.

[0007] In the process of executing computing power tasks in intelligent computing centers, computing power routing, as the core hub connecting computing power resources and task requirements, directly determines the efficiency of computing power scheduling and the quality of task execution by accurately sensing its operational status. Therefore, high-precision monitoring of computing power routing status is crucial. However, the difficulty of monitoring computing power routing is far greater than that of traditional Internet routing monitoring. Traditional Internet routing monitoring only focuses on network layer connectivity, resulting in relatively coarse monitoring granularity. Computing power routing, on the other hand, faces a multi-dimensional, highly real-time, and highly dynamic monitoring scenario involving the convergence of computing and the network. It needs to simultaneously cover the entire link of network, computing power, and tasks, and must cope with millisecond-level traffic surges and the heterogeneous complexity of tens of millions of computing and network elements, exponentially increasing the monitoring difficulty. Existing technologies have not been adapted and optimized for the core characteristics of computing power routing. Existing monitoring solutions suffer from problems such as coarse monitoring granularity, insufficient real-time performance, and single data dimensions, failing to accurately capture the real-time operational status of computing power routing, resulting in consistently low monitoring accuracy.

[0008] It is evident that since the emergence of intelligent computing centers, the monitoring accuracy of computing power routing status has been insufficient, and this problem has always been a pressing issue that needs to be addressed in this field. Summary of the Invention

[0009] This invention provides a method and apparatus for high-precision monitoring of computing power routing status through computing power usage in an intelligent computing cloud platform, thereby solving the problem of insufficient monitoring accuracy of computing power routing status in the prior art.

[0010] To solve the above problems, the present invention is implemented as follows: In a first aspect, the present invention provides a method for a smart computing cloud platform to monitor the routing status of computing power with high precision based on the use of computing power, comprising: Step S1: Through the built-in indicator exposure service of the computing power gateway, collect multi-dimensional indicator data generated by the computing power gateway in real time during the process of computing power routing and forwarding computing power running tasks. The multi-dimensional indicator data includes computing power routing indicator data, upstream service indicator data, and gateway status indicator data. Step S2: Expose multi-dimensional indicator data to the outside world, and bind corresponding dimension tag combinations to the indicator data of different dimensions when exposing them. The tag combinations corresponding to the computing power routing indicator data and the upstream service indicator data shall include at least the computing power usage tag. Step S3: Pull indicator data through an external data acquisition service at a preset frequency, and store the indicator data in a time-series format according to the tag combination; Step S4: When a query statement based on tag combination is received, the time-series stored indicator data is processed based on the query statement and the visualization results are displayed. The query statement can be a single-dimensional query statement or a multi-dimensional cross-query statement.

[0011] In one embodiment, step S1 includes: Step S1.1: Pre-configure indicator monitoring rules for the computing power gateway; Step S1.2: Through the built-in indicator exposure service of the computing power gateway, collect multi-dimensional indicator data generated by the computing power gateway in real time during the process of routing and forwarding computing power running tasks according to the indicator monitoring rules.

[0012] In one embodiment, step S2 includes: Step S2.1: For each computing power routing-related metric data and each upstream service-related metric data generated by a single computing power routing and forwarding of a single computing power running task, uniformly bind the same set of computing power-related tags corresponding to this forwarding; and, Step S2.2: Bind gateway-related tag combinations to each gateway status indicator data.

[0013] In one embodiment, step S2.1 includes: Step S2.1.1: Preset label fields for computing power routing; Step S2.1.2: Based on the routing configuration information of computing power routing, determine the fixed label values ​​corresponding to each of the label fields; Step S2.1.3: Based on the actual forwarding status of the computing power running task in this computing power routing, determine the dynamic label values ​​corresponding to each of the remaining label fields; Step S2.1.4: Fill the fixed tag value and dynamic tag value into the corresponding tag field to form a computing power related tag combination, and bind the computing power related tag combination to each computing power routing index data and each upstream service index data generated in this forwarding.

[0014] In one embodiment, the computing power-related tag combination includes at least one of the following types of tags: Routing identification related tags; Upstream service related tags; Task attribute related tags; Environmental isolation related labels; Resource specification related tags.

[0015] In one embodiment, when the computing power route corresponds to multiple candidate upstream computing power nodes, the upstream service-related tags are dynamic tags. Step S2.1.3 includes: Step S2.1.3.1: Determine the target upstream computing power node to which the computing power running task is forwarded by the computing power route. The target upstream computing power node is selected by the computing power route from the candidate upstream computing power nodes. Step S2.1.3.2: Dynamically fill in the upstream service-related tags based on the node information of the target upstream computing power node.

[0016] In one embodiment, the metrics listening port used by the metrics exposure service is independent of the task traffic port used by the computing power routing service to process computing power running tasks.

[0017] Secondly, the present invention also provides a device for a smart computing cloud platform to monitor the computing power routing status with high precision based on the use of computing power, comprising: The data collection module is used to collect multi-dimensional indicator data generated by the computing power gateway in real time during the process of routing and forwarding computing power running tasks through the built-in indicator exposure service of the computing power gateway. The multi-dimensional indicator data includes computing power routing indicator data, upstream service indicator data, and gateway status indicator data. The exposure module is used to expose multi-dimensional indicator data to the outside world, and bind corresponding dimension tag combinations to the indicator data of different dimensions during exposure. The tag combinations corresponding to the computing power routing indicator data and the upstream service indicator data shall include at least the computing power usage tag. The pull module is used to pull indicator data from an external data acquisition service at a preset frequency and store the indicator data in a time sequence according to the tag combination. The query module is used to process the time-series stored indicator data and display the visualization results when a query statement based on tag combination is received. The query statement can be a single-dimensional query statement or a multi-dimensional cross-query statement.

[0018] In one embodiment, the acquisition module is further configured to: pre-configure indicator monitoring rules for the computing power gateway; and, through the indicator exposure service built into the computing power gateway, collect multi-dimensional indicator data generated by the computing power gateway in real time during the process of routing and forwarding computing power operation tasks, in accordance with the indicator monitoring rules.

[0019] In one embodiment, the exposure module is further configured to: uniformly bind the same set of computing power-related tag combinations to each computing power routing-type indicator data and each upstream service-type indicator data generated by a single computing power routing single forwarding of a single computing power running task; and bind gateway-related tag combinations to each gateway status indicator data.

[0020] In one embodiment, the exposure module is further configured to: preset label fields for computing power routing; determine fixed label values ​​corresponding to some label fields based on the routing configuration information of computing power routing; determine dynamic label values ​​corresponding to the remaining label fields based on the actual forwarding status of the computing power running task in this instance; fill the fixed label values ​​and dynamic label values ​​into the corresponding label fields to form a computing power-related label combination, and bind the computing power-related label combination to each computing power routing-type indicator data and each upstream service-type indicator data generated in this forwarding.

[0021] In one embodiment, the combination of computing power-related tags involved in the exposure module includes at least one of the following types of tags: Routing identification related tags; Upstream service related tags; Task attribute related tags; Environmental isolation related labels; Resource specification related tags.

[0022] In one embodiment, when the computing power route corresponds to multiple candidate upstream computing power nodes, the upstream service-related tags are dynamic tags. The exposure module is further used to: determine the target upstream computing power node to which the computing power route forwards the computing power running task, the target upstream computing power node being selected by the computing power route from the candidate upstream computing power nodes; and dynamically fill in the upstream service-related tags based on the node information of the target upstream computing power node.

[0023] In one embodiment, the metric listening port used by the metric exposure service in the exposure module is independent of the task traffic port for processing computing power running tasks in the computing power routing.

[0024] Thirdly, the present invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, it implements the steps in the method for high-precision monitoring of computing power routing status by computing power usage of the intelligent computing cloud platform described in the first aspect above.

[0025] Fourthly, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the method for high-precision monitoring of computing power routing status by computing power utilization in the intelligent computing cloud platform described in the first aspect above.

[0026] Fifthly, the present invention also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps in the method for high-precision monitoring of computing power routing status by computing power utilization in the intelligent computing cloud platform described in the first aspect above.

[0027] In this invention, step S1 involves collecting multi-dimensional indicator data generated by the computing power gateway during the routing and forwarding of computing power tasks through the built-in indicator exposure service. The multi-dimensional indicator data includes computing power routing indicator data, upstream service indicator data, and gateway status indicator data. Step S2 involves exposing the multi-dimensional indicator data to the outside world and binding corresponding dimension tag combinations to the indicator data of different dimensions during exposure. Step S3 involves pulling indicator data at a preset frequency through an external data collection service and storing the indicator data in a time-series manner according to the tag combinations. Step S4 involves processing the time-series stored indicator data based on the tag combinations when a query statement is received and displaying the visualization results, wherein the query statement is a single-dimensional query statement or a multi-dimensional cross-query statement. In this way, by collecting real-time data on computing power routing metrics, upstream service metrics, and gateway status metrics, comprehensive and fine-grained monitoring of the entire computing power routing and forwarding process can be achieved. By binding corresponding dimensional tag combinations to metrics data of different dimensions, metrics from different sources and dimensions can be freely associated, aggregated, and filtered by tags, supporting single-dimensional statistics and multi-dimensional cross-analysis, greatly improving the flexibility of monitoring data. The tag mechanism can also improve the monitoring accuracy of computing power routing status. Attached Figure Description

[0028] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the description of the present invention will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 This is a flowchart of a method for high-precision monitoring of computing power routing status through computing power usage provided by the present invention; Figure 2 This is an architecture diagram of a method for high-precision monitoring of computing power routing status through computing power usage provided by the present invention; Figure 3 This is a structural diagram of a device for high-precision monitoring of computing power routing status through computing power usage provided by the present invention; Figure 4 This is a structural diagram of an electronic device provided by the present invention. Detailed Implementation

[0030] The technical solutions of this invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0031] The “computing power” mentioned in this invention refers to: the ability of computer equipment or computing / data center to process information; the ability of computer hardware and software to work together to perform a certain computing requirement; the computing power to achieve the target result output by processing information data; and a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, mainly providing services to society through computing power infrastructure.

[0032] The "computational power" (CP) described in this invention refers to the ability of a data center server to process data and output results. It is a comprehensive indicator of a data center's computing power, encompassing general computing power, supercomputing power, and intelligent computing power. The commonly used unit of measurement is floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), with higher values ​​indicating stronger overall computing power. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A supercomputers, 500,000 mainstream server CPUs, or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 +CP 智能 +CP 超级 .

[0033] The "Network Power" (NP) mentioned in this invention refers to the performance of data transmission capability of computing facilities, which includes comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, and involves network transmission within and between data centers. It is a comprehensive indicator for measuring network transmission scheduling capability.

[0034] The "Storage Power" (SP) described in this invention refers to the comprehensive capabilities of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon operation. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and internal storage devices within servers. The commonly used unit of measurement for storage capacity is exabytes (EB, 1EB = 2^60 bytes), the commonly used unit of measurement for performance is the number of read / write operations per second (IOPS / TB), and the disaster recovery ratio is an important indicator of security and reliability.

[0035] The "computing infrastructure" mentioned in this invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, enabling centralized computing, storage, transmission, and application of information.

[0036] The "new information infrastructure" mentioned in this invention refers to network infrastructure such as 5G networks, fiber optic broadband networks, backbone networks, international communication networks, and satellite internet; computing infrastructure such as data centers, general computing centers, intelligent computing centers, and supercomputing centers; and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0037] The “computing power” mentioned in this invention includes: general computing power, intelligent computing power, and supercomputing power.

[0038] The "general computing power" mentioned in this invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0039] The "intelligent computing power" mentioned in this invention refers to: a computing platform deployed on a large scale based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit) for various artificial intelligence innovative applications, such as natural language processing and machine vision.

[0040] The “supercomputing power” mentioned in this invention refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, and gene analysis.

[0041] The "intelligent computing center" described in this invention refers to a facility that, through the use of large-scale heterogeneous computing resources, including general-purpose computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.), primarily provides the necessary computing power, data, and algorithms for artificial intelligence applications (such as the development, training, and inference of deep learning models). The intelligent computing center encompasses facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enablement.

[0042] The "intelligent computing center cloud platform" mentioned in this invention, abbreviated as "intelligent computing cloud", refers to a cloud computing platform that integrates hardware and software resources based on an intelligent computing center.

[0043] The "intelligent computing center" mentioned in this invention includes, but is not limited to, "smart computing center".

[0044] The "intelligent computing center" mentioned in this invention, also known as an artificial intelligence computing center, is a type of computing infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications, based on artificial intelligence theory and adopting an artificial intelligence computing architecture.

[0045] The "computing center" mentioned in this invention refers to a facility that is mainly composed of infrastructure such as wind, thermal, hydro, and electricity, and IT hardware and software equipment, and has computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0046] The "supercomputing center" mentioned in this invention refers to a supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters. It can provide large-scale computing, storage and network services and is widely used in aerospace, defense, oil exploration, climate modeling and genome sequencing and other application scenarios.

[0047] The “computing resources” mentioned in this invention refer to the technologies and facilities required for the development of the digital society that have the ability to compute, transmit, store and apply information, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guaranteeing resources such as wind, fire, water and electricity.

[0048] The "model" mentioned in this invention includes, but is not limited to, "large language model" and "multimodal large model".

[0049] The "large language model" mentioned in this invention refers to a large-scale language model (LLM), which is a language model with a large number of parameters. It is designed to understand and generate human language, and is trained with a large amount of text data. It can perform a wide range of tasks, including text summarization, translation, and sentiment analysis.

[0050] The “Multimodal Large Models” mentioned in this invention refer to models that combine multimodal information such as text, images, videos, and audio for training, including but not limited to multimodal large language models.

[0051] The “computing power running task” mentioned in this invention refers to a specific workload or job that is executed on computing power resources and requires a certain amount of computing power support. It usually involves complex scenarios such as model inference, model training, and model fine-tuning.

[0052] The meaning of "P99" in this invention is: 99th Percentile, which means that 99% of request response times are below this value, and is used to measure the tail latency performance of computing power routing.

[0053] The meaning of "P95" in this invention is: 95th Percentile, which means that 95% of request response times are below this value, reflecting the typical latency level of computing power routing.

[0054] The meaning of "P50" in this invention is: 50th Percentile, which means that 50% of the request response times are below this value, representing the conventional latency benchmark for computing power routing.

[0055] The meaning of "qos_level" in this invention is: Quality of Service Level, which is used to distinguish the service priority (such as high priority / normal / low priority) of tasks running with different computing power, so as to realize differentiated monitoring and resource guarantee.

[0056] The "SLA" mentioned in this invention means: Service Level Agreement, which refers to the service quality standards agreed upon between the computing power service provider and the user (such as availability ≥99.9%, P99 latency ≤3s, etc.).

[0057] The “computing power purpose” mentioned in this invention refers to the specific application scenario in which the computing power runs tasks, which is used to characterize the actual task purpose when computing power is called, including but not limited to model inference, model training, model fine-tuning, etc.

[0058] The "upstream computing power node" mentioned in this invention refers to a container, virtual machine, or server located upstream of the computing power route that can provide computing power services.

[0059] In this application, "upstream" refers to an abstraction of the backend application service or node pointed to by the computing power gateway, which may include multiple actual upstream computing power nodes.

[0060] The "computing power gateway" described in this invention refers to a core forwarding device in a computing power network when transmitting computing power execution tasks. The computing power gateway determines the forwarding path of the computing power execution tasks. In this invention, the computing power gateway can be an APISIX gateway.

[0061] The "computing power network" mentioned in this invention refers to a network consisting of multiple computing power nodes and computing power gateways used to connect the computing power nodes.

[0062] The "metric exposure service" described in this invention refers to a metric collection and external exposure service integrated within the computing power gateway and implemented based on the gateway's native monitoring plugin. In this invention, the computing power gateway is an APISIX gateway, and the metric exposure service refers to the export server built into the APISIX gateway.

[0063] The "external data acquisition service" described in this invention refers to a service deployed outside the computing power gateway, used to retrieve multi-dimensional indicator data with bound tags from the exposure location of the indicator exposure service of the computing power gateway at a preset frequency, and to store the indicator data in a time-series format for subsequent querying, analysis, and visualization. In this invention, the external data acquisition service is the Prometheus service.

[0064] The "query statement" mentioned in this invention refers to a data query instruction constructed based on tag combinations and statistical functions, used to retrieve, filter, aggregate, and analyze time-series stored indicator data and return monitoring results. In this invention, the external data acquisition service is Prometheus, and the query statement specifically refers to a PromQL query statement. Aggregation operations based on PromQL query statements achieve multi-dimensional grouping statistics, and aggregation can be performed by tag dimension using the by() or without() modifiers, flexibly supporting task monitoring needs.

[0065] The "single-dimensional query statement" described in this invention refers to a query statement that constructs query conditions based on only one tag dimension, retrieves, filters, aggregates and analyzes time-series stored indicator data, and returns monitoring results.

[0066] The "multi-dimensional cross-query statement" described in this invention refers to a query statement that simultaneously performs combination matching based on two or more different labels, retrieves, filters, aggregates and analyzes time-series stored indicator data through multi-label joint constraints, and returns monitoring results.

[0067] It is important to emphasize that the difficulty of computing power routing (hereinafter referred to as "computing power routing") in this invention is far greater than the difficulty of network data routing (hereinafter referred to as "network routing") in existing technologies. Computing power routing represents a fundamental paradigm shift and a revolutionary change. The following is a detailed technical analysis of why the difficulty of computing power routing in this invention is far greater than that of network routing in existing technologies: 1. Dimensions of routing decisions: from "single objective" to "multi-objective trade-offs".

[0068] In existing network routing technologies, the primary optimization target is network metrics such as latency, bandwidth, hop count, and packet loss rate, which are relatively easy to quantify and measure.

[0069] Compared to network routing in existing technologies, the computing power routing of this invention needs to optimize the following two types of heterogeneous resources simultaneously: (1) Network resources, including latency, bandwidth, etc.

[0070] (2) Computing resources, including CPU / GPU utilization, memory size, storage input / output (I / O), specific hardware accelerators, etc.

[0071] This is a multi-objective optimization problem, and often these objectives conflict with each other (for example, the node with the strongest computing power may have high network latency), making the trade-offs and decisions extremely complex.

[0072] 2. The dynamic nature of routing states: from "relatively stable" to "constantly changing".

[0073] In existing network routing technologies, although network topology and link states may change, the frequency of change is relatively low (e.g., measured in seconds or minutes), and routing protocols have convergence time.

[0074] Compared to network routing in existing technologies, the computing node status in the computing power routing of this invention is highly dynamic. The computing power of a GPU node can change from idle to fully loaded within milliseconds. The start and end of a task can instantly change the node's load. This requires the computing power routing system to have near real-time perception and decision-making capabilities, and its update and convergence speed requirements are much higher than those of network routing.

[0075] 3. Global nature of routing information: from "local information" to "global state".

[0076] In existing technologies, network routing typically employs distributed algorithms (such as link-state routing), where each router only knows the overall network topology but does not need to know the details of every data flow in the network.

[0077] Compared to network routing in existing technologies, the computing power routing update module of this invention needs to know the real-time computing power status of all nodes in the entire network and the real-time requirements of all tasks. This creates a strong demand for centralized, global view controllers (such as software-defined networking (SDN) controllers), but it also brings the risks of scalability and single points of failure.

[0078] 4. Heterogeneity and Abstraction: From "Standard Unit" to "Various Differences".

[0079] The network routing technology in the present technology deals with standardized Internet Protocol (IP) packets, which have a uniform format and predictable behavior.

[0080] Compared to network routing in existing technologies, the computing power routing of this invention includes the following three aspects: (1) Heterogeneous computing power, with computing nodes varying greatly, from x86 CPUs to ARM CPUs, to various types of GPUs, neural network processors (NPUs), and field programmable gate arrays (FPGAs). Their architectures, instruction sets, and performance characteristics are completely different.

[0081] (2) The tasks are heterogeneous, and the computing tasks are also different. Some require high parallelism (e.g., GPU), some require high single-core performance (e.g., CPU), some are memory intensive, and some are I / O intensive.

[0082] (3) Abstraction is difficult. How to use a unified "metric" to quantify different types of computing power on different nodes is a huge challenge. It is not possible to measure all computing power simply by "floating point operations per second (FLOPS)" as bandwidth.

[0083] 5. System Coupling and Fault Handling: From “Decoupling” to “Tight Coupling”.

[0084] In existing network routing technologies, the network layer is the transport layer and is decoupled from the application layer. Network failures typically only affect communication, and applications can design retry mechanisms.

[0085] Compared to existing network routing technologies, the computing power routing of this invention tightly couples the network and computation. Once a computation task is routed to a node, if that node fails mid-computation or the network suddenly interrupts, the entire task will fail, potentially requiring the computation to start from scratch rather than simply retransmitting a few data packets. This significantly increases the system's vulnerability and the complexity of fault handling.

[0086] These fundamental differences make computing power routing a core challenge and key technology in future computing power networks, cloud-edge-device collaboration, and artificial intelligence (AI) computing networks. Essentially, it is a complex scheduling problem involving the joint optimization of multi-dimensional resources in a dynamic, heterogeneous, and distributed environment. Its difficulty is naturally far greater than that of existing network routing technologies that only concern data packet paths. The extreme difficulty in implementing computing power routing stems from the fact that it does not simply address the problem of "transporting network data packets (e.g., network data routing)," but rather the complexity of deep integration of computing power and network, dynamic balancing of multiple factors (e.g., heterogeneous indicators such as network bandwidth, computing load, and task priority), and the lack of cross-domain (e.g., cloud, edge, and device) collaboration and standardized computing power measurement. This invention addresses this pressing issue in the field of computing power.

[0087] In the process of executing computing power tasks on intelligent computing cloud platforms, computing power routing, as the core hub connecting computing power resources and task requirements, directly determines the efficiency of computing power scheduling and the quality of task execution through accurate perception of its operational status. Therefore, high-precision monitoring of computing power routing status is crucial. However, the difficulty of monitoring computing power routing is far greater than that of traditional Internet routing monitoring. Traditional Internet routing monitoring only focuses on network layer connectivity, resulting in relatively coarse monitoring granularity. Computing power routing, on the other hand, faces a multi-dimensional, highly real-time, and highly dynamic monitoring scenario involving the convergence of computing and the network. It needs to simultaneously cover the entire link of network, computing power, and tasks, and must cope with millisecond-level traffic surges and the heterogeneous complexity of tens of millions of computing and network elements, exponentially increasing the monitoring difficulty. Existing technologies have not been adapted and optimized for the core characteristics of computing power routing. Current monitoring solutions suffer from problems such as coarse monitoring granularity, insufficient real-time performance, and limited data dimensions, failing to accurately capture the real-time operational status of computing power routing, resulting in consistently low monitoring accuracy.

[0088] To address the problem of insufficient monitoring accuracy of computing power routing status mentioned above, this invention proposes a method and apparatus for intelligent computing cloud platforms to monitor computing power routing status with high precision based on computing power usage.

[0089] Please see Figure 1 , Figure 1 This is a flowchart of a method for high-precision monitoring of computing power routing status through computing power usage in an intelligent computing cloud platform, as provided by the present invention. Figure 1As shown, the method includes: Step S1: Through the built-in indicator exposure service of the computing power gateway, collect multi-dimensional indicator data generated by the computing power gateway in real time during the process of computing power routing and forwarding computing power running tasks. The multi-dimensional indicator data includes computing power routing indicator data, upstream service indicator data, and gateway status indicator data.

[0090] In this invention, the computing power gateway can be an APISIX gateway. The built-in metric exposure service of the computing power gateway is the APISIX gateway's built-in export server. The APISIX gateway natively integrates a Prometheus plugin, which can perform data tracking throughout the entire lifecycle of each computing power execution task processed by the APISIX gateway to capture multi-dimensional metric data generated when the computing power is routed and forwarded. The export server, by loading the Prometheus plugin, realizes the ability to collect and output multi-dimensional metric data.

[0091] The computing power gateway can be configured with multiple computing power routes, each with its own corresponding computing power task forwarding rules. For example, computing power route 1 can be used to forward computing power tasks of model inference to the corresponding computing power node for processing; computing power route 2 can be used to forward computing power tasks of model training to the corresponding computing power node for processing, and so on.

[0092] It is understood that when a customer initiates a computing power execution task, the task will be forwarded by the computing power route in the computing power gateway to the corresponding upstream computing power node, which will then execute the task. In this application, the upstream computing power node (i.e., the computing power node to which the computing power route forwards the task) is referred to as the upstream computing power node.

[0093] In this application, the multi-dimensional indicator data may include indicator data in three dimensions: computing power routing indicator data, upstream service indicator data, and gateway status indicator data. The following will provide a detailed introduction to the three dimensions of indicator data.

[0094] Among them, computing power routing metrics refer to metrics related to computing power routing. For example, computing power routing metrics include, but are not limited to, Requests Per Second (QPS), latency at each percentile (P99 / P95 / P50) and average latency (AVG), status code distribution (2xx / 4xx / 5xx subdivisions), error rate, throughput (bandwidth / request body size), request method distribution, client IP source, and other metrics.

[0095] Upstream service metrics refer to the metrics data of upstream computing nodes related to computing power routing, which can be used to track the health and performance of upstream computing nodes. For example, upstream service metrics include, but are not limited to, health check status (pass / fail count), number of active connections, response time percentile, number of circuit breaker triggers, rate limiting rejections, number of retries, and connection pool exhaustion events of upstream computing nodes.

[0096] The gateway status metrics data refers to the metrics data of the computing power gateway itself. For example, gateway status metrics data includes, but is not limited to, CPU, memory, worker process status, shared memory usage, ETCD connection status (ETCD is a distributed configuration center used by APISIX for unified storage of routing rules and plugin configurations), configuration synchronization latency, plugin execution time, and other metrics data.

[0097] The aforementioned indicator data includes indicator fields and the actual indicator values ​​corresponding to each indicator field.

[0098] The technical effect of step S1: By collecting real-time data on computing power routing metrics, upstream service metrics, and gateway status metrics, it is possible to achieve comprehensive and fine-grained monitoring of the entire computing power routing and forwarding process.

[0099] Step S2: Expose the multi-dimensional indicator data to the outside world, and bind the corresponding dimension tag combination to the indicator data of different dimensions when exposing it. The tag combination corresponding to the computing power routing indicator data and the upstream service indicator data shall include at least the computing power usage tag.

[0100] In this invention, after the multi-dimensional indicator data is collected in real time, it is necessary to expose the multi-dimensional indicator data to the outside world so that external data collection services can pull the indicator data from the data exposure point.

[0101] In this invention, considering that forwarding computing power execution tasks to upstream computing power nodes is a complete task process, the computing power routing index data and upstream service index data generated in this forwarding can be bound to the same set of tags. For example, the tag set may include the computing power routing identifier corresponding to the computing power routing used in this forwarding, the task name configured in the computing power routing, the node name of the upstream computing power node to which the computing power execution task is forwarded, the purpose of computing power used by the upstream computing power node to execute the computing power execution task (e.g., computing power purpose is model inference, model training, model fine-tuning, etc.), the model name used by the upstream computing power node to execute the computing power execution task, the node identifier of the upstream computing power node to which the computing power execution task is forwarded, the tenant identifier of the tenant that initiated the computing power execution task, the operating environment of the upstream computing power node (production environment / development environment / grayscale environment), the region where the upstream computing power node is located, and the service quality level corresponding to the computing power execution task, etc.

[0102] In this invention, the tag combination bound to the gateway status indicator data may include tags such as the identifier of the computing power gateway, the operating environment of the computing power gateway, and the region where the computing power gateway is located.

[0103] The technical effect of step S2: By binding the corresponding dimension's tag combination to the indicator data of different dimensions, the indicator data of each dimension can be accurately associated with its respective dimension, so that when querying indicators in the future, indicators from different sources and dimensions can be freely associated, aggregated and filtered by tags.

[0104] Step S3: Pull indicator data through an external data acquisition service at a preset frequency, and store the indicator data in a time-series manner according to the tag combination.

[0105] The external data acquisition service can be a Prometheus server located outside the computing power gateway. This service can pull the aforementioned multi-dimensional metric data from the computing power gateway's metric monitoring port.

[0106] The preset frequency refers to the frequency at which external data collection services pull exposed indicator data.

[0107] The same retrieval frequency can be set for the indicator data across different dimensions. For example, the retrieval frequency for the indicator data in the three dimensions mentioned above can all be set to retrieve indicator data once every 15 seconds.

[0108] Different retrieval frequencies can be set for different dimensions of indicator data. For example, gateway status indicator data can be set to be retrieved once every 15 seconds; upstream service indicator data can be set to be retrieved once every 30 seconds; and computing power routing indicator data can be set to be retrieved once every 60 seconds.

[0109] This invention does not impose specific restrictions on the frequency of data retrieval for each dimension, and can be adjusted according to the actual situation.

[0110] For short-lifecycle or network isolation scenarios, Pushgateway can be used as a fallback to implement the Push mode.

[0111] The technical effect of step S3 is that the indicator data is pulled by an external data collection service at a preset frequency without manual intervention, ensuring a continuous and stable inflow of indicator data; the indicator data is stored in a time sequence according to the tag combination, which facilitates efficient subsequent querying based on tags.

[0112] Step S4: When a query statement based on tag combination is received, the time-series stored indicator data is processed based on the query statement and the visualization results are displayed. The query statement can be a single-dimensional query statement or a multi-dimensional cross-query statement.

[0113] For example, when the received query statement is “sum(rate(apisix_http_status{calculation_type="reasoning"}[5m]))by(model_name)”, it indicates that a request is needed to run computing power tasks for inference purposes according to the model name. The aggregation is performed based on a 5-minute time window to realize the statistics and monitoring of real-time inference traffic for each model.

[0114] For example, when the received query statement is "histogram_quantile(0.99,sum(rate)",... When (apisix_http_latency_bucket{env="prod"}[5m]))by(le,route_name))”, it indicates that requests for computing power to run tasks in the production environment need to be aggregated based on the route name, and the P99 quantile latency of each route is statistically analyzed and monitored.

[0115] For example, when the received query statement is “sum(rate(apisix_http_requests_total[5m]))by(region, calculation_type, qos_level)”, it indicates that the request data of the computing power gateway needs to be cross-aggregated based on a 5-minute time window according to three dimensions: region, computing power purpose, and service quality level, so as to realize the statistics and monitoring of multi-dimensional traffic distribution.

[0116] In this invention, query statements based on tag combinations can be dynamically injected through Grafana variable templates. Users can select tag monitoring dimensions such as computing power routing identifier, environment, and model through drop-down boxes, and the panel will automatically refresh and present the corresponding aggregated view according to the selected dimensions.

[0117] In this invention, Grafana can be used as a dashboard to present the query results in the form of real-time charts and status panels. Even when the user is not querying, key indicators such as commonly used computing power, routing health, QPS, error rate, and latency P95 can be presented in real-time charts and status panels, allowing for a more intuitive and comprehensive overview.

[0118] The technical effect of step S4 is that when a query statement based on tag combination is received, the time-series stored indicator data is processed and the visualization results are displayed according to the single-dimensional or multi-dimensional cross-query statement. This can support flexible and efficient multi-dimensional data retrieval and visualization, meet the needs of refined monitoring and statistical analysis in different scenarios, and improve the monitoring accuracy of computing power routing status.

[0119] In this invention, step S1 involves collecting multi-dimensional indicator data generated by the computing power gateway during the routing and forwarding of computing power tasks through the built-in indicator exposure service. The multi-dimensional indicator data includes computing power routing indicator data, upstream service indicator data, and gateway status indicator data. Step S2 involves exposing the multi-dimensional indicator data to the outside world and binding corresponding dimension tag combinations to the indicator data of different dimensions during exposure. Step S3 involves pulling indicator data at a preset frequency through an external data collection service and storing the indicator data in a time-series manner according to the tag combinations. Step S4 involves processing the time-series stored indicator data based on the tag combinations when a query statement is received and displaying the visualization results, wherein the query statement is a single-dimensional query statement or a multi-dimensional cross-query statement. In this way, by collecting real-time data on computing power routing metrics, upstream service metrics, and gateway status metrics, comprehensive and fine-grained monitoring of the entire computing power routing and forwarding process can be achieved. By binding corresponding dimensional tag combinations to metrics data of different dimensions, metrics from different sources and dimensions can be freely associated, aggregated, and filtered by tags, supporting single-dimensional statistics and multi-dimensional cross-analysis, greatly improving the flexibility of monitoring data. The tag mechanism can also improve the monitoring accuracy of computing power routing status.

[0120] In one embodiment, step S1 includes: Step S1.1: Pre-configure the indicator monitoring rules for the computing power gateway.

[0121] Among them, the indicator monitoring rules are used to indicate which indicator data to collect for each dimension.

[0122] Managers can dynamically adjust the indicator monitoring rules according to different management needs. For example, they can dynamically adjust the indicator monitoring rules to increase or decrease the collection of certain indicators.

[0123] The technical effect of step S1.1: By pre-configuring the indicator monitoring rules for the computing power gateway, it supports dynamic adjustment according to management needs, enabling flexible customization and on-demand collection of multi-dimensional indicator data, avoiding resource overhead caused by invalid indicators, and improving monitoring efficiency.

[0124] Step S1.2: Through the built-in indicator exposure service of the computing power gateway, collect multi-dimensional indicator data generated by the computing power gateway in real time during the process of routing and forwarding computing power running tasks according to the indicator monitoring rules.

[0125] The built-in metrics exposure service of the computing power gateway is deployed as an independent process within the same Pod as the computing power router. The metrics exposure service reads the memory data of the computing power gateway during runtime, converts the internal state into a metric format that can be recognized by the external data collection service, and provides a metric listening port for the external data collection service to pull metric data.

[0126] Among them, refer to Figure 2 To understand, Figure 2 This application presents an architecture diagram of a method for high-precision monitoring of computing power routing status through computing power usage in an intelligent computing cloud platform. Figure 2 As shown, the client sends the computing power execution task to the computing power gateway. The corresponding computing power route in the computing power gateway forwards the computing power execution task to the upstream computing power node, and the upstream computing power node executes the computing power execution task based on the computing power. Furthermore, the computing power gateway's built-in metric exposure service reads the memory data of the computing power gateway during runtime, converts the internal state into a metric format recognizable by the external data acquisition service, and provides a metric listening port for the external data acquisition service to pull multi-dimensional metric data. The external data acquisition service pulls multi-dimensional metric data from the metric listening port at a preset frequency, and when it receives a query statement based on tag combinations, it processes the time-series stored metric data based on the query statement and displays the visualization results on a visualization panel for administrators to view.

[0127] The technical effect of step S1.2: By strictly following the monitoring rules to collect metrics through the built-in metric exposure service of the gateway, the standardization of metric collection can be guaranteed.

[0128] In one embodiment, the metric listening port used by the metric exposure service described in step S1.2 above is independent of the task traffic port used by the computing power routing to process computing power running tasks. The metric listening port refers to the port that the computing power gateway provides to the external data collection service to retrieve metrics.

[0129] The technical advantage of using an independent metric monitoring port for the metric exposure service from the task traffic port for processing computing power tasks via the computing power routing mechanism is that it enables physical isolation between monitoring metric traffic and task traffic, preventing metric retrieval operations from consuming task processing resources, thereby ensuring the stability and response performance of computing power routing forwarding. At the same time, it facilitates the separate exposure of the metric monitoring port to external data collection services within the computing cluster through a cluster mechanism, and allows the configuration of independent access control rules for the metric monitoring port (such as allowing only specific IP ranges to access), thereby improving the security of the metric monitoring port.

[0130] In one embodiment, step S2 includes: Step S2.1: For each computing power routing index and each upstream service index generated by a single computing power routing and a single computing power running task during a single forwarding, bind the same set of computing power-related tags corresponding to this forwarding.

[0131] For example, suppose computing power route 1 forwards computing power task 1 to upstream computing power node 1 for execution. During this process, a series of indicator data will be generated (such as the computing power routing indicator data and upstream service indicator data described above). In this invention, based on the routing configuration information of the computing power route used in this forwarding process and the actual forwarding situation, a set of tag combinations corresponding to this forwarding can be generated, called the computing power related tag combination. Then, this set of computing power related tag combinations is bound to each computing power routing indicator data and each upstream service indicator data generated in this forwarding process.

[0132] For example, if computing power route 1 forwards computing power task 1 to upstream computing power node 1 for execution, a total of 20 computing power routing-related indicator data and 20 upstream service-related indicator data are generated in this process. After generating the computing power-related tag combination corresponding to this forwarding based on the routing configuration information of the computing power route used in this forwarding process and the actual forwarding situation, this tag combination can be bound to these 20 computing power routing-related indicator data and 20 upstream service-related indicator data respectively.

[0133] The technical effect achieved by step S2.1 is that by uniformly binding the various indicator data generated by a single forwarding of the same computing power route to the same set of computing power-related tags, the computing power routing indicator data and the upstream service indicator data can be accurately associated in terms of tags. This facilitates aggregation query and correlation analysis by tag combination, effectively improving the efficiency of fault location and link tracing.

[0134] Step S2.2: Bind gateway-related tag combinations to each gateway status indicator data.

[0135] The tag combination that binds the gateway status indicator data may include tags such as the identifier of the computing power gateway, the operating environment of the computing power gateway, and the region where the computing power gateway is located.

[0136] The technical effect achieved by step S2.2 is as follows: There are multiple computing power gateways in the intelligent computing cloud platform. By binding the status index data of each gateway to the corresponding gateway-related tag combination, it is possible to support unified monitoring of multiple computing power gateways based on tags such as computing power gateway identifier, computing power gateway operating environment, and computing power gateway location. It can clearly distinguish the operating status of computing power gateways in different environments and regions, which facilitates unified operation and maintenance of multiple computing power gateways.

[0137] In one embodiment, the computing power-related tag combination involved in step S2.1 above includes at least one of the following types of tags: routing identifier related tags, upstream service related tags, task attribute related tags, environment isolation related tags, and resource specification related tags. These five types of tags will be described in detail below.

[0138] Among them, route identification-related tags refer to tags related to computing power routing. Route identification-related tags include, but are not limited to, the following: route_id (unique identifier of computing power route), route_name (the name of the task corresponding to the computing power route, such as llm-inference-v1), api_version (the interface version of the computing power route), and the route type of the computing power route.

[0139] Among them, upstream service-related tags refer to tags related to upstream computing power nodes. Upstream service-related tags include, but are not limited to, the following: upstream (the name of the upstream, which is an abstraction of the backend application service or node pointed to by the computing power gateway, and may contain multiple actual upstream computing power nodes), host (the address of the upstream computing power node to which the computing power running task is actually forwarded), and node_id (the instance identifier of the upstream computing power node to which the computing power running task is actually forwarded).

[0140] Among them, task attribute related tags refer to tags related to the task attributes of the computing power running task. Task attribute related tags include, but are not limited to, the following: calculation_type (the purpose of computing power, such as model inference / model training / model fine-tuning / embedding), model_name (the name of the model used to run the computing power task, such as qwen-72b), and model_family (the model family to which the model used to run the computing power task belongs).

[0141] Among them, environment isolation-related tags refer to tags related to the environment of the computing power operation task forwarding and execution link (computing power gateway + upstream computing power node) and the identifier of the tenant that initiated the computing power operation task. Environment isolation-related tags include, but are not limited to, the following: env (the operating environment where the computing power operation task forwarding and execution link is located, such as production environment / development environment / canary environment), region (the geographical region where the computing power operation task forwarding and execution link is located), cluster (the cluster identifier of the computing power gateway and upstream computing power node in the computing power operation task forwarding and execution link), and tenant_id (the identifier of the tenant that initiated the computing power operation task).

[0142] Among them, resource specification-related tags refer to the tags related to the hardware resource specifications and service quality levels on which the computing power runs tasks. Resource specification-related tags include, but are not limited to, the following: gpu_type (GPU model, such as A100, H100, L40S) and qos_level (service quality level).

[0143] The technical effects of the computing power-related tag combination, including routing identifier tags, upstream service tags, task attribute tags, environment isolation tags, and resource specification tags, are as follows: Through multiple tag types, indicators from different sources and dimensions can be freely associated, aggregated, and filtered by tags, supporting single-dimensional statistics and multi-dimensional cross-analysis, significantly improving the flexibility of monitoring data usage; through multiple tag types, multi-level perception of computing power routing, upstream computing power nodes, environment, and tasks can be achieved, enabling second-level location of single computing power node failures and accurate assessment of task impact; through task attribute tags, differentiated monitoring can be achieved for different computing power uses such as inference, training, and fine-tuning, focusing on monitoring the latency, throughput, and stability indicators of various computing power tasks, enabling refined performance control, effectively ensuring the SLA of corresponding tasks, and solving the technical shortcomings of traditional gateway monitoring that only supports the general HTTP dimension and cannot distinguish the heterogeneous characteristics of computing power tasks.

[0144] In one embodiment, step S2.1 includes: Step S2.1.1: Preset label fields for computing power routing.

[0145] The preset label fields for computing power routing can include the label fields mentioned above, such as routing identifier labels, upstream service labels, task attribute labels, environment isolation labels, and resource specification labels. Specific fields can be dynamically adjusted according to actual needs.

[0146] The technical effect achieved by step S2.1.1: By flexibly pre-setting label fields for computing power routing, it is possible to adapt to the differentiated configurations of different business scenarios and management dimensions, and meet the diverse and personalized monitoring needs of computing power routing status.

[0147] Step S2.1.2: Based on the routing configuration information of computing power routing, determine the fixed label values ​​corresponding to each of the label fields.

[0148] For a specific computing power route, the routing configuration information of that computing power route is fixed. For example, the route_id of a certain computing power route is configured as "route-001", and the computing power route is configured to forward computing power running tasks for model training. That is, the calculation_type of the computing power route is configured as "inference".

[0149] In this invention, fixed tag values ​​corresponding to certain tag fields related to the routing configuration information can be determined based on the routing configuration information of the computing power routing.

[0150] The technical effect achieved by step S2.1.2 is that by determining the fixed label values ​​corresponding to some label fields through the routing configuration information based on computing power routing, the automatic filling of some labels can be achieved, ensuring the consistency of related label values ​​under the same computing power routing.

[0151] Step S2.1.3: Based on the actual forwarding status of the computing power running task in this computing power routing, determine the dynamic label values ​​corresponding to the remaining label fields.

[0152] In addition to the label fields with fixed label values ​​mentioned above, the label fields preset for computing power routing in step S2.1.1 also include remaining label fields that need to be determined based on the actual forwarding situation of the computing power running task in this computing power routing. In this invention, it is necessary to determine the dynamic label values ​​corresponding to these remaining label fields based on the actual forwarding situation of the computing power running task in this computing power routing.

[0153] For example, if a computing power route corresponds to multiple upstream computing power nodes, and the computing power route can only forward the current computing power operation task to one of the multiple upstream computing power nodes, then the aforementioned upstream service-related tags are dynamically changing. They will be filled in according to the actual forwarding situation of the computing power route for the current computing power operation task, that is, based on the relevant information of the upstream computing power node to which the computing power route actually forwards the current computing power operation task.

[0154] For example, the tenant_id (the identifier of the tenant who initiated the computing power operation task) in the environment isolation related label can also be set as a label whose label value needs to be dynamically determined. Different tenants correspond to different tenant_ids. If the computing power operation task is initiated by tenant A, the label value corresponding to the tenant_id can be filled with "tenant_A"; if the computing power operation task is initiated by tenant B, the label value corresponding to the tenant_id can be filled with "tenant_B".

[0155] The technical effect achieved by step S2.1.3 is as follows: Based on the actual forwarding status of the computing power operation task, the dynamic label values ​​corresponding to the remaining label fields are determined, which can reflect the actual forwarding status in real time and keep the label values ​​consistent with the actual forwarding information of the task.

[0156] Step S2.1.4: Fill the fixed tag value and dynamic tag value into the corresponding tag field to form a computing power related tag combination, and bind the computing power related tag combination to each computing power routing index data and each upstream service index data generated in this forwarding.

[0157] After the computing power routing completes the forwarding of this computing power operation task, the fixed label value and dynamic label value determined above for this forwarding are filled into the corresponding label field to form a computing power related label combination. The computing power related label combination is then bound to each computing power routing category indicator data and each upstream service category indicator data generated by this forwarding.

[0158] The technical effect achieved by step S2.1.4 is as follows: By filling fixed and dynamic label values ​​into the corresponding label fields, a unified combination of computing power-related labels is formed. The combination of computing power-related labels is then bound to each computing power routing indicator data and each upstream service indicator data generated in this forwarding process. This allows for precise association between computing power routing indicator data and upstream service indicator data under the same forwarding process. This not only ensures the computing power routing configuration information but also reflects the real-time forwarding characteristics of different tasks, different tenants, and different upstream nodes. It facilitates subsequent flexible multi-dimensional querying, aggregation analysis, and fault tracing, thereby improving the refinement and intelligence level of computing power routing monitoring.

[0159] In one embodiment, referring to the above description, considering the situation where a computing power route corresponds to multiple upstream computing power nodes, and the computing power route can only forward the current computing power operation task to one of the multiple upstream computing power nodes, therefore, when the computing power route corresponds to multiple candidate upstream computing power nodes, the upstream service related label is a dynamic label. Step S2.1.3 includes: Step S2.1.3.1: Determine the target upstream computing power node to which the computing power running task is forwarded by the computing power route. The target upstream computing power node is selected by the computing power route from the candidate upstream computing power nodes.

[0160] The technical effect achieved by step S2.1.3.1 is that, in scenarios where computing power routing corresponds to multiple candidate upstream computing power nodes, the target upstream computing power node selected for actual forwarding is determined, thereby enabling the determination of accurate upstream service-related labels and making the labels bound to computing power routing index data and upstream service index data more accurate.

[0161] Step S2.1.3.2: Dynamically fill in the upstream service-related tags based on the node information of the target upstream computing power node.

[0162] For example, suppose a computing power route corresponds to three candidate upstream computing power nodes, namely upstream computing power node 1, upstream computing power node 2 and upstream computing power node 3. Each of the three candidate upstream computing power nodes has its own corresponding node information. If the computing power route selects to forward the computing power running task to upstream computing power node 1, then upstream computing power node 1 is taken as the target upstream computing power node, and the upstream service related tags are dynamically filled based on the node information of upstream computing power node 1.

[0163] The technical effect achieved by step 2.1.3.2 is as follows: In scenarios where multiple candidate upstream computing power nodes correspond to computing power routing, the upstream service-related tags are dynamically filled in according to the selected target upstream computing power node for actual forwarding. This makes the tags bound to computing power routing index data and upstream service index data more accurate. At the same time, the upstream service-related tags enable fine-grained observation at the node level, which facilitates the rapid location of abnormal nodes and the assessment of their impact on the task.

[0164] In one embodiment, to achieve standardized tag management, this invention establishes a tag naming convention and lifecycle management mechanism to ensure tag consistency during multi-team collaboration. The following tag specifications are defined for computing power routing: tag names use lowercase underscores (snake_case), tag values ​​use lowercase English letters or standard encoding (e.g., model_name=qwen-72b-instruct), and tag values ​​are prohibited from containing spaces or special characters. The validity of tags in the APISIX configuration file is verified through the CI pipeline, and illegal tags are blocked from being published. The Prometheus cardinality is regularly checked to clean up orphan tags for abandoned routes, preventing high cardinality from causing query performance degradation.

[0165] Please see Figure 3 , Figure 3 This is a structural diagram of a device provided by the present invention for a smart computing cloud platform that monitors the routing status of computing power with high precision based on the use of computing power. Figure 3 As shown, a device 300 for high-precision monitoring of computing power routing status through computing power usage in an intelligent computing cloud platform includes: The data collection module 301 is used to collect multi-dimensional indicator data generated by the computing power gateway in real time during the process of computing power routing and forwarding computing power running tasks through the built-in indicator exposure service of the computing power gateway. The multi-dimensional indicator data includes computing power routing indicator data, upstream service indicator data, and gateway status indicator data. The exposure module 302 is used to expose multi-dimensional indicator data to the outside world, and bind corresponding dimension tag combinations to the indicator data of different dimensions during exposure. The tag combinations corresponding to the computing power routing indicator data and the upstream service indicator data shall include at least the computing power usage tag. The pull module 303 is used to pull indicator data through an external data acquisition service at a preset frequency, and store the indicator data in a time sequence according to the tag combination. The query module 304 is used to process the time-series stored indicator data and display the visualization results based on the query statement when a query statement based on tag combination is received. The query statement can be a single-dimensional query statement or a multi-dimensional cross-query statement.

[0166] In one embodiment, the acquisition module 301 is further configured to: pre-configure indicator monitoring rules for the computing power gateway; and collect multi-dimensional indicator data generated by the computing power gateway in real time during the process of routing and forwarding computing power running tasks through the indicator exposure service built into the computing power gateway, in accordance with the indicator monitoring rules.

[0167] In one embodiment, the exposure module 302 is further configured to: uniformly bind the same set of computing power-related tag combinations corresponding to this forwarding to each computing power routing-type indicator data and each upstream service-type indicator data generated by a single computing power routing single forwarding of a single computing power running task; and bind the gateway-related tag combination to each gateway status indicator data.

[0168] In one embodiment, the exposure module 302 is further configured to: preset label fields for computing power routing; determine the fixed label values ​​corresponding to some label fields based on the routing configuration information of computing power routing; determine the dynamic label values ​​corresponding to the remaining label fields based on the actual forwarding situation of computing power running tasks in this instance; fill the fixed label values ​​and dynamic label values ​​into the corresponding label fields to form a computing power-related label combination, and bind the computing power-related label combination to each computing power routing-type indicator data and each upstream service-type indicator data generated in this forwarding.

[0169] In one embodiment, the combination of computing power-related tags involved in the exposure module 302 includes at least one of the following types of tags: Routing identification related tags; Upstream service related tags; Task attribute related tags; Environmental isolation related labels; Resource specification related tags.

[0170] In one embodiment, when the computing power route corresponds to multiple candidate upstream computing power nodes, the upstream service-related tags are dynamic tags. The exposure module 302 is further configured to: determine the target upstream computing power node to which the computing power route forwards the computing power running task, the target upstream computing power node being selected by the computing power route from the candidate upstream computing power nodes; and dynamically fill in the upstream service-related tags based on the node information of the target upstream computing power node.

[0171] In one embodiment, the indicator listening port used by the indicator exposure service in the exposure module 302 is independent of the task traffic port for processing computing power running tasks in the computing power routing.

[0172] The device for high-precision monitoring of computing power routing status by computing power usage provided by the present invention is capable of realizing the various processes of the above-mentioned method for high-precision monitoring of computing power routing status by computing power usage by computing power usage in the intelligent computing cloud platform. The technical features are one-to-one and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0173] It should be noted that the device for high-precision monitoring of computing power routing status by computing power usage in the intelligent computing cloud platform of this invention can be a device, or a component, integrated circuit, or chip in an electronic device.

[0174] The present invention also provides an electronic device, see below. Figure 4 , Figure 4 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. The electronic device includes a memory 401, a processor 402, and a program or instructions stored in the memory 401 that run on the memory. When the program or instructions are executed by the processor 402, they can achieve the following: Figure 1 The corresponding intelligent computing cloud platform achieves the same technical effect through any step in the method embodiment of high-precision monitoring of computing power routing status by computing power usage, and will not be elaborated here.

[0175] The processor 402 can be a CPU, ASIC, FPGA or GPU.

[0176] Those skilled in the art will understand that all or part of the steps of the above-described method embodiment for high-precision monitoring of computing power routing status through computing power usage in an intelligent computing cloud platform can be implemented by hardware related to program instructions, and the program can be stored in a readable medium.

[0177] The present invention also provides a readable storage medium on which a computer program is stored, and which, when executed by a processor, can perform the above-described functions. Figure 1 The corresponding intelligent computing cloud platform can achieve the same technical effect by using any step in the method embodiment for high-precision monitoring of computing power routing status through computing power usage. Therefore, to avoid repetition, it will not be described again here. The storage medium mentioned includes, for example, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0178] The present invention also provides a computer program product, including computer instructions that, when executed by a processor, implement the above-described... Figure 1 The corresponding intelligent computing cloud platform uses a method embodiment to monitor the routing status of computing power with high precision through the use of computing power, and can achieve the same technical effect. To avoid repetition, it will not be described in detail here.

[0179] The terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses. Additionally, the use of "and / or" in this invention indicates at least one of the connected objects, such as A and / or B and / or C, representing seven possibilities: A alone, B alone, C alone, both A and B present, both B and C present, both A and C present, and A, B, and C present.

[0180] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0181] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or second terminal device, etc.) to execute the methods of the various embodiments of the present invention.

[0182] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of the present invention.

Claims

1. A method for high-precision monitoring of computing power routing status through computing power usage in an intelligent computing cloud platform, characterized in that, include: Step S1: Through the built-in indicator exposure service of the computing power gateway, collect multi-dimensional indicator data generated by the computing power gateway in real time during the process of computing power routing and forwarding computing power running tasks. The multi-dimensional indicator data includes computing power routing indicator data, upstream service indicator data, and gateway status indicator data. Step S2: Expose the multi-dimensional indicator data to the outside world, and bind the corresponding dimension tag combination to the indicator data of different dimensions when exposing it. The tag combination corresponding to the computing power routing indicator data and the upstream service indicator data shall include at least the computing power usage tag. Step S3: Pull the indicator data through an external data acquisition service at a preset frequency, and store the indicator data in a time sequence according to the tag combination; Step S4: When a query statement based on the tag combination is received, the time-series stored indicator data is processed based on the query statement and the visualization results are displayed. The query statement is a single-dimensional query statement or a multi-dimensional cross-query statement.

2. The method as described in claim 1, characterized in that, Step S1 includes: Step S1.1: Pre-configure indicator monitoring rules for the computing power gateway; Step S1.2: Through the built-in indicator exposure service of the computing power gateway, collect multi-dimensional indicator data generated by the computing power gateway in real time during the process of routing and forwarding computing power operation tasks according to the indicator monitoring rules.

3. The method as described in claim 1 or 2, characterized in that, Step S2 includes: Step S2.1: For each computing power routing index and each upstream service index generated by a single computing power routing forwarding of a single computing power running task, uniformly bind the same set of computing power-related tags corresponding to this forwarding; and, Step S2.2: Bind a gateway-related tag combination to each of the gateway status indicator data.

4. The method as described in claim 3, characterized in that, Step S2.1 includes: Step S2.1.1: Preset label fields for the computing power routing; Step S2.1.2: Based on the routing configuration information of the computing power routing, determine the fixed label values ​​corresponding to each of the label fields; Step S2.1.3: Based on the actual forwarding status of the computing power running task in this computing power routing, determine the dynamic tag values ​​corresponding to the remaining tag fields; Step S2.1.4: Fill the fixed tag value and the dynamic tag value into the corresponding tag field to form a computing power related tag combination, and bind the computing power related tag combination to each computing power routing index data and each upstream service index data generated in this forwarding.

5. The method as described in claim 4, characterized in that, The computing power-related tag combination includes at least one of the following types of tags: Routing identification related tags; Upstream service related tags; Task attribute related tags; Environmental isolation related labels; Resource specification related tags.

6. The method according to claim 5, characterized in that, When the computing power route corresponds to multiple candidate upstream computing power nodes, the upstream service-related label is a dynamic label. Step S2.1.3 includes: Step S2.1.3.1: Determine the target upstream computing power node to which the computing power routing will forward the computing power running task, wherein the target upstream computing power node is selected by the computing power routing from the candidate upstream computing power nodes; Step S2.1.3.2: Based on the node information of the target upstream computing power node, dynamically fill in the upstream service-related tags.

7. The method as described in claim 2, characterized in that, The indicator listening port used by the indicator exposure service is independent of the task traffic port used by the computing power routing to process the computing power running tasks.

8. A device for high-precision monitoring of computing power routing status through computing power usage in an intelligent computing cloud platform, characterized in that, include: The data collection module is used to collect multi-dimensional indicator data generated by the computing power gateway in real time during the process of routing and forwarding computing power running tasks through the built-in indicator exposure service of the computing power gateway. The multi-dimensional indicator data includes computing power routing indicator data, upstream service indicator data, and gateway status indicator data. The exposure module is used to expose the multi-dimensional indicator data to the outside world, and bind the corresponding dimension tag combination to the indicator data of different dimensions when exposing it. The tag combination corresponding to the computing power routing indicator data and the upstream service indicator data shall include at least the computing power usage tag. The pull module is used to pull the indicator data through an external data acquisition service at a preset frequency, and to store the indicator data in a time sequence according to the tag combination; The query module is used to process the time-series stored indicator data and display the visualization results based on the query statement when a query statement based on the tag combination is received, wherein the query statement is a single-dimensional query statement or a multi-dimensional cross-query statement.

9. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the program implements the steps of the method for high-precision monitoring of computing power routing status by computing power usage of the intelligent computing cloud platform as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for high-precision monitoring of computing power routing status by computing power usage in the intelligent computing cloud platform as described in any one of claims 1 to 7.

11. A computer program product, characterized in that, The method includes computer instructions that, when executed by a processor, implement the steps of the method for high-precision monitoring of computing power routing status by computing power usage in an intelligent computing cloud platform as described in any one of claims 1 to 7.