Computing power operation safety monitoring platform

Through the unified identification system and flexible resource reporting method of the computing power operation security monitoring platform, the comprehensiveness of computing power resource monitoring and complex access process are solved, and efficient management of computing power resources and timely and accurate data are achieved.

CN120469895APending Publication Date: 2025-08-12CHINA ACADEMY OF INFORMATION & COMM

Patent Information

Application Number
CN202510963482.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing computing power operation monitoring platform cannot achieve comprehensive and real-time monitoring of computing power resources, resulting in inefficient resource allocation and use, and the access and reporting process is cumbersome, affecting the timeliness and accuracy of data.

Method used

Provide a computing power operation security monitoring platform, including computing power acquisition module, management module, monitoring module and access management module. It realizes accurate identification and classification of resources through a unified computing power identification system, supports the resource reporting method that actively uploads and waits for capture, and combines the data analysis module to conduct real-time monitoring and data cleaning to generate monitoring reports.

Benefits of technology

Ensure the continuity and accuracy of computing tasks, realize the rational allocation and efficient use of resources, and improve the performance of the platform and the timeliness and accuracy of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120469895A_ABST
    Figure CN120469895A_ABST
Patent Text Reader

Abstract

The invention discloses a computing power operation safety monitoring platform, and relates to the technical field of computing power internet, the computing power operation safety monitoring platform comprises a computing power acquisition module, a computing power management module, a computing power monitoring module, an access management module and a data analysis module, and the computing power management module is responsible for identification and rechecking of computing power resources, management of acquisition plug-ins and management of computing power tasks. According to the computing power operation safety monitoring platform, a unified computing power identification system is provided, so that computing power resources can be accurately identified and classified, comprehensive management of the computing power resources is realized through the computing power management module, real-time monitoring is carried out in cooperation with the computing power monitoring module, and a monitoring report is generated; potential performance bottlenecks can be found and solved in time, the continuity and accuracy of calculation tasks are ensured, and reasonable allocation and efficient use of resources are also ensured. Meanwhile, the platform supports flexible access of monitoring enterprises and computing power suppliers, two resource reporting modes of active uploading and waiting for grabbing are provided, and the reporting requirements of different enterprises are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computing power Internet technology, and specifically to a computing power operation security monitoring platform. Background Art

[0002] Computing resources refer to the capabilities and resources available within a computing system or network to perform computing tasks. These resources include, but are not limited to, various processors (CPUs, GPUs, etc.), software systems, data storage devices, and network connectivity. From individual devices to entire networks, computing resources embody data processing capabilities and are a core element of modern technological development. Computing operation refers to the process by which these computing resources operate and process tasks in real-world applications. Its efficiency and stability directly impact the quality and speed of computing tasks and are a key indicator of computing system performance.

[0003] With the rapid development of the computing internet, monitoring the security of computing power operations has become particularly important. Real-time monitoring of the operating status of computing resources helps to promptly identify and resolve potential performance bottlenecks, ensuring that computing tasks can be completed efficiently and accurately.

[0004] However, in practice, existing monitoring platforms typically focus only on specific computing resources or indicators, failing to provide comprehensive, real-time monitoring of computing resource services, computing platform services, computing application services, and computing interconnection and scheduling services. This impacts the rational allocation and use of computing resources and poses a risk to computing security. Furthermore, the access and reporting processes of existing platforms are often overly complex, making it difficult for monitoring companies and computing providers to report data. This not only affects the timeliness and accuracy of data but also reduces the efficiency of the platform. Summary of the Invention

[0005] The purpose of the present invention is to provide a computing power operation security monitoring platform to solve the problems raised in the above background technology.

[0006] To achieve the above objectives, the present invention provides the following technical solutions: a computing power operation security monitoring platform, comprising: Computing power collection module: responsible for collecting computing power identifiers and reporting them to the intermediate table, and pre-processing the collected raw data, such as data cleaning and format conversion. The computing power collection module obtains computing power identifiers by actively pulling or receiving identifiers through various system APIs, deploying probes on bare metal, and reporting through a visual interface. The intermediate table can collect monitoring indicators for four business aspects: computing power resource business operation security monitoring, computing power platform business operation security monitoring, computing power application business operation security monitoring, and computing power interconnection scheduling business operation security monitoring. Computing power management module: responsible for the identification and review of computing power resources, the management of collection plug-ins, and the management of computing power tasks, in order to achieve unified and standardized management of computing power resources and ensure the accuracy and effectiveness of computing power resources; Computing power monitoring module: Responsible for monitoring the computing power resource business, computing power platform business, computing power application business, and computing power interconnection scheduling business reported by the enterprise, as well as displaying monitoring data and generating monitoring reports. This module monitors the operating status of computing power resources in real time, promptly identifies and resolves problems, and ensures the stability and security of computing power resources. Access management module: responsible for managing the access process of monitoring companies and computing power providers, ensuring the security and effective transmission of data, and achieving effective management of monitoring companies and computing power service providers; Data analysis module: Responsible for conducting in-depth analysis of monitoring data, extracting valuable information and indicators, providing data support for the optimization and management of computing resources, helping managers better understand the operating status of computing resources, and formulating more reasonable computing resource allocation and usage strategies.

[0007] Furthermore, the computing power management module includes the following submodules: Computing power identification submodule: provides an overview query page for the computing power identification of the computing power resources reported by the enterprise, and supports queries in two dimensions: full computing power and available computing power; Computing power review submodule: This module provides a computing power review page for situations where there are significant discrepancies between the total computing power resources reported by an enterprise. It also uses an automated verification algorithm to conduct consistency audits on any repeated reporting of computing power identifiers for the same availability zone by monitored enterprises. Collection plug-in submodule: monitors the running status of the gateway plug-in in real time, including the request address, request method, authentication method, authentication data and plug-in type, and provides a unified visual interface for monitoring and management of the gateway plug-in, making it easier for administrators to monitor and manage; Computing power task submodule: Real-time viewing and management of computing power tasks, including computing power identification collection tasks, computing power identification push tasks, computing power identification reporting tasks, computing power identification verification tasks, and operation monitoring and reporting tasks. Specifically, by managing the computing power task status, you can view the progress and status of various computing power identification operations.

[0008] Furthermore, the computing power monitoring module includes the following submodules: Object Monitoring Submodule: Monitors the computing resources, computing platform, computing application, and computing interconnection scheduling services reported by enterprises to achieve comprehensive monitoring of computing resources. This module includes the following business monitoring and analysis: Computing resource business operation security monitoring: Operational security monitoring is conducted for general computing, intelligent computing, supercomputing, cloud computing and other resources from dimensions such as resource location, resource scale, resource utilization, resource availability and business response time. The reliability, stability and compliance of resources are comprehensively assessed to provide a strong basis for the optimal allocation of computing resources. More specific monitoring indicators include: the city and availability zone where the computing resources reported by the enterprise are located, as well as the number of availability zones, the number of servers and network capacity in the same availability zone, the server type of a single server, the specifications of the processor, memory, hard disk, accelerator card, network card, resource type, service type, number of replicas, storage method, computing power Internet address, chip unique number, sales status and availability status, CPU / GPU / supercomputing / cloud computing service utilization, storage service utilization of a resource pool, including the proportion of abnormal resources and the monthly resource unavailability period, and the response time of services providing resources to the outside world. Computing platform business operation security monitoring: This involves conducting operational security monitoring on computing management platforms, cloud management platforms, network management platforms, and large-scale internet platforms from dimensions such as platform business availability, platform business traffic monitoring, interface security, and business response time. This assesses the reliability, security, and stability of the business platforms and provides guidance for continuous platform optimization. More specific monitoring indicators include: monthly platform service unavailability duration, platform system disaster recovery architecture, platform upstream and downstream traffic monitoring, interface security solutions, and platform business response time. Computing power application business operation security monitoring: Conduct operational security monitoring of computing power applications such as large models, computing power cards, and cloud computers from the perspectives of business continuity, business observability, and business response time. This evaluates the continuity, observability, and responsiveness of computing power application services, and promotes the continuous improvement of computing power applications. More specific monitoring indicators include: the service recovery time (RTO) for each computing power business failure, the data loss time (RPO) for each computing power business failure recovery, the enterprise's own monitoring indicator system for computing power applications, and the response time of application services. Computing power interconnection scheduling business operation security monitoring: This monitors the operation security of the plug-ins and interfaces involved in computing power interconnection scheduling from the perspectives of consistency, interconnection scheduling efficiency, and security. This evaluates the consistency, efficiency, and security of interconnection scheduling to ensure the smooth operation of computing power interconnection scheduling. More specific monitoring indicators include: the enterprise's computing power resource opening interface and standard consistency, the peak rate of a typical availability zone using long-distance RDMA, the average transmission rate between internal intelligent computing resources, the average transmission latency between internal intelligent computing resources, interconnection identity verification, and interconnection interface security. Data monitoring submodule: monitors the monthly monitoring data reported by enterprises and displays the monitoring data on a visual page, including platform business traffic monitoring, utilization of various computing resources, abnormal resource ratio, business response time, etc. Report generation sub-module: Generate monitoring reports regularly, count and analyze the monitoring indicators reported by local enterprises from the province and enterprise dimensions, and provide decision support for management.

[0009] Furthermore, the intermediate table reporting process is as follows: After logging into the computing power operation security monitoring platform, click on computing power reporting, and then click on operation monitoring reporting. Here you can download the intermediate table template and report the intermediate table.

[0010] Furthermore, the intermediate tables are divided into the following four categories: General filling items: Enterprises should fill in according to actual conditions; Items automatically generated by the system: Enterprises do not need to fill in (this item is not included in the intermediate table). The system automatically calculates the results based on the information in the common reporting items and displays them on the platform (including: A. CPU / GPU / supercomputing / cloud computing service utilization: sales volume / total volume; B. Abnormal resource ratio: unavailable volume / total volume; C. Interconnection identity verification: the platform automatically checks whether the enterprise with the reported standard identification is present); Service response time: The enterprise fills in the accessible URL, and the platform will test the service response delay by accessing the URL; API interface to be filled in: A. Interconnection interface security: After filling in the interface API, the enterprise contacts the monitoring platform and provides API usage documentation; B. For storage service utilization, traffic monitoring, intelligent computing transmission rate, and intelligent computing transmission delay, you can choose to fill in the values directly or report via the API. Filling in the values is a normal reporting item, while filling in the API requires contacting the monitoring platform to provide API usage documentation.

[0011] Furthermore, the access process of the monitoring enterprise / computing service provider is as follows: Monitoring companies / computing power service providers register and authenticate on the computing power operation security monitoring platform; Log in with a successfully registered account and obtain the reporting certificate in the User Center; Monitoring enterprises / computing power service providers choose to report computing power identification by actively uploading (push) or waiting for crawling (pull).

[0012] Furthermore, the range of resources that can be reported by the monitoring enterprise / computing power service provider includes two categories: full resources and available resources. When reporting resources, full resources are reported first, and there is no requirement for available resources. The full resources are statistical dimensions, and users are expected to report all computing power resources as much as possible. Available resources are resource activation dimensions, and resources that users can activate and use can be considered available resources.

[0013] Furthermore, the resource reporting methods are divided into two types: active upload (push) and waiting for crawling (pull), as follows: Ⅰ. Active upload, i.e. push method, using json format API to report: 1) Fill in the reporting credentials in the request header; 2) Fill in the identification resource in the body according to the format; Ⅱ. Waiting for crawling, also known as the pull method: Provide an API interface to the computing power platform, and the platform will crawl data through the corresponding plug-in.

[0014] Furthermore, the computing power service provider reports computing power resources and obtains a corresponding computing power identifier. The computing power identifier consists of four parts: registration code, resource code, specification code, and path code, as follows: Registration code: specifies information such as computing power provider and availability zone; Resource code: describes the information of all resources in an entire availability zone; Specification code: describes the specification information within a class of servers; Path code: describes the information of a specific chip.

[0015] Furthermore, the data analysis module integrates and cleans the monitoring data reported by each enterprise to ensure the accuracy and consistency of the data, and uses statistical, machine learning and other methods to analyze and mine the monitoring data to discover patterns and trends in the data. It generates monitoring reports based on the analysis results, provides visual charts and data analysis results, and helps users better understand the monitoring data and make decisions. At the same time, it establishes an early warning mechanism to monitor and warn of abnormal data in real time to ensure that users can promptly discover and deal with potential problems.

[0016] Furthermore, the platform also includes a situational awareness screen for statistics and analysis of monitoring indicators, and presenting a three-dimensional panoramic view of computing power security monitoring, providing multi-perspective data decision support for regulatory authorities and corporate users. The situational awareness screen includes: The monitoring computing power overview and statistics screen provides users with comprehensive monitoring and statistical information on computing power resources, including core computing power resource indicators, computing power resource usage, intelligent computing resource type statistics, vendor computing power resource statistics, and enterprise storage statistics. It displays the number of computing power identifiers, total computing capacity, total number of accelerator cards, total number of servers, total number of CPU cores, total storage capacity, total number of availability zones, number of computing power enterprises, and computing power cities managed by the current monitoring platform. It also displays the computing power reported by specific monitored vendors (including the top 5 vendor computing power and vendor storage statistics rankings). The monitoring computing power type statistics screen focuses on displaying the monitoring computing power resources and usage of different computing types. It supports classification and statistics of monitoring computing power by resource type, including general computing power statistics, intelligent computing power statistics, and supercomputing power statistics. Each type of computing power statistics includes the total computing power, the total number of identifiers, the number of corresponding resources, and resource usage, providing support for users to understand the usage of different resource types such as general computing, intelligent computing, and supercomputing. National Monitoring Computing Power Map: This map focuses on displaying the distribution of monitoring computing power resources in various regions of China. It supports displaying the three dimensions of monitoring computing power calculation volume, number of computing power identifications, and number of regulated enterprises in each province and city on the map. It also provides statistics on computing power resources in each province, including the total number of computing power cards and the total amount of computing power. It also supports displaying the ranking of computing power resources in provincial and municipal administrative regions, providing intuitive data support for understanding the geographical distribution and key areas of computing power resources across the country. The present invention provides a computing power operation security monitoring platform, which has the following beneficial effects: The present invention provides a unified computing power identification system, which enables computing power resources to be accurately identified and classified. Through the computing power management module, comprehensive management of computing power resources is achieved, and the computing power monitoring module is used for real-time monitoring and generation of monitoring reports, which not only ensures the continuity and accuracy of computing tasks, but also ensures the rational allocation and efficient use of resources. At the same time, the platform supports flexible access for monitoring companies and computing power suppliers, and provides two resource reporting methods: active upload and waiting for capture, which meet the reporting needs of different companies, ensure the timeliness and accuracy of data, and further improve platform performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a schematic diagram of the module structure of the computing power operation security monitoring platform of the present invention; Figure 2 This is a schematic diagram of the access process for monitoring enterprises / computing power service providers to the computing power operation security monitoring platform of the present invention. DETAILED DESCRIPTION

[0018] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention but are not intended to limit the scope of the present invention.

[0019] like Figure 1-Figure 2 As shown in the figure, the computing power operation security monitoring platform includes: Computing power collection module: responsible for collecting computing power identifiers and reporting them to the intermediate table, and pre-processing the collected raw data, such as data cleaning and format conversion. The computing power collection module obtains computing power identifiers by actively pulling or receiving identifiers through various system APIs, deploying probes on bare metal, and reporting through a visual interface. The intermediate table can collect monitoring indicators for four business aspects: computing power resource business operation security monitoring, computing power platform business operation security monitoring, computing power application business operation security monitoring, and computing power interconnection scheduling business operation security monitoring. Computing power management module: responsible for the identification, review, management of collection plug-ins, and management of computing power tasks, in order to achieve unified and standardized management of computing power resources and ensure the accuracy and effectiveness of computing power resources. This module includes the following sub-modules: Computing power identification submodule: provides an overview query page for the computing power identification of the computing power resources reported by the enterprise, and supports queries in two dimensions: full computing power and available computing power; Computing power review submodule: This module provides a computing power review page for situations where there are significant discrepancies between the total computing power resources reported by an enterprise. It also uses an automated verification algorithm to conduct consistency audits on any repeated reporting of computing power identifiers for the same availability zone by monitored enterprises. Collection plug-in submodule: monitors the running status of the gateway plug-in in real time, including the request address, request method, authentication method, authentication data and plug-in type, and provides a unified visual interface for monitoring and management of the gateway plug-in, making it easier for administrators to monitor and manage; Computing power task submodule: Real-time viewing and management of computing power tasks, including computing power identification collection tasks, computing power identification push tasks, computing power identification reporting tasks, computing power identification verification tasks, and operation monitoring and reporting tasks. Specifically, by managing the computing power task status, you can view the progress and status of various computing power identification operations.

[0020] Computing power monitoring module: Responsible for monitoring the computing power resource business, computing power platform business, computing power application business, and computing power interconnection scheduling business reported by enterprises, as well as displaying monitoring data and generating monitoring reports, so as to monitor the operating status of computing power resources in real time, discover and solve problems in a timely manner, and ensure the stability and security of computing power resources. This module includes the following sub-modules: The object monitoring submodule monitors the computing resource services, computing platform services, computing application services, and computing interconnection scheduling services reported by enterprises to achieve comprehensive monitoring of computing resources. In this embodiment, it specifically includes the following business monitoring and analysis: Computing resource operation security monitoring: Operational security monitoring is conducted for general computing, intelligent computing, supercomputing, cloud computing and other resources from dimensions such as resource location, resource scale, resource utilization, resource availability and business response time. The reliability, stability and compliance of resources are comprehensively assessed to provide a strong basis for the optimal allocation of computing resources. More specific monitoring indicators include: the city and availability zone where the computing resources reported by the enterprise are located, as well as the number of availability zones, the number of servers and network capacity in the same availability zone, the server type of a single server, the specifications of the processor, memory, hard disk, accelerator card, network card, resource type, service type, number of replicas, storage method, computing power Internet address, chip unique number, sales status and availability status, CPU / GPU / supercomputing / cloud computing service utilization, storage service utilization of a resource pool, including the proportion of abnormal resources and the monthly resource unavailability period, and the response time of services providing resources to the outside world. Computing platform business operation security monitoring: This involves conducting operational security monitoring on computing management platforms, cloud management platforms, network management platforms, and large-scale internet platforms from dimensions such as platform business availability, platform business traffic monitoring, interface security, and business response time. This assesses the reliability, security, and stability of the business platforms and provides guidance for continuous platform optimization. More specific monitoring indicators include: monthly platform service unavailability duration, platform system disaster recovery architecture, platform upstream and downstream traffic monitoring, interface security solutions, and platform business response time. Computing power application business operation security monitoring: Conduct operational security monitoring of computing power applications such as large models, computing power cards, and cloud computers from the perspectives of business continuity, business observability, and business response time. This evaluates the continuity, observability, and responsiveness of computing power application services, and promotes the continuous improvement of computing power applications. More specific monitoring indicators include: the service recovery time (RTO) for each computing power business failure, the data loss time (RPO) for each computing power business failure recovery, the enterprise's own monitoring indicator system for computing power applications, and the response time of application services. Computing power interconnection scheduling business operation security monitoring: This monitors the operation security of the plug-ins and interfaces involved in computing power interconnection scheduling from the perspectives of consistency, interconnection scheduling efficiency, and security. This evaluates the consistency, efficiency, and security of interconnection scheduling to ensure the smooth operation of computing power interconnection scheduling. More specific monitoring indicators include: the enterprise's computing power resource opening interface and standard consistency, the peak rate of a typical availability zone using long-distance RDMA, the average transmission rate between internal intelligent computing resources, the average transmission latency between internal intelligent computing resources, interconnection identity verification, and interconnection interface security. Data monitoring submodule: monitors the monthly monitoring data reported by enterprises and displays the monitoring data on a visual page, including platform business traffic monitoring, utilization of various computing resources, abnormal resource ratio, business response time, etc. Report generation sub-module: Generate monitoring reports regularly, count and analyze the monitoring indicators reported by local enterprises from the province and enterprise dimensions, and provide decision support for management.

[0021] In this embodiment, the intermediate table reporting process is as follows: After logging into the computing power operation security monitoring platform, click on computing power reporting, and then click on operation monitoring reporting. Here you can download the intermediate table template and report the intermediate table. There are four types of intermediate tables: General filling items: Enterprises should fill in according to actual conditions; Items automatically generated by the system: Enterprises do not need to fill in (this item is not included in the intermediate table). The system automatically calculates the results based on the information in the common reporting items and displays them on the platform (including: A. CPU / GPU / supercomputing / cloud computing service utilization: sales volume / total volume; B. Abnormal resource ratio: unavailable volume / total volume; C. Interconnection identity verification: the platform automatically checks whether the enterprise with the reported standard identification is present); Service response time: The enterprise fills in the accessible URL, and the platform will test the service response delay by accessing the URL; API interface to be filled in: A. Interconnection interface security: After filling in the interface API, the enterprise contacts the monitoring platform and provides API usage documentation; B. For storage service utilization, traffic monitoring, intelligent computing transmission rate, and intelligent computing transmission delay, you can choose to fill in the values directly or report via the API. Filling in the values is a normal reporting item, while filling in the API requires contacting the monitoring platform to provide API usage documentation.

[0022] Access management module: responsible for managing the access process of monitoring enterprises and computing power suppliers, ensuring the security and effective transmission of data, and realizing effective management of monitoring enterprises and computing power service providers.

[0023] In this embodiment, the access process of monitoring enterprises / computing service providers is as follows: Monitoring companies / computing power service providers register and authenticate on the computing power operation security monitoring platform; Log in with a successfully registered account and obtain the reporting certificate in the User Center; Monitoring enterprises / computing power service providers choose to report computing power identification by actively uploading (push) or waiting for crawling (pull).

[0024] Monitoring enterprises / computing service providers can report the scope of resources, including full resources and available resources. When reporting resources, full resources should be reported first. There is no requirement for available resources. Full resources are statistical dimensions, and users are expected to report all computing resources as much as possible. Available resources are resource activation dimensions, and resources that users can activate and use can be considered available resources.

[0025] In this embodiment, resource reporting methods are divided into two types: active upload (push) and waiting for crawling (pull), as follows: Ⅰ. Active upload, i.e. push method, using json format API to report: 1) Fill in the reporting credentials in the request header; 2) Fill in the identification resource in the body according to the format; Ⅱ. Waiting for crawling, also known as the pull method: Provide an API interface to the computing power platform, and the platform will crawl data through the corresponding plug-in.

[0026] The computing power service provider reports computing power resources and obtains the corresponding computing power identification. The computing power identification consists of four parts: registration code, resource code, specification code, and path code. The details are as follows: Registration code: specifies information such as computing power provider and availability zone; Resource code: describes the information of all resources in an entire availability zone; Specification code: describes the specification information of a type of server; Path code: describes the information of a specific chip.

[0027] Data analysis module: Responsible for conducting in-depth analysis of monitoring data, extracting valuable information and indicators, providing data support for the optimization and management of computing resources, helping managers better understand the operating status of computing resources, and formulating more reasonable computing resource allocation and usage strategies.

[0028] In this embodiment, the data analysis module integrates and cleans the monitoring data reported by each enterprise to ensure the accuracy and consistency of the data, and uses statistical, machine learning and other methods to analyze and mine the monitoring data to discover patterns and trends in the data. It generates a monitoring report based on the analysis results, provides visual charts and data analysis results, and helps users better understand the monitoring data and make decisions. At the same time, an early warning mechanism is established to monitor and warn of abnormal data in real time to ensure that users can promptly discover and deal with potential problems.

[0029] Situational Awareness Screen: This screen is used to collect statistics and analyze monitoring indicators, presenting a three-dimensional overview of computing power security monitoring, and providing multi-perspective data decision support for regulators and enterprise users. The screen includes: The monitoring computing power overview and statistics screen provides users with comprehensive monitoring and statistical information on computing power resources, including core computing power resource indicators, computing power resource usage, intelligent computing resource type statistics, vendor computing power resource statistics, and enterprise storage statistics. It displays the number of computing power identifiers, total computing capacity, total number of accelerator cards, total number of servers, total number of CPU cores, total storage capacity, total number of availability zones, number of computing power enterprises, and computing power cities managed by the current monitoring platform. It also displays the computing power reported by specific monitored vendors (including the top 5 vendor computing power and vendor storage statistics rankings). The monitoring computing power type statistics screen focuses on displaying the monitoring computing power resources and usage of different computing types. It supports classification and statistics of monitoring computing power by resource type, including general computing power statistics, intelligent computing power statistics, and supercomputing power statistics. Each type of computing power statistics includes the total computing power, the total number of identifiers, the number of corresponding resources, and resource usage, providing support for users to understand the usage of different resource types such as general computing, intelligent computing, and supercomputing. National Monitoring Computing Power Map: This focuses on displaying the distribution of monitoring computing power resources in various regions of China. It supports separate display on the map based on three dimensions: the amount of monitored computing power in each province and city, the number of computing power identifiers, and the number of regulated enterprises. It also provides computing power resource statistics for each province, including the total number of computing power cards and the total amount of computing power. It also supports displaying the computing power resource rankings of provincial and municipal administrative regions, providing intuitive data support for understanding the geographical distribution and key areas of computing power resources across the country.

[0030] The embodiments of the present invention are presented for purposes of illustration and description and are not intended to be exhaustive or to limit the invention to the disclosed forms. Many modifications and variations will be apparent to those skilled in the art. The embodiments are chosen and described in order to better illustrate the principles of the invention and its practical application and to enable those skilled in the art to understand the invention and design various embodiments with various modifications as suited for specific applications.

Claims

1. The computing power operation security monitoring platform is characterized by: include: Computing power collection module: responsible for collecting computing power identifiers and reporting them to intermediate tables, and preprocessing the collected raw data; The computing power collection module acquires computing power identification by actively pulling or receiving identifications through various system APIs, deploying probes on bare metal, and reporting through a visual interface. The intermediate table can collect monitoring indicators for four business aspects: computing power resource business operation security monitoring, computing power platform business operation security monitoring, computing power application business operation security monitoring, and computing power interconnection scheduling business operation security monitoring. Computing power management module: responsible for the identification and review of computing power resources, management of collection plug-ins, and management of computing power tasks; Computing power monitoring module: responsible for monitoring the computing power resource business, computing power platform business, computing power application business, and computing power interconnection scheduling business reported by enterprises, as well as displaying monitoring data and generating monitoring reports; Access management module: responsible for managing the access process of monitoring enterprises and computing power providers to ensure the security and effective transmission of data; Data analysis module: responsible for in-depth analysis of monitoring data, extracting valuable information and indicators, and providing data support for the optimization and management of computing resources; The computing power monitoring module includes the following submodules: Object monitoring submodule: monitors the computing power resource business, computing power platform business, computing power application business, and computing power interconnection scheduling business reported by enterprises, achieving comprehensive monitoring of computing power resources; Data monitoring submodule: monitors the monthly monitoring data reported by enterprises and displays the monitoring data on a visual page, including platform business traffic monitoring, utilization of various computing resources, abnormal resource ratio, and business response time; Report generation sub-module: Generate monitoring reports regularly, count and analyze the monitoring indicators reported by local enterprises from the province and enterprise dimensions, and provide decision support for management.

2. The computing power operation safety monitoring platform according to claim 1 is characterized in that: The computing power management module includes the following submodules: Computing power identification submodule: provides a computing power identification overview query page, and supports queries in two dimensions: total computing power and available computing power; Computing power review submodule: This module provides a computing power review page for situations where there are significant discrepancies between the total computing power resources reported by an enterprise. It also uses an automated verification algorithm to conduct consistency audits on any repeated reporting of computing power identifiers for the same availability zone by monitored enterprises. Collection plug-in submodule: monitors the running status of the gateway plug-in in real time, including the request address, request method, authentication method, authentication data and plug-in type, and provides a unified visual interface for monitoring and management of the gateway plug-in; Computing power task submodule: Real-time viewing and management of computing power tasks, including computing power identification collection tasks, computing power identification push tasks, computing power identification reporting tasks, computing power identification verification tasks, and operation monitoring and reporting tasks. Specifically, by managing the computing power task status, you can view the progress and status of various computing power identification operations.

3. The computing power operation safety monitoring platform according to claim 1 is characterized in that: The intermediate table reporting process is as follows: After logging into the computing power operation security monitoring platform, click on computing power reporting, and then click on operation monitoring reporting. Here you can download the intermediate table template and report the intermediate table.

4. The computing power operation safety monitoring platform according to claim 3 is characterized in that: There are four types of intermediate tables: General filling items: Enterprises should fill in according to actual conditions; Items automatically generated by the system: Enterprises do not need to fill in the information. The system automatically calculates the results based on the information in the common reporting items and displays them on the platform; Service response time: The enterprise fills in the accessible URL, and the platform will test the service response delay by accessing the URL; API interface to be filled in: A. Interconnection interface security: After filling in the interface API, the enterprise contacts the monitoring platform and provides API usage documentation; B. For storage service utilization, traffic monitoring, intelligent computing transmission rate, and intelligent computing transmission delay, you can choose to fill in the values directly or report via the API. Filling in the values is a normal reporting item, while filling in the API requires contacting the monitoring platform to provide API usage documentation.

5. The computing power operation safety monitoring platform according to claim 1 is characterized in that: The access process of the monitoring enterprise / computing service provider: Monitoring companies / computing power service providers register and authenticate on the computing power operation security monitoring platform; Log in with a successfully registered account and obtain the reporting certificate in the User Center; Monitoring enterprises / computing power service providers can choose to report computing power identification by actively uploading or waiting for crawling.

6. The computing power operation safety monitoring platform according to claim 5, characterized in that: The scope of resources that can be reported by the monitoring enterprise / computing power service provider includes two categories: full resources and available resources. When reporting resources, full resources should be reported first, and there is no requirement for available resources.

7. The computing power operation safety monitoring platform according to claim 5, characterized in that: There are two ways to report resources: active upload and waiting for crawling, as follows: Ⅰ. Active upload, i.e. push method, using json format API to report: 1) Fill in the reporting credentials in the request header; 2) Fill in the identification resource in the body according to the format; Ⅱ. Waiting for crawling, also known as the pull method: Provide an API interface to the computing power platform, and the platform will crawl data through the corresponding plug-in.

8. The computing power operation safety monitoring platform according to claim 7, characterized in that: The computing power service provider reports computing power resources and obtains the corresponding computing power identification. The computing power identification consists of four parts: registration code, resource code, specification code, and path code. The details are as follows: Registration code: specifies the computing power provider and availability zone information; Resource code: describes the information of all resources in an entire availability zone; Specification code: describes the specification information within a class of servers; Path code: describes the information of a specific chip.

9. The computing power operation safety monitoring platform according to claim 7, characterized in that: The data analysis module integrates and cleans the monitoring data reported by each enterprise, and then uses statistical and machine learning methods to analyze and mine the monitoring data. It generates a monitoring report based on the analysis results, provides visual charts and data analysis results, and establishes an early warning mechanism to conduct real-time monitoring and early warning of abnormal data.

10. The computing power operation safety monitoring platform according to claim 1, characterized in that: The platform also includes a situational awareness screen for statistics and analysis of monitoring indicators, presenting a three-dimensional panoramic view of computing power security monitoring, and providing multi-perspective data decision support for regulatory authorities and enterprise users. The situational awareness screen includes: The monitoring computing power overview and statistics screen provides users with comprehensive monitoring and statistical information on computing power resources, including core computing power resource indicators, computing power resource usage, intelligent computing resource type statistics, vendor computing power resource statistics, and enterprise storage statistics. It displays the number of computing power identifiers, total computing capacity, total number of accelerator cards, total number of servers, total number of CPU cores, total storage capacity, total number of available zones, number of computing power enterprises, and computing power cities for the computing power resources managed by the current monitoring platform, and displays the computing power reported by specific monitoring vendors. The monitoring computing power type statistics screen displays the monitoring computing power resources and usage of different computing types. It supports classification and statistics of monitoring computing power by resource type, including general computing power statistics, intelligent computing power statistics, and super computing power statistics. Each type of computing power statistics includes the total computing power, the total number of identifiers, the number of corresponding resources, and resource usage. National Monitoring Computing Power Map: This displays the distribution of monitoring computing power resources across China. It supports displaying each province and city's monitoring computing power based on three dimensions: computing power volume, number of computing power identifiers, and number of regulated enterprises. It also provides statistics on computing power resources for each province, including the total number of computing power cards and total computing power. It also supports displaying the computing power resource rankings of provincial and municipal administrative regions.

Citation Information

Patent Citations

  • Computing power network resource center, computing power service system and data processing method

    CN116910144A

  • Identifier analysis-based heterogeneous computing power sharing platform

    CN117971467A

  • Computing power identifier management method and device, terminal equipment, storage medium and product

    CN118228978A

  • Distributed computing power resource scheduling method and system

    WO2025076899A1

Cited By

  • Enterprise side computing power identification gateway

    CN121077854A

  • Computing power identification method oriented to computing power internet

    CN121173689A

  • Nationwide integrated computing power network monitoring platform

    CN121486210A