An artificial intelligence platform resource metering management method, system, medium and processor

By detecting and evaluating computing power points, storage points, and algorithm models, and optimizing resource combinations, the problem of low resource utilization in artificial intelligence platforms has been solved, achieving efficient resource management and unified scheduling of network resources, thereby improving computing performance.

CN119690639BActive Publication Date: 2026-02-10CHINA SOUTHERN POWER GRID COMPANY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411525218.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-30
Publication Date
2026-02-10
Estimated Expiration
2044-10-30

AI Technical Summary

Technical Problem

Existing AI platforms have low resource utilization rates and lack effective measurement and optimization methods, leading to unbalanced resource management and waste of computing resources.

Method used

By detecting the performance indicators of computing power points and storage points, evaluating algorithm models and business data, and forming task combinations to optimize resource utilization, including the selection and combination of computing power points, storage points, algorithm models and business data, the optimal utilization of resources can be achieved.

Benefits of technology

It has improved the utilization rate of computing power and storage resources, reduced computing costs, achieved efficient management and optimization of resources, formed an efficient and scalable ubiquitous computing power network, and broken through the performance limits of single-point computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119690639B_ABST
    Figure CN119690639B_ABST
Patent Text Reader

Abstract

The application provides an artificial intelligence platform resource metering management method, specifically comprising the following steps: S1: detecting and metering the computing power surplus of each computing power point in the platform system, and dividing the computing power points into training points and reasoning points according to the idle computing power surplus and the computing power point type; S2: detecting and metering the storage performance indicators of each storage point in the platform system to obtain the parameters of each storage point; S3: evaluating the algorithm model in the platform system; S4: detecting and evaluating the business data in the platform system; S5: selecting the computing power points, storage points, algorithm model and business data to form a task combination, so that the utilization rate and benefit of various resources are optimal. In the application scheme, the computing power and network resources are integrated into a whole to form an efficient, scalable and flexible ubiquitous computing power network, the cluster advantage of the computing power is exerted, the scale efficiency of the computing power is improved, and the global intelligent scheduling and optimization of the computing network resources are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence platform optimization technology, and in particular to an artificial intelligence platform resource metering management method, system, medium and processor. Background Technology

[0002] In the technological wave of the 21st century, Artificial Intelligence (AI), as a key force leading future technological development, is changing our lives, work, and even society as a whole at an unprecedented pace. The vigorous development of AI relies on the support of three core engines: algorithms, computing power, and data. These three complement each other, jointly driving continuous breakthroughs and innovations in AI technology. This article will delve into the role, current status, and future trends of these three engines in the development of AI.

[0003] The relationship between computing power and algorithms is one of mutual promotion and constraint. Computing power provides the foundation for algorithm execution, while excellent algorithms can utilize computing resources more efficiently. Algorithms are the soul supporting the entire intelligent world, and computing power is the physical foundation that carries algorithms.

[0004] Unified management and scheduling of computing resources with different characteristics, how to balance the load of each node, and how to dynamically allocate resources to achieve finer-grained resource management, improve computing resource utilization, and achieve the goal of cost reduction and efficiency improvement. With the current new trend of artificial intelligence platform development, there is still little research on the metering and optimization of platform resources, and the resource utilization rate of platforms is low. Most existing research is still in the small-scale optimization within a single system.

[0005] Therefore, there is a need for a method, system, medium, and processor for resource metering management on an artificial intelligence platform. Summary of the Invention

[0006] There is a lack of research on resource metering and optimization in existing technologies, resulting in low resource utilization rates. This invention provides a method, system, medium, and processor for resource metering and management of artificial intelligence platforms, enabling optimal utilization and efficiency of IT resources (computing resources, storage resources, etc.) and atomic capabilities (datasets, model libraries, etc.). The specific technical solution is as follows:

[0007] A method for resource metering and management on an artificial intelligence platform, specifically including the following steps:

[0008] S1: Detect and measure the surplus computing power of each computing point in the platform system, and divide the computing points into training points and inference points according to the idle surplus computing power and the type of computing points.

[0009] S2: Detect and measure the storage performance indicators of each storage point in the platform system to obtain the parameters of each storage point;

[0010] S3: Evaluate the algorithm models within the platform system to determine the type, performance metrics, and status of the algorithm models;

[0011] S4: Detect and evaluate the data types, properties, uses, location distances, or data volumes of business data within the platform system;

[0012] S5: Based on the needs of the business scenario and the evaluation status of various resources, select computing power points, storage points, algorithm models and business data to form task combinations, so as to optimize the utilization and efficiency of various resources.

[0013] Furthermore, the computing power points include homogeneous computing power points and heterogeneous computing power points; the calculation formula for detecting the computing power surplus of homogeneous computing power points is as follows:

[0014] A = a * b * (1 - c);

[0015] In the above formula, A represents the computing power surplus of a certain type of processor; a represents the performance of a single processor; b represents the number of processors; and c represents the utilization rate of each processor.

[0016] Furthermore, the formula for measuring the surplus computing power of the heterogeneous computing power points is as follows:

[0017] ;

[0018] In the above formula, This represents the computing power surplus of the i-th type of processor; denoted as the surplus computing power of the heterogeneous computing power points; n represents the number of processor types in the heterogeneous computing power points.

[0019] Furthermore, the storage performance metrics of the storage point include IOPS, throughput, latency, bandwidth, data security, and storage capacity.

[0020] Furthermore, the step of selecting computing power points, storage points, algorithm models, and business data to form a task combination based on the needs of the business scenario and the evaluation status of various resources includes the following steps:

[0021] S51: Decompose the workload of each task according to the needs of the business scenario and clarify the specific task requirements;

[0022] S52: Select a suitable algorithm model based on the decomposed tasks and the characteristics of the data to be used;

[0023] S53: Select computing power points based on the state of the algorithm model;

[0024] S54: Based on the available storage points and business data, and taking into account factors such as bandwidth, distance, and data throughput, select the storage points and business data.

[0025] Furthermore, step S54 specifically includes the following steps:

[0026] First, determine the available business data based on the nature of the task;

[0027] Secondly, based on the scale of data computation, determine the available storage locations;

[0028] Finally, the specific storage location and business data are selected based on the distance, bandwidth speed, and quantity.

[0029] Furthermore, the calculation formula for selecting specific storage points and business data is as follows:

[0030] ;

[0031] In the above formula, X represents the selection weight; L represents the distance between the storage point or business data and the computing point. is the average distance between available storage points or service data and computing power points; D is the bandwidth between storage points or service data and computing power points. The average bandwidth between available storage points or business data and computing power points; T is the data throughput of storage points or business data. This represents the average throughput of storage points or business data.

[0032] Furthermore, when selecting storage locations, to consider the impact of storage capacity on business expansion—that is, after the current business ends, when the same business is started again, it is highly likely that the stored data generated in this instance will be used again, and new and more data will also be generated, or as the computation progresses, there is a certain probability that new data will be generated—considering data access at the same storage location can improve the reliability of network data usage. The specific calculation formula is as follows:

[0033] ;

[0034] In the above formula, Select the proportion of storage points; L is the distance between the storage point or business data and the computing power point; is the average distance between available storage points or service data and computing power points; D is the bandwidth between storage points or service data and computing power points. The average bandwidth between available storage points or business data and computing power points; T is the data throughput of storage points or business data. C represents the average throughput of storage points or business data; S represents the storage space remaining at the storage point; and S represents the amount of business data used in the calculation.

[0035] An artificial intelligence platform resource metering and management system, applied to the above-described artificial intelligence platform resource metering and management method, includes:

[0036] The computing power point measurement module is used to detect and measure the computing power surplus of computing power points at various locations within the platform system, and to divide computing power points into training points and inference points according to the idle computing power surplus and computing power point type.

[0037] The storage point metering module is used to detect and measure the storage performance indicators of storage points at various locations within the platform system, and obtain the parameters of each storage point.

[0038] The model measurement module is used to evaluate the algorithm models within the platform system and determine the type, performance indicators, and status of the algorithm models.

[0039] The business data measurement module is used to detect and evaluate the data types, properties, uses, locations, or volumes of business data within the platform system.

[0040] The task combination module is used to select computing power points, storage points, algorithm models and business data to form task combinations based on the needs of business scenarios and the evaluation status of various resources, so as to optimize the utilization and efficiency of various resources.

[0041] A computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device where the computer-readable storage medium is located to perform the above-described artificial intelligence platform resource metering management method.

[0042] A processor for running a program, wherein the program executes the above-described artificial intelligence platform resource metering management method during runtime.

[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0044] 1. In this application, computing power and network resources are integrated into a whole to form an efficient, scalable and flexible ubiquitous computing network, enabling the platform to evolve towards the convergence of ubiquitous interconnection and ubiquitous computing with unified scheduling and control. By driving the network to sense the location of computing power, an integrated platform service of local distribution and computing-network convergence is realized.

[0045] 2. By connecting ubiquitous computing power through the network, the performance limits of single-point computing power can be broken through, the cluster advantages of computing power can be brought into play, and the scale efficiency of computing power can be improved. Through global intelligent scheduling and optimization of computing network resources, the "flow" of computing power can be effectively promoted to meet the business's demand for computing power on demand. Attached Figure Description

[0046] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.

[0047] Figure 1 This is a flowchart illustrating a resource metering and management method for an artificial intelligence platform.

[0048] Figure 2 This is a schematic diagram of the structure of a resource metering and management system for an artificial intelligence platform. Detailed Implementation

[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0050] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0051] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0052] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0053] Improving the utilization rate of computing resources has a significant impact on the scale and cost of computing power. A McKinsey report shows that the average daily utilization rate of servers worldwide is typically only 6% at most; according to Gartner, the utilization rate of global data centers is less than 12%. These figures indicate a huge "waste" in server costs and resource consumption in data centers. If the overall utilization rate of computing resources could be increased from 6% to 90%, it would mean an immediate 15-fold increase in macroscopic computing power, while simultaneously reducing the unit cost of computing power to 1 / 15.

[0054] Example 1

[0055] like Figure 1 The diagram shows a flowchart of a resource metering management method for an artificial intelligence platform, which includes the following steps:

[0056] S1: Detect and measure the surplus computing power of each computing point in the platform system, divide the computing points into training points and inference points according to the idle surplus computing power and the type of computing points, and form computing power resource pools respectively, and obtain the location distance of each computing point.

[0057] Furthermore, the computing power types include CPU, GPU, FPGA, DSA, NPU, or DPU; for example, GPUs are more suitable for graphics visualization, rendering, AI inference, and AI training computation, FPGAs are more suitable for chip design testing, gene sequencing, and transcoding computation, and NPUs are more suitable for machine learning, etc. Simultaneously, computing power types also include, for example, combining FPGA and CPU computing architectures to form CPU+FPGA heterogeneous computing power points, or CPU+GPU+DPU forming heterogeneous computing power points that provide integrated rendering and inference computation, allowing the most suitable architecture to perform its best computation and acceleration, thereby achieving system optimization.

[0058] The performance of numerous individual chips is aggregated into a large pool of computing resources. Conversely, if the performance of individual chips cannot be aggregated into a huge pool of computing resources and instead forms isolated islands, then even the highest performance of a single chip is meaningless; it's like a scattered pile of sand, and its utilization rate is difficult to improve.

[0059] To connect individual resources into a vast resource pool, we need to:

[0060] The hardware itself needs to support (hardware) virtualization, such as Intel's VT-x / VT-d technology. I / O devices need to support full hardware virtualization based on technologies such as SR-IOV. The accelerator itself also needs to support virtualized logical processing channels.

[0061] Furthermore, virtualization technology improves the utilization of computing and other resources of individual processing chips, and the software migration function within virtualization allows upper-layer business software to easily select different physical resources (within the entire resource pool) to run. This enables the partitioning of single hardware resources and the pooling of numerous resources across multiple hardware components.

[0062] The formula for measuring the surplus computing power of the same type of computing power points is as follows:

[0063] A = a * b * (1 - c);

[0064] In the above formula, A represents the computing power surplus of a certain type of processor; a represents the performance of a single processor; b represents the number of processors; and c represents the utilization rate of each processor.

[0065] The formula for measuring the computing power surplus of heterogeneous computing power points is as follows:

[0066] ;

[0067] In the above formula, This represents the computing power surplus of the i-th type of processor; denoted as the surplus computing power of the heterogeneous computing power points; n represents the number of processor types in the heterogeneous computing power points.

[0068] S2: Measure and analyze the storage performance of each storage point within the platform system to obtain parameters for each storage point, including location distance, storage capacity, latency, bandwidth, throughput, and IOPS.

[0069] The main performance indicators of the storage resources include:

[0070] IOPS (Input / Output Operations Per Second): IOPS is a metric for measuring the random read / write performance of a storage system, representing the number of input / output operations the system can process per second. Higher IOPS indicates better random read / write performance.

[0071] Throughput: Throughput is a metric for measuring the sequential read / write performance of a storage system, representing the amount of data the storage system can transfer per second. Higher throughput indicates better sequential read / write performance.

[0072] Latency: Latency is a metric that measures the read and write response time of a storage system, representing the time required for the storage system to process read and write requests. The lower the latency, the shorter the read and write response time of the storage system.

[0073] Bandwidth: Bandwidth is a metric that measures the data transfer speed of a storage system. Higher bandwidth means faster data transfer speeds.

[0074] Data security: Data security is an indicator of a storage system's data protection capabilities. Using technologies such as RAID can improve the data protection capabilities of a storage system, thereby improving the stability and reliability of storage performance.

[0075] Storage margin: Storage margin is an indicator that measures the remaining storage space of a storage system. The larger the storage margin, the lower the utilization rate of the storage system and the larger the available space.

[0076] Furthermore, the following evaluation methods can be used to measure and measure storage performance indicators:

[0077] Use professional tools for testing: You can use professional performance testing tools such as Iometer, IOzone, and FIO, which can help measure performance metrics such as IOPS, throughput, and latency.

[0078] Simulate real-world application scenarios: Conduct tests in real-world application scenarios to more accurately evaluate the performance of the storage system. For example, test the actual performance of the storage system by simulating real-world application scenarios such as database operations and file system access.

[0079] Consider system configuration and external environmental factors: When testing, it is necessary to consider the impact of system configuration and external environmental factors on performance. For example, the ratio of reads to writes, access queue depth, and data block size will all affect the IOPS results.

[0080] By using the above methods and indicators, the performance of storage resources can be comprehensively evaluated to ensure that they meet the needs of practical applications.

[0081] S3: Evaluate the algorithm models within the platform system to determine the type, performance metrics, and status of the algorithm models.

[0082] Furthermore, the state includes, for example, whether it is a model to be trained or a trained and usable inference model.

[0083] Furthermore, a model library refers to configuring different models into a unified whole, and effectively managing and using each model through a model library management system. The models in the model library are archived and ready for use; different models with varying performance levels can be called for training and use according to different user needs. For example, algorithm models include the following types:

[0084] Linear regression models are used to predict the values ​​of continuous variables. They predict the value of the target variable by fitting a straight line or hyperplane. Linear regression models are simple to understand and applicable to most datasets, but they are weak in modeling non-linear relationships and are sensitive to outliers and noise.

[0085] Logistic regression is a model used to solve binary classification problems. It predicts the class of a sample by mapping the output of linear regression to probability values ​​between 0 and 1. It is suitable for binary classification problems such as spam classification and disease diagnosis, and its output is a probability, which can intuitively explain the likelihood of a sample belonging to the positive class.

[0086] Decision tree models are used for classification and regression tasks. Decision trees divide a dataset into different categories or predicted values ​​through a series of feature splits. Decision trees are easy to understand and interpret, and are suitable for handling non-linear relationships and missing values, but they are prone to overfitting and underfitting, and are sensitive to outliers and noise.

[0087] Random forest models are ensemble learning algorithms based on decision trees. They construct multiple decision trees and make predictions through voting or averaging. Random forests improve the accuracy and robustness of models and are suitable for fields such as image classification and credit scoring, exhibiting good generalization ability and resistance to overfitting.

[0088] Support Vector Machines (SVMs) are models used for classification and regression tasks. SVMs find an optimal hyperplane to maximize the margin between different classes by mapping data to a high-dimensional space. SVMs are suitable for handling non-linear relationships and high-dimensional data, exhibiting high accuracy and robustness, but their ability to handle large-scale datasets is relatively poor.

[0089] The Naive Bayes model is a classification algorithm based on Bayes' theorem. Naive Bayes assumes that features are independent and classifies data by calculating posterior probabilities. It is suitable for handling missing values ​​and noisy data, but performs poorly when there is a strong correlation between features.

[0090] Furthermore, the following metrics are used to measure the performance of the model:

[0091] Accuracy: Accuracy is a fundamental metric for measuring model performance. It is calculated by dividing the number of correctly predicted instances by the model into the total number of instances in the dataset, thus providing a direct reflection of the model's accuracy. However, when dealing with imbalanced datasets, accuracy may appear inaccurate due to the large number of classes.

[0092] Precision: Precision focuses on the accuracy of a model's predictions for positive samples, calculating the proportion of instances that the model predicts to be positive but are actually positive. Precision is particularly important in applications such as medical diagnostics or fraud detection because it can reduce the serious consequences of false positives.

[0093] Recall: Recall measures a model's ability to identify positive samples; that is, the proportion of instances that are actually positive that the model correctly predicts as positive. In scenarios where the model's recall rate is important, recall is a crucial metric.

[0094] F1 Score: The F1 score is the harmonic mean of precision and recall, taking into account both aspects and suitable for scenarios where both are important. A high F1 score indicates that the model has achieved a good balance between precision and recall.

[0095] Confusion Matrix: A confusion matrix is ​​a two-dimensional table that shows the relationship between the actual class and the model's predicted class, including true positives (TP), false positives (FP), true negatives (TN), and false negatives (FN). Precision, recall, and other metrics can be calculated using the confusion matrix.

[0096] ROC curve (Receiver Operating Characteristic Curve): The ROC curve plots the false positive rate (FPR) on the horizontal axis and the true positive rate (TPR) on the vertical axis, showing the performance at different thresholds. The larger the AUC (area under the ROC curve), the better the model's performance.

[0097] Mean Squared Error (MSE) and Root Mean Squared Error (RMSE): In regression problems, MSE and RMSE are commonly used evaluation metrics that measure the difference between the model's predicted values ​​and the actual values.

[0098] R-squared (R²): R² is used for regression problems and represents the model's ability to explain variance. The closer R² is to 1, the stronger the model's explanatory power.

[0099] S4: Detect and evaluate the data types, properties, uses, locations, amounts, and quality of business data within the platform system.

[0100] The assessment of data quality includes the following key indicators:

[0101] Accuracy: Data accuracy refers to the degree of closeness between the collected or observed data and the true value. The smaller the error, the higher the accuracy of the data.

[0102] Accuracy: Data accuracy refers to the degree of similarity between different observations obtained from repeated measurements of the same object. Accuracy is related to the precision of data acquisition; the higher the precision, the higher the accuracy of the data.

[0103] Consistency: Consistency refers to whether data follows a unified standard and whether the dataset maintains a consistent format. Inconsistent data can lead to errors in data analysis.

[0104] Completeness: Completeness refers to whether there are any missing data or information. Missing data can affect the accuracy and reliability of data analysis.

[0105] Timeliness: Timeliness refers to the time interval between data generation and its availability for viewing. Timely data is crucial for decision support.

[0106] Validity: The values ​​and formats of the data must conform to the requirements of the data definition or business definition. For example, phone numbers and email addresses must conform to a specific format.

[0107] Uniqueness: The uniqueness of data means that a data item or a group of data has no duplicate values. For example, ID data must be unique.

[0108] Security: Data security refers to whether data is protected during storage and transmission to prevent unauthorized access and tampering.

[0109] Accessibility: Data accessibility refers to the ease with which users can obtain and use data. Good accessibility can improve the efficiency of data utilization.

[0110] Authenticity: Data authenticity refers to the high degree of controllability and traceability of the data collection process, ensuring the accuracy of the data.

[0111] These metrics collectively form a comprehensive framework for data quality assessment, helping to ensure data reliability, consistency, and availability. For different business needs, one or more metrics can be selected for evaluation, and the decision to adopt them can be made based on the evaluation results.

[0112] S5: Based on the needs of the business scenario and the evaluation status of various resources, select computing power points, storage points, algorithm models and business data to form task combinations, so as to optimize the utilization and efficiency of various resources.

[0113] Furthermore, the step of selecting computing power points, storage points, algorithm models, and business data to form a task combination based on the needs of the business scenario and the evaluation status of various resources includes the following steps:

[0114] S51: Decompose the workload of each task according to the needs of the business scenario and clarify the specific task requirements; for example, see if the task requirement is a classification problem, a regression problem, or other types of problem solving.

[0115] S52: Select an appropriate algorithm model based on the decomposed tasks and the characteristics of the data to be used. The appropriateness of the algorithm model selection determines whether the task can be completed. Therefore, prioritize selecting an appropriate algorithm model based on the nature of the task.

[0116] S53: Select computing power points based on the state of the algorithm model (e.g., whether it is in a state of waiting to be trained or in a state of trained and ready for inference). In the management of artificial intelligence platforms, the computing time of computing power points often occupies most of the entire business time. Among them, the training time of algorithm models is often long and costly. Therefore, selecting computing power points based on the state of the algorithm model can significantly reduce the overall business completion time. If it is a training task, choose computing power points with a large surplus of computing power. For algorithm models that have been trained and are ready for inference, allocate them to other computing power points.

[0117] S54: Based on the available storage points and business data, and taking into account factors such as bandwidth, distance, and data throughput, select the appropriate storage points and business data.

[0118] First, determine the available business data based on the nature of the task. For example, if the task is to predict the electricity consumption of a certain region, the available business data would include historical electricity consumption data for that region, as well as data on weather, date, and other influencing factors that affect electricity consumption. If the task is to predict traffic flow on a certain road segment, then historical traffic flow data for that road segment, along with corresponding weather, date, and other influencing factors, can be selected.

[0119] Secondly, based on the scale of the data computation, determine the available storage locations. For example, based on the available storage capacity, bandwidth, and other indicators of the storage locations, determine which storage locations can meet the usage requirements.

[0120] Finally, based on factors such as distance, bandwidth speed, and quantity, specific storage locations and business data are selected.

[0121] The specific selection calculation formula is as follows:

[0122] ;

[0123] In the above formula, X represents the selection weight; L represents the distance between the storage point or business data and the computing point. is the average distance between available storage points or service data and computing power points; D is the bandwidth between storage points or service data and computing power points. The average bandwidth between available storage points or business data and computing power points; T is the data throughput of storage points or business data. This represents the average throughput of storage points or business data. Under certain conditions, if a certain factor (such as jitter or packet loss) has a significant impact on data transmission, the calculation weight of that factor (the value at a certain point divided by the average of all available values) can be added, referring to a formula-like structure.

[0124] Furthermore, when selecting storage locations, to consider the impact of storage capacity on business expansion—that is, after the current business ends, when the same business is started again, it is highly likely that the stored data generated in this instance will be used again, and new and more data will also be generated, or as the computation progresses, there is a certain probability that new data will be generated—considering data access at the same storage location can improve the reliability of network data usage. The specific calculation formula is as follows:

[0125] ;

[0126] In the above formula, Select the proportion of storage points; L is the distance between the storage point or business data and the computing power point; is the average distance between available storage points or service data and computing power points; D is the bandwidth between storage points or service data and computing power points. The average bandwidth between available storage points or business data and computing power points; T is the data throughput of storage points or business data. C represents the average throughput of storage points or business data; S represents the storage space remaining of storage points; C represents the amount of business data used in the calculation; the ratio of C to S can show which storage point has a larger relative space remaining.

[0127] In practice, the selection of storage points or business data should be based on the one with the highest weight.

[0128] The transformation of the platform management architecture in this application aims at improving the efficiency of data processing. By driving computation through data flow, it processes underlying data on demand and at the fastest speed, greatly improving data processing efficiency and thus achieving a qualitative leap in platform performance.

[0129] As production-driven traffic surpasses consumption-driven traffic in modern society, high-frequency, rich media, primarily based on real-time interaction, is latency-sensitive and has high bandwidth requirements. The modal development of generative AI will place network bandwidth demands exceeding Moore's Law. Computing power deployment will gradually evolve into an architecture where training is centrally deployed on certain hub nodes, while inference is distributed. The centrally trained models need to be synchronized to the distributed inference endpoints, leading to increased backbone network traffic, user data isolation, and significant traffic fluctuations in metropolitan area networks.

[0130] In this application, computing power and network resources are integrated into a whole to form an efficient, scalable, and flexible ubiquitous computing network. This enables the platform to evolve towards a convergence of ubiquitous interconnection and unified scheduling and control of ubiquitous computing. By driving the network to sense the location of computing power, an integrated platform service of local distribution and computing-network convergence is achieved.

[0131] By connecting ubiquitous computing power through the network, the performance limits of single-point computing power can be broken through, the cluster advantages of computing power can be leveraged, and the scale efficiency of computing power can be improved. Through global intelligent scheduling and optimization of computing network resources, the "flow" of computing power can be effectively promoted, meeting the business's demand for on-demand computing power.

[0132] Metered management of AI platform resources is the process of quantifying and evaluating the performance and effectiveness of the AI ​​platform. Metering objects cover computing resources, storage resources, dataset access, model library access, etc., achieving user-level / code-level metering to optimize the utilization and efficiency of IT resources (computing resources, storage resources, etc.) and atomic capabilities (datasets, model libraries, etc.). Based on the different service targets, it is further divided into core metering for users and core metering for platform operators. Core metering topics for users include: historical cost metering at the instance or task level, and future usage estimates. Core metering topics for platform operators include: research resource allocation, resource service consumption ranking, and capacity level early warning.

[0133] Example 2

[0134] like Figure 2 The diagram shows a structural schematic of an artificial intelligence platform resource metering management system, applied to the aforementioned artificial intelligence platform resource metering management method, including:

[0135] The computing power point measurement module is used to detect and measure the computing power surplus of computing power points at various locations within the platform system, and to divide computing power points into training points and inference points according to the idle computing power surplus and computing power point type.

[0136] The storage point metering module is used to detect and measure the storage performance indicators of storage points at various locations within the platform system, and obtain the parameters of each storage point.

[0137] The model measurement module is used to evaluate the algorithm models within the platform system and determine the type, performance indicators, and status of the algorithm models.

[0138] The business data measurement module is used to detect and evaluate the data types, properties, uses, locations, or volumes of business data within the platform system.

[0139] The task combination module is used to select computing power points, storage points, algorithm models and business data to form task combinations based on the needs of business scenarios and the evaluation status of various resources, so as to optimize the utilization and efficiency of various resources.

[0140] Example 3

[0141] A computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device where the computer-readable storage medium is located to perform the above-described artificial intelligence platform resource metering management method.

[0142] Example 4

[0143] A processor for running a program, wherein the program executes the above-described artificial intelligence platform resource metering management method during runtime.

[0144] This application provides a resource metering and management method for an artificial intelligence platform, specifically including the following steps: S1: Detecting and measuring the surplus computing power of computing power points at various locations within the platform system, and dividing the computing power points into training points and inference points based on the idle surplus computing power and the type of computing power points; S2: Detecting and measuring the storage performance indicators of storage points at various locations within the platform system to obtain the parameters of each storage point; S3: Evaluating the algorithm models within the platform system to determine the type, performance indicators, and status of the algorithm models; S4: Detecting and evaluating the data types, properties, uses, location distances, or data volumes of business data within the platform system; S5: Selecting computing power points, storage points, algorithm models, and business data to form task combinations based on the needs of the business scenario and the evaluation status of each resource, so as to optimize the utilization rate and efficiency of various resources. In this application's solution, computing power and network resources are integrated into a whole to form an efficient, scalable, and flexible ubiquitous computing power network, enabling the platform to evolve towards the convergence of ubiquitous interconnection and ubiquitous computing unified scheduling and control. By driving the network to perceive the location of computing power, an integrated platform service of proximity distribution and computing-network convergence is achieved. By connecting ubiquitous computing power through the network, the performance limits of single-point computing power can be broken through, the cluster advantages of computing power can be leveraged, and the scale efficiency of computing power can be improved. Through global intelligent scheduling and optimization of computing network resources, the "flow" of computing power can be effectively promoted, meeting the business's demand for on-demand computing power.

[0145] Those skilled in the art will recognize that the units of the various examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of the invention.

[0146] In the embodiments provided by the present invention, it should be understood that the division of units is only a logical functional division. In actual implementation, there may be other division methods, such as multiple units can be combined into one unit, one unit can be split into multiple units, or some features can be ignored.

[0147] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0148] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0149] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.

Claims

1. A method for resource metering and management on an artificial intelligence platform, characterized in that, Specifically, the following steps are included: S1: Detect and measure the surplus computing power of each computing point in the platform system, and divide the computing points into training points and inference points according to the idle surplus computing power and the type of computing points. S2: Detect and measure the storage performance indicators of each storage point in the platform system to obtain the parameters of each storage point; S3: Evaluate the algorithm models within the platform system to determine the type, performance metrics, and status of the algorithm models; S4: Detect and evaluate the data types, properties, uses, location distances, or data volumes of business data within the platform system; S5: Based on the needs of the business scenario and the evaluation status of various resources, select computing power points, storage points, algorithm models and business data to form task combinations, so as to optimize the utilization and efficiency of various resources. The process of selecting computing power points, storage points, algorithm models, and business data to form a task combination based on the needs of the business scenario and the assessment status of various resources includes the following steps: S51: Decompose the workload of each task according to the needs of the business scenario and clarify the specific task requirements; S52: Select a suitable algorithm model based on the decomposed tasks and the characteristics of the data to be used; S53: Select computing power points based on the state of the algorithm model; S54: Based on the available storage points and business data, and taking into account factors such as bandwidth, distance and data throughput, select the storage points and business data. Step S54 specifically includes the following steps: First, determine the available business data based on the nature of the task; Secondly, based on the scale of data computation, determine the available storage locations; Finally, based on the distance, combined with bandwidth speed and quantity, the specific storage location and business data are selected. The calculation formula for selecting specific storage points and business data is as follows: ; In the above formula, X represents the selection weight; L represents the distance between the storage point or business data and the computing point. is the average distance between available storage points or service data and computing power points; D is the bandwidth between storage points or service data and computing power points. The average bandwidth between available storage points or business data and computing power points; T is the data throughput of storage points or business data. This represents the average throughput of storage points or business data.

2. The method for resource metering and management of an artificial intelligence platform according to claim 1, characterized in that, The computing power points include homogeneous computing power points and heterogeneous computing power points; the formula for measuring the computing power surplus of homogeneous computing power points is as follows: A = a * b * (1 - c); In the above formula, A represents the computing power surplus of a certain type of processor; a represents the performance of a single processor; b represents the number of processors; and c represents the utilization rate of each processor.

3. The method for resource metering and management of an artificial intelligence platform according to claim 2, characterized in that, The formula for measuring the computing power surplus of heterogeneous computing power points is as follows: ; In the above formula, This represents the computing power surplus of the i-th type of processor; denoted as the surplus computing power of the heterogeneous computing power points; n represents the number of processor types in the heterogeneous computing power points.

4. The method for resource metering and management of an artificial intelligence platform according to claim 1, characterized in that, The storage performance metrics of the storage points include IOPS, throughput, latency, bandwidth, data security, and storage capacity.

5. A resource metering and management system for an artificial intelligence platform, characterized in that, The method for measuring and managing resources on an artificial intelligence platform as described in any one of claims 1 to 4 includes: The computing power point measurement module is used to detect and measure the computing power surplus of computing power points at various locations within the platform system, and to divide computing power points into training points and inference points according to the idle computing power surplus and computing power point type. The storage point metering module is used to detect and measure the storage performance indicators of storage points at various locations within the platform system, and obtain the parameters of each storage point. The model measurement module is used to evaluate the algorithm models within the platform system and determine the type, performance indicators, and status of the algorithm models. The business data measurement module is used to detect and evaluate the data types, properties, uses, locations, or volumes of business data within the platform system. The task combination module is used to select computing power points, storage points, algorithm models and business data to form task combinations based on the needs of business scenarios and the evaluation status of various resources, so as to optimize the utilization and efficiency of various resources.

6. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device where the computer-readable storage medium is located to perform the artificial intelligence platform resource metering management method according to any one of claims 1 to 4.

7. A processor, characterized in that, The processor is used to run a program, wherein the program executes the artificial intelligence platform resource metering management method according to any one of claims 1 to 4 when it runs.

Citation Information

Patent Citations

  • Computing task unloading method, computing device and storage medium

    CN116541106A

  • Identifier analysis-based heterogeneous computing power sharing platform

    CN117971467A