Resource evaluation method and system for inference service model of AI self-evolution technology

Through the inference business model resource evaluation method of AI self-evolution technology, the problems of manual dependence and inefficiency in traditional evaluation methods are solved, efficient and accurate resource evaluation is achieved, and the automation and intelligence level of evaluation is improved through historical data analysis.

CN120216308APending Publication Date: 2025-06-27SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510216746.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The traditional inference business model resource evaluation method relies on manual experience and subjective judgment, resulting in inaccurate evaluation and inefficient evaluation, and the historical resource evaluation plan is not electronically preserved, and lacks the ability to analyze and mine historical data.

Method used

The inference business model resource evaluation method using AI self-evolution technology is adopted, and the inference business model resources are efficient and accurate evaluation of inference business model resources through modules such as computing power indicator management, resource evaluation, result optimization, result output and historical result archiving.

Benefits of technology

It improves the automation and intelligence level of resource evaluation, reduces the time-consuming of manual calculations, enhances the accuracy and efficiency of evaluation, and provides data analysis and model accumulation capabilities through archives of historical results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216308A_ABST
    Figure CN120216308A_ABST
Patent Text Reader

Abstract

The invention discloses a resource evaluation method and system for an inference service model of an AI self-evolution technology, and belongs to the technical field of artificial intelligence, and the method comprises the steps: carrying out the management and maintenance of key indexes, including adding, modifying, customizing, deleting and cancelling; xPU resources required by output are evaluated according to the input demand indexes and historical business scene data; in the resource evaluation process, when the requirement is not met or the evaluation deviation is large, the evaluation result is adjusted and optimized, unreasonability of input requirement data is recognized and analyzed based on the fault-tolerant capability, and an optimized result and suggestion can be given based on empirical data; displaying and outputting a resource evaluation processing result or an optimized result in a visual chart form; and archiving and storing the resource evaluation result of each time. According to the method, the problems of high time consumption and low accuracy of manual calculation are solved, the historical data analysis and mining capability is provided, and the automation and intelligence level of an evaluation scheme is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and more particularly to a method and system for evaluating inference service model resources of AI self-evolution technology. Background Art

[0002] With the rapid development of artificial intelligence technology, inference service models have been widely used in various fields. However, traditional methods for evaluating inference service model resources often rely on manual experience and subjective judgment, resulting in problems such as inaccurate evaluation and low efficiency. Therefore, how to achieve efficient and accurate evaluation of inference service models has become an urgent problem to be solved. Summary of the Invention

[0003] The technical task of the present invention is to address the above deficiencies by providing a method and system for evaluating inference service model resources of AI self-evolution technology, solving the problems of time-consuming manual calculation and low accuracy, solving the problem of non-electronic preservation of historical resource evaluation schemes, providing the ability to analyze and mine historical data, and improving the automation and intelligence level of the evaluation scheme.

[0004] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0005] A method for evaluating inference service model resources of AI self-evolution technology, the implementation of which includes:

[0006] Computing power index management: managing and maintaining key indicators, including adding, modifying, customizing, deleting, and invalidating;

[0007] Resource evaluation: evaluating and outputting the required xPU resources based on the input demand indicators and historical business scenario data;

[0008] Result optimization: During the resource evaluation process, when the requirements are not met or the evaluation deviation is large, adjusting and optimizing the evaluation result to meet the actual requirements. In particular, based on the fault tolerance ability, the unreasonable input demand data can be identified and analyzed, and the optimized result and suggestions can be given based on the empirical data;

[0009] Result output: presenting and outputting the result of the resource evaluation process or the optimized result in the form of an intuitive chart;

[0010] Historical result archiving: In order to facilitate subsequent analysis and research, it is necessary to archive and save the resource evaluation results of each time. In this way, when it is necessary to review the past evaluation results, they can be directly searched in the historical result archiving module, and the internal evaluation model can be accumulated to more conveniently and quickly output a reasonable resource evaluation scheme.

[0011] Furthermore, for the computing power index management, the key indicators include the number of users, model size, model accuracy, inference speed, memory occupancy, computing power, video memory size, video memory bit width, heat dissipation form, power consumption, etc.

[0012] Furthermore, for the computing power index management, the input demand data is cleaned, denoised and standardized; feature extraction is performed, and by constructing a self-evolving neural network model, features closely related to the performance of the inference service model are automatically extracted from the preprocessed data;

[0013] The resource evaluation step evaluates the resource consumption of the inference service model based on the extracted features and generates an evaluation result.

[0014] Furthermore, for the resource evaluation, the resources include various combinations of computing power resources, memory resources, bandwidth resources, storage resources, and budgets. By reasonably allocating these resource plans, it is ensured that the resource supply meets the normal operation of the business.

[0015] Furthermore, for the resource evaluation, data analysis techniques and algorithms of multi-factor aggregation classification are used to process the data and extract valuable information from it.

[0016] Furthermore, for the resource evaluation, custom metrics are supported, allowing users to adjust the evaluation criteria according to their own needs and preferences. This flexibility enables the resource evaluation module to meet the needs of users in different industries and fields.

[0017] Furthermore, the result output includes the ability comparison, cost comparison, recommended solutions, etc. using different acceleration cards.

[0018] The present invention also claims to protect an inference service model resource evaluation system for AI self-evolving technology, including:

[0019] A computing power index management module for managing and maintaining key indicators;

[0020] A resource evaluation module for evaluating and outputting the required xPU resources according to the input demand indicators and historical business scenario data;

[0021] A result optimization module for adjusting and optimizing the evaluation result during the resource evaluation process. When the demand is not met or the evaluation deviation is large, it identifies and analyzes the irrationality of the input demand data based on the fault tolerance ability and gives the optimized result and suggestions based on the empirical data;

[0022] A result output module for presenting and outputting the result of the resource evaluation process or the optimized result in the form of an intuitive chart;

[0023] A historical result archiving module for archiving and storing the resource evaluation results each time.

[0024] The system realizes the inference service model resource evaluation through the above method.

[0025] The present invention also claims to protect an inference service model resource evaluation device for AI self-evolution technology, including: at least one memory and at least one processor;

[0026] The at least one memory is used to store machine-readable programs;

[0027] The at least one processor is used to call the machine-readable program to implement the above method.

[0028] The present invention also claims to protect a computer-readable medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the processor is caused to execute the above method.

[0029] Compared with the prior art, an inference service model resource evaluation method and system for AI self-evolution technology of the present invention have the following beneficial effects:

[0030] Through the inference service model resource evaluation method based on AI self-evolution technology provided by this method, technicians or builders can conveniently and efficiently evaluate the supply schemes of computing power resources required for inference service models in different scenarios; reduce the complexity of resource evaluation schemes, improve the diversity and accuracy of resource evaluation schemes, and have practical significance for accelerating the popularization and implementation of inference service models. Description of the Drawings

[0031] Figure 1 is a schematic diagram of the principle of an inference service model resource evaluation method for AI self-evolution technology provided by an embodiment of the present invention;

[0032] Figure 2 is a schematic diagram of the process of an inference service model resource evaluation method for AI self-evolution technology provided by an embodiment of the present invention. Detailed Embodiments

[0033] The following further describes the present invention in conjunction with the drawings and specific embodiments.

[0034] An embodiment of the present invention provides an inference service model resource evaluation method for AI self-evolution technology, and the implementation of this method includes:

[0035] 1. Computing power index management: including key performance indicators, brands, models, etc. of mass-produced xPU devices at home and abroad; supporting the management and maintenance of key indicators such as the number of users, model size, model accuracy, inference speed, memory occupancy, computing power, video memory size, video memory bit width, heat dissipation form, power consumption, etc., including addition, modification, customization, deletion, and invalidation.

[0036] 2. Resource Evaluation: Evaluate the required xPU resources for output according to the input demand indicators and historical business scenario data; these resources include various combinations of computing power resources, memory resources, bandwidth resources, storage resources, budget, etc. By reasonably allocating these resource plans, it can ensure that the resource supply meets the normal operation of the business.

[0037] 3. Result Optimization: During the resource evaluation process, for situations where the requirements are not met or the evaluation deviation is large, adjust and optimize the evaluation results to meet the actual requirements. Especially based on the fault tolerance ability, it can identify and analyze the irrationality of the input demand data, and can give the optimized results and suggestions based on empirical data.

[0038] 4. Result Output: Present and output the results of resource evaluation processing or the optimized results in the form of intuitive charts: including the ability comparison, cost comparison, recommended solutions, etc. using different acceleration cards.

[0039] 5. Historical Result Archiving: For the convenience of subsequent analysis and research, it is necessary to archive and save the resource evaluation results each time. In this way, when it is necessary to review the past evaluation results, they can be directly retrieved from the historical result archiving module, and the internal evaluation model can be accumulated to output a reasonable resource evaluation plan more conveniently and quickly.

[0040] Among them, resource evaluation has multi-dimensional index analysis and modeling. This module provides users with a comprehensive, accurate and efficient resource evaluation tool. It comprehensively evaluates computing power resources by considering various factors, such as the performance indicators, quantity, cost, risk, etc. of computing power resources. In the resource evaluation module, multi-dimensional index analysis plays a crucial role. These indicators not only include general technical performance indicators, such as floating-point operation ability, single-precision, double-precision, video memory size, etc., but also include financial indicators such as cost and revenue, and other indicators such as the supply chain security, energy consumption, sustainability, maintainability, etc. of computing power resources. Through the comprehensive analysis of these indicators, users can more comprehensively understand the advantages and disadvantages of resources and thus make more informed decisions. To achieve multi-dimensional index analysis, the resource evaluation module of this method adopts data analysis techniques and algorithms of multi-factor aggregation and classification. These techniques can process a large amount of data and extract valuable information from it. In addition, the module also supports custom indicators, allowing users to adjust the evaluation criteria according to their own needs and preferences. This flexibility enables the resource evaluation module to meet the needs of users in different industries and fields. In addition, the scalability of xPU resources also needs to be considered so that in the future, as the number of users increases or the model size grows, it can meet the performance requirements of resource horizontal expansion and resource prediction.

[0041] An inference business scenario usually involves predicting, classifying, or making decisions on new and unseen data or situations based on a pre-trained model. Inference typically involves feeding new data into the pre-trained model and then obtaining the output results of the model. This process can be real-time, such as in applications like image recognition or speech recognition. The computational resource requirements for inference are generally lower than those for the training scenario. It mainly involves one or a few forward propagations of the model. However, for larger models or larger input data, the requirements for computing power resources are also diverse.

[0042] The main aspects involved in an inference model consuming computing power resources include:

[0043] Model structure and number of parameters: Modern deep learning models, especially large language models (LLMs), usually have complex structures and a large number of parameters. For example, the Self-Attention mechanism in the Transformer model structure needs to calculate all attention scores at each position, which requires multiple linear scans of the entire input sequence and the calculation of products and accumulations between each element. This is itself a computationally intensive task.

[0044] Parallel computing requirements: To accelerate the inference process, it is usually necessary to distribute computational tasks across multiple GPUs or TPUs for parallel processing. This means not only dealing with a large amount of data but also coordinating communication and synchronization between various computing nodes, which also requires a large amount of computing resources.

[0045] Memory bandwidth limitations: During the inference process, the model needs to frequently access and update a large amount of parameter and gradient information, which places high demands on memory bandwidth. When the memory bandwidth cannot keep up with the computing speed, a bottleneck will occur, limiting the inference speed of the model.

[0046] Real-time performance and response time: For real-time or interactive applications, such as online customer service robots and games, the response time of the inference model must be fast enough to provide a good user experience. This usually requires high computing performance and fast response capabilities, and thus more computing power resources.

[0047] The core resource metric in the inference scenario is: video memory occupancy. Video memory occupancy = video memory occupied by model parameters + batch_size × video memory per sample (output gradient and momentum); video memory occupied by model parameters = number of parameters × N (float32: n = 4; float16: n = 2; double64: n = 8). Float 16 is also called half-precision and represents a number with 16 bits, that is, 2 bytes; float32 is also called single-precision and represents a number with 32 bits, that is, 4 bytes; double64 is also called double-precision floating-point number and represents a number with 64 bits, that is, 8 bytes.

[0048] In addition, the throughput of the large model can be obtained by the ratio of the number of inferences that can be performed per second to the total number of inferences required. The number of inferences per second represents the number of inference operations that the model can complete per second, which depends on the computational complexity of the model and the hardware performance.

[0049] As Figure 1 shown, it describes the core principle of the inference service model resource evaluation model based on the AI self-evolution technology. The computing power index management module provides data support for the resource evaluation engine by maintaining multi-dimensional evaluation index data; the resource evaluation engine processes the index data through model algorithms and combinations to output the resource usage requirements based on the business scenario.

[0050] The deployment and application of the inference service model need to be based on key reference indicators such as the scale of user objects served by the inference service model, the size of the model, the model accuracy, the inference speed, and the memory occupancy. Through the resource evaluation engine device, multiple resource selection schemes required for deploying and applying the inference service model are output. Specifically, each scheme includes: xPU card type, number of xPU cards, video memory size, CPU recommendation, memory capacity recommendation, model type recommendation, network bandwidth, concurrency, power consumption, budget, etc.

[0051] When conducting the evaluation, key parameters are classified and weighted to determine the key influencing factors of different indicators. First, the size of the inference model needs to be considered. The size of the inference model is measured by the number of parameters, which directly affects the required xPU resources. At the same time, for larger models, more computing resources and memory capacity need to be considered in association to effectively ensure the smooth operation of the inference model and guarantee the user experience.

[0052] The number of users is the most direct manifestation of the application of the inference service model. The number of users that need to be supported needs to be considered to meet the basic requirements of resource evaluation. The number of users refers to the number of users who simultaneously use the model for inference. The increase in the number of users will lead to an increase in the demand for GPU resources, that is, each user needs a certain amount of xPU resources to execute the inference task. The number of users is divided into range ladders to meet the usage requirements of small, medium, and large-scale users.

[0053] Figure 2 It describes the composition of the inference service model resource evaluation process based on the AI self-evolution technology, which is mainly composed of 5 core parts: computing power index management, resource evaluation engine, result optimization, result output, and historical result archiving.

[0054] This method realizes automatic feature extraction and resource evaluation of the inference service model by constructing a self-evolving evaluation model, avoiding the deficiencies of traditional methods that rely on manual experience and subjective judgment, and improving the evaluation efficiency and accuracy. At the same time, the device can realize the automated and intelligent operation of the above evaluation method, which is convenient for users to use.

[0055] An embodiment of the present invention also provides a resource evaluation system for an inference service model of AI self-evolving technology. This system realizes the resource evaluation of the inference service model through the resource evaluation method of the AI self-evolving technology described in the above embodiment.

[0056] The system includes:

[0057] 1. Computing power index management module (including key performance indicators, brands, models, etc. of mass-produced xPU devices at home and abroad): Supports the management and maintenance of key indicators such as the number of users, model size, model accuracy, inference speed, memory occupancy, computing power, video memory size, video memory bit width, heat dissipation form, power consumption, etc., including addition, modification, customization, deletion, and invalidation.

[0058] 2. Resource evaluation module: This module is responsible for evaluating and outputting the required xPU resources according to the input demand indicators and historical business scenario data. These resources include various solution combinations such as computing power resources, memory resources, bandwidth resources, storage resources, budget, etc. By reasonably allocating these resource solutions, it can ensure that the resource supply meets the normal operation of the business.

[0059] 3. Result optimization module. During the resource evaluation process, when the requirements are not met or the evaluation deviation is large, this module adjusts and optimizes the evaluation results, identifies and analyzes the irrationality of the input demand data based on the fault tolerance ability, and gives the optimized results and suggestions based on the empirical data.

[0060] 4. Result output module. This module presents and outputs the results of the resource evaluation process or the optimized results in the form of intuitive charts: including the ability comparison, cost comparison, recommended solutions, etc. using different acceleration cards.

[0061] 5. Historical result archiving module. In order to facilitate subsequent analysis and research, it is necessary to archive and save the resource evaluation results each time. In this way, when it is necessary to review the past evaluation results, they can be directly retrieved from the historical result archiving module, and the internal evaluation model can be accumulated to more conveniently and quickly output reasonable resource evaluation solutions.

[0062] An embodiment of the present invention also provides a resource evaluation device for an inference service model of AI self-evolving technology, including: at least one memory and at least one processor;

[0063] The at least one memory is used to store machine-readable programs;

[0064] The at least one processor is configured to call the machine-readable program to implement the inference service model resource evaluation method of the AI self-evolution technology described in the above embodiments.

[0065] An embodiment of the present invention further provides a computer-readable medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the processor is caused to execute the inference service model resource evaluation method of the AI self-evolution technology described in the above embodiments. Specifically, a system or device equipped with a storage medium can be provided, on which software program code for implementing the functions of any one of the above embodiments is stored, and the computer (or CPU or MPU) of the system or device is caused to read and execute the program code stored in the storage medium.

[0066] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments. Therefore, the program code and the storage medium storing the program code constitute a part of the present invention.

[0067] Embodiments of the storage medium for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code can be downloaded from a server computer via a communication network.

[0068] In addition, it should be clear that not only can the functions of any one of the above embodiments be implemented by executing the program code read by a computer, but also by causing an operating system or the like operating on the computer based on the instructions of the program code to complete part or all of the actual operations.

[0069] In addition, it can be understood that the program code read from the storage medium is written into a memory provided in an expansion board inserted into the computer or a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, the CPU or the like installed on the expansion board or the expansion unit is caused to execute part and all of the actual operations, thereby implementing the functions of any one of the above embodiments.

[0070] The present invention has been described in detail above through the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above-mentioned multiple embodiments, those skilled in the art can know that more embodiments of the present invention can be obtained by combining the code review means in the above different embodiments, and these embodiments are also within the protection scope of the present invention.

Claims

1. A method for evaluating resource of inference business model based on AI self-evolution technology, characterized in that: The implementation of this method includes: Computing power index management: manage and maintain key indicators, including adding, modifying, customizing, deleting, and invalidating; Resource evaluation: Evaluate the xPU resources required for output based on input demand indicators and historical business scenario data; Result optimization: During the resource evaluation process, when the requirements are not met or the evaluation deviation is large, the evaluation results are adjusted and optimized. Based on the fault tolerance capability, the unreasonable input demand data can be identified and analyzed, and optimized results and suggestions can be given based on empirical data; Result output: Display and output the resource assessment results or optimized results in the form of intuitive charts; Archiving of historical results: Archive each resource assessment result.

2. According to the method for evaluating the resource of the inference business model of the AI ​​self-evolution technology in claim 1, it is characterized in that: The key indicators of computing power indicator management include the number of users, model size, model accuracy, inference speed, memory occupancy, computing power, video memory size, video memory bit width, heat dissipation form, and power consumption.

3. According to the method for evaluating the resource of the inference business model of the AI ​​self-evolution technology in claim 1, it is characterized in that: The computing power indicator management cleans, denoises and standardizes the input demand data; Perform feature extraction by building a self-evolving neural network model to automatically extract features that are closely related to the performance of the inference business model from the preprocessed data; The resource evaluation step evaluates the resource consumption of the inference business model based on the extracted features and generates an evaluation result.

4. The method for evaluating resource of inference business model of AI self-evolution technology according to claim 1 or 3, characterized in that: The resource assessment includes a combination of computing power resources, memory resources, bandwidth resources, storage resources, and budget. Through the reasonable allocation of these resource solutions, it is ensured that the resource supply meets the normal operation of the business.

5. The method for evaluating resource of inference business model of AI self-evolution technology according to claim 1 or 3, characterized in that: The resource assessment uses multi-factor aggregation and classification data analysis techniques and algorithms to process data and extract valuable information from them.

6. The method for evaluating resource of inference business model of AI self-evolution technology according to claim 1 or 3, characterized in that: The resource evaluation supports custom indicators, allowing users to adjust the evaluation criteria according to their own needs and preferences.

7. According to claim 1, a method for evaluating resources of an inference business model of AI self-evolution technology is characterized in that: The result output includes capability comparison, cost comparison, and recommended solutions of using different accelerator cards.

8. A reasoning business model resource evaluation system based on AI self-evolution technology, characterized in that: include: The computing power index management module is used to manage and maintain key indicators; Resource evaluation module, used to evaluate the required xPU resources for output based on input demand indicators and historical business scenario data; The result optimization module is used to adjust and optimize the evaluation results when the requirements are not met or the evaluation deviation is large during the resource evaluation process. It identifies and analyzes the unreasonableness of the input demand data based on fault tolerance and gives optimized results and suggestions based on empirical data. The result output module is used to display and output the results of resource evaluation processing or optimization in the form of intuitive charts; The historical result archiving module is used to archive and save each resource assessment result; The system implements inference business model resource evaluation through the method described in any one of claims 1 to 7.

9. A resource evaluation device for inference business model based on AI self-evolution technology, characterized in that: include: at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is used to call the machine-readable program to implement the method described in any one of claims 1 to 7.

10. A computer-readable medium, characterized in that The computer readable medium stores computer instructions, which, when executed by a processor, cause the processor to execute the method according to any one of claims 1 to 7.