Large model computing power sharing method
By monitoring and optimizing the data of large-model computing power resources, the sharing of large-model computing power resources is solved, the problems of low resource utilization efficiency and high cost are improved, the stability and utilization of resources are improved, and sustainable development is promoted.
Patent Information
- Application Number
- CN202510166560.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, the use of large-model computing resources is inefficient, resulting in waste of resources, and users need to register multiple accounts to use models of different manufacturers, resulting in high costs.
By monitoring the model call situation of large model manufacturers, obtaining computing resource data, using weight coefficients for weight calculation and optimization, performing resource priority sorting and actual weight analysis, and finally allocating resource response requests to realize computing power sharing of large model.
It realizes efficient sharing of large-model computing power resources, reduces user usage costs, improves resource stability and reliability, improves resource utilization, reduces energy consumption, and promotes sustainable development.
Smart Images

Figure CN120045329A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the technical field of resource scheduling, and particularly to a method for sharing the computing power of large models. Background Art
[0002] In the era of large models, subscribing to a certain model is usually a way to provide large model services. However, as the number of large model manufacturers increases and the number of models provided by each large model manufacturer grows, the single subscription model monopolizes the computing power resources and cannot efficiently utilize large models, even causing waste of large model computing power resources. Moreover, if using models from different manufacturers, users often need to register different accounts, resulting in high costs. Summary of the Invention
[0003] Aiming at the above deficiencies in the prior art, a method for sharing the computing power of large models provided by the present invention solves the problems of low utilization efficiency of computing power resources and high costs for users to use large models.
[0004] To achieve the above invention purpose, the technical solution adopted by the present invention is: A method for sharing the computing power of large models, including: S1: Monitor the model call situation of large model manufacturers to obtain large model computing power resource data; S2: Based on the large model computing power resource data, use weight coefficients to calculate weights to obtain a weight allocation result; S3: Based on the weight allocation result, perform priority sorting on the large model computing power resource data to obtain a resource index result; S4: Based on the resource index result, analyze the resource response time and status code to obtain the actual resource weight; S5: Based on the actual resource weight, allocate resource response requests to obtain a large model computing power sharing result, and complete the sharing of large model computing power.
[0005] Further, the S2 includes: Based on the large model computing power resource data, construct an initial resource weight; Use the objective function to optimize the weight coefficients to obtain the target weight coefficients; Use the target weight coefficients to adjust the initial resource weight to obtain a weight optimization result; Perform multiple linear regression on the weight optimization result, and use the minimization of the loss function for linear regression estimation to obtain the weight allocation result.
[0006] Further, the expression of the target weight coefficient is: ; ; ; Among them, represents the target weight coefficient of the response time, represents the weight coefficient of the response time, represents the learning rate, represents the objective function, represents the target weight coefficient of the status code, represents the weight coefficient of the status code, represents the average response time, represents the hyperparameter, represents the error rate.
[0007] Furthermore, the expression of the weight optimization result is: ; Among them, represents the weight optimization result, represents the weight coefficient of the response time, represents the response time of each resource, represents the weight coefficient of the status code, represents the status code of each resource.
[0008] The beneficial effects of the present invention are as follows: By using the large model computing power sharing method, the large model computing power sharing result is obtained. (1) Sharing computing power resources can greatly save the usage cost of users; (2) Improve the stability and reliability of users' access to large model resources. Through the sharing mode, it does not depend on a single computing power resource, and can effectively ensure the stability of users' use in the case of invalidation of a single resource; (3) Improve the utilization rate of the entire social resources, reduce energy consumption, and promote sustainable development. Through the sharing mode, it can ensure that the computing power resources are always available and not idle, ensure the fairness and effectiveness of the distribution of computing power resources, and reduce energy consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] This specification will be further described in the form of exemplary embodiments, and these exemplary embodiments will be described in detail through the drawings. These embodiments are not restrictive. In these embodiments, the same numbers represent the same structures, where: Figure 1 is an exemplary flowchart of a method for sharing large model computing power according to some embodiments of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0010] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
[0011] Embodiment Figure 1 is an exemplary flowchart of a method for sharing large model computing power according to some embodiments of this specification. As Figure 1 shown, the process includes the following steps. In some embodiments, the process can be executed by a processor.
[0012] S1: Monitor the model call situation of the large model manufacturers and obtain the large model computing power resource data.
[0013] The large model computing power resource data is the computing power resource data provided by the large model to users. For example, the large model computing power resource data may include the number of manufacturers, the number of models used by the manufacturers, the resource quantity of the models, the response time of each resource of the manufacturers, and the status code of each resource of the manufacturers.
[0014] In some embodiments, the processor can share the large model computing power with the end users through the AI cloud box device. The end users can access the privatized AI assistant application based on edge computing built in the AI cloud box, and can save the generated data during the conversation process and the personal sensitive data of the users during the registration process inside the AI cloud box, protecting personal data and ensuring privacy security. And through the AI cloud box, it isolates the personal data from the large model to play the role of a privacy sandbox. When using, the AI cloud box accesses the large model API interface provided by the opendatsky artificial intelligence open platform to access the large model computing power API gateway, and then coordinates the computing power resources of the large model computing power resource pool through the API gateway. Through the independently developed resource automatic allocation algorithm built in the large model SDK software upstream of the gateway, it can reliably and effectively dynamically access all the computing power in the large model computing power resource pool, providing a secure shared large model computing power service for the end users. The end users can dynamically and flexibly choose the large model to be used according to their own needs. For example, there are N manufacturers, each manufacturer has M models behind it, and each model has M1 resources. First, a load balancing algorithm needs to be designed to design the weights of each resource. The heavier the weight, the higher the priority of use, to ensure a stable response when requesting a certain model. Assume that each time a certain resource is requested, the resource will return the time T of the resource request and the resource status code S. The longer T is, the less timely the resource response is, and the lower the weight. When the resource status code S is -1, it means there is an error in the request, and the weight is negative, indicating that the resource is unavailable. When the status code S is 1, it means the request is normal and the resource is available.
[0015] S2: Based on the large model computing power resource data, use the weight coefficients to calculate weights and obtain a weight allocation result.
[0016] The weight allocation result is the result reflecting the weights corresponding to each computing power resource.
[0017] In some embodiments, the processor may implement S2 based on the following steps: Based on the large model computing power resource data, construct an initial resource weight; use the objective function to optimize the weight coefficients to obtain target weight coefficients; use the target weight coefficients to adjust the initial resource weight to obtain a weight optimization result; perform multiple linear regression on the weight optimization result, and use the minimization of the loss function for linear regression estimation to obtain the weight allocation result.
[0018] The initial resource weight is the data reflecting the initial weight situation of each resource of the manufacturer.
[0019] In some embodiments, the processor may calculate based on the large model computing power resource data to obtain the initial resource weight: .
[0020] The target weight coefficient is the weight coefficient after optimizing the weight coefficients.
[0021] In some embodiments, the processor may perform a preliminary screening of the resources based on the status code of each resource. For example, when When , it indicates that the resource is unavailable; when When .
[0022] In some embodiments, the processor may, when a resource is requested, update the weight according to and . For example, the processor may, based on a response time threshold, when exceeds the threshold, increase the value of , and when becomes -1, the weight is adjusted to a negative number, indicating that the resource is unavailable.
[0023] In some embodiments, the processor may construct an objective function using corresponding metrics. For example, the objective function metrics may include average response time: the smaller the better, resource utilization rate: the higher the better, error rate: the lower the better.
[0024] In some embodiments, the processor may train the weight coefficients based on the objective function, using historical data and initial parameters to obtain the target weight coefficients. For example, the processor may use the weight coefficient of the current response time and the weight coefficient of the current status code to calculate the corresponding resource weight and record the performance metrics of the system (such as average response time, error rate, etc.), calculate the objective function based on the performance metrics, and use the gradient descent method to update and , when the objective function converges or reaches the maximum number of iterations, the training is completed, and the target weight coefficients are obtained.
[0025] In some embodiments, the expression of the target weight coefficients may be: ; ; ; where, represents the target weight coefficient of the response time, represents the weight coefficient of the response time, represents the learning rate, represents the objective function, represents the target weight coefficient of the status code, represents the weight coefficient of the status code, represents the average response time, represents the hyperparameter, represents the error rate.
[0026] In some embodiments, the processor may evaluate the performance of the optimized and using the validation dataset and validation metrics. For example, the validation metrics may include average response time, resource utilization, and error rate.
[0027] In some embodiments, when the validation result does not meet the requirements, the processor may adjust the hyperparameters in the objective function, change the optimization algorithm (such as Adam or L-BFGS), and add more features (such as historical usage frequency, load, etc. of the resource), return to step 2 to retrain, and obtain the target weight coefficients.
[0028] The weight optimization result is the result of recalculating the resource weight using the target weight coefficients.
[0029] In some embodiments, the expression of the weight optimization result may be: ; where, represents the weight optimization result, represents the target weight coefficient of the response time, Indicates the response time of each resource, Indicates the target weight coefficient of the status code, Indicates the status code of each resource.
[0030] In some embodiments, the processor may perform multiple linear regression on the weight optimization result, use the minimization of the loss function for linear regression estimation, and obtain the weight allocation result.
[0031] S3: Based on the weight allocation result, prioritize the large model computing power resource data to obtain a resource index result.
[0032] The resource index result is a result reflecting the resource index priority situation.
[0033] In some embodiments, the processor may query the weight value of the resource based on the weight allocation result, sort it from largest to smallest according to the weight size, obtain and return the resource index with the highest weight, and obtain the resource index result.
[0034] S4: Based on the resource index result, analyze the resource response time and status code to obtain the actual resource weight.
[0035] The actual resource weight is data reflecting the actual resource weight situation.
[0036] In some embodiments, the processor may schedule the resource with the highest weight based on the resource index result, obtain the response time and status code of the resource this time, calculate the actual weight this time, and save the weight this time in the database to obtain the actual resource weight.
[0037] S5: Based on the actual resource weight, allocate resource response requests to obtain the large model computing power sharing result, and complete the sharing of the large model computing power.
[0038] The large model computing power sharing result is a result reflecting the resource scheduling and sharing situation of the large model computing power.
[0039] In some embodiments, the processor may allocate resource response requests based on the actual resource weight to obtain the large model computing power sharing result.
[0040] In some embodiments of this specification, the large model computing power sharing method is utilized to obtain the large model computing power sharing result. (1) Sharing computing power resources can greatly save the usage cost of users; (2) improve the stability and reliability of users' access to large model resources. Through the sharing mode, it does not rely on a single computing power resource, and can effectively ensure the stability of users' usage in the case of invalidation of a single resource; (3) improve the utilization rate of the entire social resources, reduce energy consumption, and promote sustainable development. Through the sharing mode, it can ensure that the computing power resources are always available and not idle, ensure the fair and effective allocation of computing power resources, and reduce energy consumption.
Claims
1. A method for sharing computing power of a large model, characterized in that: include: S1: Monitor the model calling status of large model manufacturers and obtain large model computing resource data; S2: Based on the large model computing power resource data, weight calculation is performed using the weight coefficient to obtain a weight distribution result; S3: Based on the weight allocation result, the large model computing power resource data is prioritized to obtain a resource index result; S4: Based on the resource index result, analyze the resource response time and status code to obtain the actual weight of the resource; S5: Based on the actual weight of the resources, resources are allocated to respond to the request, and the large model computing power sharing result is obtained to complete the sharing of the large model computing power.
2. The method for sharing computing power of a large model according to claim 1, characterized in that: The S2 includes: Based on the large model computing power resource data, construct initial resource weights; Using the objective function, the weight coefficient is optimized to obtain the target weight coefficient; Using the target weight coefficient, adjusting the initial resource weight to obtain a weight optimization result; The weight optimization result is subjected to multivariate linear regression, and linear regression estimation is performed using a minimization loss function to obtain a weight distribution result.
3. The method for sharing computing power of a large model according to claim 2, characterized in that: The expression of the target weight coefficient is: ; ; ; in, represents the target weight coefficient of response time, Represents the weight coefficient of response time, represents the learning rate, represents the objective function, Indicates the target weight coefficient of the status code, Indicates the weight coefficient of the status code, represents the average response time, represents the hyperparameter, Indicates the error rate.
4. The method for sharing computing power of a large model according to claim 2, characterized in that: The expression of the weight optimization result is: ; in, represents the weight optimization result, Represents the weight coefficient of response time, Indicates the response time of each resource, Indicates the weight coefficient of the status code, Indicates the status code of each resource.