A Method for Large Model Resource Allocation and Transaction Settlement Oriented to Application Calls
By counting and storing the participation calculation data of the large model and allocating computing resources in combination with the application's call frequency, the problem of providing better service quality when the configuration resources are limited is solved, and the rational allocation of resources and the improvement of service quality of the large model is achieved.
Patent Information
- Application Number
- CN202410092969.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-22
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-01-22
AI Technical Summary
When a large model is called by an application, how to provide better quality of service when configuration resources are limited, especially under the high-performance requirements of computing resources and storage resources.
By counting the participating calculation data in the target model, counting the total number of references for each application, and classifying the model data according to these count values, and providing different computing resources according to the frequency of the application's call.
It realizes the provision of different storage services and computing resources when different applications call large models, ensuring that frequently called large model applications can obtain better storage service quality, avoid idle and waste of resources, and improve the overall service quality of the large model.
Smart Images

Figure CN117909078B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence technology, and particularly relates to a method for large model resource allocation and transaction settlement for application program calls. Background Art
[0002] A large model can be called by an application program as a service. When being called by an application program, the model requires certain computing resources and storage resources to complete the application program call task. The characteristic of a large model is a large number of parameters, that is, a large amount of data. Therefore, the cost of configuration resources required for both computing and storage is not low.
[0003] Currently, in order to provide better service quality when being called by an application program, in terms of computing, a large model requires high-performance computing resources in an inference application program; in terms of storage, high-performance and low-latency storage resources can provide better service quality compared to low-performance and high-latency storage resources. However, the use of high-performance computing resources and high-performance and low-latency storage resources will lead to an increase in the corresponding computing resource cost and storage resource cost.
[0004] Therefore, a method needs to be proposed to reasonably allocate the resources owned by a large model for application program calls, so as to provide better service quality under limited configuration resources. Summary of the Invention
[0005] The embodiments of this application provide a method for large model resource allocation and transaction settlement for application program calls, which can solve the problem of requiring higher configuration resources when improving the model service quality.
[0006] In a first aspect, the embodiments of this application provide a method for large model resource allocation for application program calls, including:
[0007] When a target model is called by an application program, perform reference counting on the calculation data participating in the target model to obtain a reference count value, and record the application program making the call to obtain a recording result; wherein, the calculation data participating in the calculation is the data participating in the feedforward neural network calculation in the target model;
[0008] According to the reference count value and the recording result, respectively count the total reference times of the calculation data corresponding to each application program;
[0009] According to the total reference times of the calculation data corresponding to each application program, classify all model data of the target model and store them in different storage resources respectively;
[0010] According to the total call count of each application program calling the large model, classify the application programs and provide different computing resources respectively.
[0011] When the above method is used in an application to call a large model, by counting the data participating in the calculation, the storage service provided by the model when called by the application is determined. By recording the application that makes the call, and respectively counting the total reference times of the data participating in the calculation corresponding to each application based on the reference count value and the recording result, the data reading requirements of the large model when different applications call it are respectively counted, and the model data is classified accordingly. The model data that different applications need to reference when calling the large model is stored in different storage resources, so as to provide different storage services when different programs call, enabling applications that frequently call the large model to obtain better storage service quality.
[0012] Moreover, the above method realizes the reasonable allocation of resources by providing different computing resources when different applications call the large model according to the total call count of the application calling the large model, avoids the idle and waste of resources, and improves the overall service quality of the large model when facing application calls.
[0013] In a possible implementation manner of the first aspect, the step of classifying all model data of the target model according to the total reference times of the data participating in the calculation corresponding to each application and storing them in different storage resources respectively includes:
[0014] Classify all model data according to the total reference times of the data participating in the calculation corresponding to each application to obtain classified data; the classified data includes first-class data and second-class data, and the total reference times of the data participating in the calculation corresponding to the application corresponding to any model data in the first-class data is greater than the total reference times of the data participating in the calculation corresponding to the application corresponding to any model data in the second-class data;
[0015] Store the first-class data in the first-class storage resource; store the second-class data in the second-class storage resource;
[0016] Among them, the reading performance of the first-class storage resource is higher than that of the second-class storage resource, and the network latency of the first-class storage resource is lower than that of the second-class storage resource.
[0017] The above method classifies the model data corresponding to the application according to the total reference count of the calculation data participated by the application, and stores all the model data corresponding to the application with a high total reference count in the storage resource with higher performance and lower latency, and stores all the model data corresponding to the application with a low total reference count in the storage resource with lower performance and higher latency. When the model references the data, it can ensure that the frequently called application can read at high speed when calling the model, meeting the reading speed requirements of the model data in most call cases. And the model data corresponding to the application that does not need to be frequently called is stored in the storage resource with lower performance and higher latency, reducing the data storage cost while meeting the reading requirements of the model data.
[0018] In a possible implementation manner of the first aspect, the classified data further includes a third type of data. The total reference count of the calculation data participated by the application corresponding to any model data in the third type of data is less than the total reference count of the calculation data participated by the application corresponding to any model data in the second type of data. The method further includes:
[0019] Storing the third type of data in the third type of storage resource; the reading performance of the second type of storage resource is higher than that of the third type of storage resource, and the network latency of the second type of storage resource is lower than the network latency of the third type of data;
[0020] When the storage duration of the third type of data is greater than the second preset duration, the third type of data is discarded, and the second preset duration is greater than the first preset duration.
[0021] The above method further finely divides the model data according to the total reference count of the calculation data participated by the application, and discards the third type of data after the third type of data corresponding to the application with a relatively small total reference count is stored for a long time, further reducing the storage cost.
[0022] In a possible implementation manner of the first aspect, the above method further includes:
[0023] Storing the unused reference data as an object at the end of the period. The unused reference data is the model data that is not referenced when the target model is called by the application;
[0024] When the storage duration of the unused reference data as an object is greater than the second preset duration, the unused reference data is discarded.
[0025] The above method processes redundant model data by storing uncalled reference data as objects. Taking advantage of the characteristics that object storage does not require maintaining complex directory and file structures and does not require purchasing expensive storage devices, it reduces the storage cost of this part of the model data. After the uncalled reference data has been stored as an object for a long time, the data is discarded to optimize the model, that is, to process the unimportant data in the large model, reducing the size of the model while avoiding affecting the model accuracy and precision.
[0026] In a possible implementation of the first aspect, the steps of classifying application programs according to the total call count of each application program's invocation of the large model and respectively providing different computing resources include:
[0027] When the target model is invoked by an application program, the application program is counted for invocations to obtain the total call count of the application program;
[0028] Classify the application programs according to the call count of the application programs to obtain a classification result; the classification result includes first-class application programs and second-class application programs, and the total call count of any application program in the first-class application programs is greater than that of any application program in the second-class application programs;
[0029] Allocate computing resources according to the classification result, so that the computing power provided by the computing resources allocated to the first-class application programs is greater than the computing power provided by the computing resources allocated to the second-class application programs.
[0030] The above method classifies application programs according to the total call count of the application programs and provides different service qualities for different application programs. For the first-class application programs, they are called more frequently. Relatively speaking, the call frequency is also high. Providing them with higher computing power can meet their large computational requirements when the model is frequently called, reduce the latency when calling data, and generally improve the service quality. For the second-class application programs with relatively lower usage times, allocating them second-class computing resources with less computing power can, under the same computing resources, enable the model inference quality to reach a high level in most cases.
[0031] In a possible implementation of the first aspect, before allocating computing resources according to the classification result, it includes:
[0032] According to the classification result, add a first-class mark to the call entry corresponding to the first-class application programs and add a second-class mark to the call entry corresponding to the second-class application programs;
[0033] Correspondingly, allocating computing resources according to the classification result includes:
[0034] Provide corresponding computing resources to the first type of application and the second type of application according to the first type of tag and the second type of tag respectively.
[0035] By adding tags to the call entry of the application, the above method enables direct service provision according to the tags when providing services to the application, avoiding the problem of recalculation required for each call, improving convenience, and facilitating the direct identification and determination of the computing power required by the computing power provider.
[0036] In a possible implementation manner of the first aspect, the above method further includes:
[0037] Take off the shelf the application with a total call count of zero.
[0038] In a second aspect, an embodiment of the present application provides a large model resource trading and settlement method for application calls, including:
[0039] When the target model is called by the application, perform a reference count on the data participating in the calculation in the target model to obtain a reference count value, and record the application making the call to obtain a recording result;
[0040] Among them, the data participating in the calculation is the data participating in the feedforward neural network calculation in the target model;
[0041] According to the reference count value and the recording result, respectively count the total reference times of the data participating in the calculation corresponding to each application;
[0042] Obtain the call duration of the application calling the target model;
[0043] In response to a resource settlement request, use the call duration and the total reference times as the settlement basis for resource settlement.
[0044] By recording the referenced model data and the application when the large model is called, the above method obtains the total reference times of the data participating in the calculation corresponding to the application, quantifies the services provided by the model to the application, and uses the total reference times of the data participating in the calculation corresponding to the application and the call duration as settlement vouchers, combining the service volume provided with the service duration, and further subdividing the services provided from two aspects of the call times and the call duration, improving the credibility and accuracy of the trading and settlement method.
[0045] In a possible implementation manner of the second aspect, the above method further includes:
[0046] Upload the call duration and the total reference times to the consortium blockchain. Description of the Drawings
[0047] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0048] Figure 1 is a schematic flowchart of a large model resource allocation method for application program calls provided by an embodiment of the present application;
[0049] Figure 2 is a schematic diagram of the large model encoding and decoding process;
[0050] Figure 3 is a schematic diagram of the reference counting process provided by an embodiment of the present application;
[0051] Figure 4 is a schematic diagram of the application program classification and marking process provided by an embodiment of the present application;
[0052] Figure 5 is a schematic diagram of the large model resource allocation method for application program calls provided by an embodiment of the present application;
[0053] Figure 6 is a schematic diagram of the computing resource allocation process provided by an embodiment of the present application;
[0054] Figure 7 is a schematic diagram of the storage resource allocation process provided by an embodiment of the present application;
[0055] Figure 8 is a schematic flowchart of a large model resource trading and settlement method for application program calls provided by an embodiment of the present application;
[0056] Figure 9 is a schematic flowchart of a large model resource trading and settlement method for application program calls provided by an embodiment of the present application. Detailed implementation manners
[0057] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0058] It should be understood that, as used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups.
[0059] It should also be understood that the term "and / or" as used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0060] As used in the specification of the present application and the appended claims, the term "if" may be construed, depending on the context, as "when" or "once" or "in response to determining" or "in response to detecting". Similarly, the phrases "if determined" or "if [the described condition or event] is detected" may be construed, depending on the context, as meaning "once determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]".
[0061] In addition, in the description of the specification of the present application and the appended claims, the terms "first", "second", "third", etc. are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance.
[0062] Reference to "one embodiment" or "some embodiments" or the like described in the specification of the present application means that a particular feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.
[0063] Regarding the reduction of model configuration costs, methods such as distillation, quantization, and pruning have been mainly proposed to reduce redundancy and resource occupancy. Among them, pruning techniques can be used to delete redundant parameters in the model, quantization techniques can be used to reduce the precision of parameters in the model, and knowledge distillation techniques can be used to transfer the weights of large models to small models, thereby reducing storage requirements and accelerating calculations. In practice, these methods have been widely applied to large models, leveraging their strengths and avoiding their weaknesses, and are selected for use in specific application tasks. For example, methods such as distillation, quantization, and pruning are applied to comprehension tasks, and the quantization method is applied to generative tasks. However, although these methods can reduce resource requirements to a certain extent, they will also reduce the service quality in specific application scenarios. For example, adopting the quantization method will reduce the precision of model parameters. If a large amount of information is lost during the quantization process, it may lead to a decrease in the quality or resolution of the generated images.
[0064] In a first aspect, referring to Figure 1 , this embodiment provides a large model resource allocation method for application calls, including:
[0065] Step S102: When the target model is called by an application, perform a reference count on the calculation data involved in the target model to obtain a reference count value, and record the application that makes the call to obtain a recording result; wherein, the calculation data involved is the data involved in the feed-forward neural network calculation in the target model;
[0066] Step S104: According to the reference count value and the recording result, respectively count the total reference times of the calculation data corresponding to each application.
[0067] Step S106: Classify all the model data of the target model according to the total reference times of the calculation data corresponding to each application, and store them in different storage resources respectively.
[0068] Step S108: Classify the applications according to the total call count of each application calling the large model, and provide different computing resources respectively.
[0069] In this embodiment, the target model is the large model provided for application programs to call, including the model during training and the model obtained after training for subsequent inference. It is a model with a huge number of parameters, generally more than ten billion parameters, including but not limited to large language models, general large models, vertical large models, AIGC (AI-Generated Content), and specific implementations of more large models in the future. The large model is not limited to a single model file and can be a collection of multiple model files, jointly providing inference application services externally. For example, the inference program in the large model system loads multiple model files and integrates large model data to provide services externally; the training program in the large model system loads multiple model files through multiple computing nodes or NPUs (Neural Processing Units), or multiple model files obtained by splitting a single model file, performs parallel computing, and aggregates weight data to save as a large model.
[0070] The data in the large model, that is, the model parameters, is a collection of tensor data related to the model structure. For example, a transformer model using the decoder-only framework consists of multiple identical layers, and each layer includes a self-attention block and a multi-layer perceptron (MLP) block. The model parameters of the self-attention block are the weight matrices and biases of Q, K, V, and the weight matrix and bias of the output O; the multi-layer perceptron block consists of 2 linear layers, including the weight matrices and biases of these linear layers. Each of the self-attention block and the multi-layer perceptron block has a layer normalization, which contains 2 trainable model parameters. In addition, the word embedding matrix also has a relatively large number of parameters.
[0071] The data participating in the calculation includes the data calculated by the feed-forward neural network (FFN), that is, the prompt words, the feed-forward neural network parameters, and the generated results. In the Transformer architecture, the MLP is actually the FFN. The feed-forward neural network parameters here are the weight matrices and biases of the MLP linear layers; the prompt words are used as the input of the feed-forward neural network in the form of vector data; the generated results are used as the output of the feed-forward neural network in the form of vector data.
[0072] See Figure 2, In the inference process of a general large - data model, the prompt is used as the input. Through m times of encoder calculations and n times of decoder calculations, the inference data is obtained. For large models, especially the Transformer (a deep - learning network based on the attention mechanism) large model, its model architecture includes Feed - Forward Neural Network (FFN) calculations in both the encoder and the decoder. The vector data of each calculation result describes the process of prompt prediction results. When the data of the target model is called by an application, the calculation data participating in the FFN calculation in the target model is recorded, and the reference count of the model data is associated with the actual application situation of the model data in the target model, improving the correlation between the total reference count obtained and the data reading requirements of the target model when it is called by the application.
[0073] Reference counting refers to counting the number of times the calculation data in the model is referenced. When the data is accessed, a count is made, and each time it is used, the count is incremented by 1. In an optional implementation, when the data is accessed, the data can be recorded, and the access count of the data can be determined, and then the reference count value of the data can be determined.
[0074] Alternatively, after the data undergoes FFN calculation, the data participating in the FFN calculation and the resulting data are recorded. A data can be recorded multiple times, and finally the reference count value is obtained based on the number of records. Exemplarily, see the attached Figure 3 specification. When the FFN vector data n is referenced and undergoes FFN calculation, the reference count of the FFN vector data n changes from Cn to Cn + 1, while the reference counts of other vector data remain unchanged.
[0075] In an optional implementation, the process of the application calling the large - model prompt results can be traced and recorded through a vector database. When each application calls the large - model, the calculation data participating in the calculation and the application making the reference are recorded to determine the situation of the data in the target model being referenced.
[0076] For step S104, the total reference count of the calculation data corresponding to the application refers to the sum of the number of times the calculation data corresponding to the application is referenced when the application calls the large - model. Exemplarily, see Table 1. When application A calls the large - model, the large - model references vector data a 10 times, vector data c 7 times, vector data d 15 times, and vector data g 7 times. Then the total reference count of the calculation data corresponding to application A is 39.
[0077] Optionally, when determining the total reference count of the calculation data participated by the application according to the recording result, the calculation data participated by each application can be determined according to the recording result; the total reference count of the calculation data participated by the application can be determined according to the sum of the reference count values of each calculation data participated by the application.
[0078] Exemplarily, when the data participates in the FFN calculation, the calculation data participated is counted, and the application that calls the large model at this time is recorded. After a period of time, the sum of the reference count values of the model data called by the same application is calculated to obtain the total reference count of the calculation data participated by the application. Optionally, when the large model finishes calling one application and is called by another application, the counting of the model data starts again. During a period of time, if the large model is called by the same application multiple times, the total reference count of the application is obtained according to the sum of the reference count values of the calculation data participated corresponding to the multiple calls of the application.
[0079] For step S106, in an optional implementation manner, the step of classifying all the model data of the target model according to the total reference count of the calculation data participated by each application and storing them in different storage resources respectively includes:
[0080] Classify all the model data according to the total reference count of the calculation data participated by each application to obtain classified data; the classified data includes first-class data and second-class data, and the total reference count of the calculation data participated by the application corresponding to any model data in the first-class data is greater than the total reference count of the calculation data participated by the application corresponding to any model data in the second-class data;
[0081] Store the first-class data in the first-class storage resource; store the second-class data in the second-class storage resource;
[0082] Wherein, the reading performance of the first-class storage resource is higher than that of the second-class storage resource, and the network latency of the first-class storage resource is lower than that of the second-class storage resource.
[0083] It should be noted that in this implementation manner, the first-class data and the second-class data are only two different types of data, not specifically referring to a certain type of data, and the first storage resource and the second storage resource are only two storage resources with different performances and latencies, not specifically referring to a certain type of storage resource.
[0084] In an optional implementation manner, it can be allocated according to the total reference count of the calculation data participated by all applications and the total amount of storage resources, so that the total reference count of the calculation data participated by the application and the total amount of the allocated model configuration resources are positively correlated.
[0085] Alternatively, according to the total reference count of the data involved in each application for calculation, the data involved in the application for calculation can be stored in different storage resources. Multiple intervals of the total reference count of the data can be set, and different storage resources can be allocated to the model data corresponding to the applications in different intervals; or, according to the total reference count of all recorded applications, the applications can be classified, and their corresponding model data can be classified and stored in different storage resources.
[0086] Exemplarily, referring to Table 1, the total reference counts of the data of Application A, Application B, Application C, Application D, Application E, and Application F are 39, 29, 9, 14, 16, and 22 respectively. Then, the median 19 of this group of data can be used as the first preset value. Application A, Application B, and Application F are the first type of applications, and Application C, Application D, and Application E are the second type of applications. The vector data a, vector data b, vector data c, vector data d, vector data f, and vector data g corresponding to Application A, Application B, and Application F are stored in the first type of storage resources, and the vector data a, vector data b, vector data c, vector data d, and vector data f corresponding to Application C, Application D, and Application E are stored in the second type of storage resources.
[0087] The storage resource refers to the hardware and software infrastructure for storing and processing large-scale data and models. These resources include but are not limited to storage devices, storage networks, storage management systems, distributed storage systems, etc. The access speed of the first type of storage resource is greater than that of the second type of storage resource, its performance is higher than that of the second type of storage resource, and the latency is lower than that of the second type of storage resource. Among them, high performance and low latency of the storage resource are relative concepts. For example, Storage Resource A: is configured with a solid-state drive with relatively high read and write performance, and at the same time, when classifying and storing, technologies such as Smart NIC (Smart Network Interface Card) need to be used to reduce communication overhead; Storage Resource B: is configured with a mechanical hard drive with relatively low read and write performance, and at the same time, when performing distributed storage, there is no need to consider reducing communication overhead. By comparing the two, Storage Resource A belongs to the storage resource with high performance and low latency, and the cost of configuration and deployment is also relatively high. Among them, the read performance of the storage resource is related to the size of the I / O (read / write) request, IOPS (Input / Output Per Second, the number of read and write operations per second), bandwidth, and throughput. The storage performance of the storage resource can be quantified from the above aspects.
[0088] In the above embodiments, by classifying the model data corresponding to the application according to the total number of references of the calculation data participated by the application, and storing all the model data corresponding to the application with a high total number of references in the storage resource with higher performance and lower latency, and storing all the model data corresponding to the application with a low total number of references in the storage resource with lower performance and higher latency, when the model references the data, it can ensure that the frequently called application can read at high speed when calling the model, meeting the reading speed requirements of the model data in most call cases. And storing the model data corresponding to the application that does not need to be frequently called in the storage resource with lower performance and higher latency, while meeting the reading requirements of the model data, reducing the cost of data storage.
[0089] In a possible implementation manner, the step of classifying all the model data of the target model according to the total number of references of the calculation data participated by each application further includes:
[0090] Taking the first preset duration as a period, classifying the data of the target model periodically.
[0091] In a possible implementation manner, the classified data further includes a third type of data. The total number of references of the calculation data participated by the application corresponding to any model data in the third type of data is less than that of the application corresponding to any model data in the second type of data. The above method further includes:
[0092] Storing the third type of data in the third type of storage resource; the reading performance of the second type of storage resource is higher than that of the third type of storage resource, and the network latency of the second type of storage resource is lower than the network latency of the third type of data;
[0093] When the storage duration of the third type of data is greater than the second preset duration, discarding the third type of data, and the second preset duration is greater than the first preset duration.
[0094] Exemplarily, the first preset duration is one month, and the second preset duration is one quarter; or, the first preset duration is one quarter, and the second preset duration is one year.
[0095] In an alternative implementation manner, a large model resource allocation method for application program calls provided in this embodiment further includes:
[0096] Storing the uninvoked reference data as an object at the end of the period. The uninvoked reference data is the model data that is not referenced when the target model is called by the application program;
[0097] When the storage duration of the uninvoked reference data as an object is greater than the second preset duration, discarding the uninvoked reference data.
[0098] In this embodiment, object storage is performed on the unused reference data, and redundant model data is processed. By taking advantage of the characteristics of object storage that do not require maintaining a complex directory and file structure and do not require purchasing expensive storage devices, the storage cost of this part of the model data is reduced. After the unused reference data has been stored as an object for a long time, the data is discarded and the model is optimized, that is, the unimportant data in the large model is processed, reducing the size of the model while avoiding affecting the model accuracy and precision.
[0099] In a specific embodiment, by way of example, refer to Figure 4 , during the process of analyzing the database called by the APP, the total call counts of all APPs are analyzed every once in a while to mark the call entry of the APPs as Hot APP, Normal APP, and Cold APP in sequence. The vector data called by the application program is uniformly marked. That is, if the total reference times of the data reach the application program corresponding to Hot Storage, all the participating calculation data is marked as Hot Storage. For the vector data of Hot Storage, since it is often used, high-performance and low-latency storage resources are required, while for Normal Storage and Cold Storage, storage resources with lower performance and higher latency can be selected in turn. As Figure 5 and Figure 6 shown, for the vector data marked as Hot Storage, high-performance and low-latency storage resources are configured. For the vector data marked as Normal Storage and Cold Storage, storage resources with lower performance and higher latency are selected. For the unmarked vector data, aging processing is performed, and the storage method can be object storage to further reduce the cost of storage resources. This method of automatically optimizing the large model using a vector database can not only process the redundant data of the large model but also reasonably allocate storage resources. That is to say, this solution provides different levels of data storage service quality for different APPs, which can effectively reduce the cost of storage resources.
[0100] The beneficial effects of this embodiment are as follows:
[0101] When an application calls a large model, by counting the data participating in the calculation, the storage service provided by the model when called by the application is determined. By recording the calling application and respectively counting the total reference times of the data participating in the calculation corresponding to each application according to the reference count value and the recording result, the data reading requirements of the large model when different applications call it are respectively counted, and the model data is classified accordingly. The model data that different applications need to reference when calling the large model is stored in different storage resources, so as to provide different storage services when different programs call, enabling applications that frequently call the large model to obtain better storage service quality.
[0102] Moreover, in this embodiment, by according to the total call count of the application calling the large model, different computing resources are provided when different applications call the large model, realizing the reasonable allocation of resources, avoiding the idle and waste of resources, and improving the overall service quality of the large model when facing application calls.
[0103] According to the above embodiment, in another embodiment:
[0104] The steps of classifying the applications according to the total call count of each application calling the large model and respectively providing different computing resources include:
[0105] When the target model is called by an application, the application is counted for the call count to obtain the total call count of the application;
[0106] Classify the applications according to the call count of the applications to obtain a classification result; the classification result includes a first type of application and a second type of application, and the total call count of any application in the first type of application is greater than the total call count of any application in the second type of application;
[0107] Allocate the computing resources according to the classification result, so that the computing power provided by the computing resources allocated to the first type of application is greater than the computing power provided by the computing resources allocated to the second type of application.
[0108] Computing resources include a central processing unit, memory, hard disk, GPU (Graphic Processing Unit), and the higher the computing power of the GPU graphics card, the faster the training and inference speed, the better the service quality of the provided scenario application, and the greater the corresponding computing resource cost. Storage resources refer to the hardware and software infrastructure for storing and processing large-scale data and models, and these resources include but are not limited to storage devices, storage networks, storage management systems, distributed storage systems, etc.
[0109] The performance of computing resources is a relative concept. For example, computing resource A: a graphics card array with high computing performance (a computing cluster composed of multiple graphics cards, which accelerate the model training and inference processes through parallel computing), providing high computing power, such as configuring 64 NVIDIA Tesla A100 graphics cards; computing resource B: a graphics card with low computing performance, providing low computing power, such as configuring 1 NVIDIA Tesla V100 graphics card. By comparison, computing resource A belongs to high-performance computing resources, and the configuration and deployment cost is also relatively high.
[0110] For the allocation of computing resources, computing resources such as the central processing unit, memory, hard disk, and graphics card (GPU) can be aggregated and, through software, form a virtual and infinitely scalable "computing power resource pool", and then, according to the total reference count of the computing data corresponding to the application program, perform dynamic allocation. For storage resources, the performance of the storage resources and the corresponding network overhead can be quantified and then allocated.
[0111] In the above implementation, by classifying application programs according to the total call count of the application program and providing different service qualities for different application programs. For the first type of application program, the number of calls is relatively large, and relatively speaking, the call frequency is also high. Providing it with high computing power can meet its demand for large computing volume when the model is frequently called, reduce the latency when calling data, and generally improve the service quality. For the second type of application program with relatively low usage times, allocating the second type of computing resources with less computing power can, under the same computing resources, make the model inference quality reach a high level in most cases.
[0112] In a specific implementation, application programs can be classified and divided according to a preset multiple total call count intervals. Among them, the total call count intervals can be set by users or technicians themselves, or determined by the total call count of all application programs. For example, referring to the computing data record form (Table 1), the total call counts of application program A, application program B, application program C, application program D, application program E, and application program F are 39, 29, 9, 14, 16, and 22 respectively. Then, the median 19 of this set of data can be used as the interval division point. Application program A, application program B, and application program F are type I application programs, and application program C, application program D, and application program E are type II application programs. Then, provide the first type of computing resources to application program A, application program B, and application program F, and provide the second type of computing resources to application program C, application program D, and application program E.
[0113] Application A Application B Application C Application D Application E Application F Vector data a 10 6 5 7 6 Vector data b 9 6 1 Vector data c 7 7 4 4 Vector data d 15 1 4 Vector data e 4 7 Vector data f 3 10 Vector data g 7 4 2 Total call count 39 29 9 14 16 22
[0114] Table 1
[0115] In an alternative embodiment, the total call count of the application is periodically determined to classify the application.
[0116] In an alternative embodiment, after classifying the application, the call entry of the application can be marked according to the classification result. Add a first type of mark to the call entry corresponding to the first type of application, and add a second type of mark to the call entry corresponding to the second type of application; provide corresponding computing resources to the first type of application and the second type of application respectively according to the first type of mark and the second type of mark.
[0117] See Figure 7 , provide computing resources with high computing power to the application marked as Hot APP, provide computing resources with relatively lower computing power to the application marked as NormalAPP, provide computing resources with relatively the lowest computing power to the application marked as Cold APP, and take off the shelf the unused APP.
[0118] In the above embodiment, by adding marks to the call entry of the application, when providing services to the application, services can be directly provided according to the marks, avoiding the problem of recalculation required for each call, improving convenience, and facilitating the computing power provider to directly identify and determine the required computing power.
[0119] In an alternative embodiment, for an APP without a mark, it indicates that the application is not used by anyone, and then it enters the process of taking off the shelf to further reduce the cost of computing resources.
[0120] The data resources of the industry large model are not provided to the APP for free. The application program large model data needs to be charged, and the large model data can be used for trading as a kind of data asset. At present, there is a lack of reasonable and credible technical means for the settlement of large model data transactions. This results in a single method for the settlement of fees for using data assets (for example, monthly charging), which cannot support a more precise charging model and brings challenges to the implementation of the large model in application programs.
[0121] According to the above technical problems, see Figure 8 , on the second aspect, the embodiments of the present application provide a large model resource trading settlement method for application program calls, including:
[0122] Step S802: When the target model is called by the application program, perform a reference count on the data participating in the calculation in the target model to obtain a reference count value, and record the application program that makes the call to obtain a recording result, where the data participating in the calculation is the data participating in the feedforward neural network calculation in the target model;
[0123] Step S804: According to the reference count value and the recording result, respectively count the total reference times of each application corresponding to the data participating in the calculation.
[0124] Step S806: Obtain the call duration of the application calling the target model.
[0125] Step S808: In response to the resource settlement request, use the call duration and the total reference times as the settlement basis for resource settlement.
[0126] Among them, the application refers to the application that needs to be settled, and the target model refers to the model that provides services and is called. The call duration can be timed when the call request of the target model is made and end the timing to obtain when the call ends.
[0127] In a specific implementation, upload the usage duration of the storage and computing resources called by the target APP, the total reference count of the model vector data, and the total reference times of the data participating in the calculation corresponding to the target application to the consortium blockchain as a credible voucher for settlement between the model operator and the APP manufacturer. As Figure 9 shown, use the APP resource usage measurement server to collect the usage data of the APP data resources (i.e., the total reference count of the APP referring to the model vector data and the APP call count), the usage data of the APP storage resources, and the usage data of the APP computing resources, and upload these measurement data to the consortium blockchain as a credible voucher for transaction settlement between the model operator and the APP manufacturer.
[0128] The beneficial effect of this embodiment is as follows:
[0129] In this embodiment, when the large model is called, record the referenced model data and the application, obtain the total reference times of the data participating in the calculation corresponding to the application, quantify the services provided by the model to the application, and use the total reference times of the data participating in the calculation corresponding to the application and the call duration as the settlement vouchers, combine the service volume provided with the service duration, and further subdivide the provided services from two aspects of the call times and the call duration, improving the credibility and accuracy of the transaction settlement method.
[0130] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution is prior or posterior. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0131] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in one or more computer-readable storage media. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0132] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0133] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0134] In the embodiments provided in this application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0135] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0136] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A large model resource allocation method for application calls, characterized in that: include: When the target model is called by an application, reference counting is performed on the data involved in the calculation in the target model to obtain the reference count value, and the calling application is recorded to obtain the recording result; wherein the data involved in the calculation is the data involved in the feedforward neural network calculation in the target model, and when the target model ends the call of one application and is called by another application, the reference counting of the data involved in the calculation in the target model is restarted; According to the reference count value and the recording result, the total reference count of the participating calculation data corresponding to each application is respectively obtained. Within a period of time, if the target model is called by the same application multiple times, the total reference count of the application is obtained according to the sum of the reference count values of the participating calculation data corresponding to the multiple calls of the application; Classify the applications according to the total number of references of the participating computing data corresponding to each application, classify all model data of the target model, and store the participating computing data corresponding to different types of applications in different storage resources; According to the total call count of each application calling the large model, the applications are classified and different computing resources are provided to them. The method of classifying the applications according to the total call count of each application calling the large model and providing different computing resources respectively includes: When the target model is called by the application, the call count of the application is performed to obtain the total call count of the application; Classifying the application programs according to the call counts of the application programs to obtain a classification result; the classification result includes a first category of application programs and a second category of application programs, and the total call count of any application program in the first category of application programs is greater than the total call count of any application program in the second category of application programs; The computing resources are allocated according to the classification result so that the computing power provided by the computing resources allocated to the first category of applications is greater than the computing power provided by the computing resources allocated to the second category of applications.
2. The large model resource allocation method for application calls according to claim 1 is characterized in that: The step of classifying the applications according to the total number of references of the participating computing data corresponding to each application, classifying all the model data of the target model, and storing the participating computing data corresponding to different types of applications in different storage resources respectively comprises: The applications are classified according to the total number of citations of the participating computing data corresponding to each application, and all the model data are classified to obtain classified data; the classified data includes first-category data corresponding to the first-category applications and second-category data corresponding to the second-category applications, and the total number of citations of the participating computing data corresponding to the first-category applications corresponding to any model data in the first-category data is greater than the total number of citations of the participating computing data corresponding to the second-category applications corresponding to any model data in the second-category data; storing the first category data in a first category storage resource; storing the second category data in a second category storage resource; Among them, the reading performance corresponding to the first type of storage resources is higher than the reading performance of the second type of storage resources, and the network delay corresponding to the first type of storage resources is lower than the network delay corresponding to the second type of storage resources.
3. The large model resource allocation method for application program calls according to claim 2 is characterized in that: The classifying all model data of the target model includes: All model data of the target model are periodically classified with a first preset time length as a period.
4. The large model resource allocation method for application program calls according to claim 3 is characterized in that: The classified data further includes third-category data, wherein the total number of citations of the participating computing data corresponding to the application corresponding to any model data in the third-category data is less than the total number of citations of the participating computing data corresponding to the application corresponding to any model data in the second-category data, and the method further includes: The third type of data is stored in a third type of storage resource; the read performance corresponding to the second type of storage resource is higher than the read performance of the third type of storage resource, and the network delay corresponding to the second type of storage resource is lower than the network delay corresponding to the third type of data; When the storage time of the third type of data is longer than a second preset time, the third type of data is discarded, and the second preset time is longer than the first preset time.
5. The large model resource allocation method for application program calls according to claim 4 is characterized in that: Also includes: At the end of the cycle, the uncalled reference data is stored as an object, wherein the uncalled reference data is model data that is not referenced when the target model is called by the application program; When the duration of object storage of the uncalled reference data is greater than a second preset duration, the uncalled reference data is discarded.
6. The large model resource allocation method for application program calls according to claim 1 is characterized in that: Before the step of allocating the computing resources according to the classification result, the method includes: According to the classification result, adding a first type of mark to the call entry corresponding to the first type of application, and adding a second type of mark to the call entry corresponding to the second type of application; Correspondingly, the step of allocating the computing resources according to the classification result includes: Corresponding computing resources are provided to the first category of applications and the second category of applications according to the first category of tags and the second category of tags respectively.
7. The large model resource allocation method for application program calls according to claim 1 is characterized in that: Also includes: Applications with a total call count of zero will be removed from the shelves.
8. A large model resource transaction settlement method for application program calls, characterized in that: include: When the target model is called by an application, reference counting is performed on the data involved in the calculation in the target model to obtain the reference count value, and the calling application is recorded to obtain the recording result; wherein the data involved in the calculation is the data involved in the feedforward neural network calculation in the target model, and when the target model ends the call of one application and is called by another application, the reference counting of the data involved in the calculation in the target model is restarted; According to the reference count value and the recording result, the total reference count of the participating calculation data corresponding to each application is respectively obtained. Within a period of time, if the target model is called by the same application multiple times, the total reference count of the application is obtained according to the sum of the reference count values of the participating calculation data corresponding to the multiple calls of the application; Get the calling time of the target model called by the application; In response to the resource settlement request, resource settlement is performed using the call duration and the total number of references as a settlement basis.
9. The large model resource transaction settlement method for application program invocation according to claim 8 is characterized in that: Also includes: The calling duration and the total number of citations are uploaded to the alliance chain.
Citation Information
Patent Citations
Method, device, medium and program product for storage management
CN114924696A