Large model hierarchical storage method and device oriented to Pleashippability intelligent computing center
By adopting a large-model hierarchical storage method in the intelligent computing center, disassembling the user's big model and using hash value to retrieve the index library, the resource waste caused by the intelligent computing center's full storage of the big model is solved, the effect of saving storage space and reducing economic costs is achieved, and the application of universal computing power is promoted.
Patent Information
- Application Number
- CN202510688661.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-27
AI Technical Summary
The large model of intelligent computing centers that fully store users has led to wasting of computing power resources and network traffic, which is economically costly and is difficult to achieve widespread application of universal computing power.
The large-model hierarchical storage method is used to disassemble the user's large model, calculate the hash value of each model hierarchical to retrieve whether there is a corresponding index record in the index library. If it exists, it will not be stored repeatedly. If it does not exist, it will be stored and generated an index record.
Effectively save storage space, reduce the storage and network traffic waste of computing power resources, optimize the computing power resources of the intelligent computing center, reduce economic costs, and promote the widespread application of universal computing power.
Smart Images

Figure CN120216462A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers and computing power infrastructure, and particularly relates to a large model hierarchical storage method and device for an intelligent computing center for inclusive computing power. Background Art
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.
[0003] An "intelligent computing center" refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, and mainly provides the required computing power, data and algorithms for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training and model inference, etc.). The intelligent computing center covers facilities, hardware, software, and can provide full-stack capabilities from the underlying computing power to the top-level application enabling.
[0004] The "intelligent computing center" includes but is not limited to the "intelligent computing center".
[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services and algorithm services required for artificial intelligence applications.
[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers". It is the ability of computer devices or computing / data centers to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of target results by processing information data, and a new type of productive force integrating information computing power, network carrying capacity and data storage capacity, and mainly provides services to the society through computing power infrastructure.
[0007] At present, intelligent computing centers can provide data storage capabilities for storing large models of users of tenant services. The large models are as small as several Gs and as large as hundreds of Gs. Moreover, the large models of a large number of users are large models fine-tuned from basic large models, and more than 90% of the content of the fine-tuned large models is repeated. Storing the large models of users in full will cause waste of storage of computing power resources and network traffic, resulting in high economic costs and making it difficult to achieve the wide application of inclusive computing power. Summary of the Invention
[0008] The present invention provides a large model hierarchical storage method and device for an intelligent computing center for inclusive computing power, which is used to solve the problem that the intelligent computing center stores the large models of users in full, resulting in waste of storage of computing power resources and network traffic, high economic costs, and difficulty in achieving the wide application of inclusive computing power.
[0009] To solve the above technical problems, the present invention is implemented as follows: In a first aspect, the present invention provides a large model hierarchical storage method for an inclusive computing power intelligent computing center, including: Step S1: In response to an instruction to save a large model triggered by a user, disassemble the user's large model to obtain multiple model hierarchies; Step S2: Calculate the hash value of the content of each model hierarchy; Step S3: Retrieve whether there is an index record corresponding to the hash value of the model hierarchy in the index library. If it exists, do not save the model hierarchy; if it does not exist, save the model hierarchy, and generate an index record according to the hash value of the model hierarchy and the storage path of the model hierarchy and store it in the index library; Step S4: Record the index values of all model hierarchies of the large model.
[0010] Optionally, the step S1 includes: Step S11: If the large model is a model obtained by fine-tuning a basic model, disassemble the user's large model to obtain two model hierarchies, and the two model hierarchies include: a basic model parameter layer and a fine-tuning parameter layer; And / or, the step S1 includes: Step S12: Disassemble the user's large model based on the model structure to obtain multiple model hierarchies.
[0011] Optionally, the step S3 includes: Step S32: If there is no index record corresponding to the hash value of the model hierarchy in the index library, save the model hierarchy to a private database, generate an index record corresponding to the hash value of the model hierarchy and store it in the index library, and record the tenant to which the user belongs in the index record; And / or, the step S3 includes: If there is an index record corresponding to the hash value of the model hierarchy in the index library, determine whether the tenant to which the user belongs is recorded in the index record; Step S34: If the tenant to which the user belongs is not recorded in the index record, increment the shared count of the tenant in the index record by 1, and record the tenant to which the user belongs in the index record; determine whether the model hierarchy is saved in the private database. If the model hierarchy is saved in the private database, move the model hierarchy to the public database; Step S35: If the tenant to which the user belongs is already recorded in the index record, there is no need to increment the shared count of the tenant in the index record.
[0012] Optionally, step S3 further includes: Step S36: If there is an index record corresponding to the hash value of the model layer in the index library, increment the user usage count recorded in the index record by 1.
[0013] Optionally, the method further includes: Step S5: Determine the storage cost for each tenant according to the model type, tenant sharing times, and / or file size of the stored model layer; Among them, for the model layer shared by multiple tenants, according to the tenant sharing times of the model layer, the storage cost of the model layer is evenly divided among the multiple tenants; or, for the model layer shared by multiple tenants, if the model layer is a basic model, no storage cost is charged to the tenant for the basic model; And / or, for the model layer private to the tenant, calculate the storage cost of the model layer according to the file size of the model layer.
[0014] Optionally, step S4 includes: Step S41: Use the hash value of the model layer as the index value, and record the index values of all model layers of the large model.
[0015] Optionally, the method further includes: Step S6: In response to an instruction triggered by the user to read the large model, obtain the index values of all model layers of the large model; Step S7: Query the storage path of the model layer from the index library according to the index value of the model layer, and read the model layer according to the storage path; Step S8: Merge all the read model layers to obtain the large model.
[0016] In a second aspect, the present invention provides a large model hierarchical storage device for an inclusive computing power intelligent computing center, including: A disassembly module, configured to disassemble the user's large model in response to an instruction triggered by the user to save the large model, to obtain multiple model layers; A hash value calculation module, configured to calculate the hash value of the content of each model layer; A storage module, configured to retrieve whether there is an index record corresponding to the hash value of the model layer in the index library. If there is, do not save the model layer; if not, save the model layer, and generate an index record according to the hash value of the model layer and the storage path of the model layer, and store it in the index library; A recording module, configured to record the index values of all model layers of the large model.
[0017] In a third aspect, the present invention provides a computing power device, including: a processor, a memory, and a program stored on the memory and executable on the processor. When the program is executed by the processor, it implements the steps of the large model hierarchical storage method for the inclusive computing power intelligent computing center as described in the first aspect above.
[0018] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the large model hierarchical storage method for the inclusive computing power intelligent computing center as described in the first aspect above.
[0019] In a fifth aspect, the present invention provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement the steps of the large model hierarchical storage method for the inclusive computing power intelligent computing center as described in the first aspect above.
[0020] In the present invention, when a user needs to store a large model in the intelligent computing center, the user's large model is disassembled to obtain multiple model layers, the hash values of the contents of each model layer are calculated, and it is retrieved whether there is an index record corresponding to the hash value of the model layer in the index library. If it exists, it means that the content of this model layer has been stored before, so there is no need to store it repeatedly. If it does not exist, it means that the content of this model layer has not been stored before, then this model layer is stored, and an index record is generated according to the hash value of the model layer and the storage path of the model layer and stored in the index library. When the content of the model layer of the large model has been stored before, there is no need to store it repeatedly, that is, there is no need to store the large model in full volume, thus effectively saving storage space, reducing the waste of storage and network traffic of computing power resources, optimizing the computing power resources of the intelligent computing center, greatly reducing the waste of computing power resources, greatly reducing the economic cost, and promoting the wide application of inclusive computing power. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings: Figure 1 is a flowchart of the large model hierarchical storage method for the inclusive computing power intelligent computing center of the present invention; Figure 2 is a structural diagram of the large model hierarchical storage device for the inclusive computing power intelligent computing center of the present invention; Figure 3 is a structural diagram of the computing power device of the present invention. Specific Embodiments
[0022] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts fall within the protection scope of the present invention.
[0023] First, the technical terms related to the present invention will be briefly described below.
[0024] The "computing power" referred to in the present invention means: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to process information data and output a target result, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly providing services to society through computing power infrastructure.
[0025] The "computational power" (Computational Power, CP) referred to in the present invention means: the ability of a data center server to process data and output results, a comprehensive index to measure the computing ability of a data center, including general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 +CP 智能 +CP 超级 .
[0026] The "carrying capacity" (Network Power, NP) referred to in the present invention means: the performance of the data transmission ability of computing power facilities, a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., involving network transmission inside and between data centers, and a comprehensive index to measure network transmission scheduling ability.
[0027] The "Storage Power" (SP) described in the present invention refers to the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and built-in storage devices of servers. The commonly used measurement unit for storage capacity is exabyte (EB, 1EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.
[0028] The "computing power infrastructure" described in the present invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage power, and can realize the centralized computing, storage, transmission, and application of information.
[0029] The "new type of information infrastructure" described in the present invention mainly includes network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, and satellite Internet, computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, and supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.
[0030] The "computing power" described in the present invention includes: general computing power, intelligent computing power, and super computing power.
[0031] The "general computing power" described in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0032] The "intelligent computing power" described in the present invention refers to a computing platform that is scaled for various artificial intelligence innovation applications based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing, machine vision, etc.
[0033] The "super computing power" described in the present invention mainly refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.
[0034] The "Intelligent Computing Center" as described in the present invention refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.), and mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0035] The "Intelligent Computing Center" as described in the present invention includes, but is not limited to, the "Intelligent Computing Center".
[0036] The "Intelligent Computing Center" as described in the present invention, namely the artificial intelligence computing center, is a type of computing power infrastructure that is based on artificial intelligence theory, adopts an artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.
[0037] The "Computing Power Center" as described in the present invention refers to a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, and has computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0038] The "Supercomputing Center" as described in the present invention, namely the supercomputing data center, is a data center based on supercomputers or large-scale computing clusters, and can provide functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.
[0039] The "Computing Power Resources" as described in the present invention refers to technologies and facilities with information computing, transmission, storage, and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPU and GPU, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.
[0040] The "Large Model" as described in the present invention includes, but is not limited to, the "Large Language Model" and the "Multimodal Large Model".
[0041] The "Large Language Model" as described in the present invention refers to a large-scale language model (LLM), which is a language model with a relatively large number of parameters, aiming to understand and generate human language, trained with a large amount of text data, and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.
[0042] The "Multimodal Large Models" described in the present invention refer to models that jointly train multimodal information such as text, images, videos, and audio, including but not limited to multimodal large language models.
[0043] The "inclusive computing power" described in the present invention refers to providing appropriate and effective computing power services to all social strata and groups with computing power service needs at an affordable cost based on the requirements of equal opportunity and the principle of commercial sustainability.
[0044] To solve the problem that the existing large models of intelligent computing centers storing all users' data waste computing power resources, storage, and network traffic, resulting in high economic costs and making it difficult to widely apply inclusive computing power, please refer to Figure 1 , the present invention provides a hierarchical storage method for large models in an intelligent computing center for inclusive computing power, which includes: Step S1: In response to an instruction from a user to save a large model, disassemble the large model of the user to obtain multiple model layers; Step S2: Calculate the hash value of the content of each model layer; In the present invention, the hash algorithm (Hash Algorithm) is used to calculate the hash value of the content of the model layer. The hash algorithm is a function that maps input data of any length to an output of a fixed length. The hash algorithm has the following characteristics: Determinism: The same input will always produce the same output, ensuring that the result is verifiable.
[0045] Irreversibility: It is impossible to deduce the original data from the hash value, ensuring security.
[0046] Collision resistance: The probability that different inputs produce the same output is extremely low, reducing the risk of data tampering.
[0047] Avalanche effect: A small change in the input leads to a significant difference in the output, enhancing sensitivity.
[0048] Efficiency: It can be calculated quickly and is suitable for large-scale data or real-time scenarios.
[0049] Step S3: Retrieve whether there is an index record corresponding to the hash value of the model layer in the index library. If it exists, do not save the model layer; if it does not exist, save the model layer and generate an index record according to the hash value of the model layer and the storage path of the model layer and store it in the index library; In the present invention, multiple index records can be stored in the index library, and each index record includes the hash value of a model layer and the storage path of this model layer.
[0050] For example, the index records stored in the index library can be as follows: index1={Hash1, path1} index2={Hash2, path2} …… Among them, index1={Hash1, path1} is an index record, and index1 is the index value of the index record.
[0051] In the present invention, optionally, the hash value can also be used as the index value, and the index records stored in the index library can be as follows: index1={path1} index2={path2} …… In the present invention, it should be noted that when the hash value is not used as the index value, but the index value is set separately, each hash value uniquely corresponds to an index value.
[0052] In the present invention, optionally, when saving the model layering, it can be stored in a file manner. Optionally, the index record can also include the file size (size) of the model layering.
[0053] In the present invention, optionally, the index record can also include the number of times the model layering is shared by tenants (or referred to as the number of times the tenant references, count), and the number of times the model layering is shared by tenants is used to indicate how many tenants share the model layering.
[0054] For example, the index records stored in the index library can be as follows: index1={path1, size1, count1} index2={path2, size2, count2} …… In the present invention, optionally, the index record can also include the number of times the model layering is used by users, and the number of times the model layering is used by users is used to indicate how many users use the model layering.
[0055] In the present invention, when the index record includes the file size, the number of times the model layering is shared by tenants, and / or the number of times the model layering is used by users, an index record is generated according to the hash value of the model layering, the storage path of the model layering, and the file size, the number of times the model layering is shared by tenants, and / or the number of times the model layering is used by users of the model layering, and the generated index record is stored in the index library.
[0056] It should be noted that in the present invention, a tenant refers to a tenant who leases the computing power resource storage space of the intelligent computing center. There can be multiple tenants, and a user is a user served by a tenant. The user can store a large model in the computing power resource storage space of the tenant serving him / her, and one tenant can serve multiple users.
[0057] Step S4: Record the index values of all model layers of the large model.
[0058] Optionally, the index value of a model layer is the hash value of the model layer.
[0059] For example, a certain large model can be recorded as follows: Model1: {index1, index2}. That is, Model1 is disassembled into two model layers, where the index values of the two model layers are index1 and index2 respectively.
[0060] Optionally, while recording the index values of all model layers of the large model, the storage path of the model layer and / or the file size of the model layer can also be recorded simultaneously.
[0061] For example, a certain large model can be recorded as follows: Model1 { index1= / c: / publish / model_level1,50G index2= / c: / private / model_level2,1G } Among them, / c: / publish / model_level1 is the storage path of the model layer, and 50G is the file size of the model layer.
[0062] In the present invention, when a user needs to store a large model in the intelligent computing center, the user's large model is disassembled to obtain multiple model layers, the hash values of the contents of each model layer are calculated, and it is retrieved whether there is an index record corresponding to the hash value of the model layer in the index library. If it exists, it means that the content of this model layer has been stored before, so there is no need to store it repeatedly. If it does not exist, it means that the content of this model layer has not been stored before, so this model layer is stored, and an index record is generated according to the hash value of the model layer and the storage path of the model layer and stored in the index library. When the content of the model layer of the large model has been stored before, there is no need to store it repeatedly, that is, there is no need to store the large model in full volume, thus effectively saving storage space, reducing the waste of storage and network traffic of computing power resources, optimizing the computing power resources of the intelligent computing center, greatly reducing the waste of computing power resources, greatly reducing economic costs, and promoting the wide application of inclusive computing power.
[0063] In the present invention, optionally, step S1 includes: Step S11: If the large model is a model obtained by fine-tuning a base model, disassemble the user's large model to obtain two model layers, where the two model layers include: a base model parameter layer and a fine-tuning parameter layer.
[0064] The base model refers to a pre-trained large model, such as GPT or BERT, etc.
[0065] Large model fine-tuning refers to a technical process of secondary training based on a pre-trained large model with data in a specific domain or task, so that the model adapts to a specific application scenario.
[0066] In the present invention, the large model can be a large model using LoRA fine-tuning (Low-Rank Adaptation). LoRA fine-tuning is an efficient model fine-tuning technology. Its core idea is to introduce low-rank matrices on the basis of the pre-trained model to reduce the number of parameters required for fine-tuning, thereby improving the training efficiency and avoiding overfitting.
[0067] In the present invention, optionally, step S1 includes: Step S12: Disassemble the user's large model based on the model structure to obtain multiple model layers.
[0068] For example, assuming the large model is a Transformer model architecture, the large model can be divided into two model layers: an encoding component and a decoding component. Or, more specifically, it can be divided into the self-attention layer and the feed-forward network of the encoding component, the self-attention layer, the attention layer, and the feed-forward network of the decoding component, for a total of 5 model layers.
[0069] In the present invention, optionally, step S3 includes: Step S31: Use the hash value of the model layer as the index value, generate an index record according to the index value of the model layer and the storage path of the model layer, and store the generated index record in the index library.
[0070] Using the hash value of the model layer directly as the index value can eliminate the need to introduce a separate index value, and the retrieval of the index library is more convenient.
[0071] In the present invention, optionally, step S3 further includes: Step S32: If there is no index record corresponding to the hash value of the model layer in the index library, save the model layer to the private database, generate an index record corresponding to the hash value of the model layer and store it in the index library, and the index record records the tenant to which the user belongs.
[0072] Optionally, in the present invention, step S3 further includes: Step S33: If there is an index record corresponding to the hash value of the model layering in the index library, determine whether the tenant to which the user belongs is recorded in the index record; Step S34: If the tenant to which the user belongs is not recorded in the index record, increment the tenant sharing count in the index record by 1, and record the tenant to which the user belongs in the index record; determine whether the model layering is stored in the private database. If the model layering is stored in the private database, move the model layering to the public database; Step S35: If the tenant to which the user belongs is already recorded in the index record, there is no need to increment the tenant sharing count in the index record by 1.
[0073] In the present invention, during the above determination process, if it is determined that the tenant to which the user belongs is recorded in the index record, there is no need to increment the tenant sharing count in the index record by 1.
[0074] In the present invention, during the above determination process, if it is determined that the model layering is stored in the public database, there is no need to perform further storage processing.
[0075] Optionally, in the present invention, each tenant has its own private database, and the public database is shared by multiple tenants.
[0076] In the present invention, the intelligent computing center provides a private database and a public database. If there is no index record corresponding to the hash value of a certain model layering in the index library, it means that the model layering has not been stored before, and the model layering can be stored in the private database.
[0077] If there is an index record corresponding to the hash value of a certain model layering in the index library, it means that the model layering has been stored before. At this time, it is also necessary to determine whether the tenant currently storing the model layering is recorded in the index record. If the tenant has been recorded in the index record, there is no need to increase the tenant sharing count of the index record.
[0078] If the tenant is not recorded in the index record, the tenant needs to be recorded in the index record, and the tenant sharing count of the index record is incremented by 1. Then, it is also necessary to determine whether the model layering is currently stored in the private database. If so, the model layering needs to be moved to the public database. If the model layering is already stored in the public database, there is no need to save it again.
[0079] Optionally, in the present invention, step S3 further includes: Step S36: If there is an index record corresponding to the hash value of the model layer in the index library, increment the user usage count in the index record by 1.
[0080] In the present invention, optionally, the method for hierarchical storage of large models for the inclusive computing power intelligent computing center further includes: Step S5: Determine the storage cost for each tenant according to the model type, tenant sharing times, and / or file size of the stored model layer. Among them, for the model layer shared by multiple tenants, according to the tenant sharing times of the model layer, the storage cost of the model layer is evenly divided among the multiple tenants; or, for the model layer shared by multiple tenants, if the model layer is a basic model, no storage cost is charged for the basic model from the tenant. And / or, for the model layer private to the tenant, calculate the storage cost of the model layer according to the file size of the model layer.
[0081] In the present invention, optionally, the model layer private to the tenant refers to the model layer stored in a private database.
[0082] In the present invention, the large model storage platform provided by the intelligent computing center can have multiple tenants, and each tenant serves multiple users. For each tenant, the storage cost of the large model stored by it can be calculated.
[0083] In some examples, it is not excluded to calculate the storage cost for each user.
[0084] In the present invention, optionally, the step S4 includes: Step S41: Use the hash value of the model layer as the index value, and record the index values of all model layers of the large model.
[0085] In the present invention, each large model can have a unique identifier, such as Model1. The intelligent computing center can separately provide an identifier library for recording the index values of all model layers of the large model.
[0086] For example, the record corresponding to a certain large model is as follows: Model1 {index1= / c: / publish / model_level1,50G index2= / c: / private / model_level2,1G } In the present invention, optionally, the method for hierarchical storage of large models for the inclusive computing power intelligent computing center further includes: Step S6: In response to the instruction triggered by the user to read the large model, obtain the index values of all model layers of the large model; Step S7: According to the index values of the model layers, query the storage paths of the model layers from the index library, and read the model layers according to the storage paths; Step S8: Merge all the read model layers to obtain the large model.
[0087] For example, if the user needs to use a certain stored large model, the identifier of the large model (such as Model1) can be used to retrieve the large model in the identifier library, and the index values of all model layers of the large model can be read from the record corresponding to the identifier of the large model in the identifier library. For example, if the large model includes two model layers, the index values are index1 and index2 respectively.
[0088] In some examples, if the storage paths of the model layers are stored in the identifier library, the model layers can be directly read according to the storage paths.
[0089] In some examples, if the storage paths of the model layers are not stored in the identifier library, it is necessary to retrieve the index library according to the index values of each model layer to obtain the storage paths of the model layers stored in the index library.
[0090] Please refer to Figure 2 , the present invention also provides a large model hierarchical storage device 10 for a general-purpose computing power intelligent computing center, including: A disassembling module 11, configured to disassemble the user's large model in response to the instruction triggered by the user to save the large model, to obtain a plurality of model layers; A hash value calculation module 12, configured to calculate the hash value of the content of each model layer; A storage module 13, configured to retrieve whether there is an index record corresponding to the hash value of the model layer in the index library. If it exists, do not save the model layer; if it does not exist, save the model layer, and generate an index record according to the hash value of the model layer and the storage path of the model layer, and store it in the index library; A recording module 14, configured to record the index values of all model layers of the large model.
[0091] In the present invention, when a user needs to store a large model in an intelligent computing center, the user's large model is disassembled to obtain multiple model layers. The hash values of the contents of each model layer are calculated, and it is retrieved whether there is an index record corresponding to the hash value of the model layer in the index library. If it exists, it means that the content of this model layer has been stored before, so there is no need to store it repeatedly. If it does not exist, it means that the content of this model layer has not been stored before, so this model layer is stored, and an index record is generated according to the hash value of the model layer and the storage path of the model layer and stored in the index library. When the content of the model layer of the large model has been stored before, there is no need to store it repeatedly, that is, there is no need to store the large model in full, thereby effectively saving storage space, reducing waste of storage and network traffic, optimizing the computing power resources of the intelligent computing center, greatly reducing waste of computing power resources, greatly reducing economic costs, and promoting the wide application of inclusive computing power.
[0092] In the present invention, optionally, the disassembling module 11 is configured to, if the large model is a model obtained by fine-tuning a basic model, disassemble the user's large model to obtain two model layers, and the two model layers include: a basic model parameter layer and a fine-tuning parameter layer; In the present invention, optionally, the disassembling module 11 is configured to disassemble the user's large model based on the model structure to obtain multiple model layers.
[0093] In the present invention, optionally, the storage module 13 is configured to use the hash value of the model layer as an index value, generate an index record according to the index value of the model layer and the storage path of the model layer, and store the generated index record in the index library.
[0094] In the present invention, optionally, the storage module 13 is configured to, if there is no index record corresponding to the hash value of the model layer in the index library, save the model layer to a private database, generate an index record corresponding to the hash value of the model layer and store it in the index library, and the tenant to which the user belongs is recorded in the index record; In the present invention, optionally, the storage module 13 is configured to, if there is an index record corresponding to the hash value of the model layer in the index library, determine whether the tenant to which the user belongs is recorded in the index record; if the tenant to which the user belongs is not recorded in the index record, increment the sharing times of the tenant in the index record by 1, and record the tenant to which the user belongs in the index record; determine whether the model layer is saved in the private database, and if the model layer is saved in the private database, move the model layer to the public database; if the tenant to which the user belongs has already been recorded in the index record, there is no need to increment the sharing times of the tenant in the index record by 1.
[0095] In the present invention, optionally, the storage module 13 is configured to, if there is an index record corresponding to the hash value of the model layer in the index library, increment by 1 the number of user usages recorded in the index record.
[0096] In the present invention, optionally, the large model layer storage device 10 for the inclusive computing power intelligent computing center further includes: A charging module, configured to determine the storage cost of each tenant according to the model type, tenant sharing times, and / or file size of the stored model layer; Among them, for the model layer shared by multiple tenants, according to the tenant sharing times of the model layer, the storage cost of the model layer is evenly divided among the multiple tenants; or, for the model layer shared by multiple tenants, if the model layer is a basic model, no storage cost for the basic model is charged to the tenant; And / or, for the model layer private to the tenant, the storage cost of the model layer is calculated according to the file size of the model layer.
[0097] In the present invention, optionally, the recording module 14 is configured to use the hash value of the model layer as an index value and record the index values of all model layers of the large model.
[0098] In the present invention, optionally, the large model layer storage device 10 for the inclusive computing power intelligent computing center further includes: An acquisition module, configured to, in response to an instruction triggered by a user to read the large model, acquire the index values of all model layers of the large model; A query module, configured to query the storage path of the model layer from the index library according to the index value of the model layer, and read the model layer according to the storage path; A merging module, configured to merge all the read model layers to obtain the large model.
[0099] Please refer to Figure 3 , the present invention also provides a computing power device 20, including a processor 21, a memory 22, and a computer program stored on the memory 22 and executable on the processor 21. When the computer program is executed by the processor 21, it implements each process of the above-mentioned embodiment of the large model layer storage method for the inclusive computing power intelligent computing center, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0100] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above-mentioned embodiment of the large model hierarchical storage method for the inclusive computing power intelligent computing center, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0101] The present invention also provides a computer program product, including computer instructions, which when executed by a processor implement the above-mentioned Figure 1 each process of the embodiment of the large model hierarchical storage method for the inclusive computing power intelligent computing center shown, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0102] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including that element.
[0103] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in various embodiments of the present invention.
[0104] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose of the present invention and the scope protected by the claims, and all of them belong to the protection scope of the present invention.
Claims
1. A large model hierarchical storage method for an inclusive computing power intelligent computing center, characterized in that, Including: Step S1: In response to an instruction triggered by the user to save the large model, disassemble the user's large model to obtain multiple model layers; Step S2: Calculate the hash value of the content of each model layer; Step S3: Retrieve whether there is an index record corresponding to the hash value of the model layer in the index library. If it exists, do not save the model layer; If not, save the model layer, and generate an index record according to the hash value of the model layer and the storage path of the model layer and store it in the index library; Step S4: Record the index values of all model layers of the large model.
2. The method according to claim 1, characterized in that The step S1 includes: Step S11: If the large model is a model obtained by fine-tuning the base model, disassemble the user's large model to obtain two model layers, and the two model layers include: the base model parameter layer and the fine-tuning parameter layer; And / or, the step S1 includes: Step S12: Disassemble the user's large model based on the model structure to obtain multiple model layers.
3. The method according to claim 1, characterized in that The step S3 includes: Step S31: Use the hash value of the model layer as the index value, generate an index record according to the index value of the model layer and the storage path of the model layer, and store the generated index record in the index library.
4. The method according to claim 1, wherein The step S3 further includes: Step S32: If there is no index record corresponding to the hash value of the model layer in the index library, save the model layer to the private database, generate an index record corresponding to the hash value of the model layer and store it in the index library, and record the tenant to which the user belongs in the index record; And / or, the step S3 further includes: Step S33: If there is an index record corresponding to the hash value of the model layer in the index library, determine whether the tenant to which the user belongs is recorded in the index record; Step S34: If the tenant to which the user belongs is not recorded in the index record, increment the tenant sharing times in the index record by 1, and record the tenant to which the user belongs in the index record; determine whether the model layer is saved in the private database. If the model layer is saved in the private database, move the model layer to the public database; Step S35: If the tenant to which the user belongs has been recorded in the index record, there is no need to increment the tenant sharing times in the index record by 1.
5. The method according to claim 4, characterized in that The step S3 further includes: Step S36: If there is an index record corresponding to the hash value of the model layer in the index library, increment the user usage times recorded in the index record by 1.
6. The method according to claim 4, characterized in that It further includes: Step S5: Determine the storage cost of each tenant according to the model type, tenant sharing times, and / or file size of the stored model layer; Wherein, for the model layer shared by multiple tenants, according to the tenant sharing times of the model layer, the storage cost of the model layer is evenly shared by the multiple tenants; or, for the model layer shared by multiple tenants, if the model layer is the base model, no storage cost is charged to the tenant; And / or, for the model layering private to the tenant, calculate the storage cost of the model layering according to the file size of the model layering.
7. The method according to claim 1, characterized in that, The step S4 includes: Step S41: Use the hash value of the model layering as the index value, and record the index values of all model layerings of the large model.
8. The method according to claim 1, characterized in that It further includes: Step S6: In response to an instruction for reading a large model triggered by a user, obtain the index values of all model layerings of the large model; Step S7: According to the index values of the model layering, query the storage path of the model layering from the index library, and read the model layering according to the storage path; Step S8: Merge all the read model layerings to obtain the large model.
9. A large model hierarchical storage device for an inclusive computing power intelligent computing center, characterized in that, It includes: A disassembling module, configured to disassemble the large model of the user in response to an instruction for saving a large model triggered by the user, to obtain a plurality of model layerings; A hash value calculation module, configured to calculate the hash value of the content of each model layering; A storage module, configured to retrieve whether there is an index record corresponding to the hash value of the model layering in the index library. If it exists, do not save the model layering; If it does not exist, save the model layering, and generate an index record according to the hash value of the model layering and the storage path of the model layering, and store it in the index library; A recording module, configured to record the index values of all model layerings of the large model.
10. A computing power device, characterized in that, It includes: A processor, a memory, and a program stored on the memory and executable on the processor. When the program is executed by the processor, it implements the steps of the large model layering storage method for a general-purpose computing power intelligent computing center as described in any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements the steps of the large model layering storage method for a general-purpose computing power intelligent computing center as described in any one of claims 1 to 8.
12. A computer program product, characterized in that, It includes computer instructions. When the computer instructions are executed by a processor, it implements the steps of the large model layering storage method for a general-purpose computing power intelligent computing center as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Image blocking deduplication method and system based on image edge detection
CN112200740A
Metadata management method and system based on multi-tenant SaaS architecture and electronic equipment
CN112860948A
Data processing method and storage equipment
CN113806341A
Multi-party sharing system and method for deep learning large model
CN116611471A
Model reasoning method and device based on calculation unit deployment, equipment and medium
CN117494816A